Search PubMed⌕ Search

Biomedical subjects

Wing K Fung

Publications and source records attributed to Wing K Fung.

At least 19 recordsLinked to original sources

Assessing local influence in principal component analysis with application to haematology study data.

In many medical and health studies, high-dimensional data are often encountered. Principal component analysis (PCA) is a commonly used technique to reduce such data to a few components that includes most of the information provided by the original data. However, PCA is known to be very sensitive to some abnormal observations. Therefore, it is essential to assess such sensitivity in PCA. In this paper, the assessments of local influence based on generalized influence function are developed under the case-weights and additive perturbation schemes, along with a discussion of the perturbation scheme and the generalized influence function approach. When perturbing different variables of the data, it is noted that the directions of the largest joint local influence for the eigenvalues are all the same. Moreover, these directions are completely determined by the score values of the observations, to which an approximate cut-off point is given. The proposed methods are applied to analyse a set of haematology study data for illustration. Results add new insights in finding influential observations in the studied data set.

Health Status↗

An extension of the transmission disequilibrium test incorporating imprinting.

The recombination rates in meioses of females and males are often different. Some genes that affect development and behavior in mammals are known to be imprinted, and >1% of all mammalian genes are believed to be imprinted. When the gene is imprinted and the recombination fractions are sex specific, the conventional transmission disequilibrium test (TDT) is shown to be still valid for testing for linkage. The power function of the TDT is derived, and the effect of the degree of imprinting on the power of the TDT is investigated. It is learned that imprinting has little effect on the power when the female and male recombination rates are equal. On the basis of case-parents trios, the transmissions from the heterozygous fathers/mothers to their affected children are separated as paternal and maternal, and two TDT-like statistics, TDT(p) and TDT(m), are consequently constructed. It is found that the TDT(p) possesses a higher power than the TDT for maternal imprinting genes, and the TDT(m) is more powerful than the TDT for paternal imprinting genes. On the basis of the parent-of-origin effects test statistic (POET), a novel statistic, TDT incorporating imprinting (TDTI) is proposed to test for linkage in the presence of linkage disequilibrium, which is shown to be more powerful than the TDT when parent-of-origin effects are significant but slightly less powerful than the TDT when parent-of-origin effects are negligible. The validity of the TDT and TDTI is assessed by simulation. The power approximation formulas for the TDT and TDTI are derived and the simulation results show that they are accurate. The simulation study on power comparison shows that the TDTI outperforms the TDT for imprinted genes. The improvement can be substantial in the case of complete paternal/maternal imprinting.

Animals↗

On statistical analysis of forensic DNA: theory, methods and computer programs.

Statistics plays an important role in evaluating the evidential weight of forensic DNA. In this paper, general statistical principles for forensic DNA analysis are presented. We introduce the theory and methods for the statistical assessment in kinship determination and DNA mixture evaluation. In particular, analytical formulas for testing for biological relationship among three individuals and for assessing the DNA mixture evidence in the case of multiple subdivided ethnic groups are developed. Two user-friendly computer programs are demonstrated to exhibit their wide applicability in tackling with complex kinship/paternity and mixture problems. The EasyDNA program can solve a complicated paternity case in 1 min.

DNA↗

On the formula for admixture linkage disequilibrium.

Admixture linkage disequilibrium (ALD), a phenomenon created by gene flow between genetically distinct populations, has for some time been used as a tool in gene mapping. It is therefore important to analyze the pattern of ALD over generations. In this study we explore two models of admixture: the gradual admixture (GA) model, in which admixture occurs at a variable rate in every generation; and the immediate admixture (IA) model, a special case of the GA model, in which admixture occurs in a single generation. In the case of ALD, the well-known formula of linkage disequilibrium (Delta(t)=(1-r)t Delta(0)) is not applicable under these two models. We note the effect of a random mating population (RMP) to the gametic frequencies from the parental population to the offspring population, and provide the correct formula for ALD.

Chromosome Mapping↗

Influence diagnostics for two-component Poisson mixture regression models: applications in public health.

In many medical and health applications, Poisson mixture regression models are commonly used to analyse heterogeneous count data. Motivated by two data sets drawn from public health studies, influence diagnostics are proposed for assessing the sensitivity of the fitted two-component Poisson mixture regression models. Under various perturbations of the observed data or model assumptions, influence assessments based on the local influence approach are developed for detecting clusters and/or individual observations that impact on the estimation of model parameters. Results from studies on recurrent urinary tract infections and maternity length of stay illustrate the usefulness of the influence diagnostics.

Cluster Analysis↗

Evaluation of DNA mixtures involving two pairs of relatives.

This paper considers the statistical evaluation of DNA mixtures in the following situations: (1) two unknown contributors are related respectively to two typed persons, (2) two of the unknown or untyped contributors are related and the third unknown contributor is related to a typed person, or (3) there are two pairs of related unknown contributors to the DNA mixture. The corresponding formulas for evaluating the likelihood ratios on the strength of DNA evidence are derived and the kinship coefficients for the related persons are incorporated into the calculations. Two examples are analyzed for illustration.

DNA Fingerprinting↗

Detecting differentially expressed genes by relative entropy.

DNA microarray experiments have generated large amount of gene expression measurements across different conditions. One crucial step in the analysis of these data is to detect differentially expressed genes. Some parametric methods, including the two-sample t-test (T-test) and variations of it, have been used. Alternatively, a class of non-parametric algorithms, such as the Wilcoxon rank sum test (WRST), significance analysis of microarrays (SAM) of Tusher et al. (2001), the empirical Bayesian (EB) method of Efron et al. (2001), etc., have been proposed. Most available popular methods are based on t-statistic. Due to the quality of the statistic that they used to describe the difference between groups of data, there are situations when these methods are inefficient, especially when the data follows multi-modal distributions. For example, some genes may display different expression patterns in the same cell type, say, tumor or normal, to form some subtypes. Most available methods are likely to miss these genes. We developed a new non-parametric method for selecting differentially expressed genes by relative entropy, called SDEGRE, to detect differentially expressed genes by combining relative entropy and kernel density estimation, which can detect all types of differences between two groups of samples. The significance of whether a gene is differentially expressed or not can be estimated by resampling-based permutations. We illustrate our method on two data sets from Golub et al. (1999) and Alon et al. (1999). Comparing the results with those of the T-test, the WRST and the SAM, we identified novel differentially expressed genes which are of biological significance through previous biological studies while they were not detected by the other three methods. The results also show that the genes selected by SDEGRE have a better capability to distinguish the two cell types.

Acute Disease↗

Combining the case-control methodology with the small size transmission/disequilibrium test for multiallelic markers.

Case-control studies compare marker-allele distributions in affected and unaffected individuals, and significant results may be due to linkage but can also simply reflect population structure. To test for linkage after obtaining a significant case-control finding, within-family analysis can be performed. In a transmission/disequilibrium test (TDT), genotypes of cases are compared to those of their parents to explore whether a specific allele, or marker, at a locus of interest is transmitted to a greater degree than Mendelian inheritance would warrant. For multiallelic markers, several authors have proposed extensions to the TDT. In this article, we propose a TDT test, utilizing the available information of a case-control study in the grouping of alleles for multiallelic markers, and thereby increase the statistical power of a TDT test with a small sample size.

Case-Control Studies↗

Sensitivity of score tests for zero-inflation in count data.

In many biomedical applications, count data have a large proportion of zeros and the zero-inflated Poisson regression (ZIP) model may be appropriate. A popular score test for zero-inflation, comparing the ZIP model to a standard Poisson regression model, was given by van den Broek. Similarly, for count data that exhibit extra zeros and are simultaneously overdispersed, a score test for testing the ZIP model against a zero-inflated negative binomial alternative was proposed by Ridout, Hinde and Demétrio. However, these test statistics are sensitive to anomalous cases in the data, and incorrect inferences concerning the choice of model may be drawn. In this paper, diagnostic measures are derived to assess the influence of observations on the score statistics. Two examples that motivated the application of zero-inflated regression models are considered to illustrate the importance of sensitivity analysis of the zero-inflation tests.

Accidents, Occupational↗

Full siblings impersonating parent/child prove most difficult to discredit with DNA profiling alone.

DNA profiling is currently the most widely used method for parentage verification, although many forms of it have limitations of some sort. In this paper, a general formula is derived to depict a simple relationship between the probability that a random man and the probably that a male relative of the child, other than the child's father, is excluded from paternity, when the phenotype of the child's mother is unavailable. With this, the possible limitations of a finite set of STR loci in excluding close relatives of the child from paternity are illustrated. Genetically, among the commonly encountered biologic relationships, to exclude a full sibling of the child from paternity if they pose themselves as father and child remains the most difficult.

China↗

Median regression for longitudinal data.

We review and compare three estimators of median regression in linear models with longitudinal data. The estimators are constructed based on well-known ideas of weighting, decorrelating, and the working assumption of independence. Both asymptotic efficiency calculations and finite-sample Monte Carlo studies are used to assess the performance of these estimators. We find that their relative performances depend on the nature of covariates. The estimator under the working assumption of independence is computationally simple and yet has good relative performance when the covariates are invariant over time or when the within-subject correlations are small. Its relative performance in finite samples is also found to be more favourable than suggested by the asymptotic comparisons.

Analgesics↗

User-friendly programs for easy calculations in paternity testing and kinship determinations.

Four computer programs are developed for easy calculations of likelihood ratios (LRs) for paternity testing and kinship determinations. The programs are constructed based on the concepts of Bayes Theorem, conditional probability and pedigree analysis. Computer enumeration is used for handling the calculations. The programs have wide applicability, and users can save and check the results easily. The programs employ the distinctive pull-down manual, making them very easy to use. In this paper, we explain the theory and describe various features of the software.

Alleles↗

Testing for kinship in a subdivided population.

The effect of population subdivision on the determination of kinship of any two persons is investigated in this paper. Expressions of the joint genotype probabilities and likelihood ratios on kinship testing are reported. Two real cases are analysed using the Hong Kong Chinese population data and the Spanish data. Various kinds of relationships are investigated for illustration.

Child↗

Evaluating forensic DNA mixtures with contributors of different structured ethnic origins: a computer software.

The effect of a structured population on the likelihood ratio of a DNA mixture has been studied by the current authors and others. In practice, contributors of a DNA mixture may belong to different ethnic/racial origins, a situation especially common in multi-racial countries such as the USA and Singapore. We have developed a computer software which is available on the web for evaluating DNA mixtures in multi-structured populations. The software can deal with various DNA mixture problems that cannot be handled by the methods given in a recent article of Fung and Hu.

Alleles↗

Analyzing hospital length of stay: mean or median regression?

BACKGROUND: Length of stay (LOS) is an important measure of hospital activity and health care utilization, but its empirical distribution is often positively skewed. OBJECTIVE: This study reviews the mean and median regression approaches for analyzing LOS, which have implications for service planning, resource allocation, and bed utilization. METHODS: The two approaches are applied to analyze hospital discharge data on cesarean delivery. Both models adjust for patient and health-related characteristics, and for the dependency of LOS outcomes nested within hospitals. The estimation methods are also compared in a simulation study. RESULTS: For the empirical application, the mean regression results are somewhat sensitive to the magnitude of trimming chosen. The identified factors from median regression, namely number of diagnoses, number of procedures, and payment classification, are robust to high-LOS outliers. The simulation experiment shows that median regression can outperform mean regression even when the response variable is moderately positively skewed. CONCLUSION: Median regression appears to be a suitable alternative to analyze the clustered and positively skewed LOS, without transforming and trimming the data arbitrarily.

Cesarean Section↗