Search PubMed⌕ Search

Biomedical subjects

Youngjo Lee

Publications and source records attributed to Youngjo Lee.

9 recordsLinked to original sources

Sparse Logistic Regression on Genomic Data for Prediction of Tumour Pathological Subtype.

The correct prediction of tumour subtype is critical for the treatment of cancer patients to maximise the chance of survival. The patients' genomic information, such as copy number alterations (CNA) profile, has increasingly become an important factor in the prediction to supplement the traditional pathological subtyping. The incorporation of the CNA information in a prediction model, such as logistic regression, faces two major statistical challenges: first, how to estimate the model parameters in the thousands and, second, how to deal with the correlation of CNA between genomic regions. To address them, we propose a sparse logistic regression model with random effects where some of its parameters are estimated to zero while the other parameters are non-zero. In effect, a variable selection is embedded in the modelling. To deal with the correlation of CNA across genomic regions, we extend further the model to incorporate an additional penalty in the corresponding likelihood function in the logistic regression. The results show that we can identify selected genomic regions that are informative to distinguish different tumour subtypes, while giving a good prediction ability. We illustrate the methodology using CNA dataset from a lung cancer cohort.

Journal Article↗

Robust estimation in mixed linear models with non-monotone missingness.

We introduce a model to account for abrupt changes among repeated measures with non-monotone missingness. Development of likelihood inferences for such models is hard because it involves intractable integration to obtain the marginal likelihood. We use hierarchical likelihood to overcome such difficulty. Abrupt changes among repeated measures can be well described by introducing random effects in the dispersion. A simulation study shows that the resulting estimator is efficient, robust against misspecification of fatness of tails. For illustration we use a schizophrenic behaviour data presented by Rubin and Wu.

Computer Simulation↗

Determinants of hospital closure in South Korea: use of a hierarchical generalized linear model.

Understanding causes of hospital closure is important if hospitals are to survive and continue to fulfill their missions as the center for health care in their neighborhoods. Knowing which hospitals are most susceptible to closure can be of great use for hospital administrators and others interested in hospital performance. Although prior studies have identified a range of factors associated with increased risk of hospital closure, most are US-based and do not directly relate to health care systems in other countries. We examined determinants of hospital closure in a nationally representative sample: 805 hospitals established in South Korea before 1996 were examined-hospitals established in 1996 or after were excluded. Major organizational changes (survival vs. closure) were followed for all South Korean hospitals from 1996 through 2002. With the use of a hierarchical generalized linear model, a frailty model was used to control correlation among repeated measurements for risk factors for hospital closure. Results showed that ownership and hospital size were significantly associated with hospital closure. Urban hospitals were less likely to close than rural hospitals. However, the urban location of a hospital was not associated with hospital closure after adjustment for the proportion of elderly. Two measures for hospital competition (competitive beds and 1-Hirshman--Herfindalh index) were positively associated with risk of hospital closure before and after adjustment for confounders. In addition, annual 10% change in competitive beds was significantly predictive of hospital closure. In conclusion, yearly trends in hospital competition as well as the level of hospital competition each year affected hospital survival. Future studies need to examine the contribution of internal factors such as management strategies and financial status to hospital closure in South Korea.

Economics, Hospital↗

Dispersion frailty models and HGLMs.

In medical research recurrent event times can be analysed using a frailty model in which the frailties for different individuals are independent and identically distributed. However, such a homogeneous assumption about frailties could sometimes be suspect. For modelling heterogeneity in frailties we describe dispersion frailty models arising from a new class of models, namely hierarchical generalized linear models. Using the kidney infection data we illustrate how to detect and model heterogeneity among frailties. Stratification of frailty models is also investigated.

Adult↗

HGLM versus conditional estimators for the analysis of clustered binary data.

Clustered binary data arise frequently in medical research such as cross-over clinical trials and twin studies. For the analysis of such data either a random-effects model or a conditional likelihood approach can be used. In this paper, we compare numerically the random-effects model estimator and the conditional likelihood estimator and discuss their relative merits for the analysis of binary data.

Age Factors↗

Robust ascertainment-adjusted parameter estimation.

Nonrandom ascertainment is commonly used in genetic studies of rare diseases, since this design is often more convenient than the random-sampling design. When there is an underlying latent heterogeneity, Epstein et al. ([2002] Am. J. Hum. Genet. 70:886-895) showed that it is possible to get unbiased or consistent estimation of population parameters under ascertainment adjustment, but Glidden and Liang ([2002] Genet. Epidemiol. 23:201-208) showed in a simulation study that the resulting estimates are highly sensitive to misspecification of the latent components. To overcome this difficulty, we consider a heavy-tailed model for latent variables that allows a robust estimation of the parameters. We describe a hierarchical-likelihood approach that avoids the integration used in the standard marginal likelihood approach. We revisit and extend the previous simulation, and show that the resulting estimator is efficient and robust against misspecification of the distribution of latent variables.

Genetic Diseases, Inborn↗

Multilevel mixed linear models for survival data.

For the analysis of correlated survival data mixed linear models are useful alternatives to frailty models. By their use the survival times can be directly modelled, so that the interpretation of the fixed and random effects is straightforward. However, because of intractable integration involved with the use of marginal likelihood the class of models in use has been severely restricted. Such a difficulty can be avoided by using hierarchical-likelihood, which provides a statistically efficient and fast fitting algorithm for multilevel models. The proposed method is illustrated using the chronic granulomatous disease data. A simulation study is carried out to evaluate the performance.

Case-Control Studies↗

Analysis of ulcer data using hierarchical generalized linear models.

In multi-centre clinical trials, heterogeneities in individual hospital treatment effects can be modelled as random effects. Estimates of the individual hospital treatment effects and estimate of the mean treatment effect, allowing for the presence of overall hospital differences, are required, together with some measure of their uncertainty. Systematic inferences from the hierarchical-likelihood are now possible, using hierarchical generalized linear models. We show how to construct profile likelihoods for the treatment effects of individual hospitals.

Data Interpretation, Statistical↗

Hierarchical-likelihood approach for mixed linear models with censored data.

Mixed linear models describe the dependence via random effects in multivariate normal survival data. Recently they have received considerable attention in the biomedical literature. They model the conditional survival times, whereas the alternative frailty model uses the conditional hazard rate. We develop an inferential method for the mixed linear model via Lee and Nelder's (1996) hierarchical-likelihood (h-likelihood). Simulation and a practical example are presented to illustrate the new method.

Animals↗