Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “High-dimensional regression”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

59 records · Page 4Linked to original sources

Boosting proportional hazards models using smoothing splines, with applications to high-dimensional microarray data.

MOTIVATION: An important area of research in the postgenomics era is to relate high-dimensional genetic or genomic data to various clinical phenotypes of patients. Due to large variability in time to certain clinical events among patients, studying possibly censored survival phenotypes can be more informative than treating the phenotypes as categorical variables. Due to high dimensionality and censoring, building a predictive model for time to event is more difficult than the classification/linear regression problem. We propose to develop a boosting procedure using smoothing splines for estimating the general proportional hazards models. Such a procedure can potentially be used for identifying non-linear effects of genes on the risk of developing an event. RESULTS: Our empirical simulation studies showed that the procedure can indeed recover the true functional forms of the covariates and can identify important variables that are related to the risk of an event. Results from predicting survival after chemotherapy for patients with diffuse large B-cell lymphoma demonstrate that the proposed method can be used for identifying important genes that are related to time to death due to cancer and for building a parsimonious model for predicting the survival of future patients. In addition, there is clear evidence of non-linear effects of some genes on survival time.

Algorithms↗

Is breathing in infants chaotic? Dimension estimates for respiratory patterns during quiet sleep.

We describe an analysis of dynamic behavior apparent in times-series recordings of infant breathing during sleep. Three principal techniques were used: estimation of correlation dimension, surrogate data analysis, and reduced linear (autoregressive) modeling (RARM). Correlation dimension can be used to quantify the complexity of time series and has been applied to a variety of physiological and biological measurements. However, the methods most commonly used to estimate correlation dimension suffer from some technical problems that can produce misleading results if not correctly applied. We used a new technique of estimating correlation dimension that has fewer problems. We tested the significance of dimension estimates by comparing estimates with artificial data sets (surrogate data). On the basis of the analysis, we conclude that the dynamics of infant breathing during quiet sleep can best be described as a nonlinear dynamic system with large-scale, low-dimensional and small-scale, high-dimensional behavior; more specifically, a noise-driven nonlinear system with a two-dimensional periodic orbit. Using our RARM technique, we identified the second period as cyclic amplitude modulation of the same period as periodic breathing. We conclude that our data are consistent with respiration being chaotic.

Algorithms↗

Supervised dimension reduction of intrinsically low-dimensional data.

High-dimensional data generated by a system with limited degrees of freedom are often constrained in low-dimensional manifolds in the original space. In this article, we investigate dimension-reduction methods for such intrinsically low-dimensional data through linear projections that preserve the manifold structure of the data. For intrinsically one dimensional data, this implies projecting to a curve on the plane with as few intersections as possible. We are proposing a supervised projection pursuit method that can be regarded as an extension of the single-index model for nonparametric regression. We show results from a toy and two robotic applications.

Journal Article↗

A framework for block-wise missing data in multi-omics.

High-throughput technologies have generated vast amounts of omic data. It is a consensus that the integration of diverse omics sources improves predictive models and biomarker discovery. However, managing multiple omics data poses challenges such as data heterogeneity, noise, high-dimensionality and missing data, especially in block-wise patterns. This study addresses the challenges of high dimensionality and block-wise missing data through a regularization and constrained-based approach. The methodology is implemented in the R package bwm for binary and continuous response variables, and applied to breast cancer and exposome multi-omics datasets, achieving strong performance even in scenarios with missing data present in all omics. In binary classification task, our proposed model achieves accuracy in the range of 86% to 92%, and F1 in the range of 68% to 79%. And, in regression task the correlation between true and predicted responses is in the range of 72% to 76%. However, there is a slight decline in performance metrics as the percentage of missing data increases. In scenarios where block-wise missing data affects multiple omics, the model performance actually surpasses that of scenarios where missing data is present in only one omics. One possible explanation for this might be that the other scenarios introduce a greater diversity of observation profiles, leading to a more robust model. Depending on the specific omics being studied, there is greater consistency in feature selection when comparing block-wise missing data scenarios.

Humans↗

Research issues and strategies for genomic and proteomic biomarker discovery and validation: a statistical perspective.

The development and validation of clinically useful biomarkers from high-dimensional genomic and proteomic information pose great research challenges. Present bottlenecks include: that few of the biomarkers showing promise in initial discovery were found to warrant subsequent validation; and biomarker validation is expensive and time consuming. Biomarker evaluation should proceed in an orderly fashion to enhance rigor and efficiency. A molecular profiling approach, although promising, has a high chance of yielding biased results and overfitted models. Specimens from cohorts or intervention trials are essential to eliminate biases. The high cost for biomarker validation motivates some novel study design features, including sequential filtering and DNA pooling. For data analysis, logistic regression (in particular, boosting logistic regression) has features of robustness against model misspecification, and has resistance to model overfitting. Model assessment and cross-validation are critical components of data analysis. Having an independent test set is a vital feature of study design.

Biomarkers, Tumor↗