Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “High-dimensional regression”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Sparse polygenic risk score inference with the spike-and-slab LASSO.

MOTIVATION: Large-scale biobanks, with rich phenotypic and genomic data across hundreds of thousands of samples, provide ample opportunities to elucidate the genetics of complex traits and diseases. Consequently, there is growing demand for robust and scalable methods for disease risk prediction from genotype data. Inference in this setting is challenging due to the high-dimensionality of genomic data, especially when coupled with smaller sample sizes. Popular Polygenic Risk Score (PRS) inference methods address this challenge by adopting sparse Bayesian priors or penalized regression techniques, such as the Least Absolute Shrinkage and Selection Operator (LASSO). However, the former class of methods are not as scalable and do not produce exact sparsity, while the latter tends to over-shrink large coefficients. RESULTS: In this study, we present SSLPRS, a novel PRS method based on the Spike-and-Slab LASSO (SSL) prior, which offers a theoretical bridge between the two frameworks. We extend previous work to derive a coordinate-ascent inference algorithm that operates on GWAS summary statistics, which is orders-of-magnitude more efficient than corresponding individual-level-based implementations. To illustrate the statistical properties of the proposed model, we conducted experiments involving nine simulation configurations and nine quantitative phenotypes from the UK Biobank. Our results demonstrate that SSLPRS is competitive with state-of-the-art methods in terms of prediction accuracy and exhibits superior variable selection performance, especially in sparse genetic architectures. In simulations, this translates to upwards of 50% improvement in positive predictive value. In analysis of real phenotypes, we show that selected variants are highly enriched for meaningful genomic annotations and have better replication rates in larger meta-analyses. AVAILABILITY AND IMPLEMENTATION: SSLPRS is available in the open-source package https://github.com/li-lab-mcgill/penprs.

Multifactorial Inheritance↗

Using nuclear morphometry to discriminate the tumorigenic potential of cells: a comparison of statistical methods.

Despite interest in the use of nuclear morphometry for cancer diagnosis and prognosis as well as to monitor changes in cancer risk, no generally accepted statistical method has emerged for the analysis of these data. To evaluate different statistical approaches, Feulgen-stained nuclei from a human lung epithelial cell line, BEAS-2B, and a human lung adenocarcinoma (non-small cell) cancer cell line, NCI-H522, were subjected to morphometric analysis using a CAS-200 imaging system. The morphometric characteristics of these two cell lines differed significantly. Therefore, we proceeded to address the question of which statistical approach was most effective in classifying individual cells into the cell lines from which they were derived. The statistical techniques evaluated ranged from simple, traditional, parametric approaches to newer machine learning techniques. The multivariate techniques were compared based on a systematic cross-validation approach using 10 fixed partitions of the data to compute the misclassification rate for each method. For comparisons across cell lines at the level of each morphometric feature, we found little to distinguish nonparametric from parametric approaches. Among the linear models applied, logistic regression had the highest percentage of correct classifications; among the nonlinear and nonparametric methods applied, the Classification and Regression Trees model provided the highest percentage of correct classifications. Classification and Regression Trees has appealing characteristics: there are no assumptions about the distribution of the variables to be used, there is no need to specify which interactions to test, and there is no difficulty in handling complex, high-dimensional data sets containing mixed data types.

Adenocarcinoma↗

Structural interpretation of a topological index. 1. External factor variable connectivity index (EFVCI).

The external factor variable connectivity index (EFVCI) is interpreted by mining out the structural features hidden in the space spanned by the EFVCI indices through projection pursuit combining with number-theory net (NT-net) on the unit sphere U(Us). Projection pursuit is concerned with "interesting" projections of high-dimensional data sets to machine-pick "interesting" low-dimensional projections of a high-dimensional point cloud by numerically maximizing a certain objective function or projection index. At first, the optimal EFVCI index reaches to -0.80 in the correlation with a retention index of 207 hydrocarbons produced by insects. The EFVCI indices, with regression results of R = 0.99998, s = 3.49, RMSECV = 3.90, and F = 7.9560e+005, obtain high regression quality. The model is proven valid by leave-one-out cross validation. Second, the EFVCI index is interpreted by the structure information, that is, size, branch number, graph center, and branching position of topological structures, which is searched out on the unit sphere U(Us) by projection pursuit. Finally, the interpretation information is used to discover some chemical knowledge concerning the variation of the retention index with the change in chemical structures.

Journal Article↗

Changes in hippocampal volume and shape across time distinguish dementia of the Alzheimer type from healthy aging.

Rates of hippocampal volume loss have been shown to distinguish subjects with dementia of the Alzheimer type (DAT) from nondemented controls. In this study, we obtained magnetic resonance scans in 18 subjects with very mild DAT (CDR 0.5) and 26 age-matched nondemented controls (CDR 0) 2 years apart. Large-deformation high-dimensional brain mapping was used to quantify and compare changes in hippocampal shape as well as volume in the two groups of subjects. Hippocampal volume loss over time was significantly greater in the CDR 0.5 subjects (left = 8.3%, right = 10.2%) than in the CDR 0 subjects (left = 4.0%, right = 5.5%) (ANOVA, F = 7.81, P = 0.0078). We used singular-value decomposition and logistic regression models to quantify hippocampal shape change across time within individuals, and this shape change in the CDR 0.5 and CDR 0 subjects was found to be significantly different (Wilks's lambda, P = 0.014). Further, at baseline, CDR 0.5 subjects, in comparison to CDR 0 subjects, showed inward deformation over 38% of the hippocampal surface; after 2 years this difference grew to 47%. Also, within the CDR 0 subjects, shape change between baseline and follow-up was largely confined to the head of the hippocampus and subiculum, while in the CDR 0.5 subjects, shape change involved the lateral body of the hippocampus as well as the head region and subiculum. These results suggest that different patterns of hippocampal shape change in time as well as different rates of hippocampal volume loss distinguish very mild DAT from healthy aging.

Aged↗

LimROTS: a hybrid method integrating empirical Bayes and reproducibility-optimized statistics for robust differential expression analysis.

MOTIVATION: Differential expression analysis plays a vital role in omics research enabling precise identification of features that associate with different phenotypes. This process is critical for uncovering biological differences between conditions, such as disease versus healthy states. In proteomics, several statistical methods have been used, ranging from simple t-tests to more advanced methods like DEqMS, limma and ROTS. However, a flexible method for reproducibility-optimized statistics tailored for clinical omics data has been lacking. RESULTS: In this study, we developed LimROTS, a hybrid method that integrates a linear regression model and the empirical Bayes approach with reproducibility optimized statistics, to create a novel moderated ranking statistic, for robust and flexible analysis of proteomics data. We validated its performance using twenty-one proteomics gold standard spike-in datasets with different protein mixtures, MS instruments, and techniques for benchmarking. This hybrid approach improves accuracy and reproducibility of complex proteomics data, making LimROTS a powerful tool for high-dimensional omics data analysis. AVAILABILITY AND IMPLEMENTATION: LimROTS has been implemented as an R/Bioconductor package, available at https://doi.org/doi:10.18129/B9.bioc.LimROTS. Additionally, the code used in this study is available in GitHub repository https://github.com/AliYoussef96/LimROTSmanuscript.

Bayes Theorem↗

Hidden space support vector machines.

Hidden space support vector machines (HSSVMs) are presented in this paper. The input patterns are mapped into a high-dimensional hidden space by a set of hidden nonlinear functions and then the structural risk is introduced into the hidden space to construct HSSVMs. Moreover, the conditions for the nonlinear kernel function in HSSVMs are more relaxed, and even differentiability is not required. Compared with support vector machines (SVMs), HSSVMs can adopt more kinds of kernel functions because the positive definite property of the kernel function is not a necessary condition. The performance of HSSVMs for pattern recognition and regression estimation is also analyzed. Experiments on artificial and real-world domains confirm the feasibility and the validity of our algorithms.

Algorithms↗

Bayesian framework for least-squares support vector machine classifiers, gaussian processes, and kernel Fisher discriminant analysis.

The Bayesian evidence framework has been successfully applied to the design of multilayer perceptrons (MLPs) in the work of MacKay. Nevertheless, the training of MLPs suffers from drawbacks like the nonconvex optimization problem and the choice of the number of hidden units. In support vector machines (SVMs) for classification, as introduced by Vapnik, a nonlinear decision boundary is obtained by mapping the input vector first in a nonlinear way to a high-dimensional kernel-induced feature space in which a linear large margin classifier is constructed. Practical expressions are formulated in the dual space in terms of the related kernel function, and the solution follows from a (convex) quadratic programming (QP) problem. In least-squares SVMs (LS-SVMs), the SVM problem formulation is modified by introducing a least-squares cost function and equality instead of inequality constraints, and the solution follows from a linear system in the dual space. Implicitly, the least-squares formulation corresponds to a regression formulation and is also related to kernel Fisher discriminant analysis. The least-squares regression formulation has advantages for deriving analytic expressions in a Bayesian evidence framework, in contrast to the classification formulations used, for example, in gaussian processes (GPs). The LS-SVM formulation has clear primal-dual interpretations, and without the bias term, one explicitly constructs a model that yields the same expressions as have been obtained with GPs for regression. In this article, the Bayesian evidence framework is combined with the LS-SVM classifier formulation. Starting from the feature space formulation, analytic expressions are obtained in the dual space on the different levels of Bayesian inference, while posterior class probabilities are obtained by marginalizing over the model parameters. Empirical results obtained on 10 public domain data sets show that the LS-SVM classifier designed within the Bayesian evidence framework consistently yields good generalization performances.

Artificial Intelligence↗

Diagnosis of Ovarian Cancer Using Decision Tree Classification of Mass Spectral Data.

Recent reports from our laboratory and others support the SELDI ProteinChip technology as a potential clinical diagnostic tool when combined with $n$ -dimensional analyses algorithms. The objective of this study was to determine if the commercially available classification algorithm biomarker patterns software (BPS), which is based on a classification and regression tree (CART), would be effective in discriminating ovarian cancer from benign diseases and healthy controls. Serum protein mass spectrum profiles from 139 patients with either ovarian cancer, benign pelvic diseases, or healthy women were analyzed using the BPS software. A decision tree, using five protein peaks resulted in an accuracy of 81.5% in the cross-validation analysis and 80%in a blinded set of samples in differentiating the ovarian cancer from the control groups. The potential, advantages, and drawbacks of the BPS system as a bioinformatic tool for the analysis of the SELDI high-dimensional proteomic data are discussed.

Journal Article↗

A weakly supervised deep learning-based recurrence prediction and risk stratification of lung adenocarcinoma from pathology whole-slide images.

BACKGROUND: Accurate prediction of postoperative recurrence in lung adenocarcinoma (LUAD) is essential for guiding clinical decision-making and improving patient outcomes. Although various predictive models have been developed, most rely on complex genomic analyses and high-dimensional clinical data. The complexity of these approaches substantially limits their feasibility for routine clinical use. To address this clinical challenge, this study aims to predict postoperative recurrence using routinely available hematoxylin and eosin (H&E)-stained images and characterize the associated biological features. METHODS: A total of 329 patients who underwent curative resection at the First Affiliated Hospital of Wenzhou Medical University (FHWMU) were retrospectively enrolled and randomly assigned to training and internal validation cohorts in a 7:3 ratio. An independent external validation cohort comprising 70 patients from the Clinical Proteomic Tumor Analysis Consortium (CPTAC) was included. Three patch-level feature extractors (Inception_V3, ResNet18, and DenseNet121) were evaluated within a weakly supervised multiple-instance learning (MIL) framework incorporating automated region-of-interest (ROI) detection on segmented whole-slide images (WSIs). Model performance was assessed using the area under the receiver operating characteristic curve (AUC), Kaplan-Meier (KM) survival analysis, and multivariable Cox proportional hazards regression. Transcriptomic profiling and gene set enrichment analysis (GSEA) were conducted to investigate biological differences between risk groups. RESULTS: The model achieved AUCs of 0.923 in the training cohort, 0.891 in the internal validation cohort, and 0.847 in the external validation cohort. The model effectively stratified patients into high- and low-risk groups with significantly different recurrence-free survival (RFS) across all cohorts (all P&#x2009;<&#x2009;0.001) and retained prognostic value within AJCC stages I-III. Transcriptomic analyses revealed consistent enrichment of cell cycle-related pathways and neutrophil extracellular trap (NET) formation in high-risk patients across both institutional and CPTAC cohorts, aligning with distinct biological profiles of the model-derived risk stratification. CONCLUSIONS: This weakly supervised deep learning framework enables accurate and externally validated prediction of postoperative recurrence in LUAD using routinely available histopathological images, and integration of histopathological features with molecular analyses enhances biological interpretability. This work provides a clinically accessible and cost-effective tool for postoperative risk assessment in LUAD patients.

Humans↗

Using high-dimensional environmental covariates to study genotype by environment interaction for reproductive traits in Duroc boars.

We investigated the potential of incorporating grid-cell-based environmental covariates (ECs) in the genetic evaluation of total sperm count (TSC), sperm motility (MOT), and sperm morphology (MOR) for Duroc boars. A total of 188,665 records derived from 3,684 genotyped boars, born between December 2018 and October 2024 and raised in three stud farms located in different U.S. states, were analyzed using multi-trait linear-threshold repeatability models. To account for genotype by environment interactions (GE), we constructed an interaction matrix as the Hadamard product of the genomic relationship matrix and an environmental (co)variance matrix. The environmental groups were defined in three ways: farm, farm-season, and farm-year-season. The (co)variance matrix was constructed based on daily ECs obtained from the NASA POWER database for each environmental group. Of all available ECs, those significantly associated with TSC, MOT, and MOR (temperature, relative humidity, atmospheric pressure, and wind speed and direction) were retained. We evaluated five models with different GE structures: M1 represented the baseline without accounting for GE, in M2 the GE included farm as environmental groups, in M3 the GE included farm-season as environmental groups, in M4 the GE included farm-year-season as environmental groups, and M5 involved M3 with an additional random effect of the farm-season. Estimates of heritability for TSC, MOT, and MOR ranged from 0.03 to 0.04, 0.05 to 0.08, and 0.04 to 0.08, respectively. Corresponding repeatability ranged from 0.15 to 0.23, 0.28 to 0.49, and 0.28 to 0.49. The proportion of phenotypic variance attributed to GE variance ranged from 0.00 to 0.32, 0.00 to 0.44, and 0.00 to 0.44. Lastly, estimates of genetic correlation, TSC-MOT, TSC-MOR, and MOT-MOR ranged from 0.27 to 0.31, 0.24 to 0.31, and 0.98 to 0.99, respectively, with minor differences across models. We assessed the predictive ability of models using the linear regression validation. Across traits and models, bias ranged from -0.05 to 0.02 standard deviations, slope varied from 0.88 to 0.99, the correlation ranged from 0.75 to 0.84, and accuracy from 0.41 to 0.53. Overall, building the GE matrix considering grid-cell-based ECs helped to account for GE, thereby reducing the proportion of phenotypic variance attributed to genetic components; however, it did not improve the validation metrics. Additional on-farm records for ECs may improve the model performance.

Animals↗

Penalized Cumulative Probability Model for a Continuous Outcome Subject to Detection Limits.

Mixed-type outcome data occur when the outcome variable's distribution is a mixture of both continuous and discrete ordinal variables. Such mixed-type outcomes are common in biomedical, psychological, and the health sciences, particularly for variables having either a detection or quantitation limit. When interest lies in identifying a combination of genomic features associated with a mixed-type outcome, any method used would require a variable selection strategy for high-dimensional data. Unfortunately, few variable selection methods exist for modeling a mixed-type outcome when the covariate space is high dimensional. This study develops a high-dimensional penalized cumulative probability model (CPM), to allow for the identification of genomic features associated with mixed-type outcome of interest. We demonstrated how such model may be estimated using the iterative penalization procedure-the generalized monotone incremental forward stagewise (GMIFS) algorithm. The Model-X knockoffs procedure was combined with the estimation algorithm to control the false discovery rates (FDR) when performing variable selection. Through extensive simulation studies, our penalized CPM was shown to outperform alternative methods in terms of controlled variable selection performance by achieving high statistical power with the FDR being controlled at the target level. We demonstrate the utility of our method by applying it to predict estimated glomeruli filtration rate (eGFR) in kidney transplant recipients at 24&#x2009;months post-transplant using baseline gene expression data as predictors. Our CPM model identified five genes associated with this mixed-type outcome which have important links to renal disease, which may provide prognostic guidance for kidney transplantation recipients.

Models, Statistical↗

Partial least squares dimension reduction for microarray gene expression data with a censored response.

An important application of DNA microarray technologies involves monitoring the global state of transcriptional program in tumor cells. One goal in cancer microarray studies is to compare the clinical outcome, such as relapse-free or overall survival, for subgroups of patients defined by global gene expression patterns. A method of comparing patient survival, as a function of gene expression, was recently proposed in [Bioinformatics 18 (2002) 1625] by Nguyen and Rocke. Due to the (a) high-dimensionality of microarray gene expression data and (b) censored survival times, a two-stage procedure was proposed to relate survival times to gene expression profiles. The first stage involves dimensionality reduction of the gene expression data by partial least squares (PLS) and the second stage involves prediction of survival probability using proportional hazard regression. In this paper, we provide a systematic assessment of the performance of this two-stage procedure. PLS dimension reduction involves complex non-linear functions of both the predictors and the response data, rendering exact analytical study intractable. Thus, we assess the methodology under a simulation model for gene expression data with a censored response variable. In particular, we compare the performance of PLS dimension reduction relative to dimension reduction via principal components analysis (PCA) and to a modified PLS (MPLS) approach. PLS performed substantially better relative to dimension reduction via PCA when the total predictor variance explained is low to moderate (e.g. 40%-60%). It performed similar to MPLS and slightly better in some cases. Additionally, we examine the effect of censoring on dimension reduction stage. The performance of all methods deteriorates for a high censoring rate, although PLS-PH performed relatively best overall.

Algorithms↗

Instantaneous multivariate EEG coherence analysis by means of adaptive high-dimensional autoregressive models.

This study presents an efficient algorithm for the fitting of multivariate autoregressive models (MVAR) with time-dependent parameters to multidimensional signals. Thereby, the dimension of the model may be chosen to equal the number of signal channels. The autoregressive (AR) parameter matrices are estimated by an extension of the recursive least squares (RLS) algorithm with forgetting factor. The estimation procedure includes a single trial as well as an ensemble mean approach. The latter approach allows the simultaneous fit of one mean MVAR model to a set of single trials, each of them representing the measurement of the same task. A particular advantage of this ensemble mean approach is that it requires only a low computation effort in comparison to well known procedures applied to single trials. Furthermore, the ensemble mean approach is linked with a high adaptation capability. The properties of the estimator are investigated using simulated time series. It can be demonstrated that the adaptation capability of the estimation (measured by its adaptation speed and variance) does not depend on the model dimension. The mean MVAR fit is applied to 19-dimensional EEG data, recorded during an elementary comparison procedure. The calculation of ordinary and multiple coherence is discussed. The sensitivity of the multiple instantaneous EEG coherence will be demonstrated.

Algorithms↗

Nonparametric regression applied to quantitative structure-activity relationships

Several nonparametric regressors have been applied to modeling quantitative structure-activity relationship (QSAR) data. The simplest regressor, the Nadaraya-Watson, was assessed in a genuine multivariate setting. Other regressors, the local linear and the shifted Nadaraya-Watson, were implemented within additive models--a computationally more expedient approach, better suited for low-density designs. Performances were benchmarked against the nonlinear method of smoothing splines. A linear reference point was provided by multilinear regression (MLR). Variable selection was explored using systematic combinations of different variables and combinations of principal components. For the data set examined, 47 inhibitors of dopamine beta-hydroxylase, the additive nonparametric regressors have greater predictive accuracy (as measured by the mean absolute error of the predictions or the Pearson correlation in cross-validation trails) than MLR. The use of principal components did not improve the performance of the nonparametric regressors over use of the original descriptors, since the original descriptors are not strongly correlated. It remains to be seen if the nonparametric regressors can be successfully coupled with better variable selection and dimensionality reduction in the context of high-dimensional QSARs.

Journal Article↗

Supervised machine learning techniques for the classification of metabolic disorders in newborns.

MOTIVATION: During the Bavarian newborn screening programme all newborns have been tested for about 20 inherited metabolic disorders. Owing to the amount and complexity of the generated experimental data, machine learning techniques provide a promising approach to investigate novel patterns in high-dimensional metabolic data which form the source for constructing classification rules with high discriminatory power. RESULTS: Six machine learning techniques have been investigated for their classification accuracy focusing on two metabolic disorders, phenylketo nuria (PKU) and medium-chain acyl-CoA dehydrogenase deficiency (MCADD). Logistic regression analysis led to superior classification rules (sensitivity >96.8%, specificity >99.98%) compared to all investigated algorithms. Including novel constellations of metabolites into the models, the positive predictive value could be strongly increased (PKU 71.9% versus 16.2%, MCADD 88.4% versus 54.6% compared to the established diagnostic markers). Our results clearly prove that the mined data confirm the known and indicate some novel metabolic patterns which may contribute to a better understanding of newborn metabolism.

Algorithms↗

Is breathing in infants chaotic? Dimension estimates for respiratory patterns during quiet sleep.

We describe an analysis of dynamic behavior apparent in times-series recordings of infant breathing during sleep. Three principal techniques were used: estimation of correlation dimension, surrogate data analysis, and reduced linear (autoregressive) modeling (RARM). Correlation dimension can be used to quantify the complexity of time series and has been applied to a variety of physiological and biological measurements. However, the methods most commonly used to estimate correlation dimension suffer from some technical problems that can produce misleading results if not correctly applied. We used a new technique of estimating correlation dimension that has fewer problems. We tested the significance of dimension estimates by comparing estimates with artificial data sets (surrogate data). On the basis of the analysis, we conclude that the dynamics of infant breathing during quiet sleep can best be described as a nonlinear dynamic system with large-scale, low-dimensional and small-scale, high-dimensional behavior; more specifically, a noise-driven nonlinear system with a two-dimensional periodic orbit. Using our RARM technique, we identified the second period as cyclic amplitude modulation of the same period as periodic breathing. We conclude that our data are consistent with respiration being chaotic.

Algorithms↗

Supervised dimension reduction of intrinsically low-dimensional data.

High-dimensional data generated by a system with limited degrees of freedom are often constrained in low-dimensional manifolds in the original space. In this article, we investigate dimension-reduction methods for such intrinsically low-dimensional data through linear projections that preserve the manifold structure of the data. For intrinsically one dimensional data, this implies projecting to a curve on the plane with as few intersections as possible. We are proposing a supervised projection pursuit method that can be regarded as an extension of the single-index model for nonparametric regression. We show results from a toy and two robotic applications.

Journal Article↗

A framework for block-wise missing data in multi-omics.

High-throughput technologies have generated vast amounts of omic data. It is a consensus that the integration of diverse omics sources improves predictive models and biomarker discovery. However, managing multiple omics data poses challenges such as data heterogeneity, noise, high-dimensionality and missing data, especially in block-wise patterns. This study addresses the challenges of high dimensionality and block-wise missing data through a regularization and constrained-based approach. The methodology is implemented in the R package bwm for binary and continuous response variables, and applied to breast cancer and exposome multi-omics datasets, achieving strong performance even in scenarios with missing data present in all omics. In binary classification task, our proposed model achieves accuracy in the range of 86% to 92%, and F1 in the range of 68% to 79%. And, in regression task the correlation between true and predicted responses is in the range of 72% to 76%. However, there is a slight decline in performance metrics as the percentage of missing data increases. In scenarios where block-wise missing data affects multiple omics, the model performance actually surpasses that of scenarios where missing data is present in only one omics. One possible explanation for this might be that the other scenarios introduce a greater diversity of observation profiles, leading to a more robust model. Depending on the specific omics being studied, there is greater consistency in feature selection when comparing block-wise missing data scenarios.

Humans↗