Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “High-dimensional regression”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Survival prediction of diffuse large-B-cell lymphoma based on both clinical and gene expression information.

MOTIVATION: It is important to predict the outcome of patients with diffuse large-B-cell lymphoma after chemotherapy, since the survival rate after treatment of this common lymphoma disease is <50%. Both clinically based outcome predictors and the gene expression-based molecular factors have been proposed independently in disease prognosis. However combining the high-dimensional genomic data and the clinically relevant information to predict disease outcome is challenging. RESULTS: We describe an integrated clinicogenomic modeling approach that combines gene expression profiles and the clinically based International Prognostic Index (IPI) for personalized prediction in disease outcome. Dimension reduction methods are proposed to produce linear combinations of gene expressions, while taking into account clinical IPI information. The extracted summary measures capture all the regression information of the censored survival phenotype given both genomic and clinical data, and are employed as covariates in the subsequent survival model formulation. A case study of diffuse large-B-cell lymphoma data, as well as Monte Carlo simulations, both demonstrate that the proposed integrative modeling improves the prediction accuracy, delivering predictions more accurate than those achieved by using either clinical data or molecular predictors alone.

Biomarkers, Tumor↗

LimROTS: a hybrid method integrating empirical Bayes and reproducibility-optimized statistics for robust differential expression analysis.

MOTIVATION: Differential expression analysis plays a vital role in omics research enabling precise identification of features that associate with different phenotypes. This process is critical for uncovering biological differences between conditions, such as disease versus healthy states. In proteomics, several statistical methods have been used, ranging from simple t-tests to more advanced methods like DEqMS, limma and ROTS. However, a flexible method for reproducibility-optimized statistics tailored for clinical omics data has been lacking. RESULTS: In this study, we developed LimROTS, a hybrid method that integrates a linear regression model and the empirical Bayes approach with reproducibility optimized statistics, to create a novel moderated ranking statistic, for robust and flexible analysis of proteomics data. We validated its performance using twenty-one proteomics gold standard spike-in datasets with different protein mixtures, MS instruments, and techniques for benchmarking. This hybrid approach improves accuracy and reproducibility of complex proteomics data, making LimROTS a powerful tool for high-dimensional omics data analysis. AVAILABILITY AND IMPLEMENTATION: LimROTS has been implemented as an R/Bioconductor package, available at https://doi.org/doi:10.18129/B9.bioc.LimROTS. Additionally, the code used in this study is available in GitHub repository https://github.com/AliYoussef96/LimROTSmanuscript.

Bayes Theorem↗

Predictive value of hippocampal MR imaging-based high-dimensional mapping in mesial temporal epilepsy: preliminary findings.

BACKGROUND AND PURPOSE: We objectively assessed surface structural changes of the hippocampus in mesial temporal sclerosis (MTS) and assessed the ability of large-deformation high-dimensional mapping (HDM-LD) to demonstrate hippocampal surface symmetry and predict group classification of MTS in right and left MTS groups compared with control subjects. METHODS: Using eigenvector field analysis of HDM-LD segmentations of the hippocampus, we compared the symmetry of changes in the right and left MTS groups with a group of 15 matched controls. To assess the ability of HDM-LD to predict group classification, eigenvectors were selected by a logistic regression procedure when comparing the MTS group with control subjects. RESULTS: Multivariate analysis of variance on the coefficients from the first 9 eigenvectors accounted for 75% of the total variance between groups. The first 3 eigenvectors showed the largest differences between the control group and each of the MTS groups, but with eigenvector 2 showing the greatest difference in the MTS groups. Reconstruction of the hippocampal deformation vector fields due solely to eigenvector 2 shows symmetrical patterns in the right and left MTS groups. A "leave-one-out" (jackknife) procedure correctly predicted group classification in 14 of 15 (93.3%) left MTS subjects and all 15 right MTS subjects. CONCLUSION: Analysis of principal dimensions of hippocampal shape change suggests that MTS, after accounting for normal right-left asymmetries, affects the right and left hippocampal surface structure very symmetrically. Preliminary analysis using HDM-LD shows it can predict group classification of MTS and control hippocampi in this well-defined population of patients with MTS and mesial temporal lobe epilepsy (MTLE).

Adult↗

Hidden space support vector machines.

Hidden space support vector machines (HSSVMs) are presented in this paper. The input patterns are mapped into a high-dimensional hidden space by a set of hidden nonlinear functions and then the structural risk is introduced into the hidden space to construct HSSVMs. Moreover, the conditions for the nonlinear kernel function in HSSVMs are more relaxed, and even differentiability is not required. Compared with support vector machines (SVMs), HSSVMs can adopt more kinds of kernel functions because the positive definite property of the kernel function is not a necessary condition. The performance of HSSVMs for pattern recognition and regression estimation is also analyzed. Experiments on artificial and real-world domains confirm the feasibility and the validity of our algorithms.

Algorithms↗

Bayesian framework for least-squares support vector machine classifiers, gaussian processes, and kernel Fisher discriminant analysis.

The Bayesian evidence framework has been successfully applied to the design of multilayer perceptrons (MLPs) in the work of MacKay. Nevertheless, the training of MLPs suffers from drawbacks like the nonconvex optimization problem and the choice of the number of hidden units. In support vector machines (SVMs) for classification, as introduced by Vapnik, a nonlinear decision boundary is obtained by mapping the input vector first in a nonlinear way to a high-dimensional kernel-induced feature space in which a linear large margin classifier is constructed. Practical expressions are formulated in the dual space in terms of the related kernel function, and the solution follows from a (convex) quadratic programming (QP) problem. In least-squares SVMs (LS-SVMs), the SVM problem formulation is modified by introducing a least-squares cost function and equality instead of inequality constraints, and the solution follows from a linear system in the dual space. Implicitly, the least-squares formulation corresponds to a regression formulation and is also related to kernel Fisher discriminant analysis. The least-squares regression formulation has advantages for deriving analytic expressions in a Bayesian evidence framework, in contrast to the classification formulations used, for example, in gaussian processes (GPs). The LS-SVM formulation has clear primal-dual interpretations, and without the bias term, one explicitly constructs a model that yields the same expressions as have been obtained with GPs for regression. In this article, the Bayesian evidence framework is combined with the LS-SVM classifier formulation. Starting from the feature space formulation, analytic expressions are obtained in the dual space on the different levels of Bayesian inference, while posterior class probabilities are obtained by marginalizing over the model parameters. Empirical results obtained on 10 public domain data sets show that the LS-SVM classifier designed within the Bayesian evidence framework consistently yields good generalization performances.

Artificial Intelligence↗

Diagnosis of Ovarian Cancer Using Decision Tree Classification of Mass Spectral Data.

Recent reports from our laboratory and others support the SELDI ProteinChip technology as a potential clinical diagnostic tool when combined with $n$ -dimensional analyses algorithms. The objective of this study was to determine if the commercially available classification algorithm biomarker patterns software (BPS), which is based on a classification and regression tree (CART), would be effective in discriminating ovarian cancer from benign diseases and healthy controls. Serum protein mass spectrum profiles from 139 patients with either ovarian cancer, benign pelvic diseases, or healthy women were analyzed using the BPS software. A decision tree, using five protein peaks resulted in an accuracy of 81.5% in the cross-validation analysis and 80%in a blinded set of samples in differentiating the ovarian cancer from the control groups. The potential, advantages, and drawbacks of the BPS system as a bioinformatic tool for the analysis of the SELDI high-dimensional proteomic data are discussed.

Journal Article↗

A weakly supervised deep learning-based recurrence prediction and risk stratification of lung adenocarcinoma from pathology whole-slide images.

BACKGROUND: Accurate prediction of postoperative recurrence in lung adenocarcinoma (LUAD) is essential for guiding clinical decision-making and improving patient outcomes. Although various predictive models have been developed, most rely on complex genomic analyses and high-dimensional clinical data. The complexity of these approaches substantially limits their feasibility for routine clinical use. To address this clinical challenge, this study aims to predict postoperative recurrence using routinely available hematoxylin and eosin (H&E)-stained images and characterize the associated biological features. METHODS: A total of 329 patients who underwent curative resection at the First Affiliated Hospital of Wenzhou Medical University (FHWMU) were retrospectively enrolled and randomly assigned to training and internal validation cohorts in a 7:3 ratio. An independent external validation cohort comprising 70 patients from the Clinical Proteomic Tumor Analysis Consortium (CPTAC) was included. Three patch-level feature extractors (Inception_V3, ResNet18, and DenseNet121) were evaluated within a weakly supervised multiple-instance learning (MIL) framework incorporating automated region-of-interest (ROI) detection on segmented whole-slide images (WSIs). Model performance was assessed using the area under the receiver operating characteristic curve (AUC), Kaplan-Meier (KM) survival analysis, and multivariable Cox proportional hazards regression. Transcriptomic profiling and gene set enrichment analysis (GSEA) were conducted to investigate biological differences between risk groups. RESULTS: The model achieved AUCs of 0.923 in the training cohort, 0.891 in the internal validation cohort, and 0.847 in the external validation cohort. The model effectively stratified patients into high- and low-risk groups with significantly different recurrence-free survival (RFS) across all cohorts (all P&#x2009;<&#x2009;0.001) and retained prognostic value within AJCC stages I-III. Transcriptomic analyses revealed consistent enrichment of cell cycle-related pathways and neutrophil extracellular trap (NET) formation in high-risk patients across both institutional and CPTAC cohorts, aligning with distinct biological profiles of the model-derived risk stratification. CONCLUSIONS: This weakly supervised deep learning framework enables accurate and externally validated prediction of postoperative recurrence in LUAD using routinely available histopathological images, and integration of histopathological features with molecular analyses enhances biological interpretability. This work provides a clinically accessible and cost-effective tool for postoperative risk assessment in LUAD patients.

Humans↗

Using high-dimensional environmental covariates to study genotype by environment interaction for reproductive traits in Duroc boars.

We investigated the potential of incorporating grid-cell-based environmental covariates (ECs) in the genetic evaluation of total sperm count (TSC), sperm motility (MOT), and sperm morphology (MOR) for Duroc boars. A total of 188,665 records derived from 3,684 genotyped boars, born between December 2018 and October 2024 and raised in three stud farms located in different U.S. states, were analyzed using multi-trait linear-threshold repeatability models. To account for genotype by environment interactions (GE), we constructed an interaction matrix as the Hadamard product of the genomic relationship matrix and an environmental (co)variance matrix. The environmental groups were defined in three ways: farm, farm-season, and farm-year-season. The (co)variance matrix was constructed based on daily ECs obtained from the NASA POWER database for each environmental group. Of all available ECs, those significantly associated with TSC, MOT, and MOR (temperature, relative humidity, atmospheric pressure, and wind speed and direction) were retained. We evaluated five models with different GE structures: M1 represented the baseline without accounting for GE, in M2 the GE included farm as environmental groups, in M3 the GE included farm-season as environmental groups, in M4 the GE included farm-year-season as environmental groups, and M5 involved M3 with an additional random effect of the farm-season. Estimates of heritability for TSC, MOT, and MOR ranged from 0.03 to 0.04, 0.05 to 0.08, and 0.04 to 0.08, respectively. Corresponding repeatability ranged from 0.15 to 0.23, 0.28 to 0.49, and 0.28 to 0.49. The proportion of phenotypic variance attributed to GE variance ranged from 0.00 to 0.32, 0.00 to 0.44, and 0.00 to 0.44. Lastly, estimates of genetic correlation, TSC-MOT, TSC-MOR, and MOT-MOR ranged from 0.27 to 0.31, 0.24 to 0.31, and 0.98 to 0.99, respectively, with minor differences across models. We assessed the predictive ability of models using the linear regression validation. Across traits and models, bias ranged from -0.05 to 0.02 standard deviations, slope varied from 0.88 to 0.99, the correlation ranged from 0.75 to 0.84, and accuracy from 0.41 to 0.53. Overall, building the GE matrix considering grid-cell-based ECs helped to account for GE, thereby reducing the proportion of phenotypic variance attributed to genetic components; however, it did not improve the validation metrics. Additional on-farm records for ECs may improve the model performance.

Animals↗

Cross-validated Cox regression on microarray gene expression data.

This paper describes how penalized Cox regression, in combination with cross-validated partial likelihood can be employed to obtain reliable survival prediction models for high dimensional microarray data. The suggested procedure is demonstrated on a breast cancer survival data set consisting of 295 tumours as collected in the National Cancer Institute in Amsterdam and previously reported in more general papers. The main aim of this paper it to show how generally accepted biostatistical procedures can be employed to analyse high-dimensional data.

Breast Neoplasms↗

Penalized Cumulative Probability Model for a Continuous Outcome Subject to Detection Limits.

Mixed-type outcome data occur when the outcome variable's distribution is a mixture of both continuous and discrete ordinal variables. Such mixed-type outcomes are common in biomedical, psychological, and the health sciences, particularly for variables having either a detection or quantitation limit. When interest lies in identifying a combination of genomic features associated with a mixed-type outcome, any method used would require a variable selection strategy for high-dimensional data. Unfortunately, few variable selection methods exist for modeling a mixed-type outcome when the covariate space is high dimensional. This study develops a high-dimensional penalized cumulative probability model (CPM), to allow for the identification of genomic features associated with mixed-type outcome of interest. We demonstrated how such model may be estimated using the iterative penalization procedure-the generalized monotone incremental forward stagewise (GMIFS) algorithm. The Model-X knockoffs procedure was combined with the estimation algorithm to control the false discovery rates (FDR) when performing variable selection. Through extensive simulation studies, our penalized CPM was shown to outperform alternative methods in terms of controlled variable selection performance by achieving high statistical power with the FDR being controlled at the target level. We demonstrate the utility of our method by applying it to predict estimated glomeruli filtration rate (eGFR) in kidney transplant recipients at 24&#x2009;months post-transplant using baseline gene expression data as predictors. Our CPM model identified five genes associated with this mixed-type outcome which have important links to renal disease, which may provide prognostic guidance for kidney transplantation recipients.

Models, Statistical↗

Partial least squares dimension reduction for microarray gene expression data with a censored response.

An important application of DNA microarray technologies involves monitoring the global state of transcriptional program in tumor cells. One goal in cancer microarray studies is to compare the clinical outcome, such as relapse-free or overall survival, for subgroups of patients defined by global gene expression patterns. A method of comparing patient survival, as a function of gene expression, was recently proposed in [Bioinformatics 18 (2002) 1625] by Nguyen and Rocke. Due to the (a) high-dimensionality of microarray gene expression data and (b) censored survival times, a two-stage procedure was proposed to relate survival times to gene expression profiles. The first stage involves dimensionality reduction of the gene expression data by partial least squares (PLS) and the second stage involves prediction of survival probability using proportional hazard regression. In this paper, we provide a systematic assessment of the performance of this two-stage procedure. PLS dimension reduction involves complex non-linear functions of both the predictors and the response data, rendering exact analytical study intractable. Thus, we assess the methodology under a simulation model for gene expression data with a censored response variable. In particular, we compare the performance of PLS dimension reduction relative to dimension reduction via principal components analysis (PCA) and to a modified PLS (MPLS) approach. PLS performed substantially better relative to dimension reduction via PCA when the total predictor variance explained is low to moderate (e.g. 40%-60%). It performed similar to MPLS and slightly better in some cases. Additionally, we examine the effect of censoring on dimension reduction stage. The performance of all methods deteriorates for a high censoring rate, although PLS-PH performed relatively best overall.

Algorithms↗

Instantaneous multivariate EEG coherence analysis by means of adaptive high-dimensional autoregressive models.

This study presents an efficient algorithm for the fitting of multivariate autoregressive models (MVAR) with time-dependent parameters to multidimensional signals. Thereby, the dimension of the model may be chosen to equal the number of signal channels. The autoregressive (AR) parameter matrices are estimated by an extension of the recursive least squares (RLS) algorithm with forgetting factor. The estimation procedure includes a single trial as well as an ensemble mean approach. The latter approach allows the simultaneous fit of one mean MVAR model to a set of single trials, each of them representing the measurement of the same task. A particular advantage of this ensemble mean approach is that it requires only a low computation effort in comparison to well known procedures applied to single trials. Furthermore, the ensemble mean approach is linked with a high adaptation capability. The properties of the estimator are investigated using simulated time series. It can be demonstrated that the adaptation capability of the estimation (measured by its adaptation speed and variance) does not depend on the model dimension. The mean MVAR fit is applied to 19-dimensional EEG data, recorded during an elementary comparison procedure. The calculation of ordinary and multiple coherence is discussed. The sensitivity of the multiple instantaneous EEG coherence will be demonstrated.

Algorithms↗

Nonparametric regression applied to quantitative structure-activity relationships

Several nonparametric regressors have been applied to modeling quantitative structure-activity relationship (QSAR) data. The simplest regressor, the Nadaraya-Watson, was assessed in a genuine multivariate setting. Other regressors, the local linear and the shifted Nadaraya-Watson, were implemented within additive models--a computationally more expedient approach, better suited for low-density designs. Performances were benchmarked against the nonlinear method of smoothing splines. A linear reference point was provided by multilinear regression (MLR). Variable selection was explored using systematic combinations of different variables and combinations of principal components. For the data set examined, 47 inhibitors of dopamine beta-hydroxylase, the additive nonparametric regressors have greater predictive accuracy (as measured by the mean absolute error of the predictions or the Pearson correlation in cross-validation trails) than MLR. The use of principal components did not improve the performance of the nonparametric regressors over use of the original descriptors, since the original descriptors are not strongly correlated. It remains to be seen if the nonparametric regressors can be successfully coupled with better variable selection and dimensionality reduction in the context of high-dimensional QSARs.

Journal Article↗

Supervised machine learning techniques for the classification of metabolic disorders in newborns.

MOTIVATION: During the Bavarian newborn screening programme all newborns have been tested for about 20 inherited metabolic disorders. Owing to the amount and complexity of the generated experimental data, machine learning techniques provide a promising approach to investigate novel patterns in high-dimensional metabolic data which form the source for constructing classification rules with high discriminatory power. RESULTS: Six machine learning techniques have been investigated for their classification accuracy focusing on two metabolic disorders, phenylketo nuria (PKU) and medium-chain acyl-CoA dehydrogenase deficiency (MCADD). Logistic regression analysis led to superior classification rules (sensitivity >96.8%, specificity >99.98%) compared to all investigated algorithms. Including novel constellations of metabolites into the models, the positive predictive value could be strongly increased (PKU 71.9% versus 16.2%, MCADD 88.4% versus 54.6% compared to the established diagnostic markers). Our results clearly prove that the mined data confirm the known and indicate some novel metabolic patterns which may contribute to a better understanding of newborn metabolism.

Algorithms↗

Boosting proportional hazards models using smoothing splines, with applications to high-dimensional microarray data.

MOTIVATION: An important area of research in the postgenomics era is to relate high-dimensional genetic or genomic data to various clinical phenotypes of patients. Due to large variability in time to certain clinical events among patients, studying possibly censored survival phenotypes can be more informative than treating the phenotypes as categorical variables. Due to high dimensionality and censoring, building a predictive model for time to event is more difficult than the classification/linear regression problem. We propose to develop a boosting procedure using smoothing splines for estimating the general proportional hazards models. Such a procedure can potentially be used for identifying non-linear effects of genes on the risk of developing an event. RESULTS: Our empirical simulation studies showed that the procedure can indeed recover the true functional forms of the covariates and can identify important variables that are related to the risk of an event. Results from predicting survival after chemotherapy for patients with diffuse large B-cell lymphoma demonstrate that the proposed method can be used for identifying important genes that are related to time to death due to cancer and for building a parsimonious model for predicting the survival of future patients. In addition, there is clear evidence of non-linear effects of some genes on survival time.

Algorithms↗

Is breathing in infants chaotic? Dimension estimates for respiratory patterns during quiet sleep.

We describe an analysis of dynamic behavior apparent in times-series recordings of infant breathing during sleep. Three principal techniques were used: estimation of correlation dimension, surrogate data analysis, and reduced linear (autoregressive) modeling (RARM). Correlation dimension can be used to quantify the complexity of time series and has been applied to a variety of physiological and biological measurements. However, the methods most commonly used to estimate correlation dimension suffer from some technical problems that can produce misleading results if not correctly applied. We used a new technique of estimating correlation dimension that has fewer problems. We tested the significance of dimension estimates by comparing estimates with artificial data sets (surrogate data). On the basis of the analysis, we conclude that the dynamics of infant breathing during quiet sleep can best be described as a nonlinear dynamic system with large-scale, low-dimensional and small-scale, high-dimensional behavior; more specifically, a noise-driven nonlinear system with a two-dimensional periodic orbit. Using our RARM technique, we identified the second period as cyclic amplitude modulation of the same period as periodic breathing. We conclude that our data are consistent with respiration being chaotic.

Algorithms↗

Supervised dimension reduction of intrinsically low-dimensional data.

High-dimensional data generated by a system with limited degrees of freedom are often constrained in low-dimensional manifolds in the original space. In this article, we investigate dimension-reduction methods for such intrinsically low-dimensional data through linear projections that preserve the manifold structure of the data. For intrinsically one dimensional data, this implies projecting to a curve on the plane with as few intersections as possible. We are proposing a supervised projection pursuit method that can be regarded as an extension of the single-index model for nonparametric regression. We show results from a toy and two robotic applications.

Journal Article↗

A framework for block-wise missing data in multi-omics.

High-throughput technologies have generated vast amounts of omic data. It is a consensus that the integration of diverse omics sources improves predictive models and biomarker discovery. However, managing multiple omics data poses challenges such as data heterogeneity, noise, high-dimensionality and missing data, especially in block-wise patterns. This study addresses the challenges of high dimensionality and block-wise missing data through a regularization and constrained-based approach. The methodology is implemented in the R package bwm for binary and continuous response variables, and applied to breast cancer and exposome multi-omics datasets, achieving strong performance even in scenarios with missing data present in all omics. In binary classification task, our proposed model achieves accuracy in the range of 86% to 92%, and F1 in the range of 68% to 79%. And, in regression task the correlation between true and predicted responses is in the range of 72% to 76%. However, there is a slight decline in performance metrics as the percentage of missing data increases. In scenarios where block-wise missing data affects multiple omics, the model performance actually surpasses that of scenarios where missing data is present in only one omics. One possible explanation for this might be that the other scenarios introduce a greater diversity of observation profiles, leading to a more robust model. Depending on the specific omics being studied, there is greater consistency in feature selection when comparing block-wise missing data scenarios.

Humans↗