Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Using classification tree and logistic regression methods to diagnose myocardial infarction.

Early and accurate diagnosis of myocardial infarction (MI) in patients who present to the Emergency Room (ER) complaining of chest pain is an important problem in emergency medicine. A number of decision aids have been developed to assist with this problem but have not achieved general use. Machine learning techniques, including classification tree and logistic regression (LR) methods, have the potential to create simple but accurate decision aids. Both a classification tree (FT Tree) and an LR model (FT LR) have been developed to predict the probability that a patient with chest pain is having an MI based solely upon data available at time of presentation to the ER. Training data came from a data set collected in Edinburgh, Scotland. Each model was then tested on a separate Edinburgh data set, as well as on a data set from a different hospital in Sheffield, England. Previously published models, the Goldman classification tree[1] and Kennedy LR equation[2], were evaluated on the same test data sets. On the Edinburgh test set, results showed that the FT Tree, FT LR, and Kennedy LR performed equally well, with ROC curve areas of 94.04%, 94.28%, and 94.30%, respectively, while the Goldman Tree's performance was significantly poorer, with an area of 84.03%. The difference in ROC areas between the first three models and the Goldman model is significant beyond the 0.0001 level. On the Sheffield test set, results showed that the FT Tree, FT LR, and Kennedy LR ROC areas were not significantly different (p > = 0.17), while the FT Tree again outperformed the Goldman Tree (p = 0.006). Unlike previous work[3], this study indicates that classification trees, which have certain advantages over LR models, may perform as well as LR models in the diagnosis of patients with MI.

Algorithms↗

Transcriptome Analysis, Machine Learning, and Experimental Identification of CDK7 Affecting the Progression of Pregnancy-induced Hypertension by Influencing Macrophage Polarization.

INTRODUCTION: Pregnancy-induced hypertension (PIH) is a severe pregnancy complication characterized by placental insufficiency, abnormal vascular remodeling, and immune dysregulation, but personalized therapeutic markers remain unclear. This study aimed to identify key genes and explore immune mechanisms in PIH using transcriptome analysis, machine learning, and experimental validation. METHODS: We analyzed the GSE204835 transcriptomic dataset to screen differentially expressed genes (DEGs) and performed Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), Reactome, and Gene Set Enrichment Analysis (GSEA) for functional annotation. Immune infiltration analysis was also performed to examine the immune landscape in PIH. Least Absolute Shrinkage and Selection Operator (LASSO) regression identified key genes, which were validated in a PIH cell model. Flow cytometry and immunofluorescence assays assessed the effect of CDK7 knockdown on macrophage polarization. RESULTS: A total of 1,598 DEGs (1,123 upregulated, 475 downregulated) were identified. Enrichment analyses highlighted associations with embryonic organ development, oxidative phosphorylation, angiogenesis, and oxidative stress. Immune infiltration analysis revealed altered eosinophil and macrophage polarization in PIH. LASSO regression selected 12 key genes, with CDK7 showing the most significant upregulation in the PIH model. CDK7 knockdown promoted macrophage polarization toward the anti-inflammatory M2 phenotype. DISCUSSION: These findings link CDK7 to immune dysregulation in PIH by modulating macrophage polarization, expanding our understanding of PIH's molecular mechanisms. The study's limitations include reliance on public datasets and in vitro models, warranting in vivo validation. CONCLUSION: CDK7 emerges as a potential therapeutic target for PIH, offering new insights into immunoregulatory interventions for this complication.

Female↗

MegaPlantTF: a machine learning framework for comprehensive identification and classification of plant transcription factors.

MOTIVATION: Understanding the role of transcription factors (TFs) in plants is essential for the study of gene regulation and various biological processes. However, both TF detection and classification remain challenging due to the great diversity and complexity of these proteins. Conventional approaches, such as BLAST, often suffer from high computational complexity and limited performance on less common TF families. RESULTS: We introduce MegaPlantTF, the first comprehensive machine learning and deep learning framework for the prediction (TF versus non-TF) and classification (family-level) of plant TFs. Our method employs k-mer-based protein representations and a two-stage architecture combining a deep feed-forward neural network with a stacking ensemble classifier. To ensure robust performance assessment, we report micro-, macro-, and weighted-average performance metrics, providing a holistic evaluation of both frequent and underrepresented TF families. Additionally, we employ threshold-based evaluation to calibrate confidence in TF detection. The results show that MegaPlantTF achieves strong accuracy and precision, particularly with a k-mer size of 3 and a classification threshold of 0.5, and maintains stable performance even under stringent thresholds. In addition to the standard cross-validation tests, a use case study on Sorghum bicolor confirms that our method performs strongly in the genome-wide analysis, making it highly suitable for large-scale TF identification and classification tasks. MegaPlantTF represents a novel contribution by integrating k-mer encoding, binary family-specific classifiers, and a two-stage stacking ensemble into a unified, reproducible framework for large-scale plant TF identification and classification. AVAILABILITY AND IMPLEMENTATION: MegaPlantTF is freely accessible through a public web server available at https://bioinformatics.um6p.ma/MegaPlantTF. The complete source code, including pretrained models and example datasets, is available at https://github.com/Bioinformatics-UM6P/MegaPlantTF.

Transcription Factors↗

Alzheimer's subtypes A supervised, unsupervised, multimodal, multilayered embedded recursive (SUMMER) AI study.

Since Alzheimer's disease (AD) is a heterogeneous disease, different subtypes may have distinct biological, genetic, and clinical characteristics, requiring tailored interventions. While several proposed subtypes of AD exist, there is still no clear consensus on a definitive classification. By leveraging complementary AI approaches, including supervised and unsupervised learning, within a recursive pipeline (SUMMER) that integrates multimodal datasets encompassing MRI measurements, phenotypes, and genetic data, our goal was to generate robust scientific evidence for identifying AD subtypes. Data was downloaded from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database and included neuroimaging data (MRI), genetics (SNPs), clinical diagnosis, and demographics. 1133 European American participants' images, aged 55-95, were included in this study. The analysis was multi-fold, where the first step involved applying an unsupervised application to a subset of the MRI sample (AD + cognitively normal (CN) aged matched groups, 100 men aged 68-85 years, and 76 women aged 68-85 years). The MRI brain gray matter was segmented into 44 regions of interest (ROIs) according to a standard atlas, and 618 features were extracted, including ROI voxel intensity measurements such as minimum, maximum, and histogram variables. Results identified a cluster of subtype AD men and a cluster of subtype AD women that were distinct from the rest of their respective samples. In the next step, the integrity of the identified subtype AD clusters was investigated using the XGBoost supervised machine learning application with genetic features (SNPs, N=36,724) and labels: the identified subtype AD cluster vs. the rest of the sample, stratified by sex. A significant AD subtype men model (accuracy=0.85, F1=0.72, AUC=0.83) and a significant women AD subtype model (accuracy=0.81, F1=0.81, AUC=0.81) were built, confirming the homogeneity of the isolated AD subtype clusters. Discriminative biomarkers were extracted from the significant models, including selected ROIs and SNPs. Finally, the subtype models were tested on an unseen subset of ADNI data. The genetic-based models identified clusters of AD subtype participants consisting of 34% of the men AD group and 47% of the women AD group. Phenotypic analysis indicates that lower body weight was associated with the women's AD subtype. Complex diseases like AD demand a sophisticated, multimodal approach for precise diagnosis. Effectively identifying disease subtypes enhances the potential for personalized treatment, ultimately improving patient outcomes.

Journal Article↗

Large-scale predictions of secretory proteins from mammalian genomic and EST sequences.

Machine learning techniques have improved predictions of secretory proteins from protein, genomic and expressed sequence tag (EST) sequences. Artificial neural networks, physical sequence analysis using high-performance optimization, and hidden Markov models identify extremely variable signal peptides (the vehicles of protein transport across the endoplasmic reticulum membrane), transmembrane segments, and specific extracellular and intracellular domains as indicators of possible roles in the intercellular and intracellular chemical signaling pathways. The major role of peptide hormones, blood coagulation factors, carcinogenesis agents, and other secretory proteins in orchestrating multicellular life indicates pharmacological potential in the cure of major diseases and numerous biotechnological applications.

Animals↗

Exploration of predictive and prognostic alternative splicing signatures in lung adenocarcinoma using machine learning methods.

BACKGROUND: Alternative splicing (AS) plays critical roles in generating protein diversity and complexity. Dysregulation of AS underlies the initiation and progression of tumors. Machine learning approaches have emerged as efficient tools to identify promising biomarkers. It is meaningful to explore pivotal AS events (ASEs) to deepen understanding and improve prognostic assessments of lung adenocarcinoma (LUAD) via machine learning algorithms. METHOD: RNA sequencing data and AS data were extracted from The Cancer Genome Atlas (TCGA) database and TCGA SpliceSeq database. Using several machine learning methods, we identified 24 pairs of LUAD-related ASEs implicated in splicing switches and a random forest-based classifiers for identifying lymph node metastasis (LNM) consisting of 12 ASEs. Furthermore, we identified key prognosis-related ASEs and established a 16-ASE-based prognostic model to predict overall survival for LUAD patients using Cox regression model, random survival forest analysis, and forward selection model. Bioinformatics analyses were also applied to identify underlying mechanisms and associated upstream splicing factors (SFs). RESULTS: Each pair of ASEs was spliced from the same parent gene, and exhibited perfect inverse intrapair correlation (correlation coefficient = - 1). The 12-ASE-based classifier showed robust ability to evaluate LNM status of LUAD patients with the area under the receiver operating characteristic (ROC) curve (AUC) more than 0.7 in fivefold cross-validation. The prognostic model performed well at 1, 3, 5, and 10 years in both the training cohort and internal test cohort. Univariate and multivariate Cox regression indicated the prognostic model could be used as an independent prognostic factor for patients with LUAD. Further analysis revealed correlations between the prognostic model and American Joint Committee on Cancer stage, T stage, N stage, and living status. The splicing network constructed of survival-related SFs and ASEs depicts regulatory relationships between them. CONCLUSION: In summary, our study provides insight into LUAD researches and managements based on these AS biomarkers.

Adenocarcinoma of Lung↗

Natural discriminant analysis using interactive Potts models.

Natural discriminant analysis based on interactive Potts models is developed in this work. A generative model composed of piece-wise multivariate gaussian distributions is used to characterize the input space, exploring the embedded clustering and mixing structures and developing proper internal representations of input parameters. The maximization of a log-likelihood function measuring the fitness of all input parameters to the generative model, and the minimization of a design cost summing up square errors between posterior outputs and desired outputs constitutes a mathematical framework for discriminant analysis. We apply a hybrid of the mean-field annealing and the gradient-descent methods to the optimization of this framework and obtain multiple sets of interactive dynamics, which realize coupled Potts models for discriminant analysis. The new learning process is a whole process of component analysis, clustering analysis, and labeling analysis. Its major improvement compared to the radial basis function and the support vector machine is described by using some artificial examples and a real-world application to breast cancer diagnosis.

Journal Article↗

Machine learning on multiple epigenetic features reveals H3K27Ac as a driver of gene expression prediction across patients with glioblastoma.

Epigenetic mechanisms play a crucial role in driving transcript expression and shaping the phenotypic plasticity of glioblastoma stem cells (GSCs), contributing to tumor heterogeneity and therapeutic resistance. These mechanisms dynamically regulate the expression of key oncogenic and stemness-associated genes, enabling GSCs to adapt to environmental cues and evade targeted therapies. Importantly, epigenetic reprogramming allows GSCs to transition between cellular states, including therapy-resistant mesenchymal-like phenotypes, underscoring the need for epigenetic-targeting strategies to disrupt these adaptive processes. Understanding these epigenetic drivers of gene expression provides a foundation for novel therapeutic interventions aimed at eradicating GSCs and improving glioblastoma outcomes. Using machine learning (ML), we employ cross-patient prediction of transcript expression in GSCs by combining epigenetic features from various sources, including ATAC-seq, CTCF ChIP-seq, RNAPII ChIP-seq, H3K27Ac ChIP-seq, and RNA-seq. We investigate different ML and deep learning (DL) models for this task and ultimately build our final pipeline using XGBoost. The model trained on one patient generalizes to other 11 patients with high performance. Notably, H3K27Ac alone from a single patient is sufficient to predict gene expression in all 11 patients. Furthermore, the distribution of H3K27Ac peaks across the genomes of all patients is remarkably similar. These findings suggest that GSCs share a common distributional pattern of enhancer activity characterized by H3K27Ac, which can be utilized to predict gene expression in GSCs across patients. In summary, while GSCs are known for their transcriptomic and phenotypic heterogeneity, we propose that they share a common epigenetic pattern of enhancer activation that defines their underlying transcriptomic expression pattern. This pattern can predict gene expression across patient samples, providing valuable insights into the biology of GSCs.

Glioblastoma↗

Feedback error learning neural network for trans-femoral prosthesis.

Feedback-error learning (FEL) neural network was developed for control of a powered trans-femoral prosthesis. Nonlinearities and time-variations of the dynamics of the plant, in addition to redundancy and dynamic uncertainty during the double support phase of walking, makes conventional control methods very difficult to use. Rule-based control, which uses a knowledge base determined by machine learning and finite automata method is limited since it does not respond well to perturbations and environmental changes. FEL can be regarded as a hybrid control, because it combines nonparametric identification with parametric modeling and control. This paper presents simulation of a powered trans-femoral prosthesis controlled by a FEL neural network. Results suggest that FEL can be used to identify inverse dynamics of an arbitrary trans-femoral prosthesis during simple single joint movements (e.g., sinusoidal oscillations). The identified inverse dynamics then allows the tracking of an arbitrary trajectory such as a desired walking pattern within a multijoint structure. Simulation shows that the identified controller responds correctly when the leg motion is exposed to a perturbation such as a frequent change of the ground reaction force or the hip joint torque generated by the user. FEL eliminates the need for precise, tedious, and complex identification of model parameters.

Activities of Daily Living↗

Honest assessments of automatic learning algorithm performance.

OBJECTIVE: To compare methods of evaluating probabilistic predictors in systems that learn from examples. STUDY DESIGN: The performance of four automatic learning algorithms, representing current machine learning technology, were assessed using four methodologies in the task of separating normal squamous intermediate cervical cells from all other segmented objects in digital images. Two of the methodologies were carefully constructed to model sources of variation associated with the choice of training and test sets. These assessments were statistically compared with assessments using both standard and a modified version of cross-validation. RESULTS: The investigation illustrates the tradeoffs involved in obtaining statistical rigor as compared with the cost of collecting data. While cross-validation makes frugal use of data, it can produce misleading assessments of algorithm performance in terms of both bias and variance. The modified version produces more reliable assessments but in some cases may also be misleading. CONCLUSION: We suggest that users of learning algorithms should exercise judicious care in evaluating learning algorithm performance in order to avoid unnecessary bias and large variance in their assessments.

Algorithms↗

CORA--a knowledge-based system for the analysis of case-control studies.

Carrying out a statistical analysis, the researcher is concerned with the problem of choosing an appropriate statistical technique from a large number of competing methods. Most common statistical software offer different methods for analysing the data without giving any support regarding the adequacy of a method for a particular data set. This paper outlines the main features of the computer system CORA which provides a statistical analysis of stratified contingency tables and additionally supports the researcher at the different steps of this analysis. Here, the support given by the system consists of two different aspects. On the one hand, the help system of CORA contains general information on the implemented statistical methods which can be obtained on request. On the other hand, an advice tool recommends an adequate statistical method which depends on the actual empirical case-control data to be analysed. To build up the advice tool, a set of rules being discovered by machine learning from simulation studies is integrated into the system CORA.

Case-Control Studies↗

Genome-wide association, polygenic risk scores, and machine learning for chronic post-surgical pain risk stratification: A UK biobank study.

Chronic post-surgical pain is a prevalent and debilitating complication following surgery, representing a clinical challenge. Despite the established heritability of pain phenotypes, large-scale genetic studies remain limited. This study aimed to identify genetic variants associated with chronic post-surgical pain, develop polygenic risk scores, and integrate these with clinical features for risk prediction. UK Biobank data from 47,836 participants (2490 cases and 45,346 controls) were split into training (80%; n = 38,268) and validation (20%; n = 9568) sets prior to analysis. A genome-wide association study was conducted on the training set only, across 19 million variants, and polygenic risk scores were constructed and integrated with clinical features in a logistic regression framework. Two close, rare, imputed signals crossed the genome-wide significance threshold but lacked local linkage-disequilibrium support, while 220 variants crossed the suggestive threshold. In the held-out validation set, cases had higher mean polygenic risk scores than controls (0.138 vs. -0.021; Cohen's d = 0.16, p < 0.001). A logistic regression model integrating clinical features and polygenic risk scores achieved an area under the curve of 0.639 (95% CI: 0.583-0.693), higher than models using either feature set alone. The polygenic risk score for chronic post-surgical pain was among the most important predictors. Risk stratification revealed the top quartile had 3.84-fold higher odds of chronic post-surgical pain than the bottom quartile (95% CI: 2.00-7.37). These findings suggest a possible modest genetic contribution to chronic post-surgical pain. Polygenic risk scores may complement clinical factors in surgical risk stratification. PERSPECTIVE: Chronic post-surgical pain may have a modest genetic contribution. This UK Biobank study identified over 220 variants at suggestive significance and constructed a polygenic risk score that was significantly elevated in cases. A combined clinical-genomic model achieved a 3.84-fold difference in odds across predicted-risk quartiles.

Chronic post-surgical pain↗

Landscape of essential growth and fluconazole-resistance genes in the human fungal pathogen Cryptococcus neoformans.

Fungi can cause devastating invasive infections, typically in immunocompromised patients. Treatment is complicated both by the evolutionary similarity between humans and fungi and by the frequent emergence of drug resistance. Studies in fungal pathogens have long been slowed by a lack of high-throughput tools and community resources that are common in model organisms. Here we demonstrate a high-throughput transposon mutagenesis and sequencing (TN-seq) system in Cryptococcus neoformans that enables genome-wide determination of gene essentiality. We employed a random forest machine learning approach to classify the C. neoformans genome as essential or nonessential, predicting 1,465 essential genes, including 302 that lack human orthologs. These genes are ideal targets for new antifungal drug development. TN-seq also enables genome-wide measurement of the fitness contribution of genes to phenotypes of interest. As proof of principle, we demonstrate the genome-wide contribution of genes to growth in fluconazole, a clinically used antifungal. We show a novel role for the well-studied RIM101 pathway in fluconazole susceptibility. We also show that insertions of transposons into the 5' upstream region can drive sensitization of essential genes, enabling screenlike assays of both essential and nonessential components of the genome. Using this approach, we demonstrate a role for mitochondrial function in fluconazole sensitivity, such that tuning down many essential mitochondrial genes via 5' insertions can drive resistance to fluconazole. Our assay system will be valuable in future studies of C. neoformans, particularly in examining the consequences of genotypic diversity.

Cryptococcus neoformans↗

Machine learning approaches for cancer prognosis and diagnosis via non-coding RNA: a comprehensive review.

Non-coding RNAs (ncRNAs), once considered genomic dark matter, are now established as key regulators of gene expression with widespread roles in cellular homeostasis and disease. In cancer, ncRNA expression is frequently and systematically dysregulated, and many of these molecules circulate in stable, protected form within biofluids, offering a compelling basis for non-invasive or minimally invasive diagnostic strategies. However, their clinical translation remains substantially hindered to date due to biological complexity, technical noise, and high dimensionality inherent to ncRNA expression datasets. In this context, machine learning (ML) has emerged as a powerful analytical tool to address these challenges, enabling the identification of subtle, reproducible ncRNA signatures predictive of diverse malignancies. This review critically evaluates ML-driven frameworks for cancer diagnosis and prognosis across four ncRNA subclasses, namely miRNAs, lncRNAs, circRNAs, and piRNAs, while also acknowledging the biophysical and thermodynamic models that reinforce ncRNA bioinformatics. Despite substantial methodological progress in ML-based cancer diagnosis and prognosis, key challenges persist, including tumor biological heterogeneity, limited multicenter validation, and the lack of widely adopted standardized protocols for preprocessing, normalization, and reporting workflows. Furthermore, many current ML models lack interpretability in biological or clinical context, constraining their translational utility. By synthesizing recent advances and identifying unresolved barriers, this review charts a roadmap for developing a robust, clinically actionable ncRNA biomarker platform for cancer detection. With global cancer incidence projected to exceed 35 million annual cases by 2050, validated ncRNA-ML-driven frameworks hold potential to revolutionize early-stage detection and personalized therapeutic strategies, thereby reducing the escalating socio-economic burden of cancer worldwide.

Humans↗

N6-methyladenine identification using deep learning and discriminative feature integration.

N6-methyladenine (6&#xa0;mA) is a pivotal DNA modification that plays a crucial role in epigenetic regulation, gene expression, and various biological processes. With advancements in sequencing technologies and computational biology, there is an increasing focus on developing accurate methods for 6&#xa0;mA site identification to enhance early detection and understand its biological significance. Despite the rapid progress of machine learning in bioinformatics, accurately detecting 6&#xa0;mA sites remains a challenge due to the limited generalizability and efficiency of existing approaches. In this study, we present Deep-N6mA, a novel Deep Neural Network (DNN) model incorporating optimal hybrid features for precise 6&#xa0;mA site identification. The proposed framework captures complex patterns from DNA sequences through a comprehensive feature extraction process, leveraging k-mer, Dinucleotide-based Cross Covariance (DCC), Trinucleotide-based Auto Covariance (TAC), Pseudo Single Nucleotide Composition (PseSNC), Pseudo Dinucleotide Composition (PseDNC), and Pseudo Trinucleotide Composition (PseTNC). To optimize computational efficiency and eliminate irrelevant or noisy features, an unsupervised Principal Component Analysis (PCA) algorithm is employed, ensuring the selection of the most informative features. A multilayer DNN serves as the classification algorithm to identify N6-methyladenine sites accurately. The robustness and generalizability of Deep-N6mA were rigorously validated using fivefold cross-validation on two benchmark datasets. Experimental results reveal that Deep-N6mA achieves an average accuracy of 97.70% on the F. vesca dataset and 95.75% on the R. chinensis dataset, outperforming existing methods by 4.12% and 4.55%, respectively. These findings underscore the effectiveness of Deep-N6mA as a reliable tool for early 6&#xa0;mA site detection, contributing to epigenetic research and advancing the field of computational biology.

Deep Learning↗

Patient-specific modeling identifies metabolic interventions for reversing glucose use reprogramming in alcohol-associated hepatitis.

Alcoholic hepatitis (AH) is an acute form of alcohol-associated liver disease with very few treatment options. Recent studies highlighted liver metabolic reprogramming in AH as an indicator of severity. We aim at identifying new intervention points to reverse liver metabolic dysregulation across varying degrees of AH. We develop 89 personalized genome-scale metabolic models by integrating a generic human cellular metabolic model with liver transcriptomics data from AH patients with varying disease severity and healthy controls. We grade the AH patients based on the model-predicted level of glycolysis reprogramming and validate the results using published metabolomics data. We test in silico gene knockdown interventions to reverse the aberrant metabolic reprogramming in AH. Knockdown of two glycolytic genes, Hkdc1 and Pkm, significantly rebalance the metabolic fluxes toward a healthy liver metabolic phenotype. We use machine learning on the glycolysis fluxes to develop a quantitative glucose use reprogramming score, which correlates with AH severity and patient-specific responses to in silico gene knockdown interventions. The score was independently validated using a published AH liver transcriptomics dataset. We propose a cellular metabolism-based therapy targeting Hkdc1 and Pkm in the glycolysis pathway as a potential treatment for reversing the aberrant glucose metabolism in AH.

Humans↗

Analysis of molecular profile data using generative and discriminative methods.

A modular framework is proposed for modeling and understanding the relationships between molecular profile data and other domain knowledge using a combination of generative (here, graphical models) and discriminative [Support Vector Machines (SVMs)] methods. As illustration, naive Bayes models, simple graphical models, and SVMs were applied to published transcription profile data for 1,988 genes in 62 colon adenocarcinoma tissue specimens labeled as tumor or nontumor. These unsupervised and supervised learning methods identified three classes or subtypes of specimens, assigned tumor or nontumor labels to new specimens and detected six potentially mislabeled specimens. The probability parameters of the three classes were utilized to develop a novel gene relevance, ranking, and selection method. SVMs trained to discriminate nontumor from tumor specimens using only the 50-200 top-ranked genes had the same or better generalization performance than the full repertoire of 1,988 genes. Approximately 90 marker genes were pinpointed for use in understanding the basic biology of colon adenocarcinoma, defining targets for therapeutic intervention and developing diagnostic tools. These potential markers highlight the importance of tissue biology in the etiology of cancer. Comparative analysis of molecular profile data is proposed as a mechanism for predicting the physiological function of genes in instances when comparative sequence analysis proves uninformative, such as with human and yeast translationally controlled tumour protein. Graphical models and SVMs hold promise as the foundations for developing decision support systems for diagnosis, prognosis, and monitoring as well as inferring biological networks.

Bayes Theorem↗

Benchmark of biomarker identification and prognostic modeling methods on diverse censored data.

The practices of identifying biomarkers and developing prognostic models using genomic data has become increasingly prevalent. Such data often features characteristics that make these practices difficult, namely high dimensionality, correlations between predictors, and sparsity. Many modern methods have been developed to address these problematic characteristics while performing feature selection and prognostic modeling, but a large-scale comparison of their performances in these tasks on diverse right-censored time to event data (aka survival time data) is much needed. We have compiled many existing methods, including some machine learning methods, several which have performed well in previous benchmarks, primarily for comparison in regards to variable selection capability, and secondarily for survival time prediction on many synthetic datasets with varying levels of sparsity, correlation between predictors, and signal strength of informative predictors. For illustration, we have also performed multiple analyses on a publicly available and widely used cancer cohort from The Cancer Genome Atlas using these methods. We evaluated the methods through extensive simulation studies in terms of the false discovery rate, F1-score, concordance index, Brier score, root mean square error, and computation time. Of the methods compared, CoxBoost and the Adaptive LASSO performed well in all metrics, and the LASSO and elastic net excelled when evaluating concordance index and F1-score. The Benjamini-Hoschberg and q-value procedures showed volatile performances in controlling the false discovery rate. Some methods' performances were greatly affected by differences in the data characteristics. With our extensive numerical study, we have identified the best performing methods for a plethora of data characteristics using informative metrics. This will help cancer researchers in choosing the best approach for their needs when working with genomic data.

Humans↗