Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Pan-cancer multi-omics machine learning defines a lactylation-associated immune-excluded tumor state with proteomic and experimental corroboration.

BACKGROUND: Histone lactylation links lactate metabolism to chromatin regulation, but whether lactylation-program-associated transcriptional patterns delineate recurrent pan-cancer tumor states remains unclear. METHODS: We integrated mRNA, lncRNA, and miRNA profiles from 9712 TCGA tumors across 33 cancer types with GTEx references, six GEO cohorts, IMvigor210, and an institutional clear-cell renal cell carcinoma (ccRCC) cohort used for exploratory DIA-NN proteomic corroboration. Random-effects co-expression meta-analysis, multi-omics consensus clustering, regulon inference, immune deconvolution, TIDE, oncoPredict, and SHAP-based machine learning were applied. hsa-miR-431-5p was functionally evaluated as a proof-of-concept CS2-associated miRNA in bladder cancer models. RESULTS: LacCoEx-Atlas comprised 398,491 lactylation-related co-expression pairs across 24,667 RNA features under a random-effects framework (median I² = 88.6%). Consensus clustering identified two subtypes: CS2 showed glycolytic-mesenchymal-immune-excluded features, M2 macrophage enrichment, CD8⁺ T-cell depletion, elevated HDAC4/NSD3/KDM6B activity, and worse survival, whereas CS1 showed oxidative, sirtuin-active programs. CS2 had fewer predicted ICI responders (18.3% vs. 52.0%) and a lower observed ORR in IMvigor210 (15.3% vs. 24.0%). oncoPredict identified NU7441 as a hypothesis-generating CS2-associated sensitivity signal (Hedges' g = 1.17). DIA-NN proteomics in 50 ccRCC specimens provided exploratory support for CS2-associated hypoxia, ECM degradation, and metastasis programs. The 10-feature mRNA LARItools model achieved an apparent AUC of 0.9413, while a separate multi-omics model achieved 0.971; neither was independently validated. LARItools reproduced prognostic separation across six GEO cohorts. miR-431-5p promoted malignant phenotypes and EMT in bladder cancer cells, with concordant CMU4h expression findings. CONCLUSIONS: Lactylation-program-associated transcriptional patterns delineate a recurrent immune-excluded pan-cancer tumor state associated with adverse prognosis, reduced predicted immunotherapy responsiveness, exploratory single-cancer protein-level support, and testable DNA damage response-targeting hypotheses. LacCoEx-Atlas and LARItools provide open resources for lactylation-program-associated tumor-state stratification and future translational research.

Humans↗

Supervised machine learning techniques for the classification of metabolic disorders in newborns.

MOTIVATION: During the Bavarian newborn screening programme all newborns have been tested for about 20 inherited metabolic disorders. Owing to the amount and complexity of the generated experimental data, machine learning techniques provide a promising approach to investigate novel patterns in high-dimensional metabolic data which form the source for constructing classification rules with high discriminatory power. RESULTS: Six machine learning techniques have been investigated for their classification accuracy focusing on two metabolic disorders, phenylketo nuria (PKU) and medium-chain acyl-CoA dehydrogenase deficiency (MCADD). Logistic regression analysis led to superior classification rules (sensitivity >96.8%, specificity >99.98%) compared to all investigated algorithms. Including novel constellations of metabolites into the models, the positive predictive value could be strongly increased (PKU 71.9% versus 16.2%, MCADD 88.4% versus 54.6% compared to the established diagnostic markers). Our results clearly prove that the mined data confirm the known and indicate some novel metabolic patterns which may contribute to a better understanding of newborn metabolism.

Algorithms↗

Integrating histology and spatial transcriptomics via multimodal transformers and contrastive representation learning for accurate gene expression prediction.

Predicting spatial gene expression from Histological images is a fundamental task in understanding tissue organization and molecular phenotypes. However, existing methods often rely on single-model representations or lack effective alignment between image and transcriptomic features. To address these limitations, we propose a unified multimodal learning framework that integrates histological imaging and spatial transcriptomics through a shared latent representation space. Specifically, histological H&E images are encoded by a ResNet50-based convolutional stem and a MobileViT Transformer backbone to extract hierarchical visual representations. Both modalities are projected into a shared latent space via linear-GELU-dropout transformation blocks, enabling cross-modal alignment through a contrastive learning objective that maximizes agreement between the corresponding image and the spot embeddings. Experimental results on the 10x Genomics Visium dataset of human liver tissue demonstrate that MViTGene achieves significantly higher prediction accuracy than existing methods across multiple gene subsets, with improvements of 20%, 33%, and 12% in predicting marker genes, highly expressed genes, and highly variable genes, respectively. The significant improvement in relevance indicates that the model can more accurately capture the true correspondence between tissue morphology and gene expression, therefore enabling more reliable biological interpretation. It provides a computational tool for high-throughput spatial gene expression prediction that balances performance and interpretability.

Humans↗

Ecological models of human performance based on affordance, emotion and intuition.

This paper proposes a complementary approach to Rasmussen's taxonomy of the human skill-, rule-, and knowledge-based performance models by combining the ecological concept of affordances with the neural concepts of human emotion and intuition. The classical cognitive engineering framework is extended through the neuro-ecological approach, including personal human attributes important in exercising control over the work environment. The proposed affordance-, emotion-, and intuition-based models correspond to the three types of human performance, namely: learning, adaptive and tuning control, respectively. The new framework is not a predictive model of the operator behaviour, but rather it describes the processes of neuro-ecological control of the human environment.

Cognition↗

Divergent microbial preludes to necrotising enterocolitis defined by gut phages and bacterial resistomes.

BACKGROUND: Translating microbiome correlations into robust predictive features for complex gut disorders remains elusive, partly due to oversimplified models of pathogenesis and neglect of the virome, a key player in microbial ecosystems. Necrotising enterocolitis (NEC), a devastating disease of preterm infants with no reliable clinical predictors, exemplifies this challenge. OBJECTIVE: To determine the predictive potential of the gut prophageome and polymicrobial aetiologies for NEC. DESIGN: We applied integrated metagenomic and metatranscriptomic analyses and machine learning to 1825 longitudinal stool samples from 43 preterm infants who later developed NEC and 86 gestational age-matched and birthweight-matched controls across three US hospitals. We characterised gut prophageome acquisitions and their association with clinical exposures, including antibiotics, diet and pharmacotherapies. To predict NEC risk, we integrated pre-onset prophageome, antibacterial resistome and bacteriome profiles with neonatal pathology, stratifying the cohort by disease onset timing (early: ≤40 days; late: >40 days) for separate analysis. RESULTS: NEC cases exhibited distinct viral diversity trajectories before disease onset. Early-onset NEC was best predicted by phage-bacterial interaction signatures (75% accuracy, 81% sensitivity). Metatranscriptomics revealed increased phage DNA abundance with low gene expression, suggesting a lysogenic lifestyle that may stabilise pathobionts. These phages encode metabolic genes potentially enhancing pathobiont resilience. Late-onset NEC was best predicted by antibacterial resistome profiles (83% accuracy). CONCLUSION: The gut prophageome serves as both a source of pre-symptomatic predictive signals and an active modulator of NEC pathogenesis, with distinct microbial mechanisms driving early-onset and late-onset disease. These polymicrobial etiologies inform strategies for early detection, risk stratification and the development of microbiome-targeted preventive and therapeutic interventions.

BIOMARKERS↗

Stratifying lung adenocarcinoma: a novel prognostic model based on mitochondrial outer membrane permeabilization activity.

UNLABELLED: Mitochondrial outer membrane permeabilization (MOMP) is a core apoptotic regulatory event that dictates mitochondrial integrity, where full activation drives cell death and sublethal dysregulation contributes to tumor genomic instability. We used the Cancer Genome Atlas lung adenocarcinoma cohort (TCGA-LUAD) as the training cohort and the Gene Expression Omnibus dataset GSE42127 as the validation cohort to identify prognostic genes related to MOMP activity in lung adenocarcinoma (LUAD) and to evaluate their potential biological significance. By intersecting MOMP-related genes with differentially expressed genes, combined with survival analysis, Mendelian randomization analysis, and 101 machine-learning algorithm combinations, seven prognostic genes, namely BIRC5, PSMD11, TNFRSF13C, YWHAZ, YWHAG, CYCS, and LTB, were identified. Next, an optimal prognostic model was constructed based on the gradient boosting machine (GBM) algorithm. Based on the risk score, LUAD patients were stratified into high- and low-risk groups, and patients in the high-risk group exhibited poorer overall survival in both the training and validation cohorts. Furthermore, a nomogram integrating the risk score and clinicopathological factors was developed and showed favorable predictive performance for 1-, 3-, and 5-year survival. Meanwhile, functional and immune analyses revealed that the high-risk group was enriched in DNA replication-related pathways and demonstrated a higher tumor mutation burden (TMB). Correlation analysis indicated that TNFRSF13C was positively correlated with activated B cells, whereas BIRC5 was negatively correlated with eosinophils, suggesting that MOMP-related genes might be involved in remodeling the immune microenvironment of LUAD. Drug sensitivity analysis showed differences in predicted half-maximal inhibitory concentration (IC50) values between the risk groups, suggesting the potential value of this model in assisting therapeutic stratification. Single-cell RNA sequencing (scRNA-seq) further identified T lymphocytes as a key cell type, with numerous prognostic genes exhibiting differential expression in T cells or dynamic changes during differentiation. We suggest that the MOMP-related signature established in this study may provide a reference for prognostic stratification in LUAD and offers candidate prognostic genes for subsequent experimental and clinical validation. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13205-026-05058-6.

Lung adenocarcinoma↗

Proteomic signatures and predictive modeling of cadmium-associated anxiety in middle-aged and elderly populations: an environmental exposure association study.

BACKGROUND: Emerging evidence implicates environmental contaminants such as cadmium (Cd) as modifiable risk factors for anxiety. Despite growing recognition of heavy metal toxicity in neuropsychiatric disorders, the molecular mechanisms linking environmental exposure to anxiety pathogenesis remain poorly understood. METHODS: Based on the established cohort of individuals with cognitive impairment in cadmium-contaminated areas, this cross-sectional association study enrolled 50 middle-aged and elderly hospitalized patients from these regions, adhering to the STROBE guidelines. Blood concentrations of cadmium (Cd), lead (Pb), and mercury (Hg) were analyzed in relation to anxiety severity assessed via the Hamilton Anxiety Rating Scale (HAMA). Plasma proteomic profiling was performed using data-independent acquisition (DIA) quantitative technology with an LC-MS/MS platform (timsTOF Pro, Bruker Daltonics), systematically characterizing 2,531 proteins across all samples. Machine learning techniques, specifically XGBoost and LASSO, were employed to identify biomarkers that were subsequently validated through mediation analysis and animal experiments, allowing for the screening of key protein signatures. Finally, clinical variables were integrated to construct a comprehensive model, which was then thoroughly evaluated. RESULTS: Anxious individuals exhibited significantly higher blood Cd levels than controls (&#x3b2;&#x2009;=&#x2009;0.50, 95% CI: 0.07-0.93, p&#x2009;<&#x2009;0.01), with anxiety positively correlating with depression (r&#x2009;=&#x2009;0.62, p&#x2009;=&#x2009;0.003) and inversely with ApoE3 genotype prevalence. Proteomics identified 120 differentially expressed proteins in anxious patients, enriched in oxidative phosphorylation and neurodegenerative pathways. CCDC126 emerged as a cadmium-associated biomarker, validated in rat models exposed to Cd. Combining CCDC126, blood Cd, Pb, and hypertension, a clinical prediction model achieved robust discrimination (AUC&#x2009;=&#x2009;0.80, validation cohort). CONCLUSIONS: This first integrative environmental-proteomic study highlights cadmium's synergistic role in anxiety pathophysiology and psychiatric comorbidity. The predictive model offers translatable potential for early risk stratification, while CCDC126 provides mechanistic insights for targeted interventions in populations exposed to environmental pollutants.

Cadmium↗

Discriminating different classes of toxicants by transcript profiling.

Male rats were treated with various model compounds or the appropriate vehicle controls. Most substances were either well-known hepatotoxicants or showed hepatotoxicity during preclinical testing. The aim of the present study was to determine if biological samples from rats treated with various compounds can be classified based on gene expression profiles. In addition to gene expression analysis using microarrays, a complete serum chemistry profile and liver and kidney histopathology were performed. We analyzed hepatic gene expression profiles using a supervised learning method (support vector machines; SVMs) to generate classification rules and combined this with recursive feature elimination to improve classification performance and to identify a compact subset of probe sets with potential use as biomarkers. Two different SVM algorithms were tested, and the models obtained were validated with a compound-based external cross-validation approach. Our predictive models were able to discriminate between hepatotoxic and nonhepatotoxic compounds. Furthermore, they predicted the correct class of hepatotoxicant in most cases. We provide an example showing that a predictive model built on transcript profiles from one rat strain can successfully classify profiles from another rat strain. In addition, we demonstrate that the predictive models identify nonresponders and are able to discriminate between gene changes related to pharmacology and toxicity. This work confirms the hypothesis that compound classification based on gene expression data is feasible.

Algorithms↗

Warmr: a data mining tool for chemical data.

Data mining techniques are becoming increasingly important in chemistry as databases become too large to examine manually. Data mining methods from the field of Inductive Logic Programming (ILP) have potential advantages for structural chemical data. In this paper we present Warmr, the first ILP data mining algorithm to be applied to chemoinformatic data. We illustrate the value of Warmr by applying it to a well studied database of chemical compounds tested for carcinogenicity in rodents. Data mining was used to find all frequent substructures in the database, and knowledge of these frequent substructures is shown to add value to the database. One use of the frequent substructures was to convert them into probabilistic prediction rules relating compound description to carcinogenesis. These rules were found to be accurate on test data, and to give some insight into the relationship between structure and activity in carcinogenesis. The substructures were also used to prove that there existed no accurate rule, based purely on atom-bond substructure with less than seven conditions, that could predict carcinogenicity. This results put a lower bound on the complexity of the relationship between chemical structure and carcinogenicity. Only by using a data mining algorithm, and by doing a complete search, is it possible to prove such a result. Finally the frequent substructures were shown to add value by increasing the accuracy of statistical and machine learning programs that were trained to predict chemical carcinogenicity. We conclude that Warmr, and ILP data mining methods generally, are an important new tool for analysing chemical databases.

Algorithms↗

Metagenomic analyses reveal E. coli-derived siderophores as potential signatures for breast cancer.

BACKGROUND: Breast cancer remains a leading cause of cancer-related mortality in women. Recent evidence implicates the gut microbiome and metabolites in breast cancer pathogenesis. This study explores associations between gut microbial species, their predicted metabolites, and breast cancer to uncover potential mechanistic insights. METHODS: Comprehensive metagenomic analyses were conducted on the gut microbiome of pre- and postmenopausal breast cancer patients, where microbial species were profiled through AMPHORA2 and metabolites were predicted through antiSMASH. Multivariate association analysis was used to identify significant associations between specific microbial species, predicted metabolites, and breast cancer status. A custom ensemble machine learning classifier was developed to classify pre- and postmenopausal breast cancer cases and controls based on microbial and predicted metabolite features. Additionally, a synthetic microbiome dataset was generated through MIDASim to validate the reproducibility of the ML results. Using our results, we explored the underlying dynamics of identified taxa and metabolite in breast cancer through literature and statistical support. RESULTS: Our analysis identified 471 microbial species and predicted 40 key metabolites in the metagenomic data. Multivariate analysis identified significant positive associations (p-value&#x2009;<&#x2009;0.05) of E. coli, siderophore, and thiopeptide with breast cancer. The custom ensemble model achieved accuracy and AUC as high as 78% and 90%, respectively, in classifying pre- and postmenopausal cases and controls. The high-ranking features i.e., E. coli, siderophore, and thiopeptide were consistent with the results of the multivariate association analysis, thereby substantiating their biological significance. Using these findings, we propose a mechanistic model in which E. coli secretes siderophores under iron-limited conditions in breast cancer patients, for iron sequestration from the host, which can potentially promote angiogenesis and tumor progression. CONCLUSION: Our findings suggest that microbial iron acquisition mechanisms may play a critical role in breast cancer pathophysiology. Functional validation of these mechanisms is needed to assess therapeutic potential. This study highlights gut microbiota and their metabolites as promising targets for breast cancer research and intervention.

Breast Neoplasms↗

An accurate QSPR study of O-H bond dissociation energy in substituted phenols based on support vector machines.

The support vector machine (SVM), as a novel type of learning machine, was used to develop a Quantitative Structure-Property Relationship (QSPR) model of the O-H bond dissociation energy (BDE) of 78 substituted phenols. The six descriptors calculated solely from the molecular structures of compounds selected by forward stepwise regression were used as inputs for the SVM model. The root-mean-square (rms) errors in BDE predictions for the training, test, and overall data sets were 3.808, 3.320, and 3.713 BDE units (kJ mol(-1)), respectively. The results obtained by Gaussian-kernel SVM were much better than those obtained by multiple linear regression, radial basis function neural networks, linear-kernel SVM, and other QSPR approaches.

Journal Article↗

Mismatch string kernels for discriminative protein classification.

MOTIVATION: Classification of proteins sequences into functional and structural families based on sequence homology is a central problem in computational biology. Discriminative supervised machine learning approaches provide good performance, but simplicity and computational efficiency of training and prediction are also important concerns. RESULTS: We introduce a class of string kernels, called mismatch kernels, for use with support vector machines (SVMs) in a discriminative approach to the problem of protein classification and remote homology detection. These kernels measure sequence similarity based on shared occurrences of fixed-length patterns in the data, allowing for mutations between patterns. Thus, the kernels provide a biologically well-motivated way to compare protein sequences without relying on family-based generative models such as hidden Markov models. We compute the kernels efficiently using a mismatch tree data structure, allowing us to calculate the contributions of all patterns occurring in the data in one pass while traversing the tree. When used with an SVM, the kernels enable fast prediction on test sequences. We report experiments on two benchmark SCOP datasets, where we show that the mismatch kernel used with an SVM classifier performs competitively with state-of-the-art methods for homology detection, particularly when very few training examples are available. Examination of the highest-weighted patterns learned by the SVM classifier recovers biologically important motifs in protein families and superfamilies.

Algorithms↗

A machine learning-derived intratumoral heterogeneity-related signature predicts the prognosis for and therapeutic response in patients with skin cutaneous melanoma.

BACKGROUND: Reliable biomarkers for predicting prognosis and therapeutic response in skin cutaneous melanoma (SKCM) remain limited. This study aimed to develop an intratumoral heterogeneity (ITH)-related prognostic signature for SKCM using integrative machine learning. METHODS: RNA sequencing (RNA-seq) data from 472 SKCM patients in The Cancer Genome Atlas (TCGA) and 214 patients in the GSE65904 cohort were analyzed. ITH scores were calculated using the DEPTH2 algorithm. Differentially expressed genes (DEGs) were identified between high- and low-ITH groups [|log2fold change (FC)| &#x2265;1, false discovery rate (FDR) <0.05]. Based on 38 prognostic DEGs identified by univariate Cox regression, we employed an integrative framework of 101 machine learning algorithm combinations to construct prognostic models in the TCGA training cohort. The model with the highest average concordance index (C-index) was validated in the GSE65904 cohort and selected as the prognostic ITH-related signature (PIRS). Associations of the PIRS risk score with tumor mutational burden (TMB), immune cell infiltration, immune checkpoint gene expression, and drug sensitivity were systematically evaluated. Model performance was assessed using receiver operating characteristic (ROC) curves and Cox regression analyses. RESULTS: A 38-gene PIRS was constructed using the plsRcox algorithm. Patients with high PIRS risk scores exhibited significantly poorer overall survival (OS) in both the TCGA and Gene Expression Omnibus (GEO) cohorts. The PIRS was identified as an independent prognostic factor, with area under the curve (AUC) values of 0.779, 0.734, and 0.756 for 1-, 3-, and 5-year survival, respectively. High-risk samples displayed significantly lower TMB (P<0.05), reduced immune and stromal cell infiltration (P<0.001), downregulated immune function, and decreased expression of immune checkpoint genes. Additionally, high- and low-PIRS risk score groups exhibited distinct sensitivity patterns to different classes of targeted agents. CONCLUSIONS: The machine learning-derived PIRS robustly predicts prognosis in SKCM patients. Its clinical application is promising for optimizing patient risk stratification and treatment decisions, though further prospective validation is warranted.

Skin cutaneous melanoma (SKCM)↗

A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings.

In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime.

Humans↗

Joint learning of gene functions--a Bayesian network model approach.

In this paper, we develop a machine learning system for determining gene functions from heterogeneous data sources using a Weighted Naive Bayesian network (WNB). The knowledge of gene functions is crucial for understanding many fundamental biological mechanisms such as regulatory pathways, cell cycles and diseases. Our major goal is to accurately infer functions of putative genes or Open Reading Frames (ORFs) from existing databases using computational methods. However, this task is intrinsically difficult since the underlying biological processes represent complex interactions of multiple entities. Therefore, many functional links would be missing when only one or two sources of data are used in the prediction. Our hypothesis is that integrating evidence from multiple and complementary sources could significantly improve the prediction accuracy. In this paper, our experimental results not only suggest that the above hypothesis is valid, but also provide guidelines for using the WNB system for data collection, training and predictions. The combined training data sets contain information from gene annotations, gene expressions, clustering outputs, keyword annotations, and sequence homology from public databases. The current system is trained and tested on the genes of budding yeast Saccharomyces cerevisiae. Our WNB model can also be used to analyze the contribution of each source of information toward the prediction performance through the weight training process. The contribution analysis could potentially lead to significant scientific discovery by facilitating the interpretation and understanding of the complex relationships between biological entities.

Artificial Intelligence↗

The 2000 Olympic Games of protein structure prediction; fully automated programs are being evaluated vis-à-vis human teams in the protein structure prediction experiment CAFASP2.

In this commentary, we describe two new protein structure prediction experiments being run in parallel with the CASP experiment, which together may be regarded as the 2000 Olympic Games of structure prediction. The first new experiment is CAFASP, the Critical Assessment of Fully Automated Structure Prediction. In CAFASP, the participants are fully automated programs or Internet servers, and here the automated results of the programs are evaluated, without any human intervention. The second new experiment, named LiveBench, follows the CAFASP ideology in that it is aimed towards the evaluation of automatic servers only, while it runs on a large set of prediction targets and in a continuous fashion. Researchers will be watching the 2000 protein structure prediction Olympic Games, to be held in December, in order to learn about the advances in the classical 'human-plus-machine' CASP category, the fully automated CAFASP category, and the comparison between the two.

Amino Acid Sequence↗

Benchmark of biomarker identification and prognostic modeling methods on diverse censored data.

The practices of identifying biomarkers and developing prognostic models using genomic data has become increasingly prevalent. Such data often features characteristics that make these practices difficult, namely high dimensionality, correlations between predictors, and sparsity. Many modern methods have been developed to address these problematic characteristics while performing feature selection and prognostic modeling, but a large-scale comparison of their performances in these tasks on diverse right-censored time to event data (aka survival time data) is much needed. We have compiled many existing methods, including some machine learning methods, several which have performed well in previous benchmarks, primarily for comparison in regards to variable selection capability, and secondarily for survival time prediction on many synthetic datasets with varying levels of sparsity, correlation between predictors, and signal strength of informative predictors. For illustration, we have also performed multiple analyses on a publicly available and widely used cancer cohort from The Cancer Genome Atlas using these methods. We evaluated the methods through extensive simulation studies in terms of the false discovery rate, F1-score, concordance index, Brier score, root mean square error, and computation time. Of the methods compared, CoxBoost and the Adaptive LASSO performed well in all metrics, and the LASSO and elastic net excelled when evaluating concordance index and F1-score. The Benjamini-Hoschberg and q-value procedures showed volatile performances in controlling the false discovery rate. Some methods' performances were greatly affected by differences in the data characteristics. With our extensive numerical study, we have identified the best performing methods for a plethora of data characteristics using informative metrics. This will help cancer researchers in choosing the best approach for their needs when working with genomic data.

Humans↗

A cerebellar model for predictive motor control tested in a brain-based device.

The cerebellum is known to be critical for accurate adaptive control and motor learning. We propose here a mechanism by which the cerebellum may replace reflex control with predictive control. This mechanism is embedded in a learning rule (the delayed eligibility trace rule) in which synapses onto a Purkinje cell or onto a cell in the deep cerebellar nuclei become eligible for plasticity only after a fixed delay from the onset of suprathreshold presynaptic activity. To investigate the proposal that the cerebellum is a general-purpose predictive controller guided by a delayed eligibility trace rule, a computer model based on the anatomy and dynamics of the cerebellum was constructed. It contained components simulating cerebellar cortex and deep cerebellar nuclei, and it received input from a middle temporal visual area and the inferior olive. The model was incorporated in a real-world brain-based device (BBD) built on a Segway robotic platform that learned to traverse curved paths. The BBD learned which visual motion cues predicted impending collisions and used this experience to avoid path boundaries. During learning, the BBD adapted its velocity and turning rate to successfully traverse various curved paths. By examining neuronal activity and synaptic changes during this behavior, we found that the cerebellar circuit selectively responded to motion cues in specific receptive fields of simulated middle temporal visual areas. The system described here prompts several hypotheses about the relationship between perception and motor control and may be useful in the development of general-purpose motor learning systems for machines.

Cerebellum↗