Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Human mutations in high-confidence Tourette disorder genes affect sensorimotor behavior, reward learning, and striatal dopamine in mice.

UNLABELLED: Tourette disorder (TD) is poorly understood, despite affecting 1/160 children. A lack of animal models possessing construct, face, and predictive validity hinders progress in the field. We used CRISPR/Cas9 genome editing to generate mice with mutations orthologous to human de novo variants in two high-confidence Tourette genes, CELSR3 and WWC1 . Mice with human mutations in Celsr3 and Wwc1 exhibit cognitive and/or sensorimotor behavioral phenotypes consistent with TD. Sensorimotor gating deficits, as measured by acoustic prepulse inhibition, occur in both male and female Celsr3 TD models. Wwc1 mice show reduced prepulse inhibition only in females. Repetitive motor behaviors, common to Celsr3 mice and more pronounced in females, include vertical rearing and grooming. Sensorimotor gating deficits and rearing are attenuated by aripiprazole, a partial agonist at dopamine type II receptors. Unsupervised machine learning reveals numerous changes to spontaneous motor behavior and less predictable patterns of movement. Continuous fixed-ratio reinforcement shows Celsr3 TD mice have enhanced motor responding and reward learning. Electrically evoked striatal dopamine release, tested in one model, is greater. Brain development is otherwise grossly normal without signs of striatal interneuron loss. Altogether, mice expressing human mutations in high-confidence TD genes exhibit face and predictive validity. Reduced prepulse inhibition and repetitive motor behaviors are core behavioral phenotypes and are responsive to aripiprazole. Enhanced reward learning and motor responding occurs alongside greater evoked dopamine release. Phenotypes can also vary by sex and show stronger affection in females, an unexpected finding considering males are more frequently affected in TD. SIGNIFICANCE STATEMENT: We generated mouse models that express mutations in high-confidence genes linked to Tourette disorder (TD). These models show sensorimotor and cognitive behavioral phenotypes resembling TD-like behaviors. Sensorimotor gating deficits and repetitive motor behaviors are attenuated by drugs that act on dopamine. Reward learning and striatal dopamine is enhanced. Brain development is grossly normal, including cortical layering and patterning of major axon tracts. Further, no signs of striatal interneuron loss are detected. Interestingly, behavioral phenotypes in affected females can be more pronounced than in males, despite male sex bias in the diagnosis of TD. These novel mouse models with construct, face, and predictive validity provide a new resource to study neural substrates that cause tics and related behavioral phenotypes in TD.

Preprint↗

STRUMP-I: Structure-based machine learning approach to pMHC-I binding prediction using force field energy features.

The adaptive immune system monitors cellular integrity by recognizing short peptides from intracellular proteins presented on Major Histocompatibility Complex class I (MHC-I) molecules, collectively termed peptide-MHC complexes (pMHC), enabling detection of foreign or mutated proteins. With the rising importance of immunotherapies targeting neoantigens in cancers, the ability to accurately predict which peptides will bind to the diverse population of MHC alleles is critically important. Current computational methods for pMHC-I prediction fall broadly into sequence-based methods, which rely heavily on large training datasets, and structure-based methods that leverage structural modeling and energetics of pMHC binding. While sequence-based methods have been popularly used, their performance is dependent on the size and quality of training data. On the other hands, while structure-based approaches can generalize better across diverse MHC alleles, they traditionally depend on identifying a single global minimum energy conformation, an assumption that often fails due to the inherent binding promiscuity of MHC-I molecules. To address these limitations, we developed a STRUMP-I (STRUcture-based pMHC Prediction (for class I)), a novel pMHC binding prediction tool that directly leverages a broad set of force-field-derived energy terms as machine-learning features. STRUMP-I achieves performance comparable to state-of-the-art sequence-based models while significantly outperforming them on MHC alleles with limited representation in training data. Furthermore, STRUMP-I demonstrates strong synergy when integrated with sequence-based methods, notably enhancing prediction precision. The robustness and generalizability of STRUMP-I were confirmed by evaluating its predictive performance on independent, previously unseen datasets, including an experimentally validated cancer neoantigen dataset. This combined approach advances our capability to reliably identify clinically relevant neoantigen targets. The source code and trained models are available at https://github.com/yoonjoolab/STRUMP-I.

energy optimization↗

Human mutations in high-confidence Tourette disorder genes affect sensorimotor behavior, reward learning, and striatal dopamine in mice.

Tourette disorder (TD) is poorly understood, despite affecting 1/160 children. A lack of animal models possessing construct, face, and predictive validity hinders progress in the field. We used CRISPR/Cas9 genome editing to generate mice with mutations orthologous to human de novo variants in two high-confidence Tourette genes, CELSR3 and WWC1. Mice with human mutations in Celsr3 and Wwc1 exhibit cognitive and/or sensorimotor behavioral phenotypes consistent with TD. Sensorimotor gating deficits, as measured by acoustic prepulse inhibition, occur in both male and female Celsr3 TD models. Wwc1 mice show reduced prepulse inhibition only in females. Repetitive motor behaviors, common to Celsr3 mice and more pronounced in females, include vertical rearing and grooming. Sensorimotor gating deficits and rearing are attenuated by aripiprazole, a partial agonist at dopamine type II receptors. Unsupervised machine learning reveals numerous changes to spontaneous motor behavior and less predictable patterns of movement. Continuous fixed-ratio reinforcement shows that Celsr3 TD mice have enhanced motor responding and reward learning. Electrically evoked striatal dopamine release, tested in one model, is greater. Brain development is otherwise grossly normal without signs of striatal interneuron loss. Altogether, mice expressing human mutations in high-confidence TD genes exhibit face and predictive validity. Reduced prepulse inhibition and repetitive motor behaviors are core behavioral phenotypes and are responsive to aripiprazole. Enhanced reward learning and motor responding occur alongside greater evoked dopamine release. Phenotypes can also vary by sex and show stronger affection in females, an unexpected finding considering males are more frequently affected in TD.

Animals↗

CCDC137 knockdown suppresses bladder cancer progression by downregulating SCD.

BACKGROUND: The Coiled-coil domain-containing (CCDC) family, due to its unique protein structural domain and broad involvement in diverse biological processes, has emerged as a focus in oncology research. Nevertheless, its clinical significance and function in bladder cancer (BLCA) remain poorly defined. METHODS: Machine learning algorithms were employed to identify pivotal CCDC genes in the cancer genome atlas (TCGA), and a prognostic model was subsequently constructed. Multi-omics data encompassing pan-cancer cohorts, single-cell sequencing, and spatial transcriptomics were integrated to characterize the expression patterns and prognostic significance of Coiled-coil domain-containing 137 (CCDC137), a previously uncharacterized CCDC family member in BLCA. Tissue microarray confirmed CCDC137 abnormal expression in bladder carcinoma specimens. The effect of CCDC137 knockdown on BLCA progression was evaluated through CCK8 assay, clonogenic formation, wound healing, Transwell, and subcutaneous xenograft models. RNA sequencing, quantitative RT-PCR, and western blot were utilized to delineate its regulatory network. RESULTS: A prognostic model incorporating 10 CCDC genes was successfully established in the TCGA-BLCA cohort. Then, we found that CCDC137 exhibited pan-cancer overexpression and usually correlation with poor clinical outcomes. Immunohistochemistry further substantiated its dysregulation in bladder carcinoma. Integrated multi-omics analyses suggested associations between CCDC137 expression and a tumor immunosuppressive microenvironment. CCDC137 knockdown significantly suppressed bladder cancer cell proliferation and migratory capacity in vitro. Correspondingly, subcutaneous xenograft tumor growth was inhibited in vivo. Moreover, decreased expression of stearoyl-CoA desaturase (SCD), a key lipid metabolic enzyme, accompanied CCDC137 depletion. These findings collectively suggest a cancer-promoting role for CCDC137 in bladder carcinoma. CONCLUSIONS: This systematic investigation combining multi-omics bioinformatics analyses and experimental validation demonstrates the role of CCDC137 in bladder carcinoma progression, providing novel mechanistic insights into the pathogenesis of BLCA and offering a theoretical foundation for therapeutic targeting of CCDC137 in urothelial malignancies.

Urinary Bladder Neoplasms↗

Application of genetic search in derivation of matrix models of peptide binding to MHC molecules.

T cells of the vertebrate immune system recognise peptides bound by major histocompatibility complex (MHC) molecules on the surface of host cells. Peptide binding to MHC molecules is necessary for immune recognition, but only a subset of peptides are capable of binding to a particular MHC molecule. Common amino acid patterns (binding motifs) have been observed in sets of peptides that bind to specific MHC molecules. Recently, matrix models for peptide/MHC interaction have been reported. These encode the rules of peptide/ MHC interactions for an individual MHC molecule as a 20 x 9 matrix where the contribution to binding of each amino acid at each position within a 9-mer peptide is quantified. The artificial intelligence techniques of genetic search and machine learning have proved to be very useful in the area of biological sequence analysis. The availability of peptide/MHC binding data can facilitate derivation of binding matrices using machine learning techniques. We performed a simulation study to determine the minimum number of peptide samples required to derive matrices, given the pre-defined accuracy of the matrix model. The matrices were derived using a genetic search. In addition, matrices for peptide binding to the human class I MHC molecules, HLA-B35 and -A24, were derived, validated by independent experimental data and compared to previously-reported matrices. The results indicate that at least 150 peptide samples are required to derive matrices of acceptable accuracy. This result is based on a maximum noise content of 5%, the availability of precise affinity measurements and that acceptable accuracy is determined by an area under the Relative Operating Characteristic curve (Aroc) of > 0.8. More than 600 peptide samples are required to derive matrices of excellent accuracy (Aroc > 0.9). Finally, we derived a human HLA-B27 binding matrix using a genetic search and 404 experimentally-tested peptides, and estimated its accuracy at Aroc > 0.88. The results of this study are expected to be of practical interest to immunologists for efficient identification of peptides as candidates for immunotherapy.

Amino Acid Sequence↗

Equilibrium point control of a monkey arm simulator by a fast learning tree structured artificial neural network.

A planar 17 muscle model of the monkey's arm based on realistic biomechanical measurements was simulated on a Symbolics Lisp Machine. The simulator implements the equilibrium point hypothesis for the control of arm movements. Given initial and final desired positions, it generates a minimum-jerk desired trajectory of the hand and uses the backdriving algorithm to determine an appropriate sequence of motor commands to the muscles (Flash 1987; Mussa-Ivaldi et al. 1991; Dornay 1991b). These motor commands specify a temporal sequence of stable (attractive) equilibrium positions which lead to the desired hand movement. A strong disadvantage of the simulator is that it has no memory of previous computations. Determining the desired trajectory using the minimum-jerk model is instantaneous, but the laborious backdriving algorithm is slow, and can take up to one hour for some trajectories. The complexity of the required computations makes it a poor model for biological motor control. We propose a computationally simpler and more biologically plausible method for control which achieves the benefits of the backdriving algorithm. A fast learning, tree-structured network (Sanger 1991c) was trained to remember the knowledge obtained by the backdriving algorithm. The neural network learned the nonlinear mapping from a 2-dimensional cartesian planar hand position (x,y) to a 17-dimensional motor command space (u1, . . ., u17). Learning 20 training trajectories, each composed of 26 sample points [[x,y], [u1, . . ., u17] took only 20 min on a Sun-4 Sparc workstation. After the learning stage, new, untrained test trajectories as well as the original trajectories of the hand were given to the neural network as input. The network calculated the required motor commands for these movements. The resulting movements were close to the desired ones for both the training and test cases.

Algorithms↗

Detecting Interspecific Positive Selection Using Convolutional Neural Networks.

Traditional statistical methods using maximum likelihood and Bayesian inference can detect positive selection from an interspecific phylogeny and a codon sequence alignment based on model assumptions, but they are prone to false positives due to alignment errors and can lack power. These problems are particularly pronounced when faced with high levels of indels and divergence. To address these issues, we trained and tested convolutional neural network models on simulated data and achieved higher accuracy in detecting selection across a specific range of phylogenetic scenarios and evolutionary modes. This advantage is particularly evident when performing inference on noisy data prone to misalignments. Our method shows some ability to account for these errors, where most statistical frameworks fail to do so in a tractable manner. We explore the generalizability of our convolutional neural network models to unseen evolutionary scenarios and identify future avenues to achieve broader utility. Once trained, our convolutional neural network model is faster at test time, making it a scalable alternative to traditional statistical methods for large-scale, multigene analyses. In addition to binary classification (inference of the presence or absence of positive selection during the evolution of the sequences), we use saliency maps to understand what the model learns and observe how this could be leveraged for sitewise inference of positive selection.

Neural Networks, Computer↗

Disease candidate genes prediction using positive labeled and unlabeled instances.

Identifying disease genes and understanding their performance is critical in producing drugs for genetic diseases. Nowadays, laboratory approaches are not only used for disease gene identification but also using computational approaches like machine learning are becoming considerable for this purpose. In machine learning methods, researchers can only use two data types (disease genes and unknown genes) to predict disease candidate genes. Notably, there is no source for the negative data set. The proposed method is a two-step process: The first step is the extraction of reliable negative genes from a set of unlabeled genes by one-class learning and a filter based on distance indicators from known disease genes; this step is performed separately for each disease. The second step is the learning of a binary model using causing genes of each disease as a positive learning set and the reliable negative genes extracted from that disease. Each gene in the unlabeled gene's production and ranking step is assigned a normalized score using two filters and a learned model. Consequently, disease genes are predicted and ranked. The proposed method evaluation of various six diseases and Cancer class indicates better results than other studies.

Humans↗

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting.

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

ClinVar↗

FluxRETAP: a REaction TArget Prioritization genome-scale modeling technique for selecting genetic targets.

MOTIVATION: Metabolic engineering is rapidly evolving as a result of new advances in synthetic biology tools and automation platforms that enable high throughput strain construction, as well as the development of machine learning tools (ML) for biology. However, selecting genetic engineering targets that effectively guide the metabolic engineering process is still challenging. ML can provide predictive power for synthetic biology, but current technical limitations prevent the independent use of ML approaches without previous biological knowledge. RESULTS: Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale models for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing the production of a desired metabolite. This method can provide a list of desirable engineering targets that can be combined with current ML pipelines. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production, 50% of targets that experimentally improved taxadiene production in E. coli and ∼60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida, while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets. AVAILABILITY AND IMPLEMENTATION: FluxRETAP is implemented in python and released under the creative commons license. The implementation and code are freely available at: https://github.com/JBEI/FluxRETAP.

Escherichia coli↗

Blood-based DNA methylation and exposure risk scores predict PTSD with high accuracy in military and civilian cohorts.

BACKGROUND: Incorporating genomic data into risk prediction has become an increasingly useful approach for rapid identification of individuals most at risk for complex disorders such as PTSD. Our goal was to develop and validate Methylation Risk Scores (MRS) using machine learning to distinguish individuals who have PTSD from those who do not. METHODS: Elastic Net was used to develop three risk score models using a discovery dataset (n = 1226; 314 cases, 912 controls) comprised of 5 diverse cohorts with available blood-derived DNA methylation (DNAm) measured on the Illumina Epic BeadChip. The first risk score, exposure and methylation risk score (eMRS) used cumulative and childhood trauma exposure and DNAm variables; the second, methylation-only risk score (MoRS) was based solely on DNAm data; the third, methylation-only risk scores with adjusted exposure variables (MoRSAE) utilized DNAm data adjusted for the two exposure variables. The potential of these risk scores to predict future PTSD based on pre-deployment data was also assessed. External validation of risk scores was conducted in four independent cohorts. RESULTS: The eMRS model showed the highest accuracy (92%), precision (91%), recall (87%), and f1-score (89%) in classifying PTSD using 3730 features. While still highly accurate, the MoRS (accuracy = 89%) using 3728 features and MoRSAE (accuracy = 84%) using 4150 features showed a decline in classification power. eMRS significantly predicted PTSD in one of the four independent cohorts, the BEAR cohort (beta = 0.6839, p-0.003), but not in the remaining three cohorts. Pre-deployment risk scores from all models (eMRS, beta = 1.92; MoRS, beta = 1.99 and MoRSAE, beta = 1.77) displayed a significant (p < 0.001) predictive power for post-deployment PTSD. CONCLUSION: Results, especially those from the eMRS, reinforce earlier findings that methylation and trauma are interconnected and can be leveraged to increase the correct classification of those with vs. without PTSD. Moreover, our models can potentially be a valuable tool in predicting the future risk of developing PTSD. As more data become available, including additional molecular, environmental, and psychosocial factors in these scores may enhance their accuracy in predicting the condition and, relatedly, improve their performance in independent cohorts.

DNA methylation↗

Representation and semiautomatic acquisition of medical knowledge in CADIAG-1 and CADIAG-2.

CADIAG-1 and CADIAG-2 (Computer-Assisted DIAGnosis) are medical expert systems especially designed for ill-defined areas such as internal medicine. Both systems are being tested in the setting of a medical information system. With respect to their knowledge representation, CADIAG-1 has obvious advantages in totally ill-defined areas such as syndromes in internal medicine, whereas CADIAG-2 seems more suited for domains with basic laboratory programs, e.g., hepatology or gall bladder and bile duct diseases. The formalization of relationships between medical entities led to first-order predicate calculus formulas in the case of CADIAG-1 and to a model based on fuzzy set theory in the case of CADIAG-2. In both systems two kinds of relationships between medical entities are considered: (1) necessity of occurrence and (2) sufficiency of occurrence. Statistical interpretations using the 2 X 2 table paradigm yield a way to calculate these relationships automatically from samples of patient data. Results obtained by exploiting 3530 patient records from a rheumatological hospital are presented. The described application is a machine-learning program that allows inductive learning from examples under statistical uncertainty.

Artificial Intelligence↗

LYCEUM: learning to call copy number variants on low-coverage ancient genomes.

MOTIVATION: Copy number variants (CNVs) are pivotal in driving phenotypic variation that facilitates species adaptation. They are significant contributors to various disorders, making ancient genomes crucial for uncovering the genetic origins of disease susceptibility across populations. However, detecting CNVs in ancient DNA (aDNA) samples poses substantial challenges due to several factors: (i) aDNA is often highly degraded; (ii) contamination from microbial DNA and DNA from closely related species introduces additional noise into sequencing data; and finally, (iii) the typically low-coverage of aDNA renders accurate CNV detection particularly difficult. Conventional CNV calling algorithms, which are optimized for high-coverage read-depth signals, underperform under such conditions. RESULTS: To address these limitations, we introduce LYCEUM, the first machine learning-based CNV caller for aDNA. To overcome challenges related to data quality and scarcity, we employ a two-step training strategy. First, the model is pre-trained on whole genome sequencing data from the 1000 Genomes Project, teaching it CNV-calling capabilities similar to conventional methods. Next, the model is fine-tuned using high-confidence CNV calls derived from only a few existing high-coverage aDNA samples. During this stage, the model adapts to making CNV calls based on the downsampled read depth signals of the same aDNA samples. LYCEUM achieves accurate detection of CNVs even in typically low-coverage ancient genomes. We also observe that the segmental deletion calls made by LYCEUM show correlation with the demographic history of the samples and exhibit patterns of negative selection inline with natural selection. AVAILABILITY AND IMPLEMENTATION: LYCEUM is available at https://github.com/ciceklab/LYCEUM.

DNA Copy Number Variations↗

Blood-based DNA methylation and exposure risk scores predict PTSD with high accuracy in military and civilian cohorts.

BACKGROUND: Incorporating genomic data into risk prediction has become an increasingly popular approach for rapid identification of individuals most at risk for complex disorders such as PTSD. Our goal was to develop and validate Methylation Risk Scores (MRS) using machine learning to distinguish individuals who have PTSD from those who do not. METHODS: Elastic Net was used to develop three risk score models using a discovery dataset (n&#x2009;=&#x2009;1226; 314 cases, 912 controls) comprised of 5 diverse cohorts with available blood-derived DNA methylation (DNAm) measured on the Illumina Epic BeadChip. The first risk score, exposure and methylation risk score (eMRS) used cumulative and childhood trauma exposure and DNAm variables; the second, methylation-only risk score (MoRS) was based solely on DNAm data; the third, methylation-only risk scores with adjusted exposure variables (MoRSAE) utilized DNAm data adjusted for the two exposure variables. The potential of these risk scores to predict future PTSD based on pre-deployment data was also assessed. External validation of risk scores was conducted in four independent cohorts. RESULTS: The eMRS model showed the highest accuracy (92%), precision (91%), recall (87%), and f1-score (89%) in classifying PTSD using 3730 features. While still highly accurate, the MoRS (accuracy&#x2009;=&#x2009;89%) using 3728 features and MoRSAE (accuracy&#x2009;=&#x2009;84%) using 4150 features showed a decline in classification power. eMRS significantly predicted PTSD in one of the four independent cohorts, the BEAR cohort (beta&#x2009;=&#x2009;0.6839, p=0.006), but not in the remaining three cohorts. Pre-deployment risk scores from all models (eMRS, beta&#x2009;=&#x2009;1.92; MoRS, beta&#x2009;=&#x2009;1.99 and MoRSAE, beta&#x2009;=&#x2009;1.77) displayed a significant (p&#x2009;<&#x2009;0.001) predictive power for post-deployment PTSD. CONCLUSION: The inclusion of exposure variables adds to the predictive power of MRS. Classification-based MRS may be useful in predicting risk of future PTSD in populations with anticipated trauma exposure. As more data become available, including additional molecular, environmental, and psychosocial factors in these scores may enhance their accuracy in predicting PTSD and, relatedly, improve their performance in independent cohorts.

Humans↗

Boosted mixture of experts: an ensemble learning scheme.

We present a new supervised learning procedure for ensemble machines, in which outputs of predictors, trained on different distributions, are combined by a dynamic classifier combination model. This procedure may be viewed as either a version of mixture of experts (Jacobs, Jordan, Nowlan, & Hintnon, 1991), applied to classification, or a variant of the boosting algorithm (Schapire, 1990). As a variant of the mixture of experts, it can be made appropriate for general classification and regression problems by initializing the partition of the data set to different experts in a boostlike manner. If viewed as a variant of the boosting algorithm, its main gain is the use of a dynamic combination model for the outputs of the networks. Results are demonstrated on a synthetic example and a digit recognition task from the NIST database and compared with classifical ensemble approaches.

Algorithms↗

Comparison of classic statistical methods and machine learning approaches to classify readiness.

MOTIVATION: Predicting physical and cognitive readiness in warfighters is critical for mission success. These predictions can be improved by identifying key biomarkers using multiple omics modalities. The MASTR-E study conducted by McKetney and colleagues is one of the most comprehensive multi-omics studies of saliva samples collected from warfighters, which also applied classic linear statistical (CLS) techniques to discover key biomarkers of readiness. Aligning with McKetney et al.'s assumptions, we operationalize readiness as a binary proxy, where pre-mission samples are labeled as "ready" to reflect a rested, unstressed physiological baseline, while post-mission samples are labeled "not ready" to reflect cumulative physical and cognitive load from the mission. As such, readiness here is not a direct biological or physiological construct, but an inferred state likely dominated by stress-related physiological changes. This assumption and definition is discussed further in the Introduction and Limitations sections. Here, we apply machine learning (ML) analyses to better assess generalizability, consider hidden interactions, and identify nonlinear patterns in the data. We investigated whether ML approaches could predict readiness and identify relevant biomarkers. ML models were trained on proteomics-only or metabolomics-only datasets to classify participants as ready or not ready and important model features were considered as putative biomarkers. Training and testing datasets were curated for two objectives: (i) recognize biomolecular signatures indicative of readiness within the same donor and (ii) assess generalizability across warfighters by withholding donors for testing. RESULTS: Proteomics-based models achieved AUCs of 0.907&#x2009;&#xb1;&#x2009;0.034 and 0.860&#x2009;&#xb1;&#x2009;0.063 for Objectives 1 and 2, respectively. Metabolomics-based models achieved Objective 1 AUC of 0.994&#x2009;&#xb1;&#x2009;0.007 and Objective 2 AUC of 0.993&#x2009;&#xb1;&#x2009;0.010. Comparative analysis with existing literature validates the model's feature importances, but the identified putative biomarkers significantly differ from those discovered through CLS analyses, as only one ML-identified biomarker overlapping with those identified through CLS methods. We show that these ML models and identified features are more robust to noise and generalizable across participants than those identified using CLS methods. AVAILABILITY: The analysis pipelines are provided as Jupyter notebooks, including all code and documentation, and are available publicly on GitHub at {https://github.com/netrias/ReadinessClassification}.

Machine Learning↗

Pan-cancer multi-omics machine learning defines a lactylation-associated immune-excluded tumor state with proteomic and experimental corroboration.

BACKGROUND: Histone lactylation links lactate metabolism to chromatin regulation, but whether lactylation-program-associated transcriptional patterns delineate recurrent pan-cancer tumor states remains unclear. METHODS: We integrated mRNA, lncRNA, and miRNA profiles from 9712 TCGA tumors across 33 cancer types with GTEx references, six GEO cohorts, IMvigor210, and an institutional clear-cell renal cell carcinoma (ccRCC) cohort used for exploratory DIA-NN proteomic corroboration. Random-effects co-expression meta-analysis, multi-omics consensus clustering, regulon inference, immune deconvolution, TIDE, oncoPredict, and SHAP-based machine learning were applied. hsa-miR-431-5p was functionally evaluated as a proof-of-concept CS2-associated miRNA in bladder cancer models. RESULTS: LacCoEx-Atlas comprised 398,491 lactylation-related co-expression pairs across 24,667 RNA features under a random-effects framework (median I&#xb2; = 88.6%). Consensus clustering identified two subtypes: CS2 showed glycolytic-mesenchymal-immune-excluded features, M2 macrophage enrichment, CD8&#x207a; T-cell depletion, elevated HDAC4/NSD3/KDM6B activity, and worse survival, whereas CS1 showed oxidative, sirtuin-active programs. CS2 had fewer predicted ICI responders (18.3% vs. 52.0%) and a lower observed ORR in IMvigor210 (15.3% vs. 24.0%). oncoPredict identified NU7441 as a hypothesis-generating CS2-associated sensitivity signal (Hedges' g = 1.17). DIA-NN proteomics in 50 ccRCC specimens provided exploratory support for CS2-associated hypoxia, ECM degradation, and metastasis programs. The 10-feature mRNA LARItools model achieved an apparent AUC of 0.9413, while a separate multi-omics model achieved 0.971; neither was independently validated. LARItools reproduced prognostic separation across six GEO cohorts. miR-431-5p promoted malignant phenotypes and EMT in bladder cancer cells, with concordant CMU4h expression findings. CONCLUSIONS: Lactylation-program-associated transcriptional patterns delineate a recurrent immune-excluded pan-cancer tumor state associated with adverse prognosis, reduced predicted immunotherapy responsiveness, exploratory single-cancer protein-level support, and testable DNA damage response-targeting hypotheses. LacCoEx-Atlas and LARItools provide open resources for lactylation-program-associated tumor-state stratification and future translational research.

Humans↗

An interpretable deep learning framework uncovers features governing CRISPR-Cas9 genome-editing efficiency.

MOTIVATION: CRISPR-Cas9 genome-editing efficiency is strongly influenced by the sequence composition and positional context of single-guide RNAs (sgRNAs). Although numerous deep learning-based models have been developed to predict Cas9 efficiency from sgRNA sequences, most operate as black boxes, offering limited insight into the sequence determinants underlying Cas9 activity. In addition, previous studies often overlook how the positional context of sequence motifs within sgRNAs influences their effects on Cas9 binding or cleavage. RESULTS: We introduce DeepCC9, an interpretable machine learning framework that combines explicit sequence feature extraction with a residual block-based deep architecture to improve interpretability and identify composition- and position-based motifs governing Cas9 genome-editing efficiency. We applied this method to multiple Cas9 variant datasets, achieving superior predictive performance compared with existing methods while enabling direct interpretation of sequence motifs and their positional effects. Our analysis uncovered 74 sequence motifs enriched or depleted at specific positions within sgRNAs and strongly associated with Cas9 efficiency, providing mechanistic insight into sequence features that influence guide performance. Together, these results establish DeepCC9 as a generalizable and interpretable framework for modeling sequence-function relationships and advancing the understanding of the sequence determinants underlying CRISPR-Cas9 genome editing. AVAILABILITY AND IMPLEMENTATION: The authors have implemented their algorithm in the Python programming language (version 3.X), which is accessible using (https://zenodo.org/records/20073890).

Deep Learning↗