Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Integrated transcriptome analysis and machine learning to construct a homeostatic model of acetylation for bladder cancer and validate the key gene CES1.

BACKGROUND: Bladder cancer (BLCA) is one of the most common malignant tumors of the urinary system. Protein acetylation (PA) plays a critical role in regulating multiple biological processes (BPs), cellular homeostasis, and cancer-related signaling pathways. This study aimed to construct a homeostatic model of acetylation for BLCA using integrated transcriptome analysis and machine learning and to validate the key gene CES1. METHODS: RNA sequencing (RNA-seq) and clinical data were obtained from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases. Acetylation-related differentially expressed genes (DEGs) in BLCA were screened using differential expression analysis (DEA). An acetylation homeostatic model was constructed via univariate, machine learning-based least absolute shrinkage and selection operator (LASSO) and multivariate Cox regression analyses, followed by validation in multiple cohorts. Single-cell RNA-seq analysis was used to explore gene expression patterns in diverse cell types. Enrichment analysis (EA), immune infiltration, and drug sensitivity analysis (DSA) were performed to characterize molecular features of different risk groups. Finally, the biological function of CES1 as the key gene was verified by in vitro knockdown experiments. RESULTS: We established a robust acetylation homeostatic model consisting of five genes, which effectively predicted overall survival (OS) and served as an independent prognostic factor in BLCA. High-risk patients showed significantly poorer prognosis, distinct immune infiltration profiles, and differential drug sensitivity. CES1 was identified and validated as the key gene in this model, which was highly expressed in BLCA and associated with poor prognosis. Knockdown of CES1 markedly suppressed cell proliferation, invasion, and migration, and reduced intracellular coenzyme A (CoA) levels, thereby regulating PA homeostasis. CONCLUSIONS: We developed and validated a novel acetylation homeostatic model for survival stratification and personalized treatment guidance in BLCA, based on integrated transcriptome analysis and machine learning. CES1 is closely associated with intracellular CoA levels and the malignant progression of BLCA. Its potential association with PA homeostasis requires further mechanistic validation, and it may act as a candidate therapeutic biomarker for BLCA.

Bladder cancer (BLCA)↗

Comparison between neural networks and multiple logistic regression to predict acute coronary syndrome in the emergency room.

OBJECTIVE: Patients with suspicion of acute coronary syndrome (ACS) are difficult to diagnose and they represent a very heterogeneous group. Some require immediate treatment while others, with only minor disorders, may be sent home. Detecting ACS patients using a machine learning approach would be advantageous in many situations. METHODS AND MATERIALS: Artificial neural network (ANN) ensembles and logistic regression models were trained on data from 634 patients presenting an emergency department with chest pain. Only data immediately available at patient presentation were used, including electrocardiogram (ECG) data. The models were analyzed using receiver operating characteristics (ROC) curve analysis, calibration assessments, inter- and intra-method variations. Effective odds ratios for the ANN ensembles were compared with the odds ratios obtained from the logistic model. RESULTS: The ANN ensemble approach together with ECG data preprocessed using principal component analysis resulted in an area under the ROC curve of 80%. At the sensitivity of 95% the specificity was 41%, corresponding to a negative predictive value of 97%, given the ACS prevalence of 21%. Adding clinical data available at presentation did not improve the ANN ensemble performance. Using the area under the ROC curve and model calibration as measures of performance we found an advantage using the ANN ensemble models compared to the logistic regression models. CONCLUSION: Clinically, a prediction model of the present type, combined with the judgment of trained emergency department personnel, could be useful for the early discharge of chest pain patients in populations with a low prevalence of ACS.

Acute Disease↗

Using surrogate modeling in the prediction of fibrinogen adsorption onto polymer surfaces.

We present a Surrogate (semiempirical) Model for prediction of protein adsorption onto the surfaces of biodegradable polymers that have been designed for tissue engineering applications. The protein used in these studies, fibrinogen, is known to play a key role in blood clotting. Therefore, fibrinogen adsorption dictates the performance of implants exposed to blood. The Surrogate Model combines molecular modeling, machine learning and an Artificial Neural Network. This novel approach includes an accounting for experimental error using a Monte Carlo analysis. Briefly, measurements of human fibrinogen adsorption were obtained for 45 polymers. A total of 106 molecular descriptors were generated for each polymer. Of these, 102 descriptors were computed using the Molecular Operating Environment (MOE) software based upon the polymer chemical structures, two represented different monomer types, and two were measured experimentally. The Surrogate Model was developed in two stages. In the first stage, the three descriptors with the highest correlation to adsorption were determined by calculating the information gain of each descriptor. Here a Monte Carlo approach enabled a direct assessment of the effect of the experimental uncertainty on the results. The three highest-ranking descriptors, defined as those with the highest information gain for the sample set, were then selected as the input variables for the second stage, an Artificial Neural Network (ANN) to predict fibrinogen adsorption. The ANN was trained using one-half of the experimental data set (the training set) selected at random. The effect of experimental error on predictive capability was again explored using a Monte Carlo analysis. The accuracy of the ANN was assessed by comparison of the predicted values for fibrinogen adsorption with the experimental data for the remaining polymers (the validation set). The mean value of the Pearson correlation coefficient for the validation data sets was 0.54 +/- 0.12. The average root-mean-square (relative) error in prediction for the validation data sets is 38%. This is an order of magnitude less than the range of experimental values (i.e., 366%) and compares favorably with the average percent relative standard deviation of the experimental measurements (i.e., 17.9%). The effects of each of the user-defined parameters in the ANN were explored. None were observed to have a significant effect on the results. Thus, the Surrogate Model can be used to accurately and unambiguously identify polymers whose fibrinogen absorption is at the limits of the range (i.e., low or high) which is an essential requirement for assessing polymers for regenerative tissue applications.

Adsorption↗

Integration of Infant Metabolite, Genetic, and Islet Autoimmunity Signatures to Predict Type 1 Diabetes by Age 6 Years.

CONTEXT: Biomarkers that can accurately predict risk of type 1 diabetes (T1D) in genetically predisposed children can facilitate interventions to delay or prevent the disease. OBJECTIVE: This work aimed to determine if a combination of genetic, immunologic, and metabolic features, measured at infancy, can be used to predict the likelihood that a child will develop T1D by age 6 years. METHODS: Newborns with human leukocyte antigen (HLA) typing were enrolled in the prospective birth cohort of The Environmental Determinants of Diabetes in the Young (TEDDY). TEDDY ascertained children in Finland, Germany, Sweden, and the United States. TEDDY children were either from the general population or from families with T1D with an HLA genotype associated with T1D specific to TEDDY eligibility criteria. From the TEDDY cohort there were 702 children will all data sources measured at ages 3, 6, and 9 months, 11.4% of whom progressed to T1D by age 6 years. The main outcome measure was a diagnosis of T1D as diagnosed by American Diabetes Association criteria. RESULTS: Machine learning-based feature selection yielded classifiers based on disparate demographic, immunologic, genetic, and metabolite features. The accuracy of the model using all available data evaluated by the area under a receiver operating characteristic curve is 0.84. Reducing to only 3- and 9-month measurements did not reduce the area under the curve significantly. Metabolomics had the largest value when evaluating the accuracy at a low false-positive rate. CONCLUSION: The metabolite features identified as important for progression to T1D by age 6 years point to altered sugar metabolism in infancy. Integrating this information with classic risk factors improves prediction of the progression to T1D in early childhood.

Autoantibodies↗

Boolean matrix logic programming for active learning of gene functions in genome-scale metabolic network models.

Reasoning about hypotheses and updating knowledge through empirical observations are central to scientific discovery. In this work, we applied logic-based machine learning methods to drive biological discovery by guiding experimentation. Genome-scale metabolic network models (GEMs) - comprehensive representations of metabolic genes and reactions - are widely used to evaluate genetic engineering of biological systems. However, GEMs often fail to accurately predict the behaviour of genetically engineered cells, primarily due to incomplete annotations of gene interactions. The task of learning the intricate genetic interactions within GEMs presents computational and empirical challenges. To efficiently predict using GEM, we describe a novel approach called Boolean Matrix Logic Programming (BMLP) by leveraging Boolean matrices to evaluate large logic programs. We developed a new system, [Formula: see text], which guides cost-effective experimentation and uses interpretable logic programs to encode a state-of-the-art GEM of a model bacterial organism. Notably, [Formula: see text] successfully learned the interaction between a gene pair with fewer training examples than random experimentation, overcoming the increase in experimental design space. [Formula: see text] enables rapid optimisation of metabolic models to reliably engineer biological systems for producing useful compounds. It offers a realistic approach to creating a self-driving lab for biological discovery, which would then facilitate microbial engineering for practical applications.

Active learning↗

Feature subset selection by genetic algorithms and estimation of distribution algorithms. A case study in the survival of cirrhotic patients treated with TIPS.

The transjugular intrahepatic portosystemic shunt (TIPS) is an interventional treatment for cirrhotic patients with portal hypertension. In the light of our medical staff's experience, the consequences of TIPS are not homogeneous for all the patients and a subgroup dies in the first 6 months after TIPS placement. Actually, there is no risk indicator to identify this subgroup of patients before treatment. An investigation for predicting the survival of cirrhotic patients treated with TIPS is carried out using a clinical database with 107 cases and 77 attributes. Four supervised machine learning classifiers are applied to discriminate between both subgroups of patients. The application of several feature subset selection (FSS) techniques has significantly improved the predictive accuracy of these classifiers and considerably reduced the amount of attributes in the classification models. Among FSS techniques, FSS-TREE, a new randomized algorithm inspired on the new EDA (estimation of distribution algorithm) paradigm has obtained the best average accuracy results for each classifier.

Algorithms↗

Proteomic mass spectra classification using decision tree based ensemble methods.

MOTIVATION: Modern mass spectrometry allows the determination of proteomic fingerprints of body fluids like serum, saliva or urine. These measurements can be used in many medical applications in order to diagnose the current state or predict the evolution of a disease. Recent developments in machine learning allow one to exploit such datasets, characterized by small numbers of very high-dimensional samples. RESULTS: We propose a systematic approach based on decision tree ensemble methods, which is used to automatically determine proteomic biomarkers and predictive models. The approach is validated on two datasets of surface-enhanced laser desorption/ionization time of flight measurements, for the diagnosis of rheumatoid arthritis and inflammatory bowel diseases. The results suggest that the methodology can handle a broad class of similar problems.

Algorithms↗

Patient-specific modeling identifies metabolic interventions for reversing glucose use reprogramming in alcohol-associated hepatitis.

Alcoholic hepatitis (AH) is an acute form of alcohol-associated liver disease with very few treatment options. Recent studies highlighted liver metabolic reprogramming in AH as an indicator of severity. We aim at identifying new intervention points to reverse liver metabolic dysregulation across varying degrees of AH. We develop 89 personalized genome-scale metabolic models by integrating a generic human cellular metabolic model with liver transcriptomics data from AH patients with varying disease severity and healthy controls. We grade the AH patients based on the model-predicted level of glycolysis reprogramming and validate the results using published metabolomics data. We test in silico gene knockdown interventions to reverse the aberrant metabolic reprogramming in AH. Knockdown of two glycolytic genes, Hkdc1 and Pkm, significantly rebalance the metabolic fluxes toward a healthy liver metabolic phenotype. We use machine learning on the glycolysis fluxes to develop a quantitative glucose use reprogramming score, which correlates with AH severity and patient-specific responses to in silico gene knockdown interventions. The score was independently validated using a published AH liver transcriptomics dataset. We propose a cellular metabolism-based therapy targeting Hkdc1 and Pkm in the glycolysis pathway as a potential treatment for reversing the aberrant glucose metabolism in AH.

Humans↗

Structure-based prediction of bZIP partnering specificity.

Predicting protein interaction specificity from sequence is an important goal in computational biology. We present a model for predicting the interaction preferences of coiled-coil peptides derived from bZIP transcription factors that performs very well when tested against experimental protein microarray data. We used only sequence information to build atomic-resolution structures for 1711 dimeric complexes, and evaluated these with a variety of functions based on physics, learned empirical weights or experimental coupling energies. A purely physical model, similar to those used for protein design studies, gave reasonable performance. The results were improved significantly when helix propensities were used in place of a structurally explicit model to represent the unfolded reference state. Further improvement resulted upon accounting for residue-residue interactions in competing states in a generic way. Purely physical structure-based methods had difficulty capturing core interactions accurately, especially those involving polar residues such as asparagine. When these terms were replaced with weights from a machine-learning approach, the resulting model was able to correctly order the stabilities of over 6000 pairs of complexes with greater than 90% accuracy. The final model is physically interpretable, and suggests specific pairs of residues that are important for bZIP interaction specificity. Our results illustrate the power and potential of structural modeling as a method for predicting protein interactions and highlight obstacles that must be overcome to reach quantitative accuracy using a de novo approach. Our method shows unprecedented performance in predicting protein-protein interaction specificity accurately using structural modeling and suggests that predicting coiled-coil interactions generally may be within reach.

Basic-Leucine Zipper Transcription Factors↗

Stochastic epigenetic mutation profiles as biomarkers of clinical activity in juvenile idiopathic arthritis: a multi-omic machine learning approach for gene prioritization.

BACKGROUND: Juvenile idiopathic arthritis (JIA) is a rare autoimmune disease arising from a complex interplay between genetic and environmental factors. Epigenetic modifications such as DNA methylation (DNAm) have been described as potential mediators in gene-environment interactions, contributing to immune system dysregulation. Emerging evidence suggests that DNAm profiles also predict therapeutic responses in autoimmune diseases. This study aims to identify epigenetic biomarkers and epigenetic-driven gene expression changes associated with JIA clinical activity. METHODS: We reanalyzed a publicly available dataset of 44 JIA patients, with whole-genome DNAm and gene expression from CD4 + T cells measured at two points: at anti-TNF therapy withdrawal (T0) and eight months later (Tend). At Tend, 30 patients maintained inactive disease (ID) while 14 did not (NO ID). We investigated differences between ID and NO ID patients in the epigenetic mutation load and various epigenetic clocks through linear regression models, and prioritized genomic regions with significantly higher number of epimutations in NO ID patients through machine learning. RESULTS: We found a higher mutation load in NO ID than ID patients, both at T0 and at Tend, with the differences at Tend reaching statistical significance (p = 0.02). In contrast, we found no evidence of association between epigenetic clocks and JIA clinical activity. Using a multi-omic approach, we identified a List of candidate epigenetically-driven differentially expressed genes, 80 up-regulated and 77 down-regulated, in NO ID patients. Finally, comparing our candidate gene list with the Connectivity Map database, we identified new candidate potential therapeutic targets. Key findings were validated in independent datasets: DNAm profiles from CD4 + T cells (56 JIA patients, 57 controls) and transcriptomic data from PBMCs of JIA patients with active or inactive disease, confirming dysregulation of pathways such as TNF-α signaling via NF-kB and TGF-β signaling among others. CONCLUSIONS: We described a significant association of epigenetic mutations with JIA clinical activity, indicating that epigenetic changes might precede clinical symptoms and may serve as biomarkers for early disease monitoring. Further, our results shed light on biomolecular mechanisms of JIA, supporting the development of more effective treatments.

Humans↗

Artificial Intelligence for Diagnosis, Risk Stratification, and Prognosis of Neuroblastoma - A Systematic Review and Meta-Analysis.

PURPOSE: To synthesizes evidence on artificial intelligence (AI) performance in neuroblastoma (NB) diagnosis, risk stratification, prognosis, and genomic characterization. MATERIALS AND METHODS: A systematic review and meta-analysis was conducted following PRISMA 2020 guidelines (PROSPERO: CRD42024539475) across five databases. Meta-analyses used random-effects models with logit-transformed Area Under the Curve (AUCs) and cluster-robust standard errors. AI models were classified as Machine Learning Models (MLM) or Hybrid Nomograms (HN) based on their construction methodology. RESULTS: Of 3,742 articles identified, 53 were included. MLMs demonstrated higher point estimates than radiologists in differential diagnosis (AUC: 0.87 vs. 0.83), though this difference was not statistically significant and carried substantial uncertainty. HNs achieved stronger performance in risk stratification (AUC: 0.87). AI-derived nomograms (AUC: 0.9) and gene signatures (AUC: 0.8) outperformed conventional prognostic markers descriptively. Chemotherapy response prediction remained below clinical utility thresholds across all model types. Only 33.9% of models reported calibration and 24.5% underwent external validation. CONCLUSIONS: AI demonstrates proof-of-concept across multiple NB clinical domains. However, clinical adoption remains premature given persistent gaps in external validation, calibration, dataset size, and pediatric-specific model development. Future studies should test these models prospectively in multicenter pediatric cohorts, ideally through COG or SIOPEN, using shared definitions for diagnosis, risk group, treatment response, and survival outcomes.

Humans↗

Deep learning-based annotation of plant abiotic stress resistance genes for crops.

The declining costs of DNA sequencing have expanded genomic data, crucial for understanding plant abiotic stress responses and crop improvement. However, accurate gene annotation remains challenging. To address this limitation, we propose the PASRGA, a deep learning approach that leverages transfer learning and contrastive learning to annotate genes related to drought, salt, cold, and UV resistance. PASRGA achieves high F1-scores, area under the receiver operating characteristic (AUROC), area under the precision-recall curve (AUPRC), and Matthews correlation coefficient (MCC) in annotating stress resistance genes, significantly outperforming the general protein annotation model CLEAN, the plant phosphatase gene annotation model PF-NET, the top-ranked model in the CAFA5 challenge NetGO 4.0, and four traditional machine learning methods. Its effectiveness was further validated with a salt stress treatment experiment in Eutrema salsugineum. To facilitate crop breeding practices, we utilized PASRGA to annotate the genomes of 17 major crops. To improve accessibility and utility, we incorporated both manually curated and PASRGA-predicted gene data, together with the PASRGA tool, into the PlantASRG database (https://bioinfor.nefu.edu.cn/PlantASRG/). This comprehensive resource aims to support crop breeding initiatives and ensure food security.

Crops, Agricultural↗

Unveiling tumor heterogeneity by single cell RNA-sequencing: From basic considerations to clinical applications.

Tumor heterogeneity-encompassing diverse cellular phenotypes, genomic alterations, and microenvironmental contexts-is a principal barrier to effective cancer therapy. Single-cell RNA sequencing (scRNA-seq) has transformed our ability to resolve this complexity by capturing transcriptomes at single-cell resolution. Here, we review the technical foundations required for high-quality scRNA-seq studies. We then trace the evolution of scRNA-seq platforms from manual micromanipulation to high-throughput systems, and describe the computational pipelines that enable reliable data interpretation. The application of scRNA-seq is exemplarily shown in the context of lung cancer, where single-cell profiling has revealed (i) the clonal and sub-clonal architecture of tumors, (ii) extensive remodeling of the immune microenvironment, iii) key mechanisms underlying resistance to targeted agents and immune-checkpoint blockade, and (iv) the dynamics of neo-antigen-specific T-cell responses. Integrating machine-learning techniques-such as deep-learning classifiers and graph-based models-with single-cell transcriptomic data has markedly sped up biomarker discovery, produced more accurate risk-stratification scores, and enabled the generation of patient-specific therapeutic predictions. We surveyed the major trial registry ClinicalTrials.gov and identified ∼380 ongoing or completed studies that explicitly incorporate scRNA-seq as a correlative or pharmacodynamic endpoint. Overall, the analysis shows that scRNA-seq becomes an increasingly important component of modern trials, providing high-resolution cellular and molecular readouts that complement conventional imaging and bulk-omics endpoints. While key challenges remain, ranging from costs, scalability and need for rigorous validation before routine clinical deployment, ongoing technological advances continue to expand the potential of scRNA-seq as a cornerstone of precision medicine.

Humans↗

GANN: genetic algorithm neural networks for the detection of conserved combinations of features in DNA.

BACKGROUND: The multitude of motif detection algorithms developed to date have largely focused on the detection of patterns in primary sequence. Since sequence-dependent DNA structure and flexibility may also play a role in protein-DNA interactions, the simultaneous exploration of sequence- and structure-based hypotheses about the composition of binding sites and the ordering of features in a regulatory region should be considered as well. The consideration of structural features requires the development of new detection tools that can deal with data types other than primary sequence. RESULTS: GANN (available at http://bioinformatics.org.au/gann) is a machine learning tool for the detection of conserved features in DNA. The software suite contains programs to extract different regions of genomic DNA from flat files and convert these sequences to indices that reflect sequence and structural composition or the presence of specific protein binding sites. The machine learning component allows the classification of different types of sequences based on subsamples of these indices, and can identify the best combinations of indices and machine learning architecture for sequence discrimination. Another key feature of GANN is the replicated splitting of data into training and test sets, and the implementation of negative controls. In validation experiments, GANN successfully merged important sequence and structural features to yield good predictive models for synthetic and real regulatory regions. CONCLUSION: GANN is a flexible tool that can search through large sets of sequence and structural feature combinations to identify those that best characterize a set of sequences.

Algorithms↗

A review on the integration of artificial intelligence into coastal modeling.

With the development of computing technology, mechanistic models are often employed to simulate processes in coastal environments. However, these predictive tools are inevitably highly specialized, involving certain assumptions and/or limitations, and can be manipulated only by experienced engineers who have a thorough understanding of the underlying theories. This results in significant constraints on their manipulation as well as large gaps in understanding and expectations between the developers and practitioners of a model. The recent advancements in artificial intelligence (AI) technologies are making it possible to integrate machine learning capabilities into numerical modeling systems in order to bridge the gaps and lessen the demands on human experts. The objective of this paper is to review the state-of-the-art in the integration of different AI technologies into coastal modeling. The algorithms and methods studied include knowledge-based systems, genetic algorithms, artificial neural networks, and fuzzy inference systems. More focus is given to knowledge-based systems, which have apparent advantages over the others in allowing more transparent transfers of knowledge in the use of models and in furnishing the intelligent manipulation of calibration parameters. Of course, the other AI methods also have their individual contributions towards accurate and reliable predictions of coastal processes. The integrated model might be very powerful, since the advantages of each technique can be combined.

Artificial Intelligence↗

Prediction of estrogen receptor agonists and characterization of associated molecular descriptors by statistical learning methods.

Specific estrogen receptor (ER) agonists have been used for hormone replacement therapy, contraception, osteoporosis prevention, and prostate cancer treatment. Some ER agonists and partial-agonists induce cancer and endocrine function disruption. Methods for predicting ER agonists are useful for facilitating drug discovery and chemical safety evaluation. Structure-activity relationships and rule-based decision forest models have been derived for predicting ER binders at impressive accuracies of 87.1-97.6% for ER binders and 80.2-96.0% for ER non-binders. However, these are not designed for identifying ER agonists and they were developed from a subset of known ER binders. This work explored several statistical learning methods (support vector machines, k-nearest neighbor, probabilistic neural network and C4.5 decision tree) for predicting ER agonists from comprehensive set of known ER agonists and other compounds. The corresponding prediction systems were developed and tested by using 243 ER agonists and 463 ER non-agonists, respectively, which are significantly larger in number and structural diversity than those in previous studies. A feature selection method was used for selecting molecular descriptors responsible for distinguishing ER agonists from non-agonists, some of which are consistent with those used in other studies and the findings from X-ray crystallography data. The prediction accuracies of these methods are comparable to those of earlier studies despite the use of significantly more diverse range of compounds. SVM gives the best accuracy of 88.9% for ER agonists and 98.1% for non-agonists. Our study suggests that statistical learning methods such as SVM are potentially useful for facilitating the prediction of ER agonists and for characterizing the molecular descriptors associated with ER agonists.

Forecasting↗

Identification and Classification of Expressed Orphan Genes, Spurious Orphan Genes, and Conserved Genes in the Human Gut Microbiome.

Orphan genes (OGs)-genes lacking detectable homologs outside a species-are widespread in microbial genomes and are thought to contribute to their adaptation and molecular innovation. However, not all predicted OGs may represent novel functional coding sequences. False positive OGs, also called spurious OGs, can arise from gene prediction errors. We reason that OGs lacking detectable expression are more likely to be spurious. To test this, we combined large-scale metatranscriptomic profiling of the human gut microbiome with machine learning to distinguish expressed OGs from spurious ones and compare them with conserved genes (CGs) found in multiple species. Using nearly 5,000 metatranscriptome libraries, we identified ∼218,000 OGs supported by expression evidence, while ∼330,000 predicted OGs lacked detectable expression and were classified as spurious. We extracted 154 features for sequence, structural, and evolutionary properties for each gene and trained XGBoost classifiers while accounting for genomic representation. The models achieved an area under the receiver operating characteristic curve (AUC) of 0.82 in distinguishing expressed OGs from spurious OGs and an AUC of 0.93 in distinguishing expressed OGs from CGs. Interpretation based on SHAP (SHapley Additive exPlanations) revealed clear biological signals. Particularly, expressed orphans were present in more genomes than spurious ones, and expressed OGs were shorter than CGs. This work improves OG discovery and suggests that expressed OGs differ systematically from CGs and spurious OGs in sequence composition, structural constraints, and evolutionary signals.

Humans↗

Computer-automated dementia screening using a touch-tone telephone.

BACKGROUND: This study investigated the sensitivity and specificity of a computer-automated telephone system to evaluate cognitive impairment in elderly callers to identify signs of early dementia. METHODS: The Clinical Dementia Rating Scale was used to assess 155 subjects aged 56 to 93 years (n = 74, 27, 42, and 12, with a Clinical Dementia Rating Scale score of 0, 0.5, 1, and 2, respectively). These subjects performed a battery of tests administered by an interactive voice response system using standard Touch-Tone telephones. Seventy-four collateral informants also completed an interactive voice response version of the Symptoms of Dementia Screener. RESULTS: Sixteen cognitively impaired subjects were unable to complete the telephone call. Performances on 6 of 8 tasks were significantly influenced by Clinical Dementia Rating Scale status. The mean (SD) call length was 12 minutes 27 seconds (2 minutes 32 seconds). A subsample (n = 116) was analyzed using machine-learning methods, producing a scoring algorithm that combined performances across 4 tasks. Results indicated a potential sensitivity of 82.0% and specificity of 85.5%. The scoring model generalized to a validation subsample (n = 39), producing 85.0% sensitivity and 78.9% specificity. The kappa agreement between predicted and actual group membership was 0.64 (P<.001). Of the 16 subjects unable to complete the call, 11 provided sufficient information to permit us to classify them as impaired. Standard scoring of the interactive voice response-administered Symptoms of Dementia Screener (completed by informants) produced a screening sensitivity of 63.5% and 100% specificity. A lower criterion found a 90.4% sensitivity, without lowering specificity. CONCLUSIONS: Computer-automated telephone screening for early dementia using either informant or direct assessment is feasible. Such systems could provide wide-scale, cost-effective screening, education, and referral services to patients and caregivers.

Aged↗