Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Machine learning-ready genomic biomarkers: ATF3 polymorphisms predict postoperative analgesic demand through AI-compatible phenotyping.

PURPOSE: To determine whether ATF3 polymorphisms can serve as genetic biomarkers for machine learning-based precision analgesia by establishing a genotype-phenotype association suitable for predictive modeling of postoperative opioid requirements. METHODS: In a prospective cohort of 167 adults undergoing abdominal surgery, ATF3 SNPs rs3122721 and rs3125293 were genotyped. A structured dataset architecture was developed to represent genetic profiles as input features for supervised learning models, enabling translational analysis of genotype‑dependent opioid consumption over 72 h. RESULTS: Patients with homozygous genotypes of the ATF3 SNPs had significantly higher opioid requirements than non‑carriers, despite reporting similar subjective pain scores. This consistent genotype‑dependent pattern provided a clinically relevant phenotype suitable for integration into predictive algorithms. CONCLUSION: ATF3 genotyping offers a promising biomarker for computationally informed precision analgesia. By linking genomic variability to clinically meaningful outcomes within a structured clinical and genomic framework, this approach supports the future development of risk-stratified clinical decision-support systems to optimize postoperative pain management.Trial registration ChiCTR1900021991, registered 30 April 2019. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13755-026-00480-9.

ATF3↗

Gaussian processes for machine learning.

Gaussian processes (GPs) are natural generalisations of multivariate Gaussian random variables to infinite (countably or continuous) index sets. GPs have been applied in a large number of fields to a diverse range of ends, and very many deep theoretical analyses of various properties are available. This paper gives an introduction to Gaussian processes on a fairly elementary level with special emphasis on characteristics relevant in machine learning. It draws explicit connections to branches such as spline smoothing models and support vector machines in which similar ideas have been investigated. Gaussian process models are routinely used to solve hard machine learning problems. They are attractive because of their flexible non-parametric nature and computational simplicity. Treated within a Bayesian framework, very powerful statistical methods can be implemented which offer valid estimates of uncertainties in our predictions and generic model selection procedures cast as nonlinear optimization problems. Their main drawback of heavy computational scaling has recently been alleviated by the introduction of generic sparse approximations.13,78,31 The mathematical literature on GPs is large and often uses deep concepts which are not required to fully understand most machine learning applications. In this tutorial paper, we aim to present characteristics of GPs relevant to machine learning and to show up precise connections to other "kernel machines" popular in the community. Our focus is on a simple presentation, but references to more detailed sources are provided.

Algorithms↗

Transcriptome-based high-frequency recurrence index predicts frequent recurrence in non-muscle-invasive bladder cancer after Bacillus Calmette-Guérin therapy.

BACKGROUND: High-frequency recurrence (HfR,&#x2009;&#x2265;&#x2009;2 recurrences) in non-muscle-invasive bladder cancer (NMIBC) poses a significant clinical burden. Current risk models, such as the European Organization for Research and Treatment of Cancer (EORTC), the European Association of Urology (EAU), and the UROMOL classification, offer limited predictive accuracy for identifying patients at risk for frequent recurrence despite appropriate treatment. METHODS: A 75-gene high-frequency recurrence index (HfRI) was constructed by selecting recurrence-associated genes using differential expression and Cox regression analyses. The HfRI was computed as a weighted sum of normalized gene expression values. The model was trained on a discovery cohort and validated in multiple cohorts (n&#x2009;=&#x2009;1379) using machine-learning approaches. Clinical relevance was assessed using recurrence-free survival (RFS) and Cox models, and predictive performance was compared with that of the EORTC, EAU, and UROMOL classifications using the area under the curve (AUC) and the concordance index (c-index). RESULTS: The HfRI robustly stratified patients into high-risk and low-risk groups across six independent NMIBC cohorts. Patients classified as HfRI-high had a significantly greater likelihood of experiencing&#x2009;&#x2265;&#x2009;2 recurrences (&#x3c7;2, p&#x2009;=&#x2009;0.001) and showed markedly reduced RFS (log-rank test, p&#x2009;<&#x2009;0.001). The adverse prognostic effect of the HfRI persisted even among patients treated with BCG therapy (log-rank test, p&#x2009;=&#x2009;0.02). Multivariate analysis revealed that the HfRI was an independent predictor of HfR (HR&#x2009;=&#x2009;2.82, 95% CI&#x2009;=&#x2009;1.89-4.20, p&#x2009;<&#x2009;0.001). Compared with established clinical risk classifiers, the HfRI demonstrated superior predictive performance (AUC&#x2009;=&#x2009;0.736, c-index&#x2009;=&#x2009;0.673) in terms of the EORTC (AUC&#x2009;=&#x2009;0.594), EAU (AUC&#x2009;=&#x2009;0.557) risk groups, and UROMOL2021 (AUC&#x2009;=&#x2009;0.596) classification. Pathway analysis revealed that HfRI-high tumors were characterized by upregulation of cell cycle progression and DNA replication pathways, accompanied by suppression of immune signaling pathways. These biological features provide a mechanistic explanation for the reduced responsiveness to intravesical BCG therapy, underscoring the role of HfRI not only as a predictor of recurrence risk but also as a biomarker capable of identifying patients unlikely to benefit from standard BCG treatment. CONCLUSIONS: HfRI represents a robust, transcriptome-based tool for predicting frequent recurrence in NMIBC patients. The HfRI supports earlier identification of patients at risk of high-frequency recurrence, thereby supporting personalized treatment strategies.

Humans↗

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (&#x3a8;) represents one of the most abundant and conserved RNA modifications. &#x3a8; provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of &#x3a8; sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel &#x3a8; site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA &#x3a8;-site prediction. The &#x3a8; modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA &#x3a8;-site prediction. Meta-PseU offers a new framework for robust &#x3a8;-site identification by using long sequences.

Pseudouridine↗

Combining multi-species genomic data for microRNA identification using a Naive Bayes classifier.

MOTIVATION: Most computational methodologies for microRNA gene prediction utilize techniques based on sequence conservation and/or structural similarity. In this study we describe a new technique, which is applicable across several species, for predicting miRNA genes. This technique is based on machine learning, using the Naive Bayes classifier. It automatically generates a model from the training data, which consists of sequence and structure information of known miRNAs from a variety of species. RESULTS: Our study shows that the application of machine learning techniques, along with the integration of data from multiple species is a useful and general approach for miRNA gene prediction. Based on our experiments, we believe that this new technique is applicable to an extensive range of eukaryotes' genomes. Specific structure and sequence features are first used to identify miRNAs followed by a comparative analysis to decrease the number of false positives (FPs). The resulting algorithm exhibits higher specificity and similar sensitivity compared to currently used algorithms that rely on conserved genomic regions to decrease the rate of FPs.

Algorithms↗

Prediction of primate splice junction gene sequences with a cooperative knowledge acquisition system.

We propose a cooperative conceptual modelling environment in which two agents interact: the machine and the human expert. The former is able to extract knowledge from data using a symbolic-numeric machine learning system, and the latter is able to control the learning process by accepting and validating the machine results, or by criticizing those results or the explanation that the system produces on them. The improvement of the conceptual modelling relies on the cooperation between the two agents. Results obtained with our method on prediction of primate splice junctions sites in genetic sequences are far better than those reported in the literature with other symbolic machine learning systems, and are as better as those obtained with some artificial neural networks methods reported at present. But in opposite to neural networks which lack of argumentation, our system provides the user a plausible explanation of its prediction.

Algorithms↗

Ecologic niche modeling and differentiation of populations of Triatoma brasiliensis neiva, 1911, the most important Chagas' disease vector in northeastern Brazil (hemiptera, reduviidae, triatominae).

Ecologic niche modeling has allowed numerous advances in understanding the geographic ecology of species, including distributional predictions, distributional change and invasion, and assessment of ecologic differences. We used this tool to characterize ecologic differentiation of Triatoma brasiliensis populations, the most important Chagas' disease vector in northeastern Brazil. The species' ecologic niche was modeled based on data from the Fundação Nacional de Saúde of Brazil (1997-1999) with the Genetic Algorithm for Rule-Set Prediction (GARP). This method involves a machine-learning approach to detecting associations between occurrence points and ecologic characteristics of regions. Four independent "ecologic niche models" were developed and used to test for ecologic differences among T. brasiliensis populations. These models confirmed four ecologically distinct and differentiated populations, and allowed characterization of dimensions of niche differentiation. Patterns of ecologic similarity matched patterns of molecular differentiation, suggesting that T. brasiliensis is a complex of distinct populations at various points in the process of speciation.

Animals↗

Prediction of the tissue/blood partition coefficients of organic compounds based on the molecular structure using least-squares support vector machines.

The accurate nonlinear model for predicting the tissue/blood partition coefficients (PC) of organic compounds in different tissues was firstly developed based on least-squares support vector machines (LS-SVM), as a novel machine learning technique, by using the compounds' molecular descriptors calculated from the structure alone and the composition features of tissues. The heuristic method (HM) was used to select the appropriate molecular descriptors and build the linear model. The prediction result of the LS-SVM model is much better than that obtained by HM method and the prediction values of tissue/blood partition coefficients based on the LS-SVM model are in good agreement with the experimental values, which proved that nonlinear model can simulate the relationship between the structural descriptors, the tissue composition and the tissue/blood partition coefficients more accurately as well as LS-SVM was a powerful and promising tool in the prediction of the tissue/blood partition behaviour of compounds. Furthermore, this paper provided a new and effective method for predicting the tissue/blood partition behaviour of the compounds in the different tissues from their structures and gave some insight into structural features related to the partition process of the organic compounds in different tissues.

Least-Squares Analysis↗

Detection of antibiotic heteroresistance in clinical microbiology: current and emerging methodologies.

BACKGROUND: Antibiotic heteroresistance (HR) is characterised by the coexistence of susceptible and resistant subpopulations within an apparently isogenic bacterial isolate. Because routine antimicrobial susceptibility testing (AST) primarily assesses the dominant population, HR may escape detection, potentially leading to discrepancies between laboratory susceptibility categorisation and the underlying bacterial population structure. OBJECTIVES: To provide a critical and practice-oriented evaluation of current and emerging methodologies for HR detection and to discuss their strengths, limitations, and potential for clinical implementation. SOURCES: Narrative review based on PubMed searches, complemented by screening of key reference lists and relevant EUCAST and CLSI documents. Peer-reviewed literature was prioritised. CONTENT: Phenotypic approaches, particularly population analysis profiling, remain the reference method for HR definition, but their labour-intensive workflows, long turnaround times, and limited standardisation restrict routine implementation. Alternative strategies, including modified AST assays, metabolic assays, and single-cell platforms, offer gains in speed or throughput but require broader validation. Molecular approaches such as quantitative PCR, droplet digital PCR, targeted deep sequencing, and whole-genome sequencing improve detection of minority resistance determinants. Emerging computational frameworks, including machine learning models integrating phenotypic and genomic data, represent a promising frontier for scalable HR prediction. IMPLICATIONS: Available evidence supports the clinical relevance of HR, although its association with adverse outcomes varies across bacterial species and antibiotic classes. Harmonised methodologies and clinically validated interpretive criteria are needed to support integration of HR assessment into routine diagnostics. Prospective multicentre studies and further standardisation, including engagement with EUCAST and CLSI, will be important to advance clinical implementation.

Antimicrobial resistance↗

transFold: a web server for predicting the structure and residue contacts of transmembrane beta-barrels.

Transmembrane beta-barrel (TMB) proteins are embedded in the outer membrane of Gram-negative bacteria, mitochondria and chloroplasts. The cellular location and functional diversity of beta-barrel outer membrane proteins makes them an important protein class. At the present time, very few non-homologous TMB structures have been determined by X-ray diffraction because of the experimental difficulty encountered in crystallizing transmembrane (TM) proteins. The transFold web server uses pairwise inter-strand residue statistical potentials derived from globular (non-outer-membrane) proteins to predict the supersecondary structure of TMB. Unlike all previous approaches, transFold does not use machine learning methods such as hidden Markov models or neural networks; instead, transFold employs multi-tape S-attribute grammars to describe all potential conformations, and then applies dynamic programming to determine the global minimum energy supersecondary structure. The transFold web server not only predicts secondary structure and TMB topology, but is the only method which additionally predicts the side-chain orientation of transmembrane beta-strand residues, inter-strand residue contacts and TM beta-strand inclination with respect to the membrane. The program transFold currently outperforms all other methods for accuracy of beta-barrel structure prediction. Available at http://bioinformatics.bc.edu/clotelab/transFold.

Amino Acids↗

Multi&#x2011;omics identification of a novel signature for serous ovarian carcinoma in the context of 3P medicine and based on twelve programmed cell death patterns: a multi-cohort machine learning study.

BACKGROUND: Predictive, preventive, and personalized medicine (PPPM/3PM) is a strategy aimed at improving the prognosis of cancer, and programmed cell death (PCD) is increasingly recognized as a potential target in cancer therapy and prognosis. However, a PCD-based predictive model for serous ovarian carcinoma (SOC) is lacking. In the present study, we aimed to establish a cell death index (CDI)-based model using PCD-related genes. METHODS: We included 1254 genes from 12 PCD patterns in our analysis. Differentially expressed genes (DEGs) from the Cancer Genome Atlas (TCGA) and Genotype-Tissue Expression (GTEx) were screened. Subsequently, 14 PCD-related genes were included in the PCD-gene-based CDI model. Genomics, single-cell transcriptomes, bulk transcriptomes, spatial transcriptomes, and clinical information from TCGA-OV, GSE26193, GSE63885, and GSE140082 were collected and analyzed to verify the prediction model. RESULTS: The CDI was recognized as an independent prognostic risk factor for patients with SOC. Patients with SOC and a high CDI had lower survival rates and poorer prognoses than those with a low CDI. Specific clinical parameters and the CDI were combined to establish a nomogram that accurately assessed patient survival. We used the PCD-genes model to observe differences between high and low CDI groups. The results showed that patients with SOC and a high CDI showed immunosuppression and hardly benefited from immunotherapy; therefore, trametinib_1372 and BMS-754807 may be potential therapeutic agents for these patients. CONCLUSIONS: The CDI-based model, which was established using 14 PCD-related genes, accurately predicted the tumor microenvironment, immunotherapy response, and drug sensitivity of patients with SOC. Thus this model may help improve the diagnostic and therapeutic efficacy of PPPM.

Humans↗

Comprehensive in silico genomics analysis of global trends and host-specific emergence of aminoglycoside resistance in Staphylococcus aureus: a One-Health perspective.

BACKGROUND: Aminoglycosides remain clinically valuable against Staphylococcus aureus. Aminoglycoside resistance in S. aureus represents a critical One Health concern and is primarily driven by aminoglycoside-modifying enzymes (AMEs), which are frequently plasmid-encoded. Although regional studies have provided valuable insights, the global epidemiology of aminoglycoside resistance determinants remains poorly characterized because comprehensive data integrating human, animal, and environmental reservoirs are still lacking. This study addresses this gap by analyzing over 110,000 S. aureus genomes (2000-2025) to map the global resistome, quantify temporal and host-specific trends, and assess the association between genetic determinants and phenotypic resistance. METHODS: We performed a retrospective One Health meta-analysis of 110,309 S. aureus genomes collected between 2000 and 2025 from 128 countries. Genomes were quality-filtered and aminoglycoside resistance determinants were identified using NCBI AMRFinderPlus (v4.0.23). Multilocus sequence typing and host-source harmonization (Human, Animal, Environment, Unknown) enabled clonal and reservoir stratification. Temporal trends in gene prevalence and resistance burden were modeled with robust regression. Geographic and host-associated structuring of key genes was assessed via &#x3c7;2 and enrichment tests. Machine-learning models (elastic-net, random forests, XGBoost) were benchmarked for minimum inhibitory concentration (MIC) prediction via nested cross-validation, with performance evaluated by mean absolute error, RMSE, and SHAP-based feature importance. All analyses were conducted in R and Python using publicly available, de-identified genomic data. RESULTS: Aminoglycoside resistance-associated genes were dominated by modifying enzyme determinants, with ant(6)-Ia, ant(9)-Ia, aph(3')-IIIa, sat4, aadD1, and aac(6')-Ie/aph(2'')-Ia occurring in 14-22% of isolates worldwide. Temporal analysis revealed significant declines in several major determinants, most notably ant(9)-Ia (-2.22 percentage points per year, p&#x2009;<&#x2009;0.001), whereas apmA exhibited a non-significant decreasing trend in animal isolates. Host structuring was marked: human clinical isolates concentrated common determinants, while animal and environmental isolates harbored rare alleles (apmA, spw, str, spd). Geographic mapping confirmed near-universal distribution of common genes but focal restriction of rare ones. Publicly available phenotypic data indicated strong activity of amikacin, whereas gentamicin showed a distinct resistant subpopulation that closely corresponded with AME gene carriage. Genotype-phenotype analyses demonstrated strong concordance, with gene-rich complements predicting resistant MIC strata and absence of determinants predicting susceptibility. Analysis across different gene classes revealed frequent co-occurrence of aminoglycoside resistance genes with determinants from other classes, such as mecA, blaZ, and MLS_B, embedding them within multidrug-resistant (MDR) genomic contexts. CONCLUSION: Over 25&#xa0;years, the prevalence of aminoglycoside resistance-associated genes in S. aureus has declined for several common determinants, while rare veterinary-linked alleles are emerging in animal isolates. Strong genotype-phenotype concordance supports genomic prediction for gentamicin and amikacin, where MIC data are available, although phenotypic confirmation remains essential. The frequent co-occurrence of aminoglycoside resistance genes with other antimicrobial resistance determinants indicates their integration within co-occurrence patterns of MDR genes, defined here as clusters of co-occurring resistance genes often carried on shared mobile genetic elements. These patterns highlight the need for integrated One Health surveillance combining clinical, veterinary, and environmental monitoring with plasmid-context resolution to anticipate emerging threats.

Aminoglycosides↗

Genomic Language Model for Predicting Enhancers and Their Allele-Specific Activity in the Human Genome.

Predicting and deciphering the regulatory logic of enhancers is a challenging problem, due to the intricate sequence features and lack of consistent genetic or epigenetic signatures that can accurately discriminate enhancers from other genomic regions. Recent machine-learning based methods have spotlighted the importance of extracting nucleotide composition of enhancers but failed to learn the sequence context and perform suboptimally. Motivated by advances in genomic language models, we developed DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. We trained two different models, using large collection of enhancers curated from the ENCODE registry of candidate cis-Regulatory Elements. The best fine-tuned model achieved 88.05% accuracy with Matthews correlation coefficient of 76% on independent set aside data. Further, we present the analysis of the predicted enhancers for all chromosomes of the human genome by comparing with the enhancer regions reported in publicly available databases. Finally, we applied DNABERT-Enhancer along with other DNABERT based regulatory genomic region prediction models to predict candidate SNPs with allele-specific enhancer and transcription factor binding activity. The genome-wide enhancer annotations and candidate loss-of-function genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies.

Journal Article↗

Multi-omics dynamic profiling reveals predictive biomarkers for first-line immunochemotherapy in extensive-stage small-cell lung cancer.

BACKGROUND: Extensive-stage small-cell lung cancer (ES-SCLC) is associated with a poor prognosis. Although first-line immunochemotherapy improves clinical outcomes, robust prognostic biomarkers for this treatment modality remain unavailable. The aim of this study was to identify non-invasive, easily accessible, and dynamically monitored biomarkers of ES-SCLC by machine learning integrating serum metabolomics, lipidomics, and proteomics at multiple time points. METHODS: A total of 816 serum samples were collected from ES-SCLC patients receiving first-line immunotherapy combined with chemotherapy or first-line chemotherapy for metabolomics, lipidomics, and proteomics analysis. The immunochemotherapy cohort was randomly divided into training and validation subsets at a 6:4 ratio. Biomarkers were identified using machine learning algorithms, and their prognostic significance was evaluated through receiver operating characteristic (ROC) analysis, Kaplan&#x2013;Meier survival analysis, and multivariate Cox regression. Potential metabolic pathways and mechanisms were further explored via integrated multi-omic analysis. RESULTS: The immunochemotherapy exhibited a prolonged median progression-free survival (PFS) and higher objective response rate (ORR) compared to the chemotherapy group. A total of 5 serum metabolites (uric acid, L-aspartate-semialdehyde, dimethisterone, xanthine, L-cysteine), 6 lipids (Cer d18:1/26:0, Cer d18:2/25:0, SM d18:1/20:1, SM d17:1/25:1, DG O-18:1_16:0, PS 18:0_24:0), and 3 proteins (ACIN1, ACSL4, PHGDH) were identified and constructed into independent prognostic models. Among patients receiving immunochemotherapy, those categorized as low-risk based on the model demonstrated significantly longer PFS compared with those in the high-risk group. These prognostic signatures also retained predictive value in patients who underwent second-line treatment with anlotinib plus immunochemotherapy. Integrated analysis revealed that glycine, serine, and threonine metabolism was the commonly enriched pathway across all three omics layers. Notably, PHGDH (protein), L-aspartate-semialdehyde and L-cysteine (metabolites), and PS (18:0_24:0) (lipid), key elements in this pathway, were all incorporated in the predictive model. In addition, models of the composition of these substances after one cycle of treatment can still predict the prognosis of patients. CONCLUSION: In this study, we constructed and validated a set of non-invasive, dynamically monitorable prognostic models (containing 5 metabolites, 6 lipids, and 3 proteins) using machine learning by integrating multiple time point data from the serum metabolome, lipid panel, and proteome to accurately distinguish the prognostic risk of patients with ES-SCLC receiving immunochemotherapy. PFS was significantly prolonged in patients in the low-risk group, and this model remains predictive in the subsequent second-line treatment with anlotinib in combination with immunochemotherapy. Glycine-serine-threonine metabolic pathway may be the key mechanism, of which PHGDH, L-aspartate semialdehyde, L-cysteine and PS (18:0_24:0) are the core predictors. This study provides the first multi-omics dynamic prognostic tool for ES-SCLC immunochemotherapy and reveals potential therapeutic targets.

Humans↗

A development environment for predictive modelling in foods.

Waikato Environment for Knowledge Analysis (WEKA) is a comprehensive suite of Java class libraries that implement many state-of-the-art machine learning/data mining algorithms. Non-programmers interact with the software via a user interface component called the Knowledge Explorer. Applications constructed from the WEKA class libraries can be run on any computer with a web-browsing capability, allowing users to apply machine learning techniques to their own data regardless of computer platform. This paper describes the user interface component of the WEKA system in reference to previous applications in the predictive modelling of foods.

Algorithms↗

Diversity and complexity of HIV-1 drug resistance: a bioinformatics approach to predicting phenotype from genotype.

Drug resistance testing has been shown to be beneficial for clinical management of HIV type 1 infected patients. Whereas phenotypic assays directly measure drug resistance, the commonly used genotypic assays provide only indirect evidence of drug resistance, the major challenge being the interpretation of the sequence information. We analyzed the significance of sequence variations in the protease and reverse transcriptase genes for drug resistance and derived models that predict phenotypic resistance from genotypes. For 14 antiretroviral drugs, both genotypic and phenotypic resistance data from 471 clinical isolates were analyzed with a machine learning approach. Information profiles were obtained that quantify the statistical significance of each sequence position for drug resistance. For the different drugs, patterns of varying complexity were observed, including between one and nine sequence positions with substantial information content. Based on these information profiles, decision tree classifiers were generated to identify genotypic patterns characteristic of resistance or susceptibility to the different drugs. We obtained concise and easily interpretable models to predict drug resistance from sequence information. The prediction quality of the models was assessed in leave-one-out experiments in terms of the prediction error. We found prediction errors of 9.6-15.5% for all drugs except for zalcitabine, didanosine, and stavudine, with prediction errors between 25.4% and 32.0%. A prediction service is freely available at http://cartan.gmd.de/geno2pheno.html.

Computational Biology↗

Error criteria for cross validation in the context of chaotic time series prediction.

The prediction of a chaotic time series over a long horizon is commonly done by iterating one-step-ahead prediction. Prediction can be implemented using machine learning methods, such as radial basis function networks. Typically, cross validation is used to select prediction models based on mean squared error. The bias-variance dilemma dictates that there is an inevitable tradeoff between bias and variance. However, invariants of chaotic systems are unchanged by linear transformations; thus, the bias component may be irrelevant to model selection in the context of chaotic time series prediction. Hence, the use of error variance for model selection, instead of mean squared error, is examined. Clipping is introduced, as a simple way to stabilize iterated predictions. It is shown that using the error variance for model selection, in combination with clipping, may result in better models.

Journal Article↗

Feature mining and predictive model construction from severe trauma patient's data.

In management of severe trauma patients, trauma surgeons need to decide which patients are eligible for damage control. Such decision may be supported by utilizing models that predict the patient's outcome. The study described in this paper investigates the possibility to construct patient outcome prediction models from retrospective patient's data at the end of initial damage control surgery by using feature mining and machine learning techniques. As the data used comprises rather excessive number of features, special attention was paid to the problem of selecting only the most relevant features. We show that a small subset of features may carry enough information to construct reasonably accurate prognostic models. Furthermore, the techniques used in our study identified two factors, namely the pH value when admitted to ICU and the worst partial active thromboplastin time, to be of highest importance for prediction. This finding is pathophysiologically reasonable and represents two of three major problems with severe trauma patients, metabolic acidosis, hypothermia, and coagulopathy.

Algorithms↗