Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Construction of a sequence motif characteristic of aminergic G protein-coupled receptors.

An approach to discover sequence patterns characteristic of ligand classes is described and applied to aminergic G protein-coupled receptors (GPCRs). Putative ligand-binding residue positions were inferred from considering three lines of evidence: conservation in the subfamily absent or underrepresented in the superfamily, any available mutation data, and the physicochemical properties of the ligand. For aminergic GPCRs, the motif is composed of a conserved aspartic acid in the third transmembrane (TM) domain (rhodopsin position 117) and a conserved tryptophan in the seventh TM domain (rhodopsin position 293); the roles of each are readily justified by molecular modeling of ligand-receptor interactions. This minimally defined motif is an appropriate computational tool for identifying additional, potentially novel aminergic GPCRs from a set of experimentally uncharacterized "orphan" GPCRs, complementing existing sequence matching, clustering, and machine-learning techniques. Motif sensitivity stems from the stepwise addition of residues characteristic of an entire class of ligand (and not tailored for any particular biogenic amine). This sensitivity is balanced by careful consideration of residues (evidence drawn from mutation data, correlation of ligand properties to residue properties, and location with respect to the extracellular face), thereby maintaining specificity for the aminergic class. A number of orphan GPCRs assigned to the aminergic class by this motif were later discovered to be a novel subfamily of trace amine GPCRs, as well as the successful classification of the histamine H4 receptor.

Amino Acid Sequence↗

The Mycobacterium tuberculosis Transposon Sequencing Database (MtbTnDB): A Large-Scale Guide to Genetic Conditional Essentiality.

Characterizing genetic essentiality across various conditions is fundamental for understanding gene function. Transposon sequencing (TnSeq) is a powerful technique to generate genome-wide essentiality profiles in bacteria and has been extensively applied to Mycobacterium tuberculosis (Mtb). Dozens of TnSeq screens have yielded valuable insights into the biology of Mtb in vitro, inside macrophages, and in model host organisms. Despite their value, these Mtb TnSeq profiles have not been standardized or collated into a single, easily searchable database. This results in significant challenges when attempting to query and compare these resources, limiting our ability to obtain a comprehensive and consistent understanding of genetic conditional essentiality in Mtb. We address this problem by building a central repository of publicly available Mtb TnSeq screens, the Mtb transposon sequencing database (MtbTnDB). The MtbTnDB is a living resource that encompasses to date ≈150 standardized TnSeq screens, enabling open access to data, visualizations, and functional predictions through an interactive web app (www.mtbtndb.app). We conduct several statistical analyses on the complete database, such as demonstrating that (i) genes in the same genomic neighborhood have similar TnSeq profiles, and (ii) clusters of genes with similar TnSeq profiles are enriched for genes from similar functional categories. We further analyze the performance of machine learning models trained on TnSeq profiles to predict the functional annotation of orphan genes in Mtb. By facilitating the comparison of TnSeq screens across conditions, the MtbTnDB will accelerate the exploration of conditional genetic essentiality, provide insights into the functional organization of Mtb genes, and help predict gene function in this important human pathogen.

DNA Transposable Elements↗

Ecological Filtering by Tuber Compartments Shapes Stable Core Microbiomes That Underpin Potato Plant Growth Across Environments.

Harnessing plant microbiomes for sustainable agriculture requires understanding not only whether they can boost crop performance, but also how ecological processes govern their assembly, stability, and functional contributions across environments. While we previously showed that seed tuber microbiomes can predict potato vigour using machine learning, it remained unclear how ecological processes shape tuber microbiome stability and functionality across host genotypes, tuber compartments, soil types, and years. Here, we analyzed the national-scale dataset of 240 field-collected potato seedlots, spanning six genotypes, two soil types, and two growing years, with a focus on the spatially distinct heel and eye compartments of the potato tuber. By profiling over 1200 bacterial and fungal communities and linking microbiome composition to plant performance, we show that plant genotype and tuber compartment are the strongest determinants of microbial diversity and composition. Compartment-specific enrichment of functional traits revealed spatial partitioning of microbial functions, with organic compound conversion and nitrogen cycling dominant in the heel, and energy metabolism enriched in the eye. Applying a macroecological abundance-occupancy framework, we identified a stable core microbiome of bacterial and fungal taxa that persisted across all environments and years. These core members were more strongly associated with plant growth-related traits than non-core taxa, and core taxa in different tuber compartments showed distinct correlations with taxa of potential pathogenic relevance. Together, our findings demonstrate that tuber compartments act as ecological filters that structure persistent, functionally specialised microbiomes linked to plant growth-related traits across environments. By providing an ecological and functional framework for compartment-resolved, stable core microbiomes, this study advances mechanistic understanding of plant-microbe interactions and identifies stable microbial partners as promising targets for improving potato resilience and productivity.

Journal Article↗

Rhizosphere Dialogue: Microorganisms Mediated by Root Exudates Alleviate Drought Stress in Grasses.

Drought stress threatens the ecological functions and economic value of grasses, posing a major challenge to their sustainable production. Plants co-evolve with rhizosphere microbial communities, sometimes described as the plant's second genome, that can contribute to drought adaptation. Drought alters root architecture, hormonal and redox regulation and belowground carbon allocation, thereby modifying the quantity and composition of root exudation and reshaping the rhizosphere environment. This review uses the rhizosphere dialogue as an integrative framework to link these plant responses with microbial recruitment and subsequent feedback to the host. We summarise three linked stages of this dialogue: drought-induced changes in root exudation; microbial recruitment and colonisation through chemotaxis, attachment, biofilm formation, and root colonisation; and microbiome-mediated feedback that improves plant water relations, hormonal and redox homoeostasis, nutrient acquisition, and root function. We highlight microbial extracellular polymeric substances, 1-aminocyclopropane-1-carboxylate deaminase, and microbial volatile organic compounds as key mediators of drought alleviation. We then discuss how this framework may inform rational synthetic microbial community (SynCom) design, microbiome-informed breeding, artificial intelligence and machine-learning assisted strain prioritisation, rhizosphere legacy effects, and real-time monitoring. Future work should distinguish active exudate-mediated recruitment from drought-driven environmental filtering and integrate multi-omics, plant genetics, functional validation, and multi-location field trials to determine whether rhizosphere dialogue can become a predictive framework for climate-resilient grass production.

drought stress↗

Donor Microbiota Features Associated With Liver Transplant Recipient Infectious Complications: A Pilot Study Using Deep Intestinal Sampling During Liver Procurement.

BACKGROUND: The gut microbiota of living organ donors has been linked to transplant outcomes. However, little is known about the characteristics of the deceased donor gut microbiota or its potential impact on recipient outcomes. METHODS: We analyzed the deep intestinal microbiota from 24 deceased donors. Samples included luminal stool from the right and left colon as well as bile. Microbial composition was characterized using 16S V4 rRNA sequencing. &#x3b1;- and &#x3b2;-diversity analyses were performed to compare microbial communities between donor enteric sites and against stool samples from 28 healthy community controls, 14 critically ill intensive care comparators, and 12 matched liver transplant recipients. Machine learning models and logistic regression analysis were applied to explore whether features of the donor microbiota could predict recipient post-transplant complications. FINDINGS: The deceased donor microbiota showed an absence of the expected compositional variability between sampling sites, with no significant differences in either &#x3b1;- or &#x3b2;-diversity observed between bile, right and left colonic samples (all p > 0.05). Donor samples exhibited distinct microbial profiles compared with stool from both healthy and ICU comparators, including increased abundance of potential pathogens within the Enterobacteriaceae family (all p < 0.001). Features of the donor microbiota, particularly enrichment of Enterobacteriaceae, were associated with an increased risk of early post-transplant infection in recipients (&#x2264;&#xa0;30 days; p&#xa0;=&#xa0;0.011). INTERPRETATION: The deceased donor gut microbiota may represent a distinct microbial community with potential clinical relevance. Microbial profiling of donor enteric microbiota may help identify recipients at heightened risk of early post-transplant infectious complications.

Enterobacteriaceae↗

Optimized approach to decision fusion of heterogeneous data for breast cancer diagnosis.

As more diagnostic testing options become available to physicians, it becomes more difficult to combine various types of medical information together in order to optimize the overall diagnosis. To improve diagnostic performance, here we introduce an approach to optimize a decision-fusion technique to combine heterogeneous information, such as from different modalities, feature categories, or institutions. For classifier comparison we used two performance metrics: The receiving operator characteristic (ROC) area under the curve [area under the ROC curve (AUC)] and the normalized partial area under the curve (pAUC). This study used four classifiers: Linear discriminant analysis (LDA), artificial neural network (ANN), and two variants of our decision-fusion technique, AUC-optimized (DF-A) and pAUC-optimized (DF-P) decision fusion. We applied each of these classifiers with 100-fold cross-validation to two heterogeneous breast cancer data sets: One of mass lesion features and a much more challenging one of microcalcification lesion features. For the calcification data set, DF-A outperformed the other classifiers in terms of AUC (p < 0.02) and achieved AUC=0.85 +/- 0.01. The DF-P surpassed the other classifiers in terms of pAUC (p < 0.01) and reached pAUC=0.38 +/- 0.02. For the mass data set, DF-A outperformed both the ANN and the LDA (p < 0.04) and achieved AUC=0.94 +/- 0.01. Although for this data set there were no statistically significant differences among the classifiers' pAUC values (pAUC=0.57 +/- 0.07 to 0.67 +/- 0.05, p > 0.10), the DF-P did significantly improve specificity versus the LDA at both 98% and 100% sensitivity (p < 0.04). In conclusion, decision fusion directly optimized clinically significant performance measures, such as AUC and pAUC, and sometimes outperformed two well-known machine-learning techniques when applied to two different breast cancer data sets.

Algorithms↗

Causal protein-signaling networks derived from multiparameter single-cell data.

Machine learning was applied for the automated derivation of causal influences in cellular signaling networks. This derivation relied on the simultaneous measurement of multiple phosphorylated protein and phospholipid components in thousands of individual primary human immune system cells. Perturbing these cells with molecular interventions drove the ordering of connections between pathway components, wherein Bayesian network computational methods automatically elucidated most of the traditionally reported signaling relationships and predicted novel interpathway network causalities, which we verified experimentally. Reconstruction of network models from physiologically relevant primary single cells might be applied to understanding native-state tissue signaling biology, complex drug actions, and dysfunctional signaling in diseased cells.

Algorithms↗

Tenofovir resistance and resensitization.

Human immunodeficiency viruses in 321 samples from tenofovir-naïve patients were retrospectively evaluated for resistance to this nucleotide analogue. All virus strains with insertions between amino acids 67 and 70 of the reverse transcriptase (n = 6) were highly resistant. Virus strains with the Q151M mutation were divided into susceptible (n = 12) and highly resistant (n = 8) viruses. This difference was due to the absence or presence of the K65R mutation, which was confirmed by site-directed mutagenesis. Viral clones with various combinations of the mutations M41L, K70R, L210W, and T215F or T215Y were analyzed for cross-resistance induced by thymidine analogue mutations (TAMs). The levels of increased resistance induced by single, double, and triple mutations at the indicated positions could be ranked as follows: for mutants with single mutations, mutations at positions 41 > 215 > 70; for mutants with double mutations, mutations at positions 41 and 215 > 70 and 215 = 210 and 215 > 41 and 70; for mutants with triple mutations, mutations at positions 41, 210, and 215 > 41, 70, and 215. Viral clones with M184V or M184I exhibited slightly increased susceptibilities to tenofovir (0.7-fold). Almost all clones with TAM-induced resistance were resensitized when M184V was present (P < 0.001). Among the viruses in the clinical samples, the rate of tenofovir resistance significantly increased with the number of TAMs both in the samples with 184M and in those with 184V (P = 0.005 and P = 0.003, respectively). A resensitizing effect of M184V was confirmed for all samples exhibiting at least one TAM (P = 0.03). However, accumulation of at least two TAMs resulted in more than 2.0-fold reduced susceptibility to tenofovir, irrespective of the presence of M184V. Decision tree building, a classical machine learning technique, was used to generate models for the interpretation of mutations with respect to tenofovir resistance. The application of previously proposed cutoffs for a reduced response to therapy and treatment failure demonstrated the central roles of positions 215 and 65 for 1.5- and 4.0-fold reduced susceptibilities, respectively. Thus, clinically relevant resistance may be conferred by the accumulation of TAMs, and the resensitizing effect of M184V should be considered only minor.

Adenine↗

Integrated analysis of established and novel microbial and chemical methods for microbial source tracking.

Several microbes and chemicals have been considered as potential tracers to identify fecal sources in the environment. However, to date, no one approach has been shown to accurately identify the origins of fecal pollution in aquatic environments. In this multilaboratory study, different microbial and chemical indicators were analyzed in order to distinguish human fecal sources from nonhuman fecal sources using wastewaters and slurries from diverse geographical areas within Europe. Twenty-six parameters, which were later combined to form derived variables for statistical analyses, were obtained by performing methods that were achievable in all the participant laboratories: enumeration of fecal coliform bacteria, enterococci, clostridia, somatic coliphages, F-specific RNA phages, bacteriophages infecting Bacteroides fragilis RYC2056 and Bacteroides thetaiotaomicron GA17, and total and sorbitol-fermenting bifidobacteria; genotyping of F-specific RNA phages; biochemical phenotyping of fecal coliform bacteria and enterococci using miniaturized tests; specific detection of Bifidobacterium adolescentis and Bifidobacterium dentium; and measurement of four fecal sterols. A number of potentially useful source indicators were detected (bacteriophages infecting B. thetaiotaomicron, certain genotypes of F-specific bacteriophages, sorbitol-fermenting bifidobacteria, 24-ethylcoprostanol, and epycoprostanol), although no one source identifier alone provided 100% correct classification of the fecal source. Subsequently, 38 variables (both single and derived) were defined from the measured microbial and chemical parameters in order to find the best subset of variables to develop predictive models using the lowest possible number of measured parameters. To this end, several statistical or machine learning methods were evaluated and provided two successful predictive models based on just two variables, giving 100% correct classification: the ratio of the densities of somatic coliphages and phages infecting Bacteroides thetaiotaomicron to the density of somatic coliphages and the ratio of the densities of fecal coliform bacteria and phages infecting Bacteroides thetaiotaomicron to the density of fecal coliform bacteria. Other models with high rates of correct classification were developed, but in these cases, higher numbers of variables were required.

Animals↗

Discrimination of modes of action of antifungal substances by use of metabolic footprinting.

Diploid cells of Saccharomyces cerevisiae were grown under controlled conditions with a Bioscreen instrument, which permitted the essentially continuous registration of their growth via optical density measurements. Some cultures were exposed to concentrations of a number of antifungal substances with different targets or modes of action (sterol biosynthesis, respiratory chain, amino acid synthesis, and the uncoupler). Culture supernatants were taken and analyzed for their "metabolic footprints" by using direct-injection mass spectrometry. Discriminant function analysis and hierarchical cluster analysis allowed these antifungal compounds to be distinguished and classified according to their modes of action. Genetic programming, a rule-evolving machine learning strategy, allowed respiratory inhibitors to be discriminated from others by using just two masses. Metabolic footprinting thus represents a rapid, convenient, and information-rich method for classifying the modes of action of antifungal substances.

Antifungal Agents↗

Definition of the Bacillus subtilis PurR operator using genetic and bioinformatic tools and expansion of the PurR regulon with glyA, guaC, pbuG, xpt-pbuX, yqhZ-folD, and pbuO.

The expression of the pur operon, which encodes enzymes of the purine biosynthetic pathway in Bacillus subtilis, is subject to control by the purR gene product (PurR) and phosphoribosylpyrophosphate. This control is also exerted on the purA and purR genes. A consensus sequence for the binding of PurR, named the PurBox, has been suggested (M. Kilstrup, S. G. Jessing, S. B. Wichmand-Jørgensen, M. Madsen, and D. Nilsson, J. Bacteriol. 180:3900-3906, 1998). To determine whether the expression of other genes might be regulated by PurR, we performed a search for PurBox sequences in the B. subtilis genome sequence and found several candidate PurBoxes. By the use of transcriptional lacZ fusions, five selected genes or operons (glyA, yumD, yebB, xpt-pbuX, and yqhZ-folD), all having a putative PurBox in their upstream regulatory regions, were found to be regulated by PurR. Using a machine-learning algorithm developed for sequence pattern finding, we found that all of the genes identified as being PurR regulated have two PurBoxes in their upstream control regions. The two boxes are divergently oriented, forming a palindromic sequence with the inverted repeats separated by 16 or 17 nucleotides. A computerized search revealed one additional PurR-regulated gene, ytiP. The significance of the tandem PurBox motifs was demonstrated in vivo by deletion analysis and site-directed mutagenesis of the two PurBox sequences located upstream of glyA. All six genes or operons encode enzymes or transporters playing a role in purine nucleotide metabolism. Functional analysis showed that yebB encodes the previously characterized hypoxanthine-guanine permease PbuG and that ytiP encodes another guanine-hypoxanthine permease and is now named pbuO. yumD encodes a GMP reductase and is now named guaC.

Bacillus subtilis↗

Semen-specific genetic characteristics of human immunodeficiency virus type 1 env.

Human immunodeficiency virus type 1 (HIV-1) in the male genital tract may comprise virus produced locally in addition to virus transported from the circulation. Virus produced in the male genital tract may be genetically distinct, due to tissue-specific cellular characteristics and immunological pressures. HIV-1 env sequences derived from paired blood and semen samples from the Los Alamos HIV Sequence Database were analyzed to ascertain a male genital tract-specific viral signature. Machine learning algorithms could predict seminal tropism based on env sequences with accuracies exceeding 90%, suggesting that a strong genetic signature does exist for virus replicating in the male genital tract. Additionally, semen-derived viral populations exhibited constrained diversity (P < 0.05), decreased levels of positive selection (P < 0.025), decreased CXCR4 coreceptor utilization, and altered glycosylation patterns. Our analysis suggests that the male genital tract represents a distinct selective environment that contributes to the apparent genetic bottlenecks associated with the sexual transmission of HIV-1.

Computational Biology↗

Divergent microbial preludes to necrotising enterocolitis defined by gut phages and bacterial resistomes.

BACKGROUND: Translating microbiome correlations into robust predictive features for complex gut disorders remains elusive, partly due to oversimplified models of pathogenesis and neglect of the virome, a key player in microbial ecosystems. Necrotising enterocolitis (NEC), a devastating disease of preterm infants with no reliable clinical predictors, exemplifies this challenge. OBJECTIVE: To determine the predictive potential of the gut prophageome and polymicrobial aetiologies for NEC. DESIGN: We applied integrated metagenomic and metatranscriptomic analyses and machine learning to 1825 longitudinal stool samples from 43 preterm infants who later developed NEC and 86 gestational age-matched and birthweight-matched controls across three US hospitals. We characterised gut prophageome acquisitions and their association with clinical exposures, including antibiotics, diet and pharmacotherapies. To predict NEC risk, we integrated pre-onset prophageome, antibacterial resistome and bacteriome profiles with neonatal pathology, stratifying the cohort by disease onset timing (early: &#x2264;40 days; late: >40&#x2009;days) for separate analysis. RESULTS: NEC cases exhibited distinct viral diversity trajectories before disease onset. Early-onset NEC was best predicted by phage-bacterial interaction signatures (75% accuracy, 81% sensitivity). Metatranscriptomics revealed increased phage DNA abundance with low gene expression, suggesting a lysogenic lifestyle that may stabilise pathobionts. These phages encode metabolic genes potentially enhancing pathobiont resilience. Late-onset NEC was best predicted by antibacterial resistome profiles (83% accuracy). CONCLUSION: The gut prophageome serves as both a source of pre-symptomatic predictive signals and an active modulator of NEC pathogenesis, with distinct microbial mechanisms driving early-onset and late-onset disease. These polymicrobial etiologies inform strategies for early detection, risk stratification and the development of microbiome-targeted preventive and therapeutic interventions.

BIOMARKERS↗

Towards discovering structural signatures of protein folds based on logical hidden Markov models.

With the growing number of determined protein structures and the availability of classification schemes, it becomes increasingly important to develop computer methods that automatically extract structural signatures for classes of proteins. In this paper, we introduce and apply a new Machine Learning technique, Logical Hidden Markov Models (LOHMMs), to the task of finding structural signatures of folds according to the classification scheme SCOP. Our results indicate that LOHMMs are applicable to this task and possess several advantages over other approaches.

Algorithms↗

ANN-Spec: a method for discovering transcription factor binding sites with improved specificity.

This work describes ANN-Spec, a machine learning algorithm and its application to discovering un-gapped patterns in DNA sequence. The approach makes use of an Artificial Neural Network and a Gibbs sampling method to define the Specificity of a DNA-binding protein. ANN-Spec searches for the parameters of a simple network (or weight matrix) that will maximize the specificity for binding sequences of a positive set compared to a background sequence set. Binding sites in the positive data set are found with the resulting weight matrix and these sites are then used to define a local multiple sequence alignment. Training complexity is O(lN) where l is the width of the pattern and N is the size of the positive training data. A quantitative comparison of ANN-Spec and a few related programs is presented. The comparison shows that ANN-Spec finds patterns of higher specificity when training with a background data set. The program and documentation are available from the authors for UNIX systems.

Algorithms↗

Textquest: document clustering of Medline abstracts for concept discovery in molecular biology.

We present an algorithm for large-scale document clustering of biological text, obtained from Medline abstracts. The algorithm is based on statistical treatment of terms, stemming, the idea of a 'go-list', unsupervised machine learning and graph layout optimization. The method is flexible and robust, controlled by a small number of parameter values. Experiments show that the resulting document clusters are meaningful as assessed by cluster-specific terms. Despite the statistical nature of the approach, with minimal semantic analysis, the terms provide a shallow description of the document corpus and support concept discovery.

Abstracting and Indexing↗

Radical pruning: a method to construct skeleton radial basis function networks.

Trained radial basis function networks are well-suited for use in extracting rules and explanations because they contain a set of locally tuned units. However, for rule extraction to be useful, these networks must first be pruned to eliminate unnecessary weights. The pruning algorithm cannot search the network exhaustively because of the computational effort involved. It is shown that using multiple pruning methods with smart ordering of the pruning candidates, the number of weights in a radial basis function network can be reduced to a small fraction of the original number. The complexity of the pruning algorithm is quadratic (instead of exponential) in the number of network weights. Pruning performance is shown using a variety of benchmark problems from the University of California, Irvine machine learning database.

Algorithms↗

Assessing rbf networks using DELVE.

In this paper, different methods for training radial basis function (RBF) networks for regression problems are described and illustrated. Then, using data from the DELVE archive, they are empirically compared with each other and with some other well known methods for machine learning. Each of the RBF methods performs well on at least one DELVE task, but none are as consistent as the best of the other non-RBF methods.

Algorithms↗