Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

Knowledge-based computational search for genes associated with the metabolic syndrome.

MOTIVATION: A methodology to search for genes associated with multifactorial diseases by integrating the large amount of accumulated knowledge is seriously needed. A comprehensive understanding derived from a holistic view of gene relationship structures can be gained from our proposed analysis called the cross-subspace analysis (CSA). In this analysis, gene objects are generated by machine learning using their term occurrence patterns in MEDLINE abstracts and the degree of relationship between gene objects is quantified by matching these patterns. RESULTS: Structuralization of relationships of a set of genes was performed using CSA, which were retrieved using the terms, 'obesity', 'diabetes', 'hypertriglyceridemia' and 'hypertension' that refer to diseases comprising metabolic syndrome, on a 2D plane inferring important biomedical concepts from the gene distribution. Then, we prioritized the significance of 6131 well-annotated human genes in terms of the distance on the plane from the centroid of 'metabolic syndrome'-related genes distribution. The validity was confirmed by comparing the knowledge extracted by the ordering with existing medical knowledge.

Abstracting and Indexing↗

Proteomic mass spectra classification using decision tree based ensemble methods.

MOTIVATION: Modern mass spectrometry allows the determination of proteomic fingerprints of body fluids like serum, saliva or urine. These measurements can be used in many medical applications in order to diagnose the current state or predict the evolution of a disease. Recent developments in machine learning allow one to exploit such datasets, characterized by small numbers of very high-dimensional samples. RESULTS: We propose a systematic approach based on decision tree ensemble methods, which is used to automatically determine proteomic biomarkers and predictive models. The approach is validated on two datasets of surface-enhanced laser desorption/ionization time of flight measurements, for the diagnosis of rheumatoid arthritis and inflammatory bowel diseases. The results suggest that the methodology can handle a broad class of similar problems.

Algorithms↗

Analysis of differentially-regulated genes within a regulatory network by GPS genome navigation.

MOTIVATION: A critical challenge of the post-genomic era is to understand how genes are differentially regulated even when they belong to a given network. Because the fundamental mechanism controlling gene expression operates at the level of transcription initiation, computational techniques have been developed that identify cis regulatory features and map such features into expression patterns to classify genes into distinct networks. However, these methods are not focused on distinguishing between differentially regulated genes within a given network. Here we describe an unsupervised machine learning method, termed GPS for gene promoter scan, that discriminates among co-regulated promoters by simultaneously considering both cis-acting regulatory features and gene expression. GPS is particularly useful for knowledge discovery in environments with reduced datasets and high levels of uncertainty. RESULTS: Application of this method to the enteric bacteria Escherichia coli and Salmonella enterica uncovered novel members, as well as regulatory interactions in the regulon controlled by the PhoP protein that were not discovered using previous approaches. The predictions made by GPS were experimentally validated to establish that the PhoP protein uses multiple mechanisms to control gene transcription, and is a central element in a highly connected network. AVAILABILITY: The scripts and programs used in this work are accessible from the gps-tools.wustl.edu website. Data and predictions are available by request.

Algorithms↗

Application of latent semantic analysis to protein remote homology detection.

MOTIVATION: Remote homology detection between protein sequences is a central problem in computational biology. The discriminative method such as the support vector machine (SVM) is one of the most effective methods. Many of the SVM-based methods focus on finding useful representations of protein sequence, using either explicit feature vector representations or kernel functions. Such representations may suffer from the peaking phenomenon in many machine-learning methods because the features are usually very large and noise data may be introduced. Based on these observations, this research focuses on feature extraction and efficient representation of protein vectors for SVM protein classification. RESULTS: In this study, a latent semantic analysis (LSA) model, which is an efficient feature extraction technique from natural language processing, has been introduced in protein remote homology detection. Several basic building blocks of protein sequences have been investigated as the 'words' of 'protein sequence language', including N-grams, patterns and motifs. Each protein sequence is taken as a 'document' that is composed of bags-of-word. The word-document matrix is constructed first. The LSA is performed on the matrix to produce the latent semantic representation vectors of protein sequences, leading to noise-removal and smart description of protein sequences. The latent semantic representation vectors are then evaluated by SVM. The method is tested on the SCOP 1.53 database. The results show that the LSA model significantly improves the performance of remote homology detection in comparison with the basic formalisms. Furthermore, the performance of this method is comparable with that of the complex kernel methods such as SVM-LA and better than that of other sequence-based methods such as PSI-BLAST and SVM-pairwise.

Algorithms↗

Improving missing value estimation in microarray data with gene ontology.

MOTIVATION: Gene expression microarray experiments produce datasets with frequent missing expression values. Accurate estimation of missing values is an important prerequisite for efficient data analysis as many statistical and machine learning techniques either require a complete dataset or their results are significantly dependent on the quality of such estimates. A limitation of the existing estimation methods for microarray data is that they use no external information but the estimation is based solely on the expression data. We hypothesized that utilizing a priori information on functional similarities available from public databases facilitates the missing value estimation. RESULTS: We investigated whether semantic similarity originating from gene ontology (GO) annotations could improve the selection of relevant genes for missing value estimation. The relative contribution of each information source was automatically estimated from the data using an adaptive weight selection procedure. Our experimental results in yeast cDNA microarray datasets indicated that by considering GO information in the k-nearest neighbor algorithm we can enhance its performance considerably, especially when the number of experimental conditions is small and the percentage of missing values is high. The increase of performance was less evident with a more sophisticated estimation method. We conclude that even a small proportion of annotated genes can provide improvements in data quality significant for the eventual interpretation of the microarray experiments. AVAILABILITY: Java and Matlab codes are available on request from the authors. SUPPLEMENTARY MATERIAL: Available online at http://users.utu.fi/jotatu/GOImpute.html.

Algorithms↗

A multi-step approach to time series analysis and gene expression clustering.

MOTIVATION: The huge growth in gene expression data calls for the implementation of automatic tools for data processing and interpretation. RESULTS: We present a new and comprehensive machine learning data mining framework consisting in a non-linear PCA neural network for feature extraction, and probabilistic principal surfaces combined with an agglomerative approach based on Negentropy aimed at clustering gene microarray data. The method, which provides a user-friendly visualization interface, can work on noisy data with missing points and represents an automatic procedure to get, with no a priori assumptions, the number of clusters present in the data. Cell-cycle dataset and a detailed analysis confirm the biological nature of the most significant clusters. AVAILABILITY: The software described here is a subpackage part of the ASTRONEURAL package and is available upon request from the corresponding author. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Artificial Intelligence↗

Optimized multilayer perceptrons for molecular classification and diagnosis using genomic data.

MOTIVATION: Multilayer perceptrons (MLP) represent one of the widely used and effective machine learning methods currently applied to diagnostic classification based on high-dimensional genomic data. Since the dimensionalities of the existing genomic data often exceed the available sample sizes by orders of magnitude, the MLP performance may degrade owing to the curse of dimensionality and over-fitting, and may not provide acceptable prediction accuracy. RESULTS: Based on Fisher linear discriminant analysis, we designed and implemented an MLP optimization scheme for a two-layer MLP that effectively optimizes the initialization of MLP parameters and MLP architecture. The optimized MLP consistently demonstrated its ability in easing the curse of dimensionality in large microarray datasets. In comparison with a conventional MLP using random initialization, we obtained significant improvements in major performance measures including Bayes classification accuracy, convergence properties and area under the receiver operating characteristic curve (A(z)). SUPPLEMENTARY INFORMATION: The Supplementary information is available on http://www.cbil.ece.vt.edu/publications.htm

Biomarkers, Tumor↗

Automated discovery of 3D motifs for protein function annotation.

MOTIVATION: Function inference from structure is facilitated by the use of patterns of residues (3D motifs), normally identified by expert knowledge, that correlate with function. As an alternative to often limited expert knowledge, we use machine-learning techniques to identify patterns of 3-10 residues that maximize function prediction. This approach allows us to test the assumption that residues that provide function are the most informative for predicting function. RESULTS: We apply our method, GASPS, to the haloacid dehalogenase, enolase, amidohydrolase and crotonase superfamilies and to the serine proteases. The motifs found by GASPS are as good at function prediction as 3D motifs based on expert knowledge. The GASPS motifs with the greatest ability to predict protein function consist mainly of known functional residues. However, several residues with no known functional role are equally predictive. For four groups, we show that the predictive power of our 3D motifs is comparable with or better than approaches that use the entire fold (Combinatorial-Extension) or sequence profiles (PSI-BLAST). AVAILABILITY: Source code is freely available for academic use by contacting the authors. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

Functional bioinformatics for Arabidopsis thaliana.

MOTIVATION: The genome of Arabidopsis thaliana, which has the best understood plant genome, still has approximately one-third of its genes with no functional annotation at all from either MIPS or TAIR. We have applied our Data Mining Prediction (DMP) method to the problem of predicting the functional classes of these protein sequences. This method is based on using a hybrid machine-learning/data-mining method to identify patterns in the bioinformatic data about sequences that are predictive of function. We use data about sequence, predicted secondary structure, predicted structural domain, InterPro patterns, sequence similarity profile and expressions data. RESULTS: We predicted the functional class of a high percentage of the Arabidopsis genes with currently unknown function. These predictions are interpretable and have good test accuracies. We describe in detail seven of the rules produced.

Algorithms↗

A computational approach toward label-free protein quantification using predicted peptide detectability.

We propose here a new concept of peptide detectability which could be an important factor in explaining the relationship between a protein's quantity and the peptides identified from it in a high-throughput proteomics experiment. We define peptide detectability as the probability of observing a peptide in a standard sample analyzed by a standard proteomics routine and argue that it is an intrinsic property of the peptide sequence and neighboring regions in the parent protein. To test this hypothesis we first used publicly available data and data from our own synthetic samples in which quantities of model proteins were controlled. We then applied machine learning approaches to demonstrate that peptide detectability can be predicted from its sequence and the neighboring regions in the parent protein with satisfactory accuracy. The utility of this approach for protein quantification is demonstrated by peptides with higher detectability generally being identified at lower concentrations over those with lower detectability in the synthetic protein mixtures. These results establish a direct link between protein concentration and peptide detectability. We show that for each protein there exists a level of peptide detectability above which peptides are detected and below which peptides are not detected in an experiment. We call this level the minimum acceptable detectability for identified peptides (MDIP) which can be calibrated to predict protein concentration. Triplicate analysis of a biological sample showed that these MDIP values are consistent among the three data sets.

Algorithms↗

Pathway analysis using random forests classification and regression.

MOTIVATION: Although numerous methods have been developed to better capture biological information from microarray data, commonly used single gene-based methods neglect interactions among genes and leave room for other novel approaches. For example, most classification and regression methods for microarray data are based on the whole set of genes and have not made use of pathway information. Pathway-based analysis in microarray studies may lead to more informative and relevant knowledge for biological researchers. RESULTS: In this paper, we describe a pathway-based classification and regression method using Random Forests to analyze gene expression data. The proposed methods allow researchers to rank important pathways from externally available databases, discover important genes, find pathway-based outlying cases and make full use of a continuous outcome variable in the regression setting. We also compared Random Forests with other machine learning methods using several datasets and found that Random Forests classification error rates were either the lowest or the second-lowest. By combining pathway information and novel statistical methods, this procedure represents a promising computational strategy in dissecting pathways and can provide biological insight into the study of microarray data. AVAILABILITY: Source code written in R is available from http://bioinformatics.med.yale.edu/pathway-analysis/rf.htm.

Algorithms↗

Genetic attributes of cerebrospinal fluid-derived HIV-1 env.

HIV-1 often invades the CNS during primary infection, eventually resulting in neurological disorders in up to 50% of untreated patients. The CNS is a distinct viral reservoir, differing from peripheral tissues in immunological surveillance, target cell characteristics and antiretroviral penetration. Neurotropic HIV-1 likely develops distinct genotypic characteristics in response to this unique selective environment. We sought to catalogue the genetic features of CNS-derived HIV-1 by analysing 456 clonal RNA sequences of the C2-V3 env subregion generated from CSF and plasma of 18 chronically infected individuals. Neuropsychological performance of all subjects was evaluated and summarized as a global deficit score. A battery of phylogenetic, statistical and machine learning tools was applied to these data to identify genetic features associated with HIV-1 neurotropism and neurovirulence. Eleven of 18 individuals exhibited significant viral compartmentalization between blood and CSF (P < 0.01, Slatkin-Maddison test). A CSF-specific genetic signature was identified, comprising positions 9, 13 and 19 of the V3 loop. The residue at position 5 of the V3 loop was highly correlated with neurocognitive deficit (P < 0.0025, Fisher's exact test). Antibody-mediated HIV-1 neutralizing activity was significantly reduced in CSF with respect to autologous blood plasma (P < 0.042, Student's t-test). Accordingly, CSF-derived sequences exhibited constrained diversity and contained fewer glycosylated and positively selected sites. Our results suggest that there are several genetic features that distinguish CSF- and plasma-derived HIV-1 populations, probably reflecting altered cellular entry requirements and decreased immune pressure in the CNS. Furthermore, neurological impairment may be influenced by mutations within the viral V3 loop sequence.

Amino Acid Sequence↗

Improving risk indexes for Alzheimer's disease and related dementias for use in midlife.

Knowledge of a person's risk for Alzheimer's disease and related dementias (ADRDs) is required to triage candidates for preventive interventions, surveillance, and treatment trials. ADRD risk indexes exist for this purpose, but each includes only a subset of known risk factors. Information missing from published indexes could improve risk prediction. In the Dunedin Study of a population-representative New Zealand-based birth cohort followed to midlife (N&#x2009;=&#x2009;938, 49.5% female), we compared associations of four leading risk indexes with midlife antecedents of ADRD against a novel benchmark index comprised of nearly all known ADRD risk factors, the Dunedin ADRD Risk Benchmark (DunedinARB). Existing indexes included the Cardiovascular Risk Factors, Aging, and Dementia index (CAIDE), LIfestyle for BRAin health index (LIBRA), Australian National University Alzheimer's Disease Risk Index (ANU-ADRI), and risks selected by the Lancet Commission on Dementia. The Dunedin benchmark was comprised of 48 separate indicators of risk organized into 10 conceptually distinct risk domains. Midlife antecedents of ADRD treated as outcome measures included age-45 measures of brain structural integrity [magnetic resonance imaging-assessed: (i) machine-learning-algorithm-estimated brain age, (ii) log-transformed volume of white matter hyperintensities, and (iii) mean grey matter volume of the hippocampus] and measures of brain functional integrity [(i) objective cognitive function assessed via the Wechsler Adult Intelligence Scale-IV, (ii) subjective problems in everyday cognitive function, and (iii) objective cognitive decline measured as residualized change in cognitive scores from childhood to midlife on matched Weschler Intelligence scales]. All indexes were quantitatively distributed and proved informative about midlife antecedents of ADRD, including algorithm-estimated brain age (&#x3b2;'s from 0.16 to 0.22), white matter hyperintensities volume (&#x3b2;'s from 0.16 to 0.19), hippocampal volume (&#x3b2;'s from -0.08 to -0.11), tested cognitive deficits (&#x3b2;'s from -0.36 to -0.49), everyday cognitive problems (&#x3b2;'s from 0.14 to 0.38), and longitudinal cognitive decline (&#x3b2;'s from -0.18 to -0.26). Existing indexes compared favourably to the comprehensive benchmark in their association with the brain structural integrity measures but were outperformed in their association with the functional integrity measures, particularly subjective cognitive problems and tested cognitive decline. Results indicated that existing indexes could be improved with targeted additions, particularly of measures assessing socioeconomic status, physical and sensory function, epigenetic aging, and subjective overall health. Existing premorbid ADRD risk indexes perform well in identifying linear gradients of risk among members of the general population at midlife, even when they include only a small subset of potential risk factors. They could be improved, however, with targeted additions to more holistically capture the different facets of risk for this multiply determined, age-related disease.

Alzheimer&#x2019;s disease↗

Fuel spill identification using solid-phase extraction and solid-phase microextraction. 1. Aviation turbine fuels.

The water-soluble fraction of aviation jet fuels is examined using solid-phase extraction and solid-phase microextraction. Gas chromatographic profiles of solid-phase extracts and solid-phase microextracts of the water-soluble fraction of kerosene- and nonkerosene-based jet fuels reveal that each jet fuel possesses a unique profile. Pattern recognition analysis reveals fingerprint patterns within the data characteristic of fuel type. By using a novel genetic algorithm (GA) that emulates human pattern recognition through machine learning, it is possible to identify features characteristic of the chromatographic profile of each fuel class. The pattern recognition GA identifies a set of features that optimize the separation of the fuel classes in a plot of the two largest principal components of the data. Because principal components maximize variance, the bulk of the information encoded by the selected features is primarily about the differences between the fuel classes.

Journal Article↗

Sex-specific associations of the plasma-proteome with incident coronary artery disease.

AIMS: The etiology of coronary artery Disease (CAD) appears different for men and women, yet insights into underlying sex-specific biological mechanisms are limited. We integrated genomic and proteomic analyses to investigate sex-specific associations of the plasma-proteome with CAD. METHODS AND RESULTS: In 40,829 UK Biobank participants (free-of-CAD, baseline-365 days thereafter; 55% women; mean age 56.9&#x2009;&#xb1;&#x2009;8.1 years), we examined associations between 2,922 plasma proteins and incident CAD over a median follow-up of 13.7 years (IQR 13.1-14.4) using multivariable-adjusted Cox proportional hazards models. Sex-specific analyses identified 440 female exclusive and 32 male exclusive proteins associated with incident CAD (FDR-corrected p&#x2009;<&#x2009;0.05), revealing distinct pathway enrichments, including innate immune response in women and angiogenesis in men. Causality was assessed through combined and sex-stratified two-sample Mendelian randomization (MR) using inverse-variance-weighted analyses with genome wide association summary statistics from 422,108 men (61,969 cases) and 521,695 women (27,128 cases) (UK Biobank, FinnGen freeze 9). Integration of direct sex-protein interaction analyses with sex-combined MR identified 59 proteins with evidence for sex-specific causal effects. Four proteins demonstrated concordant directionality in sex-stratified MR analyses (n&#x2009;=&#x2009;943,803) and multivariable regression models, namely CDKN2D, MYH9, and SKAP2 (women), and CTSH (men). To assess translational relevance, prioritized targets were further evaluated in secondary major adverse cardiovascular events among carotid endarterectomy patients (MACE; Athero-Express) and acute myocardial infarction (AMI; MISSION!) using plasma proteomics and ELISA. After further top-target identification in the context of MACE and AMI, clinical drug candidates were identified through a machine learning framework, including CTSH (men), and TNFRSF4 (both sexes). CONCLUSIONS: We identified sex-specific associations of proteins and biological pathways with incident CAD. Whereas the majority of proteins had consistent associations in both men and women, our findings suggest a degree of sex-specific pathogenesis with evidence for potential causality, opening new alleys for tailored prevention strategies and clinical cardiovascular risk management.

Journal Article↗

Quantitative classification and natural clustering of Caenorhabditis elegans behavioral phenotypes.

Genetic analysis of nervous system function relies on the rigorous description of behavioral phenotypes. However, standard methods for classifying the behavioral patterns of mutant Caenorhabditis elegans rely on human observation and are therefore subjective and imprecise. Here we describe the application of machine learning to quantitatively define and classify the behavioral patterns of C. elegans nervous system mutants. We have used an automated tracking and image processing system to obtain measurements of a wide range of morphological and behavioral features from recordings of representative mutant types. Using principal component analysis, we represented the behavioral patterns of eight mutant types as data clouds distributed in multidimensional feature space. Cluster analysis using the k-means algorithm made it possible to quantitatively assess the relative similarities between different behavioral phenotypes and to identify natural phenotypic clusters among the data. Since the patterns of phenotypic similarity identified in this study closely paralleled the functional similarities of the mutant gene products, the complex phenotypic signatures obtained from these image data appeared to represent an effective diagnostic of the mutants' underlying molecular defects.

Animals↗

Lay person-based screening for early detection of Alzheimer's disease: development and validation of an instrument.

Symptoms of cognitive impairment reported to telephone interviewers by caregivers of 272 patients were analyzed with respect to research diagnoses of dementia. All patients received neuropsychological evaluation for establishing the research diagnoses. A data mining program that used machine learning algorithms produced an optimized binary decision tree for differentiating patient groups according to all available information. The results of this analysis were used to help four dementia experts create a dementia screening instrument amenable to application and scoring by nonclinical personnel. The validity of the resulting instrument was then evaluated in an independent sample of 103 patients administered neuropsychological testing within the previous 60 days. The psychometric properties of the empirically derived scale and its performance for discriminating control from probable or possible Alzheimer's patients indicate strong potential for use as a dementia screener for the general population.

Adult↗

Mapping ovarian cellular and molecular landscape across the lifespan of women: a scoping review.

BACKGROUND: With growing interest in ART, fertility preservation, and postmenopausal health of women, reproductive medicine is increasingly focused on characterizing oocytes and ovarian tissue composition, as well as understanding the molecular mechanisms that guide ovarian function throughout its lifecycle. High-throughput omics technologies have enabled the characterization of different molecular layers, leading to substantial advances in our understanding of their complex dynamics. However, not all molecular aspects are studied equally, and studies examining the same modalities often show inconsistencies, underscoring the need for data standardization and highlighting the potential for using transformative artificial intelligence and machine-learning (AI/ML) methods for ovary studies. OBJECTIVE AND RATIONALE: This study aims to evaluate how multi-omic studies have advanced our understanding of the ovarian lifecycle from fetal development to postmenopause. We systematically reviewed published studies that have investigated molecular/omic layers, including the genome, methylome, transcriptome, and proteome throughout ovarian development and aging. Our analysis identified key molecular and cellular patterns, highlighted inconsistencies across studies and addressed gaps in data analysis, interpretation, and reproducibility to guide future research. SEARCH METHODS: We conducted a systematic literature search of Medline (PubMed), Embase (Ovid), and Web of Science Core Collection (Clarivate) using a combination of controlled and free text terms for human ovary, oogenesis, folliculogenesis, ovary development and (epi)genome, transcriptome, proteome, and multi-omic mechanisms to find relevant articles published before August 2025. To focus the scope of the current review, studies of domesticated and farm animals, rodents and other model organisms, non-human primates, as well as those examining various human ovarian pathologies were excluded. OUTCOMES: The search identified 23 546 studies for screening, of which 637 full-text studies were assessed for eligibility. Subsequently, we extracted data from 121 studies. Most studies analyzed the transcriptome of oocytes, granulosa cells, and ovarian tissue from reproductive-age individuals (n&#x2009;=&#x2009;91), with fewer studies examining samples from individuals of advanced reproductive age (n&#x2009;=&#x2009;45) and fetal (n&#x2009;=&#x2009;16) samples. Transcriptome analyses were most common (n&#x2009;=&#x2009;103, 85%), followed by proteome (n&#x2009;=&#x2009;19, 16%) and epigenome (n&#x2009;=&#x2009;14, 12%) studies. We found substantial variation in how studies defined and reported participants' groups as well as in their sequencing technologies and data analysis methods, with a lack of standardized reporting of background clinical information, data analysis methods, and pipeline details. The key findings underscore the prevailing consensus on genes defining major ovarian cell types and their roles throughout the ovarian lifespan, from prenatal development to postmenopausal transformation. This review highlighted the underrepresentation of certain patient groups, particularly prepubertal and peri-/postmenopausal individuals, among researched populations, due to obvious clinical and ethical reasons. WIDER IMPLICATIONS: This scoping review offers a comprehensive overview and benchmark of the current state of high-throughput omics-based research on ovarian cellular composition and molecular dynamics. To address these shortcomings, we propose general recommendations for multi-omics ovary studies and emphasize the necessity for more thorough multi-omic data integration by effectively applying novel AI/ML approaches. They can potentially improve the quality of multi-omics analyses at both single-cell and tissue levels despite limited sample sizes and enable integration of molecular profiling data with clinical and radiology datasets, enabling a more comprehensive understanding of ovarian biology. Such advancements can enhance reproducibility of research findings and guide future research to deepen our understanding of ovarian biology and ultimately support the development of medical technologies for better preserving fertility and alleviating infertility. REGISTRATION NUMBER: A protocol was published a priori on the Open Science Framework (https://osf.io/z38gb/).

Female↗