Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

Genome-wide association analysis reveals specialization to hosts and niches in multiple species of the Lactobacillaceae.

The Lactobacillaceae inhabit diverse environments, but the extent of their habitat adaptation remains unclear and the colonization factors unknown. First, we applied multiple machine learning models to determine if we can distinguish strains of the same species isolated from two different habitats based on their gene content. Surprisingly, we show that no species is differentially adapted to the oral cavity versus the human gut, or food versus the human gut, while only Lactobacillus crispatus showed specialization to the human urogenital system versus human gut. We then asked which species of Lactobacillaceae are habitat-specialized and how they could be identified. Using multiple lifestyle predictors incorporated in logistic regression models, we found that Limosilactobacillus reuteri, Ligilactobacillus ruminis, L. salivarius, L. crispatus, and L. mucosae displayed the highest degrees of host specialization. Applying our microbial genome-wide association study tool, aurora, to these species identified genes encoding adhesins and bacteriocins as the strongest and most common adaptation factors. This work establishes a generalizable framework for identifying novel species-habitat pairs with strong evidence of specialization and for uncovering the genomic features underlying within-species host and habitat adaptation.

Humans↗

Application of expert systems analysis to interpretation of fatal cases involving amitriptyline.

Part I. Expert 4: The object of this study was to investigate the applicability of commercially available expert system shells to interpretation in forensic toxicology. Amitriptyline toxicology was selected as a pilot trial. Blood and tissue concentrations of amitriptyline and nortriptyline in fatal and nonfatal amitriptyline cases from the literature and from the Registry of Human Toxicology databank were entered into the expert system shell Expert 4 (Rivers, Elsevier). The statistical evaluation routines of the shell were used to search for patterns in the data. Successive changes in the data base were made to test for the influence of the data base on the conclusions. Finally the data base was refined, based on the evaluations, to strengthen the probabilities of the conclusions. The refined database was coupled with the Expert 4 inference engine to infer unknown parameters in the cases. The results of the expert system analysis were compared to known values and published expert opinions. The ratio of amitriptyline/nortriptyline and tissue levels of nortriptyline were found to be the most significant measures for interpretation of effect and time since ingestion. Part II. Computer Induction of Rules: Blood and tissue concentrations of amitriptyline and nortriptyline in fatal and nonfatal amitriptyline cases from the literature and from the Registry of Human Toxicology databank were entered into the expert system shell BEAGLE, (Forsyth, Machine Learning Research, Ltd.). The automatic rule induction routines of the shell were used to search for patterns in the data. The program expressed these patterns as numerical predictions or Boolean logic rules. The results were compared to those obtained with the Expert 4 (Rivers, Elsevier) using the same case knowledge base.(ABSTRACT TRUNCATED AT 250 WORDS)

Amitriptyline↗

Circulating tumor human papillomavirus DNA whole genome sequencing enables human papillomavirus-associated oropharynx cancer early detection.

BACKGROUND: Early detection of HPV-associated oropharyngeal squamous cell carcinoma, the most common HPV cancer in the United States, could reduce disease-related morbidity and mortality, yet currently, there are no early detection tests. HPV circulating tumor DNA (ctDNA) is a sensitive and specific biomarker for HPV-associated oropharyngeal squamous cell carcinoma at diagnosis. It is unknown if ctDNA HPV is detectable prior to diagnosis, and thus its potential as an early detection test is also unknown. METHODS: Plasma samples from the Mass General Brigham Biobank collected 1.3-10.8 years prior to diagnosis from HPV-associated oropharyngeal squamous cell carcinoma patients (n = 28) and age- and sex-matched controls (n = 28) were blinded and run on a newly developed and validated multifeature HPV whole genome sequencing liquid biopsy assay and a validated HPV antibody assay. RESULTS: HPV ctDNA results were positive in 22 of 28 prediagnostic samples from HPV-associated oropharyngeal squamous cell carcinoma cases (sensitivity 79%) with a maximum lead time of 7.8 years. HPV ctDNA results were negative in all controls (0 of 28, 100% specificity). Diagnostic accuracy was highest within 4 years of cancer diagnosis and was higher than HPV Ab detection within the same timeframe (P = .004). Application of a machine-learning model trained and tested on an independent cohort of 306 cases and controls increased the sensitivity of detection to 27 of 28 cases (overall sensitivity 96%) and the maximum lead time to 10.3 years. CONCLUSIONS: HPV ctDNA can be detected in the blood years prior to diagnosis with HPV-associated oropharyngeal squamous cell carcinoma, with high specificity, in a case-control cohort of 56 participants. HPV ctDNA detection alone, or in combination with previously identified serological biomarkers, may be a feasible approach to early detection of HPV-associated oropharyngeal squamous cell carcinoma.

Humans↗

SPINE: an integrated tracking database and data mining approach for identifying feasible targets in high-throughput structural proteomics.

High-throughput structural proteomics is expected to generate considerable amounts of data on the progress of structure determination for many proteins. For each protein this includes information about cloning, expression, purification, biophysical characterization and structure determination via NMR spectroscopy or X-ray crystallography. It will be essential to develop specifications and ontologies for standardizing this information to make it amenable to retrospective analysis. To this end we created the SPINE database and analysis system for the Northeast Structural Genomics Consortium. SPINE, which is available at bioinfo.mbb.yale.edu/nesg or nesg.org, is specifically designed to enable distributed scientific collaboration via the Internet. It was designed not just as an information repository but as an active vehicle to standardize proteomics data in a form that would enable systematic data mining. The system features an intuitive user interface for interactive retrieval and modification of expression construct data, query forms designed to track global project progress and external links to many other resources. Currently the database contains experimental data on 985 constructs, of which 740 are drawn from Methanobacterium thermoautotrophicum, 123 from Saccharomyces cerevisiae, 93 from Caenorhabditis elegans and the remainder from other organisms. We developed a comprehensive set of data mining features for each protein, including several related to experimental progress (e.g. expression level, solubility and crystallization) and 42 based on the underlying protein sequence (e.g. amino acid composition, secondary structure and occurrence of low complexity regions). We demonstrate in detail the application of a particular machine learning approach, decision trees, to the tasks of predicting a protein's solubility and propensity to crystallize based on sequence features. We are able to extract a number of key rules from our trees, in particular that soluble proteins tend to have significantly more acidic residues and fewer hydrophobic stretches than insoluble ones. One of the characteristics of proteomics data sets, currently and in the foreseeable future, is their intermediate size ( approximately 500-5000 data points). This creates a number of issues in relation to error estimation. Initially we estimate the overall error in our trees based on standard cross-validation. However, this leaves out a significant fraction of the data in model construction and does not give error estimates on individual rules. Therefore, we present alternative methods to estimate the error in particular rules.

Animals↗

Discrimination among individual Watson-Crick base pairs at the termini of single DNA hairpin molecules.

Nanoscale alpha-hemolysin pores can be used to analyze individual DNA or RNA molecules. Serial examination of hundreds to thousands of molecules per minute is possible using ionic current impedance as the measured property. In a recent report, we showed that a nanopore device coupled with machine learning algorithms could automatically discriminate among the four combinations of Watson-Crick base pairs and their orientations at the ends of individual DNA hairpin molecules. Here we use kinetic analysis to demonstrate that ionic current signatures caused by these hairpin molecules depend on the number of hydrogen bonds within the terminal base pair, stacking between the terminal base pair and its nearest neighbor, and 5' versus 3' orientation of the terminal bases independent of their nearest neighbors. This report constitutes evidence that single Watson-Crick base pairs can be identified within individual unmodified DNA hairpin molecules based on their dynamic behavior in a nanoscale pore.

Algorithms↗

Proteome Analyst: custom predictions with explanations in a web-based tool for high-throughput proteome annotations.

Proteome Analyst (PA) (http://www.cs.ualberta.ca/~bioinfo/PA/) is a publicly available, high-throughput, web-based system for predicting various properties of each protein in an entire proteome. Using machine-learned classifiers, PA can predict, for example, the GeneQuiz general function and Gene Ontology (GO) molecular function of a protein. In addition, PA is currently the most accurate and most comprehensive system for predicting subcellular localization, the location within a cell where a protein performs its main function. Two other capabilities of PA are notable. First, PA can create a custom classifier to predict a new property, without requiring any programming, based on labeled training data (i.e. a set of examples, each with the correct classification label) provided by a user. PA has been used to create custom classifiers for potassium-ion channel proteins and other general function ontologies. Second, PA provides a sophisticated explanation feature that shows why one prediction is chosen over another. The PA system produces a Naïve Bayes classifier, which is amenable to a graphical and interactive approach to explanations for its predictions; transparent predictions increase the user's confidence in, and understanding of, PA.

Internet↗

nsSNPAnalyzer: identifying disease-associated nonsynonymous single nucleotide polymorphisms.

Nonsynonymous single nucleotide polymorphisms (nsSNPs) are prevalent in genomes and are closely associated with inherited diseases. To facilitate identifying disease-associated nsSNPs from a large number of neutral nsSNPs, it is important to develop computational tools to predict the nsSNP's phenotypic effect (disease-associated versus neutral). nsSNPAnalyzer, a web-based software developed for this purpose, extracts structural and evolutionary information from a query nsSNP and uses a machine learning method called Random Forest to predict the nsSNP's phenotypic effect. nsSNPAnalyzer server is available at http://snpanalyzer.utmem.edu/.

Algorithms↗

miRNAMap: genomic maps of microRNA genes and their target genes in mammalian genomes.

Recent work has demonstrated that microRNAs (miRNAs) are involved in critical biological processes by suppressing the translation of coding genes. This work develops an integrated database, miRNAMap, to store the known miRNA genes, the putative miRNA genes, the known miRNA targets and the putative miRNA targets. The known miRNA genes in four mammalian genomes such as human, mouse, rat and dog are obtained from miRBase, and experimentally validated miRNA targets are identified in a survey of the literature. Putative miRNA precursors were identified by RNAz, which is a non-coding RNA prediction tool based on comparative sequence analysis. The mature miRNA of the putative miRNA genes is accurately determined using a machine learning approach, mmiRNA. Then, miRanda was applied to predict the miRNA targets within the conserved regions in 3'-UTR of the genes in the four mammalian genomes. The miRNAMap also provides the expression profiles of the known miRNAs, cross-species comparisons, gene annotations and cross-links to other biological databases. Both textual and graphical web interface are provided to facilitate the retrieval of data from the miRNAMap. The database is freely available at http://mirnamap.mbc.nctu.edu.tw/.

Animals↗

transFold: a web server for predicting the structure and residue contacts of transmembrane beta-barrels.

Transmembrane beta-barrel (TMB) proteins are embedded in the outer membrane of Gram-negative bacteria, mitochondria and chloroplasts. The cellular location and functional diversity of beta-barrel outer membrane proteins makes them an important protein class. At the present time, very few non-homologous TMB structures have been determined by X-ray diffraction because of the experimental difficulty encountered in crystallizing transmembrane (TM) proteins. The transFold web server uses pairwise inter-strand residue statistical potentials derived from globular (non-outer-membrane) proteins to predict the supersecondary structure of TMB. Unlike all previous approaches, transFold does not use machine learning methods such as hidden Markov models or neural networks; instead, transFold employs multi-tape S-attribute grammars to describe all potential conformations, and then applies dynamic programming to determine the global minimum energy supersecondary structure. The transFold web server not only predicts secondary structure and TMB topology, but is the only method which additionally predicts the side-chain orientation of transmembrane beta-strand residues, inter-strand residue contacts and TM beta-strand inclination with respect to the membrane. The program transFold currently outperforms all other methods for accuracy of beta-barrel structure prediction. Available at http://bioinformatics.bc.edu/clotelab/transFold.

Amino Acids↗

PLPD: reliable protein localization prediction from imbalanced and overlapped datasets.

Subcellular localization is one of the key functional characteristics of proteins. An automatic and efficient prediction method for the protein subcellular localization is highly required owing to the need for large-scale genome analysis. From a machine learning point of view, a dataset of protein localization has several characteristics: the dataset has too many classes (there are more than 10 localizations in a cell), it is a multi-label dataset (a protein may occur in several different subcellular locations), and it is too imbalanced (the number of proteins in each localization is remarkably different). Even though many previous works have been done for the prediction of protein subcellular localization, none of them tackles effectively these characteristics at the same time. Thus, a new computational method for protein localization is eventually needed for more reliable outcomes. To address the issue, we present a protein localization predictor based on D-SVDD (PLPD) for the prediction of protein localization, which can find the likelihood of a specific localization of a protein more easily and more correctly. Moreover, we introduce three measurements for the more precise evaluation of a protein localization predictor. As the results of various datasets which are made from the experiments of Huh et al. (2003), the proposed PLPD method represents a different approach that might play a complimentary role to the existing methods, such as Nearest Neighbor method and discriminate covariant method. Finally, after finding a good boundary for each localization using the 5184 classified proteins as training data, we predicted 138 proteins whose subcellular localizations could not be clearly observed by the experiments of Huh et al. (2003).

Algorithms↗

eSLDB: eukaryotic subcellular localization database.

Eukaryotic Subcellular Localization DataBase collects the annotations of subcellular localization of eukaryotic proteomes. So far five proteomes have been processed and stored: Homo sapiens, Mus musculus, Caenorhabditis elegans, Saccharomyces cerevisiae and Arabidopsis thaliana. For each sequence, the database lists localization obtained adopting three different approaches: (i) experimentally determined (when available); (ii) homology-based (when possible); and (iii) predicted. The latter is computed with a suite of machine learning based methods, developed in house. All the data are available at our website and can be searched by sequence, by protein code and/or by protein description. Furthermore, a more complex search can be performed combining different search fields and keys. All the data contained in the database can be freely downloaded in flat file format. The database is available at http://gpcr.biocomp.unibo.it/esldb/.

Animals↗

Identifying fundamental gaps in functional metagenomics: a step towards unlocking microbiome research potential.

Incomplete functional annotation limits biological interpretation in microbiome studies and their translational potential. Poor annotation arises from multiple causes, with incomplete gene-protein-reaction mapping being one tractable yet under-examined contributor. We address this gap by developing a comprehensive hierarchical framework that systematically integrates gene families in UniRef, proteins in UniProt, and metabolic reactions in MetaCyc and BioCyc through UniProtKB accession, EC number, and Pfam-domain matching. Applied to a human gut metagenome dataset via HUMAnN3, our MetaCyc-based mapping recovers up to 2.3-fold more unique reaction identifiers than the default pipeline and increases reaction prevalence across samples from ≈32% to 52% core reactions, addressing the data sparsity that limits statistical and machine-learning applications in microbiome research. Biological plausibility for the tested functions was supported by positive and negative controls: gut-microbial hormone-metabolism reactions previously linked to this dataset were recovered, while vertebrate-specific hormone-metabolism reactions remained correctly undetected. These gains derive from systematic database integration alone, without predictive algorithms, indicating that a tractable, mapping-related component of functional dark matter and data sparsity in microbiome studies is directly addressable. Because Pfam- and BioCyc-derived mappings trade specificity for coverage, confidence in any individual reaction assignment depends on the supporting evidence tier and source database.

Humans↗

IPSA-Inductive Protein Structure Analysis.

The Inductive Structure Protein Analysis (IPSA) project presents a new method for investigating protein structure. IPSA includes the creation of a new database which was designed specifically for the analysis of protein structure by statistics and machine learning. The Protein Representation Language (PRL) database includes explicit and symbolic representations of geometrical, topological and chemophysical information about secondary structures and the relationships between secondary structures. The IPSA methodology consists of: the use of PRL information to produce a new database of examples of secondary structures which associate together (examples of possible super-secondary structures); then the use of a variety of clustering techniques to produce a consensus clustering of these examples (super-secondary structures); these super-secondary structures are finally examined to uncover any biological features of significance. We have applied this method to find simple super-secondary structures consisting of pairs of alpha-helices. We found four well-defined super-secondary structures, one formed exclusively by long range interactions, and another in association with an additional element of secondary structure (alpha t alpha-motif). Examinations were carried out using homologous pairs and conformational fits which confirm our clustering.

Cluster Analysis↗

Prediction of protein secondary structure with a reliability score estimated by local sequence clustering.

Most algorithms for protein secondary structure prediction are based on machine learning techniques, e.g. neural networks. Good architectures and learning methods have improved the performance continuously. The introduction of profile methods, e.g. PSI-BLAST, has been a major breakthrough in increasing the prediction accuracy to close to 80%. In this paper, a brute-force algorithm is proposed and the reliability of each prediction is estimated by a z-score based on local sequence clustering. This algorithm is intended to perform well for those secondary structures in a protein whose formation is mainly dominated by the neighboring sequences and short-range interactions. A reliability z-score has been defined to estimate the goodness of a putative cluster found for a query sequence in a database. The database for prediction was constructed by experimentally determined, non-redundant protein structures with <25% sequence homology, a list maintained by PDBSELECT. Our test results have shown that this new algorithm, belonging to what is known as nearest neighbor methods, performed very well within the expectation of previous methods and that the reliability z-score as defined was correlated with the reliability of prediction. This led to the possibility of making very accurate predictions for a few selected residues in a protein with an accuracy measure of Q3 > 80%. The further development of this algorithm, and a nucleation mechanism for protein folding are suggested.

Algorithms↗

Analysis and prediction of leucine-rich nuclear export signals.

We present a thorough analysis of nuclear export signals and a prediction server, which we have made publicly available. The machine learning prediction method is a significant improvement over the generally used consensus patterns. Nuclear export signals (NESs) are extremely important regulators of the subcellular location of proteins. This regulation has an impact on transcription and other nuclear processes, which are fundamental to the viability of the cell. NESs are studied in relation to cancer, the cell cycle, cell differentiation and other important aspects of molecular biology. Our conclusion from this analysis is that the most important properties of NESs are accessibility and flexibility allowing relevant proteins to interact with the signal. Furthermore, we show that not only the known hydrophobic residues are important in defining a nuclear export signals. We employ both neural networks and hidden Markov models in the prediction algorithm and verify the method on the most recently discovered NESs. The NES predictor (NetNES) is made available for general use at http://www.cbs.dtu.dk/.

Active Transport, Cell Nucleus↗

A novel statistical ligand-binding site predictor: application to ATP-binding sites.

Structural genomics initiatives are leading to rapid growth in newly determined protein 3D structures, the functional characterization of which may still be inadequate. As an attempt to provide insights into the possible roles of the emerging proteins whose structures are available and/or to complement biochemical research, a variety of computational methods have been developed for the screening and prediction of ligand-binding sites in raw structural data, including statistical pattern classification techniques. In this paper, we report a novel statistical descriptor (the Oriented Shell Model) for protein ligand-binding sites, which utilizes the distance and angular position distribution of various structural and physicochemical features present in immediate proximity to the center of a binding site. Using the support vector machine (SVM) as the classifier, our model identified 69% of the ATP-binding sites in whole-protein scanning tests and in eukaryotic proteins the accuracy is particularly high. We propose that this feature extraction and machine learning procedure can screen out ligand-binding-capable protein candidates and can yield valuable biochemical information for individual proteins.

Adenosine Triphosphate↗

Efforts towards a precision medicine approach in juvenile idiopathic arthritis.

Juvenile idiopathic arthritis (JIA) is the commonest group of childhood arthritides. Despite the availability of advanced therapeutics, many children and young people (CYP) with JIA experience disease flares, and in some, chronic joint damage. Tailoring treatment based on unique biological profiles would benefit CYP with JIA given their variable clinical presentation and disease course. To date, biomarkers to predict treatment response are lacking. With advances in single cell technologies, we are now able to profile the genes and proteins of target tissues at unprecedented resolution to define the biological basis of disease and guide novel treatment approaches. The complex analyses and combination of biological and clinical outcome data from large datasets across disease phenotypes have become possible with the development of computational and machine learning methods. Here, we summarize the strategies to integrate data through multimodal based approaches to maximize precision medicine and research priorities for CYP with JIA.

Humans↗

Extracting and characterizing gene-drug relationships from the literature.

A fundamental task of pharmacogenetics is to collect and classify relationships between genes and drugs. Currently, this useful information has not been comprehensively aggregated in any database and remains scattered throughout the published literature. Although there are efforts to collect this information manually, they are limited by the size of the published literature on gene-drug relationships. Therefore, we investigated computational methods to extract and characterize pharmacogenetic relationships between genes and drugs from the literature. We first evaluated the effectiveness of the co-occurrence method in identifying related genes and drugs. We then used supervised machine learning algorithms to classify the relationships between genes and drugs from the Pharmacogenetics and Pharmacogenomics Knowledge Base (PharmGKB) into five categories that have been defined by active pharmacogenetic researchers as relevant to their work. The final co-occurrence algorithm was able to extract 78% of the related genes and drugs that were published in a review article from the literature. Our algorithm subsequently classified the relationships between genes and drugs from the PharmGKB into five categories with 74% accuracy. We have made the data available on a supplementary website at http://bionlp.stanford.edu/genedrug/ Gene-drug relationships can be accurately extracted from text and classified into categories. Although the relationships that we have identified do not capture the details and fine distinctions often made in the literature, these methods will help scientists to track the ever-growing literature and create information resources to support future discoveries.

Algorithms↗