Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

mRNA degradation. A tale of poly(A) and multiprotein machines.

The Escherichia coli RNA degradosome is a multiprotein complex containing an endoribonuclease, polynucleotide phosphorylase and a DEAD-box RNA helicase. A related complex has been described in the spinach chloroplast. The exosome and the mtEXO complex have recently been described in yeast and it is likely that related complexes also exist in animal cells. This research suggests the widespread existence of sophisticated machines for the efficient degradation of messenger RNA. The DEAD-box helicase in the degradosome can unwind regions of RNA structure that interfere with 3'-5' degradation. The polyadenylation of RNA 3' ends is also known to promote degradation by creating a 'toehold' for the degradation machinery. Much remains to be learned about the regulation of mRNA stability. The complexity of the degradation process, both in the eubacteria and in the eukaryotes, suggests that many steps are possible points of control.

Bacterial Proteins↗

A novel structure-based encoding for machine-learning applied to the inference of SH3 domain specificity.

MOTIVATION: Unravelling the rules underlying protein-protein and protein-ligand interactions is a crucial step in understanding cell machinery. Peptide recognition modules (PRMs) are globular protein domains which focus their binding targets on short protein sequences and play a key role in the frame of protein-protein interactions. High-throughput techniques permit the whole proteome scanning of each domain, but they are characterized by a high incidence of false positives. In this context, there is a pressing need for the development of in silico experiments to validate experimental results and of computational tools for the inference of domain-peptide interactions. RESULTS: We focused on the SH3 domain family and developed a machine-learning approach for inferring interaction specificity. SH3 domains are well-studied PRMs which typically bind proline-rich short sequences characterized by the PxxP consensus. The binding information is known to be held in the conformation of the domain surface and in the short sequence of the peptide. Our method relies on interaction data from high-throughput techniques and benefits from the integration of sequence and structure data of the interacting partners. Here, we propose a novel encoding technique aimed at representing binding information on the basis of the domain-peptide contact residues in complexes of known structure. Remarkably, the new encoding requires few variables to represent an interaction, thus avoiding the 'curse of dimension'. Our results display an accuracy >90% in detecting new binders of known SH3 domains, thus outperforming neural models on standard binary encodings, profile methods and recent statistical predictors. The method, moreover, shows a generalization capability, inferring specificity of unknown SH3 domains displaying some degree of similarity with the known data.

Algorithms↗

A machine learning evaluation of an artificial immune system.

ARTIS is an artificial immune system framework which contains several adaptive mechanisms. LISYS is a version of ARTIS specialized for the problem of network intrusion detection. The adaptive mechanisms of LISYS are characterized in terms of their machine-learning counterparts, and a series of experiments is described, each of which isolates a different mechanism of LISYS and studies its contribution to the system's overall performance. The experiments were conducted on a new data set, which is more recent and realistic than earlier data sets. The network intrusion detection problem is challenging because it requires one-class learning in an on-line setting with concept drift. The experiments confirm earlier experimental results with LISYS, and they study in detail how LISYS achieves success on the new data set.

Algorithms↗

Applying hybrid reasoning to mine for associative features in biological data.

We develop the means to mine for associative features in biological data. The hybrid reasoning schema for deterministic machine learning and its implementation via logic programming is presented. The methodology of mining for correlation between features is illustrated by the prediction tasks for protein secondary structure and phylogenetic profiles. The suggested methodology leads to a clearer approach to hierarchical classification of proteins and a novel way to represent evolutionary relationships. Comparative analysis of Jasmine and other statistical and deterministic systems (including Explanation-Based Learning and Inductive Logic Programming) are outlined. Advantages of using deterministic versus statistical data mining approaches for high-level exploration of correlation structure are analyzed.

Algorithms↗

Predicting protein secondary structure by a support vector machine based on a new coding scheme.

Protein structure prediction is one of the most important problems in modern computational biology. Protein secondary structure prediction is a key step in prediction of protein tertiary structure. There have emerged many methods based on machine learning techniques, such as neural networks (NN) and support vector machine (SVM) etc., to focus on the prediction of the secondary structures. In this paper, a new method was proposed based on SVM. Different from the existing methods, this method takes into account of the physical-chemical properties and structure properties of amino acids. When tested on the most popular dataset CB513, it achieved a Q(3) accuracy of 0.7844, which illustrates that it is one of the top range methods for protein of secondary structure prediction.

Algorithms↗

Predicting risk of ischemic stroke: A transformer model using genomic data.

BACKGROUND AND OBJECTIVE: Ischemic stroke is a leading cause of mortality and long-term disability worldwide. Genetic factors contribute to IS susceptibility, yet conventional polygenic risk score approaches are primarily based on additive effects and may not fully capture non-linear relationships or positional context and interactions among genetic variants. This study aimed to develop and evaluate a transformer-based genomic model incorporating position-wise genotype embedding for IS risk prediction. METHODS: We conducted a genome-wide association study using the UK Biobank dataset to identify IS-associated loci. Gene prioritisation was subsequently performed using tissue-specific expression quantitative trait locus-based Mendelian randomisation and colocalization analyses in whole blood and brain cortex. We then developed a transformer-based model that encoded genotype and SNP-position information using a position-wise embedding layer. Model performance was evaluated across three UK Biobank control definitions and externally assessed in the independent All of Us cohort. Performance metrics included the area under the receiver operating characteristic curve (AUROC), precision, recall, and F1 score. RESULTS: Across the three UK Biobank control definitions, the proposed method achieved the numerically highest discrimination among the evaluated models, with AUROCs of 0.8109, 0.7843, and 0.7468 using MRF-negative, combined, and MRF-positive controls, respectively. In the external All of Us cohort, the proposed method achieved an AUROC of 0.7251 and retained the highest AUROC among the evaluated models. In a separate incident-stroke survival analysis, medium- and high-score groups had hazard ratios of 1.13 and 1.21, respectively, relative to the low-score group. A total of 18 IS-associated loci were identified. Among the tissue-specific MR results, EDEM2 in the brain cortex remained significant after Bonferroni correction, while DCHS2 showed a nominal association. CONCLUSIONS: The proposed transformer-based framework provides a genomic modelling approach that achieved the highest discrimination among the evaluated models in this study and retained comparative performance in an independent external cohort. In further applications, integrating this genomic framework with conventional clinical, lifestyle, and environmental risk factors may support more comprehensive and personalised IS risk assessment. Prospective, population-representative, and multi-ancestry validation will be important to establish its potential role in future prevention-oriented risk management.

Genomics and bioinformatics↗

Selecting informative genes with parallel genetic algorithms in tissue classification.

Recent advances in biotechnology offer the ability to measure the levels of expression of thousands of genes in parallel. Analysis of such data can provide understanding and insight into gene function and regulatory mechanisms. Several machine learning approaches have been used to aid to understand the functions of genes. However, these tasks are made more difficult due to the noisy nature of array data and the overwhelming number of gene features. In this paper, we use the parallel genetic algorithm to filter out the informative genes relative to classification. By combing with the classification method proposed by Golub et al. and Slonim et al., we classify the data sets with tissues of different classes, and the preliminary results are presented in this paper.

Algorithms↗

Prediction of protein structural classes by support vector machines.

In this paper, we apply a new machine learning method which is called support vector machine to approach the prediction of protein structural class. The support vector machine method is performed based on the database derived from SCOP which is based upon domains of known structure and the evolutionary relationships and the principles that govern their 3D structure. As a result, high rates of both self-consistency and jackknife test are obtained. This indicates that the structural class of a protein inconsiderably correlated with its amino and composition, and the support vector machine can be referred as a powerful computational tool for predicting the structural classes of proteins.

Artificial Intelligence↗

Vessel enhancement in digital X-ray angiographic sequences by temporal statistical learning.

In this paper, we present a vessel enhancement method, SVM temporal filtering (STF), for X-ray angiographic (XA) images using Support Vector Machine (SVM). We show that the linear SVM applied to vessel enhancement can be regarded as a matched linear filter optimizing the contrast-to-noise ratio in XA images. We propose a non-linear kernel function for the SVM leading to good enhancement with noisy, varying grey-level dynamics at vessel pixels. One key advantage over the matched filters is that an optimal filter is learnt from images, not estimated at design stage. Results on clinical XA images show that learning-based enhancement achieves better results compared to simple subtraction and other image stacking methods.

Algorithms↗

Feature subset selection for splice site prediction.

MOTIVATION: The large amount of available annotated Arabidopsis thaliana sequences allows the induction of splice site prediction models with supervised learning algorithms (see Haussler (1998) for a review and references). These algorithms need information sources or features from which the models can be computed. For splice site prediction, the features we consider in this study are the presence or absence of certain nucleotides in close proximity to the splice site. Since it is not known how many and which nucleotides are relevant for splice site prediction, the set of features is chosen large enough such that the probability that all relevant information sources are in the set is very high. Using only those features that are relevant for constructing a splice site prediction system might improve the system and might also provide us with useful biological knowledge. Using fewer features will of course also improve the prediction speed of the system. RESULTS: A wrapper-based feature subset selection algorithm using a support vector machine or a naive Bayes prediction method was evaluated against the traditional method for selecting features relevant for splice site prediction. Our results show that this wrapper approach selects features that improve the performance against the use of all features and against the use of the features selected by the traditional method. AVAILABILITY: The data and additional interactive graphs on the selected feature subsets are available at http://www.psb.rug.ac.be/gps

Arabidopsis↗

Evaluation of a belief-network-based reminder system that learns from utility feedback.

PRETRIEVE is a belief-network-based, unsolicited information-retrieval system that performs machine learning based on user feedback. We report here on the document-ordering and document-retrieval performance of PRETRIEVE. We developed a test collection of 410 judgments of document utility in a simulated medical order-entry context. We characterized the validity of these judgments, which were elicited from domain experts, by measuring interrater and intrarater reproducibility. We developed a measure of the quality of document orderings similar to the ROC-curve analysis used to evaluate document-retrieval systems. We found that the ordering performance of the PRETRIEVE system was (1) substantially better than random, (2) somewhat less than ideal, and (3) superior to that of versions of the PRETRIEVE system that used relevance feedback instead of utility feedback. Under a set of assumptions, which we make explicit, we found that the documents retrieved by a version of PRETRIEVE that modeled time cost were of higher utility than those retrieved by a similar rule-based system.

Artificial Intelligence↗

Learning kernels from biological networks by maximizing entropy.

MOTIVATION: The diffusion kernel is a general method for computing pairwise distances among all nodes in a graph, based on the sum of weighted paths between each pair of nodes. This technique has been used successfully, in conjunction with kernel-based learning methods, to draw inferences from several types of biological networks. RESULTS: We show that computing the diffusion kernel is equivalent to maximizing the von Neumann entropy, subject to a global constraint on the sum of the Euclidean distances between nodes. This global constraint allows for high variance in the pairwise distances. Accordingly, we propose an alternative, locally constrained diffusion kernel, and we demonstrate that the resulting kernel allows for more accurate support vector machine prediction of protein functional classifications from metabolic and protein-protein interaction networks. AVAILABILITY: Supplementary results and data are available at noble.gs.washington.edu/proj/maxent

Algorithms↗

Towards an internal model in pilot training.

Optimal decision making requires an information seeking behavior which reflects the comprehension of the overall system dynamics. Research in the area of human monitors in man-machine systems supports the notion of an internal model with built-in expectancies. It is doubtful that the current approach to pilot training helps develop this internal model in the most efficient way. But this is crucial since the role of the pilot is changing to a systems' manager and decision maker. An extension of the behavioral framework of pilot training might help to prepare the pilot better for the increasingly complex flight environment. This extension is based on the theoretical model of schema theory, which evolved out of psychological research. The technological advances in aircraft simulators and in-flight performance measurement devices allow investigation of the still-unresolved issues.

Aerospace Medicine↗

Using pseudo-amino acid composition and support vector machine to predict protein structural class.

As a result of genome and other sequencing projects, the gap between the number of known protein sequences and the number of known protein structural classes is widening rapidly. In order to narrow this gap, it is vitally important to develop a computational prediction method for fast and accurately determining the protein structural class. In this paper, a novel predictor is developed for predicting protein structural class. It is featured by employing a support vector machine learning system and using a different pseudo-amino acid composition (PseAA), which was introduced to, to some extent, take into account the sequence-order effects to represent protein samples. As a demonstration, the jackknife cross-validation test was performed on a working dataset that contains 204 non-homologous proteins. The predicted results are very encouraging, indicating that the current predictor featured with the PseAA may play an important complementary role to the elegant covariant discriminant predictor and other existing algorithms.

Amino Acid Sequence↗

Ensemble DNA methylation clock demonstrates Immune-metabolic aging signatures associated with mortality.

Aging is a multifactorial process that is best described in terms of the progressive acquisition of multiple layers of phenotypic changes, such as epigenetic modifications, inflammation, and metabolic dysregulation. DNA methylation clocks have been extensively used to construct epigenetic clocks based on the DNAm profiles that can be used to estimate biological age and predict age-associated outcomes. Nevertheless, the vast majority of clocks constructed so far have been based on linear models, which are unlikely to fully account for the heterogeneity and non-linearity of survival-related DNAm signatures. In this work, we constructed a heterogeneous stacked ensemble survival model based on DNAm data obtained from the Framingham Heart Study. We first identified 190 CpG loci using elastic net Cox regression and subsequently constructed a survival prediction model based on the fusion of five complementary survival models by means of a neural network meta-learner. The prediction power of the survival model was evaluated in an external validation cohort, where we observed strong performance for predicting all-cause mortality that significantly exceeded PhenoAge and was statistically comparable to GrimAge. These performance estimates were derived in cohorts of European ancestry and externally validated in postmenopausal women aged 50-79 years, and should therefore be interpreted as applicable only to demographically similar populations.

Humans↗

Predicting protein structural class with AdaBoost Learner.

The structural class is an important feature in characterizing the overall topological folding type of a protein or the domains therein. Prediction of protein structural classification has attracted the attention and efforts from many investigators. In this paper a novel predictor, the AdaBoost Learner, was introduced to deal with this problem. The essence of the AdaBoost Learner is that a combination of many 'weak' learning algorithms, each performing just slightly better than a random guessing algorithm, will generate a 'strong' learning algorithm. Demonstration thru jackknife cross-validation on two working datasets constructed by previous investigators indicated that AdaBoost outperformed other predictors such as SVM (support vector machine), a powerful algorithm widely used in biological literatures. It has not escaped our notice that AdaBoost may hold a high potential for improving the quality in predicting the other protein features as well, such as subcellular location and receptor type, among many others. Or at the very least, it will play a complementary role to many of the existing algorithms in this regard.

Algorithms↗

Discovering functional gene expression patterns in the metabolic network of Escherichia coli with wavelets transforms.

BACKGROUND: Microarray technology produces gene expression data on a genomic scale for an endless variety of organisms and conditions. However, this vast amount of information needs to be extracted in a reasonable way and funneled into manageable and functionally meaningful patterns. Genes may be reasonably combined using knowledge about their interaction behaviour. On a proteomic level, biochemical research has elucidated an increasingly complete image of the metabolic architecture, especially for less complex organisms like the well studied bacterium Escherichia coli. RESULTS: We sought to discover central components of the metabolic network, regulated by the expression of associated genes under changing conditions. We mapped gene expression data from E. coli under aerobic and anaerobic conditions onto the enzymatic reaction nodes of its metabolic network. An adjacency matrix of the metabolites was created from this graph. A consecutive ones clustering method was used to obtain network clusters in the matrix. The wavelet method was applied on the adjacency matrices of these clusters to collect features for the classifier. With a feature extraction method the most discriminating features were selected. We yielded network sub-graphs from these top ranking features representing formate fermentation, in good agreement with the anaerobic response of hetero-fermentative bacteria. Furthermore, we found a switch in the starting point for NAD biosynthesis, and an adaptation of the l-aspartate metabolism, in accordance with its higher abundance under anaerobic conditions. CONCLUSION: We developed and tested a novel method, based on a combination of rationally chosen machine learning methods, to analyse gene expression data on the basis of interaction data, using a metabolic network of enzymes. As a case study, we applied our method to E. coli under oxygen deprived conditions and extracted physiologically relevant patterns that represent an adaptation of the cells to changing environmental conditions. In general, our concept may be transferred to network analyses on biological interaction data, when data for two comparable states of the associated nodes are made available.

Algorithms↗

Mapping ovarian cellular and molecular landscape across the lifespan of women: a scoping review.

BACKGROUND: With growing interest in ART, fertility preservation, and postmenopausal health of women, reproductive medicine is increasingly focused on characterizing oocytes and ovarian tissue composition, as well as understanding the molecular mechanisms that guide ovarian function throughout its lifecycle. High-throughput omics technologies have enabled the characterization of different molecular layers, leading to substantial advances in our understanding of their complex dynamics. However, not all molecular aspects are studied equally, and studies examining the same modalities often show inconsistencies, underscoring the need for data standardization and highlighting the potential for using transformative artificial intelligence and machine-learning (AI/ML) methods for ovary studies. OBJECTIVE AND RATIONALE: This study aims to evaluate how multi-omic studies have advanced our understanding of the ovarian lifecycle from fetal development to postmenopause. We systematically reviewed published studies that have investigated molecular/omic layers, including the genome, methylome, transcriptome, and proteome throughout ovarian development and aging. Our analysis identified key molecular and cellular patterns, highlighted inconsistencies across studies and addressed gaps in data analysis, interpretation, and reproducibility to guide future research. SEARCH METHODS: We conducted a systematic literature search of Medline (PubMed), Embase (Ovid), and Web of Science Core Collection (Clarivate) using a combination of controlled and free text terms for human ovary, oogenesis, folliculogenesis, ovary development and (epi)genome, transcriptome, proteome, and multi-omic mechanisms to find relevant articles published before August 2025. To focus the scope of the current review, studies of domesticated and farm animals, rodents and other model organisms, non-human primates, as well as those examining various human ovarian pathologies were excluded. OUTCOMES: The search identified 23 546 studies for screening, of which 637 full-text studies were assessed for eligibility. Subsequently, we extracted data from 121 studies. Most studies analyzed the transcriptome of oocytes, granulosa cells, and ovarian tissue from reproductive-age individuals (n = 91), with fewer studies examining samples from individuals of advanced reproductive age (n = 45) and fetal (n = 16) samples. Transcriptome analyses were most common (n = 103, 85%), followed by proteome (n = 19, 16%) and epigenome (n = 14, 12%) studies. We found substantial variation in how studies defined and reported participants' groups as well as in their sequencing technologies and data analysis methods, with a lack of standardized reporting of background clinical information, data analysis methods, and pipeline details. The key findings underscore the prevailing consensus on genes defining major ovarian cell types and their roles throughout the ovarian lifespan, from prenatal development to postmenopausal transformation. This review highlighted the underrepresentation of certain patient groups, particularly prepubertal and peri-/postmenopausal individuals, among researched populations, due to obvious clinical and ethical reasons. WIDER IMPLICATIONS: This scoping review offers a comprehensive overview and benchmark of the current state of high-throughput omics-based research on ovarian cellular composition and molecular dynamics. To address these shortcomings, we propose general recommendations for multi-omics ovary studies and emphasize the necessity for more thorough multi-omic data integration by effectively applying novel AI/ML approaches. They can potentially improve the quality of multi-omics analyses at both single-cell and tissue levels despite limited sample sizes and enable integration of molecular profiling data with clinical and radiology datasets, enabling a more comprehensive understanding of ovarian biology. Such advancements can enhance reproducibility of research findings and guide future research to deepen our understanding of ovarian biology and ultimately support the development of medical technologies for better preserving fertility and alleviating infertility. REGISTRATION NUMBER: A protocol was published a priori on the Open Science Framework (https://osf.io/z38gb/).

Female↗