Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Predicting transmembrane beta-barrels and interstrand residue interactions from sequence.

Transmembrane beta-barrel (TMB) proteins are embedded in the outer membrane of Gram-negative bacteria, mitochondria, and chloroplasts. The cellular location and functional diversity of beta-barrel outer membrane proteins (omps) makes them an important protein class. At the present time, very few nonhomologous TMB structures have been determined by X-ray diffraction because of the experimental difficulty encountered in crystallizing transmembrane proteins. A novel method using pairwise interstrand residue statistical potentials derived from globular (nonouter membrane) proteins is introduced to predict the supersecondary structure of transmembrane beta-barrel proteins. The algorithm transFold employs a generalized hidden Markov model (i.e., multitape S-attribute grammar) to describe potential beta-barrel supersecondary structures and then computes by dynamic programming the minimum free energy beta-barrel structure. Hence, the approach can be viewed as a "wrapping" component that may capture folding processes with an initiation stage followed by progressive interaction of the sequence with the already-formed motifs. This approach differs significantly from others, which use traditional machine learning to solve this problem, because it does not require a training phase on known TMB structures and is the first to explicitly capture and predict long-range interactions. TransFold outperforms previous programs for predicting TMBs on smaller (<or=200 residues) proteins and matches their performance for straightforward recognition of longer proteins. An exception is for multimeric porins where the algorithm does perform well when an important functional motif in loops is initially identified. We verify our simulations of the folding process by comparing them with experimental data on the functional folding of TMBs. A Web server running transFold is available and outputs contact predictions and locations for sequences predicted to form TMBs.

Amino Acid Sequence↗

Modeling intestinal absorption and other nutrition-related processes using PSPICE and STELLA.

In summary, SPICE models are constructed by translating a highly organized biological system into a network diagram by using a disciplined, systematic method for converting flows through barriers and chemical reactions into branches in a network connecting the compartments in the tissue according to the identity of the flowing entities. The first step in building a simulation model is essentially the same as the first step in learning about the method. Simple mechanisms are mastered first; and then as proficiency and understanding of the system grow, these can be connected and elaborated to produce simulations more closely approximating the real complexity of the living system. Other methods exist that may be easier to deal with initially, but often they cannot be utilized as generally as SPICE owing to their inherent limitations. One program available on Apple machines that has a high degree of user friendliness is STELLA. We will make some brief comparisons here, since STELLA is often an easier way to get started in simulation and often perfectly adequate for smaller problems.

Computer Simulation↗

Classifying 'drug-likeness' with kernel-based learning methods.

In this article we report about a successful application of modern machine learning technology, namely Support Vector Machines, to the problem of assessing the 'drug-likeness' of a chemical from a given set of descriptors of the substance. We were able to drastically improve the recent result by Byvatov et al. (2003) on this task and achieved an error rate of about 7% on unseen compounds using Support Vector Machines. We see a very high potential of such machine learning techniques for a variety of computational chemistry problems that occur in the drug discovery and drug design process.

Artificial Intelligence↗

Predicting protein-ligand binding affinities using novel geometrical descriptors and machine-learning methods.

Inspired by the concept of knowledge-based scoring functions, a new quantitative structure-activity relationship (QSAR) approach is introduced for scoring protein-ligand interactions. This approach considers that the strength of ligand binding is correlated with the nature of specific ligand/binding site atom pairs in a distance-dependent manner. In this technique, atom pair occurrence and distance-dependent atom pair features are used to generate an interaction score. Scoring and pattern recognition results obtained using Kernel PLS (partial least squares) modeling and a genetic algorithm-based feature selection method are discussed.

Algorithms↗

Proteome-wide structural and interaction analysis using cross-linking mass spectrometry and its applications.

Deciphering the mechanisms of protein-protein interactions (PPIs) and protein structural changes within the native cellular environment is crucial for advancing drug discovery. In vivo chemical cross-linking coupled with mass spectrometry (XL-MS) captures weak, transient, and higher-order interactions that are often dysregulated under altered physiological conditions and remain challenging to detect using conventional methods. Applications of in vivo XL-MS range from targeted mapping of PPIs to large-scale identification of interactome networks within the cells. The integration of quantitative approaches further facilitates comparison across different physiological conditions. The recent incorporation of machine learning (ML) tools into XL-MS workflows is transforming the depth and efficiency of this technology. AI-driven algorithms now enable more accurate identification of cross-linked peptides and the mapping of interaction topologies. Furthermore, the synergistic coupling of in vivo XL-MS data with AI-assisted structural modeling platforms such as AlphaFold allows dynamic and high-throughput prediction of protein networks. This review discusses the broader applications of in vivo XL-MS in complex biological samples, ranging from organelles and cells to whole tissues, and highlights how AI integration is expanding structural biology toward a systems-level understanding of proteome architecture.

Mass Spectrometry↗

Predicting the efficiency of UAG translational stop signal through studies of physicochemical properties of its composite mono- and dinucleotides.

In this study, we explored the problem of predicting the UAG stop-codon read-through efficiency. The reported nucleotide sequences were first converted into physicochemical property vectors before being presented to a machine learning algorithm. Two sets of physicochemical properties were applied: one for mononucleosides (in terms of steric bulk, hydrophobicity and electronics) and another for dinucleotides. To the best of our knowledge, this is the first report of how dinucleotides are converted into principle components derived from NMR chemical shift data. A few efficiency prediction models were then derived and a comparison between mononucleoside and dinucleotide-based models was shown. In the derived models, the coefficients of these property based predictors lend themselves to bio-physical interpretations, an advantage which is demonstrated in this study via a prediction model based on the steric bulk factor. Although it is quite simple, the steric bulk factor model explained well the effect of sequence variations surrounding the amber stop codon and the tRNA bearing UCCU anticodon. We further proposed new alternatives at position -1 and +4 of a UAG stop codon sequence to enhance the readthrough efficiency. This research may contribute to a better understanding of the readthrough mechanisms and may also help to study the normal translation termination process.

Codon, Terminator↗

Metagenomic analyses reveal E. coli-derived siderophores as potential signatures for breast cancer.

BACKGROUND: Breast cancer remains a leading cause of cancer-related mortality in women. Recent evidence implicates the gut microbiome and metabolites in breast cancer pathogenesis. This study explores associations between gut microbial species, their predicted metabolites, and breast cancer to uncover potential mechanistic insights. METHODS: Comprehensive metagenomic analyses were conducted on the gut microbiome of pre- and postmenopausal breast cancer patients, where microbial species were profiled through AMPHORA2 and metabolites were predicted through antiSMASH. Multivariate association analysis was used to identify significant associations between specific microbial species, predicted metabolites, and breast cancer status. A custom ensemble machine learning classifier was developed to classify pre- and postmenopausal breast cancer cases and controls based on microbial and predicted metabolite features. Additionally, a synthetic microbiome dataset was generated through MIDASim to validate the reproducibility of the ML results. Using our results, we explored the underlying dynamics of identified taxa and metabolite in breast cancer through literature and statistical support. RESULTS: Our analysis identified 471 microbial species and predicted 40 key metabolites in the metagenomic data. Multivariate analysis identified significant positive associations (p-value&#x2009;<&#x2009;0.05) of E. coli, siderophore, and thiopeptide with breast cancer. The custom ensemble model achieved accuracy and AUC as high as 78% and 90%, respectively, in classifying pre- and postmenopausal cases and controls. The high-ranking features i.e., E. coli, siderophore, and thiopeptide were consistent with the results of the multivariate association analysis, thereby substantiating their biological significance. Using these findings, we propose a mechanistic model in which E. coli secretes siderophores under iron-limited conditions in breast cancer patients, for iron sequestration from the host, which can potentially promote angiogenesis and tumor progression. CONCLUSION: Our findings suggest that microbial iron acquisition mechanisms may play a critical role in breast cancer pathophysiology. Functional validation of these mechanisms is needed to assess therapeutic potential. This study highlights gut microbiota and their metabolites as promising targets for breast cancer research and intervention.

Breast Neoplasms↗

Molecular classification of liver cirrhosis in a rat model by proteomics and bioinformatics.

Liver cirrhosis is a worldwide health problem. Reliable, noninvasive methods for early detection of liver cirrhosis are not available. Using a three-step approach, we classified sera from rats with liver cirrhosis following different treatment insults. The approach consisted of: (i) protein profiling using surface-enhanced laser desorption/ionization (SELDI) technology; (ii) selection of a statistically significant serum biomarker set using machine learning algorithms; and (iii) identification of selected serum biomarkers by peptide sequencing. We generated serum protein profiles from three groups of rats: (i) normal (n=8), (ii) thioacetamide-induced liver cirrhosis (n=22), and (iii) bile duct ligation-induced liver fibrosis (n=5) using a weak cation exchanger surface. Profiling data were further analyzed by a recursive support vector machine algorithm to select a panel of statistically significant biomarkers for class prediction. Sensitivity and specificity of classification using the selected protein marker set were higher than 92%. A consistently down-regulated 3495 Da protein in cirrhosis samples was one of the selected significant biomarkers. This 3495 Da protein was purified on-chip and trypsin digested. Further structural characterization of this biomarkers candidate was done by using cross-platform matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS) peptide mass fingerprinting (PMF) and matrix-assisted laser desorption/ionization time of flight/time of flight (MALDI-TOF/TOF) tandem mass spectrometry (MS/MS). Combined data from PMF and MS/MS spectra of two tryptic peptides suggested that this 3495 Da protein shared homology to a histidine-rich glycoprotein. These results demonstrated a novel approach to discovery of new biomarkers for early detection of liver cirrhosis and classification of liver diseases.

Algorithms↗

MetaChrome: An Open-Source, User-Friendly Tool for Automated Metaphase Chromosome Analysis.

DNA Fluorescence In Situ Hybridization (FISH) is an essential technique to study chromosome biology and genetics, enabling precise visualization of specific genomic loci to study structural abnormalities, gene mapping, and chromosomal rearrangements. High-Throughput Imaging (HTI) can automate the analysis of DNA-FISH chromosome images, but the accurate and automated segmentation of mitotic chromosomes and simultaneous colocalization of FISH signals remains a challenge. While several commercial automated karyotyping tools partially solve these issues, open-source software that effectively combines robust chromosome segmentation with comprehensive colocalization analysis capabilities remains necessary. To address this unmet need, we developed MetaChrome, an open-source software platform built around a graphical user interface and explicitly designed for automated metaphase chromosome analysis. MetaChrome leverages fine-tuned deep learning models to automate metaphase chromosome segmentation, together with colocalization analysis of chromosome-specific FISH probes and immunofluorescent-labeled proteins. Importantly, MetaChrome achieves enhanced segmentation accuracy compared to traditional image processing methods by adopting a Cellpose segmentation model fine-tuned with manually annotated metaphase chromosome datasets. The fine-tuned model ensures precise assignment of DNA-FISH spots to individual chromosomes in an automated manner. This facilitates rapid identification of chromosomal abnormalities, reduces human error, and advances high-throughput chromosome analysis workflows, addressing a key bottleneck in chromosome biology research.

Chromosome segmentation↗

Quantitative structure-toxicity relationships (QSTRs): a comparative study of various non linear methods. General regression neural network, radial basis function neural network and support vector machine in predicting toxicity of nitro- and cyano- aromatics to Tetrahymena pyriformis.

Prediction of toxicity of 203 nitro- and cyano-aromatic chemicals to Tetrahymena pyriformis was carried out by radial basis function neural network, general regression neural network and support vector machine, in non-linear response surface methodology. Toxicity was predicted from hydrophobicity parameter (log Kow) and maximum superdelocalizability (Amax). Special attention was drawn to prediction ability and robustness of the models, investigated both in a leave-one-out and 10-fold cross validation (CV) processes. The influence that the corresponding changes in the learning sets during these CV processes could have on a common external test set including 41 compounds was also examined. This allowed us to establish the stability of the models. The non linear results slightly outperform (as expected) multilinear relationships (MLR) and also favourably compete with various other non linear approaches recently proposed by Ren (J. Chem. Inf. Comput. Sci., 43 1679 (2003)).

Animals↗

Knowledge-enhanced protein subcellular localization prediction from 3D fluorescence microscope images.

MOTIVATION: Pinpointing the subcellular location of proteins is essential for studying protein function and related diseases. Advances in spatial proteomics have shown that automatic recognition of protein subcellular localization from images could highly facilitate protein translocation analysis and biomarker discovery, but existing machine-learning works have been mostly limited to processing 2D images. By contrast, 3D images have higher spatial resolution&#xa0;and allow researchers to observe cellular structures in their natural context, but currently, there are only a few studies of 3D image processing for protein distribution analysis due to the lack of data and complexity of modeling. RESULTS: We developed a knowledge-enhanced protein subcellular localization model, KE3DLoc, which could recognize distribution patterns in 3D fluorescence microscope images using deep learning methods. The model designs an image feature extraction module that incorporates information from 3D and 2D projected cells and implements asymmetric loss and confidence weights to address data imbalance and weak cell annotation issues. Besides, considering that the biological knowledge in the Gene Ontology (GO) database can provide valuable support for protein location understanding, the KE3DLoc model incorporates a novel knowledge enhancement module that optimizes the protein representation by related knowledge graphs derived from the GO. Since the image module and the knowledge module calculate features from different levels, KE3DLoc designs protein ID aggregation to enhance the consistency of protein features across different cells. Experimental results on three public datasets have demonstrated that the KE3DLoc significantly outperforms existing methods and provides valuable insights for spatial proteomics research. AVAILABILITY AND IMPLEMENTATION: All datasets and codes used in this study are available at GitHub: https://github.com/PRBioimages/KE3DLoc.

Microscopy, Fluorescence↗

Using nuclear morphometry to discriminate the tumorigenic potential of cells: a comparison of statistical methods.

Despite interest in the use of nuclear morphometry for cancer diagnosis and prognosis as well as to monitor changes in cancer risk, no generally accepted statistical method has emerged for the analysis of these data. To evaluate different statistical approaches, Feulgen-stained nuclei from a human lung epithelial cell line, BEAS-2B, and a human lung adenocarcinoma (non-small cell) cancer cell line, NCI-H522, were subjected to morphometric analysis using a CAS-200 imaging system. The morphometric characteristics of these two cell lines differed significantly. Therefore, we proceeded to address the question of which statistical approach was most effective in classifying individual cells into the cell lines from which they were derived. The statistical techniques evaluated ranged from simple, traditional, parametric approaches to newer machine learning techniques. The multivariate techniques were compared based on a systematic cross-validation approach using 10 fixed partitions of the data to compute the misclassification rate for each method. For comparisons across cell lines at the level of each morphometric feature, we found little to distinguish nonparametric from parametric approaches. Among the linear models applied, logistic regression had the highest percentage of correct classifications; among the nonlinear and nonparametric methods applied, the Classification and Regression Trees model provided the highest percentage of correct classifications. Classification and Regression Trees has appealing characteristics: there are no assumptions about the distribution of the variables to be used, there is no need to specify which interactions to test, and there is no difficulty in handling complex, high-dimensional data sets containing mixed data types.

Adenocarcinoma↗

Detection and analysis of statistical differences in anatomical shape.

We present a computational framework for image-based analysis and interpretation of statistical differences in anatomical shape between populations. Applications of such analysis include understanding developmental and anatomical aspects of disorders when comparing patients versus normal controls, studying morphological changes caused by aging, or even differences in normal anatomy, for example, differences between genders. Once a quantitative description of organ shape is extracted from input images, the problem of identifying differences between the two groups can be reduced to one of the classical questions in machine learning of constructing a classifier function for assigning new examples to one of the two groups while making as few misclassifications as possible. The resulting classifier must be interpreted in terms of shape differences between the two groups back in the image domain. We demonstrate a novel approach to such interpretation that allows us to argue about the identified shape differences in anatomically meaningful terms of organ deformation. Given a classifier function in the feature space, we derive a deformation that corresponds to the differences between the two classes while ignoring shape variability within each class. Based on this approach, we present a system for statistical shape analysis using distance transforms for shape representation and the support vector machines learning algorithm for the optimal classifier estimation and demonstrate it on artificially generated data sets, as well as real medical studies.

Algorithms↗

Survey of clustering algorithms.

Data analysis plays an indispensable role for understanding various phenomena. Cluster analysis, primitive exploration with little or no prior knowledge, consists of research developed across a wide variety of communities. The diversity, on one hand, equips us with many tools. On the other hand, the profusion of options causes confusion. We survey clustering algorithms for data sets appearing in statistics, computer science, and machine learning, and illustrate their applications in some benchmark data sets, the traveling salesman problem, and bioinformatics, a new field attracting intensive efforts. Several tightly related topics, proximity measure, and cluster validation, are also discussed.

Algorithms↗

Application of causal discovery of factors driving dissolved oxygen in estuarine environments.

Dissolved oxygen (DO) concentrations in estuarine bottom waters are a manifestation of multiple, interacting physical and biogeochemical processes, yet identifying their independent contributions remains challenging. Here, we analyze monthly water quality monitoring data from eight stations across Long Island Sound from 1994 to 2022 using a causal discovery framework (PCMCI+) and transformation of forcing variables. Our goal is to identify and isolate variables that causally influence bottom DO and improve predictive models by minimizing overfitting and multicollinearity. PCMCI+ reveals surface-layer temperature as the most important and consistent negative driver of bottom DO, followed by stratification. Wind events exhibit only brief relief by advection and mixing, while river discharge shows no direct causal link to DO, making it less influential than previously thought. Biogeochemical variables, including chlorophyll-a (Chl-a), nitrate and nitrite, and particulate carbon, influence DO through both contemporaneous and time-lagged pathways, often with signs that shift depending on the process. The derived models were evaluated by comparing skill scores, mean squared error, and Akaike Information Criterion. Both model types perform well, with coefficient of determination values exceeding 0.90 at multiple stations using only 3-5 predictors. Our analysis reveals that the best causal predictors are surface-layer temperature, stratification, Chl-a, and particle carbon. This approach provides a scalable framework for improving prediction models and understanding the mechanistic links that control the seasonal variability of DO in estuarine systems.

Estuaries↗

Divergent microbial preludes to necrotising enterocolitis defined by gut phages and bacterial resistomes.

BACKGROUND: Translating microbiome correlations into robust predictive features for complex gut disorders remains elusive, partly due to oversimplified models of pathogenesis and neglect of the virome, a key player in microbial ecosystems. Necrotising enterocolitis (NEC), a devastating disease of preterm infants with no reliable clinical predictors, exemplifies this challenge. OBJECTIVE: To determine the predictive potential of the gut prophageome and polymicrobial aetiologies for NEC. DESIGN: We applied integrated metagenomic and metatranscriptomic analyses and machine learning to 1825 longitudinal stool samples from 43 preterm infants who later developed NEC and 86 gestational age-matched and birthweight-matched controls across three US hospitals. We characterised gut prophageome acquisitions and their association with clinical exposures, including antibiotics, diet and pharmacotherapies. To predict NEC risk, we integrated pre-onset prophageome, antibacterial resistome and bacteriome profiles with neonatal pathology, stratifying the cohort by disease onset timing (early: &#x2264;40 days; late: >40&#x2009;days) for separate analysis. RESULTS: NEC cases exhibited distinct viral diversity trajectories before disease onset. Early-onset NEC was best predicted by phage-bacterial interaction signatures (75% accuracy, 81% sensitivity). Metatranscriptomics revealed increased phage DNA abundance with low gene expression, suggesting a lysogenic lifestyle that may stabilise pathobionts. These phages encode metabolic genes potentially enhancing pathobiont resilience. Late-onset NEC was best predicted by antibacterial resistome profiles (83% accuracy). CONCLUSION: The gut prophageome serves as both a source of pre-symptomatic predictive signals and an active modulator of NEC pathogenesis, with distinct microbial mechanisms driving early-onset and late-onset disease. These polymicrobial etiologies inform strategies for early detection, risk stratification and the development of microbiome-targeted preventive and therapeutic interventions.

BIOMARKERS↗

Interaction profile-based protein classification of death domain.

BACKGROUND: The increasing number of protein sequences and 3D structure obtained from genomic initiatives is leading many of us to focus on proteomics, and to dedicate our experimental and computational efforts on the creation and analysis of information derived from 3D structure. In particular, the high-throughput generation of protein-protein interaction data from a few organisms makes such an approach very important towards understanding the molecular recognition that make-up the entire protein-protein interaction network. Since the generation of sequences, and experimental protein-protein interactions increases faster than the 3D structure determination of protein complexes, there is tremendous interest in developing in silico methods that generate such structure for prediction and classification purposes. In this study we focused on classifying protein family members based on their protein-protein interaction distinctiveness. Structure-based classification of protein-protein interfaces has been described initially by Ponstingl et al. 1 and more recently by Valdar et al. 2 and Mintseris et al. 3, from complex structures that have been solved experimentally. However, little has been done on protein classification based on the prediction of protein-protein complexes obtained from homology modeling and docking simulation. RESULTS: We have developed an in silico classification system entitled HODOCO (Homology modeling, Docking and Classification Oracle), in which protein Residue Potential Interaction Profiles (RPIPS) are used to summarize protein-protein interaction characteristics. This system applied to a dataset of 64 proteins of the death domain superfamily was used to classify each member into its proper subfamily. Two classification methods were attempted, heuristic and support vector machine learning. Both methods were tested with a 5-fold cross-validation. The heuristic approach yielded a 61% average accuracy, while the machine learning approach yielded an 89% average accuracy. CONCLUSION: We have confirmed the reliability and potential value of classifying proteins via their predicted interactions. Our results are in the same range of accuracy as other studies that classify protein-protein interactions from 3D complex structure obtained experimentally. While our classification scheme does not take directly into account sequence information our results are in agreement with functional and sequence based classification of death domain family members.

Humans↗