Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

PostDOCK: a structural, empirical approach to scoring protein ligand complexes.

In this work we introduce a postprocessing filter (PostDOCK) that distinguishes true binding ligand-protein complexes from docking artifacts (that are created by DOCK 4.0.1). PostDOCK is a pattern recognition system that relies on (1) a database of complexes, (2) biochemical descriptors of those complexes, and (3) machine learning tools. We use the protein databank (PDB) as the structural database of complexes and create diverse training and validation sets from it based on the "families of structurally similar proteins" (FSSP) hierarchy. For the biochemical descriptors, we consider terms from the DOCK score, empirical scoring, and buried solvent accessible surface area. For the machine-learners, we use a random forest classifier and logistic regression. Our results were obtained on a test set of 44 structurally diverse protein targets. Our highest performing descriptor combinations obtained approximately 19-fold enrichment (39 of 44 binding complexes were correctly identified, while only allowing 2 of 44 decoy complexes), and our best overall accuracy was 92%.

Ligands↗

Learning about protein hydrogen bonding by minimizing contrastive divergence.

Defining the strength and geometry of hydrogen bonds in protein structures has been a challenging task since early days of structural biology. In this article, we apply a novel statistical machine learning technique, known as contrastive divergence, to efficiently estimate both the hydrogen bond strength and the geometric characteristics of strong interpeptide backbone hydrogen bonds, from a dataset of structures representing a variety of different protein folds. Despite the simplifying assumptions of the interatomic energy terms used, we determine the strength of these hydrogen bonds to be between 1.1 and 1.5 kcal/mol, in good agreement with earlier experimental estimates. The geometry of these strong backbone hydrogen bonds features an almost linear arrangement of all four atoms involved in hydrogen bond formation. We estimate that about a quarter of all hydrogen bond donors and acceptors participate in these strong interpeptide hydrogen bonds.

Amino Acids↗

Downregulated lysyl oxidase in plasma extracellular vesicles: a biomarker linked to brain metastasis risk in lung adenocarcinoma.

BACKGROUND: Brain metastasis (BrM) is a leading cause of mortality in patients with lung adenocarcinoma (LUAD). Extracellular vesicles (EVs), which carry bioactive molecules, play a critical role in tumor microenvironment remodeling and exhibit metastatic organotropism, holding promise as liquid biopsy biomarkers. This study aims to identify plasma EV-derived proteins associated with LUAD-BrM. METHODS: A multi-omics framework was applied. Plasma EVs from 59 stage IV LUAD patients (30 BrM vs 29 non-BrM) were profiled using data-independent acquisition mass spectrometry proteomics. Candidate proteins were screened via bioinformatics and machine learning (LASSO/RF/SVM). Initial validation included tissue proteomics (n = 13), single-cell transcriptomics (TISCH2), and Western blot analysis of a subset of the discovery samples. Functional experiments were conducted in vitro. The lead candidate was ultimately validated in an independent plasma cohort (n = 158) through ELISA. RESULTS: Proteomic analysis implicated collagen-containing extracellular matrix (ECM) pathways. Lysyl oxidase (LOX), a key ECM cross-linking enzyme, was identified as a lead candidate. LOX and its family member LOXL1 were consistently downregulated in BrM tissues and plasma EVs. Single-cell analysis revealed decreased LOX expression specifically in BrM-associated fibroblasts, which showed suppressed ECM-related pathways. In vitro experiments supported a PI3K/AKT-LOX-ECM regulatory axis. Plasma EV-derived LOX demonstrated strong diagnostic performance in the independent cohort, with an AUC of 0.786 (95% CI 0.713iated fi. CONCLUSIONS: Our study establishes plasma EV-derived LOX as a promising non-invasive biomarker for LUAD-BrM through a comprehensive multi-omics validation strategy. We propose a model wherein downregulation of LOX, potentially driven by PI3K/AKT signaling in tumor-associated fibroblasts, contributes to ECM degradation and may promote brain-tropic metastasis. This finding offers new insights for risk stratification and timely intervention in LUAD patients.

Humans↗

Support vector machines for prediction and analysis of beta and gamma-turns in proteins.

Tight turns have long been recognized as one of the three important features of proteins, together with alpha-helix and beta-sheet. Tight turns play an important role in globular proteins from both the structural and functional points of view. More than 90% tight turns are beta-turns and most of the rest are gamma-turns. Analysis and prediction of beta-turns and gamma-turns is very useful for design of new molecules such as drugs, pesticides, and antigens. In this paper we investigated two aspects of applying support vector machine (SVM), a promising machine learning method for bioinformatics, to prediction and analysis of beta-turns and gamma-turns. First, we developed two SVM-based methods, called BTSVM and GTSVM, which predict beta-turns and gamma-turns in a protein from its sequence. When compared with other methods, BTSVM has a superior performance and GTSVM is competitive. Second, we used SVMs with a linear kernel to estimate the support of amino acids for the formation of beta-turns and gamma-turns depending on their position in a protein. Our analysis results are more comprehensive and easier to use than the previous results in designing turns in proteins.

Algorithms↗

RASE: recognition of alternatively spliced exons in C.elegans.

MOTIVATION: Eukaryotic pre-mRNAs are spliced to form mature mRNA. Pre-mRNA alternative splicing greatly increases the complexity of gene expression. Estimates show that more than half of the human genes and at least one-third of the genes of less complex organisms, such as nematodes or flies, are alternatively spliced. In this work, we consider one major form of alternative splicing, namely the exclusion of exons from the transcript. It has been shown that alternatively spliced exons have certain properties that distinguish them from constitutively spliced exons. Although most recent computational studies on alternative splicing apply only to exons which are conserved among two species, our method only uses information that is available to the splicing machinery, i.e. the DNA sequence itself. We employ advanced machine learning techniques in order to answer the following two questions: (1) Is a certain exon alternatively spliced? (2) How can we identify yet unidentified exons within known introns? RESULTS: We designed a support vector machine (SVM) kernel well suited for the task of classifying sequences with motifs having positional preferences. In order to solve the task (1), we combine the kernel with additional local sequence information, such as lengths of the exon and the flanking introns. The resulting SVM-based classifier achieves a true positive rate of 48.5% at a false positive rate of 1%. By scanning over single EST confirmed exons we identified 215 potential alternatively spliced exons. For 10 randomly selected such exons we successfully performed biological verification experiments and confirmed three novel alternatively spliced exons. To answer question (2), we additionally used SVM-based predictions to recognize acceptor and donor splice sites. Combined with the above mentioned features we were able to identify 85.2% of skipped exons within known introns at a false positive rate of 1%. AVAILABILITY: Datasets, model selection results, our predictions and additional experimental results are available at http://www.fml.tuebingen.mpg.de/~raetsch/RASE SUPPLEMENTARY INFORMATION: http://www.fml.tuebingen.mpg.de/raetsch/RASE.

Algorithms↗

Reversal of cancer gene expression identifies repurposed drugs for diffuse intrinsic pontine glioma.

Diffuse intrinsic pontine glioma (DIPG) is an aggressive incurable brainstem tumor that targets young children. Complete resection is not possible, and chemotherapy and radiotherapy are currently only palliative. This study aimed to identify potential therapeutic agents using a computational pipeline to perform an in silico screen for novel drugs. We then tested the identified drugs against a panel of patient-derived DIPG cell lines. Using a systematic computational approach with publicly available databases of gene signature in DIPG patients and cancer cell lines treated with a library of clinically available drugs, we identified drug hits with the ability to reverse a DIPG gene signature to one that matches normal tissue background. The biological and molecular effects of drug treatment was analyzed by cell viability assay and RNA sequence. In vivo DIPG mouse model survival studies were also conducted. As a result, two of three identified drugs showed potency against the DIPG cell lines Triptolide and mycophenolate mofetil (MMF) demonstrated significant inhibition of cell viability in DIPG cell lines. Guanosine rescued reduced cell viability induced by MMF. In vivo, MMF treatment significantly inhibited tumor growth in subcutaneous xenograft mice models. In conclusion, we identified clinically available drugs with the ability to reverse DIPG gene signatures and anti-DIPG activity in vitro and in vivo. This novel approach can repurpose drugs and significantly decrease the cost and time normally required in drug discovery.

Humans↗

The application of artificial intelligence in healthcare practice: A mapping review of systematic reviews.

Artificial intelligence (AI) is rapidly transforming healthcare practice, with growing evidence supporting its use in diagnosis, prognosis, treatment planning, and operational decision-making. The proliferation of systematic reviews in recent years underscores the need for an updated synthesis of the literature to inform research, policy, and practice. We searched PubMed, Web of Science, Scopus, IEEE Xplore, and CINAHL for systematic reviews and meta-analyses published between 2019 and February 2026. Eligible reviews focused on AI applications in healthcare practice, were peer-reviewed, and written in English. A total of 368 reviews met the inclusion criteria. Publication volume increased steadily, peaking in 2025. AI research was concentrated in high-density domains, such as radiology, oncology, and critical care. Across reviews, diagnostic imaging, electronic health record (EHR) data, and biomarkers/laboratory results accounted for 68% of training data sources, though newer data types, such as wearable device and sensor data, emerged from 2022 onward. Diagnosis, prognosis, and treatment comprised over 80% of AI applications, with novel uses emerging in recent years, such as AI-assisted clinical documentation (e.g., ambient documentation tools) and patient education. Ethical concerns were reported in 78.5% of reviews, with privacy, model accuracy, data and algorithmic bias, and explainability as recurrent themes. The proportion of reviews reporting ethical concerns increased from 2021 to 2025. AI applications in healthcare are expanding in scope, diversifying in data sources, and evolving toward novel clinical and operational uses. The human-centered AI or augmented intelligence paradigm, integrating computational precision with clinical expertise, holds significant promise but will require parallel advances in governance, regulatory frameworks, and ethical oversight to ensure safe adoption.

Artificial Intelligence↗

A genetic-based machine learning system to discover the diagnostic rules for female urinary incontinence.

A machine learning system named Galactica has been developed which uses a genetic algorithm to discover the rules for an expert system from databases. Galactica devised accurate diagnostic rules for female urinary incontinence from difficult heterogeneous data. The percentages of correctly classified stress, mixed and sensory urge incontinence testing cases were 89, 86 and 87%, respectively. However, these rules were rather general, consisting of 4-6 out of 13 conditions available in the data. Diagnostic rules for stress and mixed incontinence extracted from straightforward homogeneous data were highly accurate, classifying 100% of testing cases correctly as well as being specific, having from 10 to 11 conditions. More specific, but less accurate, rules were found from heterogeneous data with a biased fitness function. All of the rules were correct, i.e. every condition in the rules had the expected value specified by the expert. Although, Galactica achieved a slightly better classification than the discriminant analysis, it is argued that the genetic approach is better than the statistical one, due to symbolic rules being comprehensible, whereas understanding a complex mathematical model requires statistical expertise.

Algorithms↗

Predicting the toxicity of complex mixtures using artificial neural networks.

Industrial and municipal wastewaters constitute major sources of contamination of the aquatic compartment and represent a threat to aquatic life. Artificial neural networks based on three different learning paradigms were studied as a means of predicting acute toxicity to trout (5 days exposure to wastewaters) using input data from two simple microbiotests requiring only 5 or 15 min of incubation. These microbiotests were 1) the chemoluminescent peroxidase (Cl-Per) assay, which can detect radical scavengers and enzyme-inhibiting substances, and 2) the luminescent bacteria toxicity test (Microtox), in which reduction of light emission by bacteria during exposure is taken as a measure of toxicity. The responses obtained with the trout bioassay, the Cl-Per and the Microtox test were analyzed through statistical correlation (Pearson product-moment correlation), unsupervised learning by a self-organizing network, and assisted learning by the backpropagation and the Boltzmann machine (probabilistic) paradigms. No significant correlation (p < 0.05) was found between the responses obtained with either the Cl-Per assay (p = 0.121) or the Microtox (p = 0.061) microbiotest and those resulting from the trout bioassay. The self-organizing network was able to identify by itself a maximum of five classes that were more or less relevant for predicting toxicity to fish: class 1 contained 2 samples that were toxic to fish, class 2 contained 2/3 samples that were toxic, class 3 showed 6/8 samples that were non toxic, class 4 contained 5/6 samples that were non-toxic and class 5 comprised one sample that was toxic. Supervised learning with backpropagation analysis yielded two kinds of networks that hold potential. The first one was able to predict the actual toxic wastewater concentration with an overall performance of 65% when fed fresh data, while the second one, which was designed to differentiate between toxic and non-toxic effluents, exhibited a much better performance (90%). However, the probabilistic network also proved to be a very good predictive model for toxicity to fish, with an overall performance of 90%. Although more data are needed, the network based on the backpropagation paradigm seems to be a better predictor or classifier of trout toxicity when used with the Cl-Per and the Microtox microbiotests.

Algorithms↗

Use of an electronic nose to diagnose bacterial sinusitis.

BACKGROUND: Having previously established that an electronic nose (enose) can distinguish among bacteria samples, between cerebrospinal fluid leak and serum, and can identify patients with ventilator-associated pneumonia, we hypothesized that bacterial sinusitis could be diagnosed by sampling exhaled gas with an enose. METHODS: Using a nasal continuous positive airway pressure mask, we sampled gas exhaled through the nose of patients with sinusitis and compared them with controls. Data were first projected onto the principal components and then classified by support vector machine (SVM), a machine learning algorithm for pattern recognition. RESULTS: SVM analysis showed good discrimination using three approaches. First, 11 samples were used to create a training set that was used to predict whether individual samples from each set were a member of the control or infected sets. The enose was correct 98.4% of the time. Second, one-half of the samples from each of the same 11 control and infected groups were used to construct a training set, which was used to predict infection in the remaining samples. The enose was correct 82% of the time. Finally, 68 samples (34 positive and 34 controls) were analyzed using a leave-one-out scheme for creating training sets and testing sets. This method, designed to reflect the generalization property of the SVM classifier, scored a classification rate of 72%. CONCLUSION: Using the enose to sample nasal exhalation from patients with suspected sinusitis, we were able to predict correctly the diagnosis of sinusitis in at least 72% of the samples. The next step will be to do forward prediction using this model.

Algorithms↗

Karyotyping of comparative genomic hybridization human metaphases using kernel nearest-neighbor algorithm.

BACKGROUND: Comparative genomic hybridization (CGH) is a relatively new molecular cytogenetic method that detects chromosomal imbalances. Automatic karyotyping is an important step in CGH analysis because the precise position of the chromosome abnormality must be located and manual karyotyping is tedious and time-consuming. In the past, computer-aided karyotyping was done by using the 4',6-diamidino-2-phenylindole, dihydrochloride (DAPI)-inverse images, which required complex image enhancement procedures. METHODS: An innovative method, kernel nearest-neighbor (K-NN) algorithm, is proposed to accomplish automatic karyotyping. The algorithm is an application of the "kernel approach," which offers an alternative solution to linear learning machines by mapping data into a high dimensional feature space. By implicitly calculating Euclidean or Mahalanobis distance in a high dimensional image feature space, two kinds of K-NN algorithms are obtained. New feature extraction methods concerning multicolor information in CGH images are used for the first time. RESULTS: Experiment results show that the feature extraction method of using multicolor information in CGH images improves greatly the classification success rate. A high success rate of about 91.5% has been achieved, which shows that the K-NN classifier efficiently accomplishes automatic chromosome classification from relatively few samples. CONCLUSIONS: The feature extraction method proposed here and K-NN classifiers offer a promising computerized intelligent system for automatic karyotyping of CGH human chromosomes.

Algorithms↗

Evolving beyond perfection: an investigation of the effects of long-term evolution on fractal gene regulatory networks.

This paper continues a theme of exploring algorithms based on principles of biological development for tasks such as pattern generation, machine learning and robot control. Previous work has investigated the use of genes expressed as fractal proteins to enable greater evolvability of gene regulatory networks (GRNs). Here, the evolution of such GRNs is investigated further to determine whether evolution exhibits natural tendencies towards efficiency and graceful degradation of developmental programs. Experiments where "perfect" GRNs are evolved for a further thousand generations without the addition of any further selection pressure, confirm this hypothesis. After further evolution, the perfect GRNs operate in a more efficient manner (using fewer proteins) and show an improved ability to function correctly with missing genes. When the algorithm is applied to applications (e.g. robot control) this equates to efficient and fault-tolerant controllers.

Algorithms↗

Reading protocol for dynamic contrast-enhanced MR images of the breast: sensitivity and specificity analysis.

PURPOSE: To prospectively determine sensitivity and specificity of breast magnetic resonance (MR) imaging in a screening and symptomatic population by using independent double reading, with histologic or cytologic results or a minimum 18-month follow-up as the standard. MATERIALS AND METHODS: Informed consent and ethical approval were obtained. Reader performance was analyzed in 44 radiologists at 18 centers from 1541 examinations, including 1441 screening examinations in 638 high-risk women aged 24-51 years (mean, 40.5 years) and 100 examinations in symptomatic women aged 23-81 years (mean, 49.2 years). A screening protocol of dynamic T1-weighted three-dimensional imaging and 0.2 mmol/kg gadolinium-based intravenous contrast agent was used. Logistic and Poisson regressions were used to analyze reader performance in relation to experience. Correlation between readers was determined with kappa statistics. Sensitivity and specificity were analyzed according to reader, field strength, machine type, and histologic results. RESULTS: The proportion of studies with lesions analyzed reduced significantly with reader experience (odds ratio, 0.84 per 6 months; P < .001), and number of regions per lesion analyzed also diminished (incidence rate ratio, 0.98 per 6 months; P = .047). The two readers for each study agreed 87% of the time, with a moderately good kappa statistic of 0.52 (95% confidence interval [CI]: 0.45, 0.58). By taking the reading with the highest score (most likely to be malignant) from each double-read study, sensitivity was 91% (95% CI: 83%, 96%) and specificity was 81% (95% CI: 79%, 83%). Single readings had 7% lower sensitivity (95% CI: 4%, 11%) and 7% higher specificity (95% CI: 6%, 7%). Sensitivity did not differ between MR imager manufacturers or between 1.0- and 1.5-T field strength, but there were significant differences in specificity for machine type (P = .001) and for field strength adjusted for manufacturer (P = .001). Specificity, but not sensitivity, was higher in women younger than 50 years (P = .02). CONCLUSION: Independent double reading by 44 radiologists blinded to mammography results showed sensitivity and specificity acceptable for screening; sensitivity was higher when two readings were used, at the cost of specificity. Interreader correlation was moderately good, and evidence of learning was seen. Equipment manufacturer, field strength, and age affected specificity but not sensitivity.

Adult↗

Descriptor-based protein remote homology identification.

Here, we report a novel protein sequence descriptor-based remote homology identification method, able to infer fold relationships without the explicit knowledge of structure. In a first phase, we have individually benchmarked 13 different descriptor types in fold identification experiments in a highly diverse set of protein sequences. The relevant descriptors were related to the fold class membership by using simple similarity measures in the descriptor spaces, such as the cosine angle. Our results revealed that the three best-performing sets of descriptors were the sequence-alignment-based descriptor using PSI-BLAST e-values, the descriptors based on the alignment of secondary structural elements (SSEA), and the descriptors based on the occurrence of PROSITE functional motifs. In a second phase, the three top-performing descriptors were combined to obtain a final method with improved performance, which we named DescFold. Class membership was predicted by Support Vector Machine (SVM) learning. In comparison with the individual PSI-BLAST-based descriptor, the rate of remote homology identification increased from 33.7% to 46.3%. We found out that the composite set of descriptors was able to identify the true remote homolog for nearly every sixth sequence at the 95% confidence level, or some 10% more than a single PSI-BLAST search. We have benchmarked the DescFold method against several other state-of-the-art fold recognition algorithms for the 172 LiveBench-8 targets, and we concluded that it was able to add value to the existing techniques by providing a confident hit for at least 10% of the sequences not identifiable by the previously known methods.

Algorithms↗

Shaping the zebrafish notochord.

Promptly after the notochord domain is specified in the vertebrate dorsal mesoderm, it undergoes dramatic morphogenesis. Beginning during gastrulation, convergence and extension movements change a squat cellular array into a narrow, elongated one that defines the primary axis of the embryo. Convergence and extension might be coupled by a highly organized cellular intermixing known as mediolateral intercalation behavior (MIB). To learn whether MIB drives early morphogenesis of the zebrafish notochord, we made 4D recordings and quantitatively analyzed both local cellular interactions and global changes in the shape of the dorsal mesodermal field. We show that MIB appears to mediate convergence and can account for extension throughout the dorsal mesoderm. Comparing the notochord and adjacent somitic mesoderm reveals that extension can be regulated separately from convergence. Moreover, mutational analysis shows that extension does not require convergence. Hence, a cellular machine separate from MIB that can drive dorsal mesodermal extension exists in the zebrafish gastrula. The likely redundant control of morphogenesis may provide for plasticity at this critical stage of early development.

Animals↗

Multi-class protein fold classification using a new ensemble machine learning approach.

Protein structure classification represents an important process in understanding the associations between sequence and structure as well as possible functional and evolutionary relationships. Recent structural genomics initiatives and other high-throughput experiments have populated the biological databases at a rapid pace. The amount of structural data has made traditional methods such as manual inspection of the protein structure become impossible. Machine learning has been widely applied to bioinformatics and has gained a lot of success in this research area. This work proposes a novel ensemble machine learning method that improves the coverage of the classifiers under the multi-class imbalanced sample sets by integrating knowledge induced from different base classifiers, and we illustrate this idea in classifying multi-class SCOP protein fold data. We have compared our approach with PART and show that our method improves the sensitivity of the classifier in protein fold classification. Furthermore, we have extended this method to learning over multiple data types, preserving the independence of their corresponding data sources, and show that our new approach performs at least as well as the traditional technique over a single joined data source. These experimental results are encouraging, and can be applied to other bioinformatics problems similarly characterised by multi-class imbalanced data sets held in multiple data sources.

Amino Acid Sequence↗

Asymmetric Boltzmann machines.

We study asymmetric stochastic networks from two points of view: combinatorial optimization and learning algorithms based on relative entropy minimization. We show that there are non trivial classes of asymmetric networks which admit a Lyapunov function L under deterministic parallel evolution and prove that the stochastic augmentation of such networks amounts to a stochastic search for global minima of L. The problem of minimizing L for a totally antisymmetric parallel network is shown to be associated to an NP-complete decision problem. The study of entropic learning for general asymmetric networks, performed in the non equilibrium, time dependent formalism, leads to a Hebbian rule based on time averages over the past history of the system. The general algorithm for asymmetric networks is tested on a feed-forward architecture.

Algorithms↗

Machine learning in prognosis of the femoral neck fracture recovery.

We compare the performance of several machine learning algorithms in the problem of prognostics of the femoral neck fracture recovery: the K-nearest neighbours algorithm, the semi-naive Bayesian classifier, backpropagation with weight elimination learning of the multilayered neural networks, the LFC (lookahead feature construction) algorithm, and the Assistant-I and Assistant-R algorithms for top down induction of decision trees using information gain and RELIEFF as search heuristics, respectively. We compare the prognostic accuracy and the explanation ability of different classifiers. Among the different algorithms the semi-naive Bayesian classifier and Assistant-R seem to be the most appropriate. We analyze the combination of decisions of several classifiers for solving prediction problems and show that the combined classifier improves both performance and the explanation ability.

Algorithms↗