Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Application of data mining approaches to drug delivery.

Computational approaches play a key role in all areas of the pharmaceutical industry from data mining, experimental and clinical data capture to pharmacoeconomics and adverse events monitoring. They will likely continue to be indispensable assets along with a growing library of software applications. This is primarily due to the increasingly massive amount of biology, chemistry and clinical data, which is now entering the public domain mainly as a result of NIH and commercially funded projects. We are therefore in need of new methods for mining this mountain of data in order to enable new hypothesis generation. The computational approaches include, but are not limited to, database compilation, quantitative structure activity relationships (QSAR), pharmacophores, network visualization models, decision trees, machine learning algorithms and multidimensional data visualization software that could be used to improve drug delivery after mining public and/or proprietary data. We will discuss some areas of unmet needs in the area of data mining for drug delivery that can be addressed with new software tools or databases of relevance to future pharmaceutical projects.

Computer Simulation↗

Genome-wide detection of human 5' UTR variants that impact protein translation.

The 5' untranslated region (5' UTR) of messenger RNAs (mRNAs) plays a central role in regulating protein synthesis initiation, particularly through the Kozak sequence and upstream open reading frames (uORFs). Genetic variants within these regulatory elements could affect translation, altering gene expression and contributing to clinical phenotypes in humans. We developed a computational method called 5ULTRA (5' Untranslated Region Annotation) for analysis of whole-exome sequencing and whole-genome sequencing data to detect, annotate, and prioritize 5' UTR variants with potential translation impact. 5ULTRA identifies single-nucleotide variants, indels, and splicing variants that affect uORFs by creating or disrupting start/stop codons and that alter Kozak sequence strength of either the uORFs or the main coding sequence. 5ULTRA incorporates recent uORF databases and provides comprehensive annotations. 5ULTRA implements a machine-learning score to prioritize candidate variants with predicted effects on translation and also provides specific mechanistic predictions. The score correlates strongly with experimentally measured protein-level effects of 5' UTR variants. We applied 5ULTRA to multiple genetics datasets across diverse disease contexts, identifying candidate variants including potential cancer-driving somatic mutations predicted to decrease ABI1 level or increase NRAS abundance; common variants associated with traits such as multiple sclerosis, lung function, and cardiovascular function, by altering protein levels of TAGAP, VRTN, and SPAAR, respectively; and rare germline variants in our cohort, including a splicing variant of RPSA leading to 5' UTR sequence alteration that causes congenital asplenia and a variant of TNF that could predispose to tuberculosis.

Humans↗

Radon in a thermal spring: identification of anomalies related to seismic activity.

Anomalies have been observed in the radon content of thermal spring water at the Italian-Slovenian border. To distinguish the anomalies caused by environmental parameters (air and water temperature, barometric and hydrostatic pressure, rainfall) from those ascribed solely to earthquakes with M(L) from 1.2 to 2.5 and epicentres, R(E), within 2R(D) (R(D)--Dobrovolsky's radius), two approaches have been used: (i) correlation between time gradients of radon concentration and hydrostatic pressure, and (ii) regression trees within machine learning programs. The regression trees approach has been improved by introducing additional environmental parameters and prolonging the measuring period.

Disasters↗

Filter versus wrapper gene selection approaches in DNA microarray domains.

DNA microarray experiments generating thousands of gene expression measurements, are used to collect information from tissue and cell samples regarding gene expression differences that could be useful for diagnosis disease, distinction of the specific tumor type, etc. One important application of gene expression microarray data is the classification of samples into known categories. As DNA microarray technology measures the gene expression en masse, this has resulted in data with the number of features (genes) far exceeding the number of samples. As the predictive accuracy of supervised classifiers that try to discriminate between the classes of the problem decays with the existence of irrelevant and redundant features, the necessity of a dimensionality reduction process is essential. We propose the application of a gene selection process, which also enables the biology researcher to focus on promising gene candidates that actively contribute to classification in these large scale microarrays. Two basic approaches for feature selection appear in machine learning and pattern recognition literature: the filter and wrapper techniques. Filter procedures are used in most of the works in the area of DNA microarrays. In this work, a comparison between a group of different filter metrics and a wrapper sequential search procedure is carried out. The comparison is performed in two well-known DNA microarray datasets by the use of four classic supervised classifiers. The study is carried out over the original-continuous and three-intervals discretized gene expression data. While two well-known filter metrics are proposed for continuous data, four classic filter measures are used over discretized data. The same wrapper approach is used for both continuous and discretized data. The application of filter and wrapper gene selection procedures leads to considerably better accuracy results in comparison to the non-gene selection approach, coupled with interesting and notable dimensionality reductions. Although the wrapper approach mainly shows a more accurate behavior than filter metrics, this improvement is coupled with considerable computer-load necessities. We note that most of the genes selected by proposed filter and wrapper procedures in discrete and continuous microarray data appear in the lists of relevant-informative genes detected by previous studies over these datasets. The aim of this work is to make contributions in the field of the gene selection task in DNA microarray datasets. By an extensive comparison with more popular filter techniques, we would like to make contributions in the expansion and study of the wrapper approach in this type of domains.

Artificial Intelligence↗

Evolving rule-based systems in two medical domains using genetic programming.

OBJECTIVE: To demonstrate and compare the application of different genetic programming (GP) based intelligent methodologies for the construction of rule-based systems in two medical domains: the diagnosis of aphasia's subtypes and the classification of pap-smear examinations. MATERIAL: Past data representing (a) successful diagnosis of aphasia's subtypes from collaborating medical experts through a free interview per patient, and (b) correctly classified smears (images of cells) by cyto-technologists, previously stained using the Papanicolaou method. METHODS: Initially a hybrid approach is proposed, which combines standard genetic programming and heuristic hierarchical crisp rule-base construction. Then, genetic programming for the production of crisp rule based systems is attempted. Finally, another hybrid intelligent model is composed by a grammar driven genetic programming system for the generation of fuzzy rule-based systems. RESULTS: Results denote the effectiveness of the proposed systems, while they are also compared for their efficiency, accuracy and comprehensibility, to those of an inductive machine learning approach as well as to those of a standard genetic programming symbolic expression approach. CONCLUSION: The proposed GP-based intelligent methodologies are able to produce accurate and comprehensible results for medical experts performing competitive to other intelligent approaches. The aim of the authors was the production of accurate but also sensible decision rules that could potentially help medical doctors to extract conclusions, even at the expense of a higher classification score achievement.

Aphasia↗

A spatio-temporal Bayesian network classifier for understanding visual field deterioration.

OBJECTIVE: Progressive loss of the field of vision is characteristic of a number of eye diseases such as glaucoma which is a leading cause of irreversible blindness in the world. Recently, there has been an explosion in the amount of data being stored on patients who suffer from visual deterioration including field test data, retinal image data and patient demographic data. However, there has been relatively little work in modelling the spatial and temporal relationships common to such data. In this paper we introduce a novel method for classifying visual field (VF) data that explicitly models these spatial and temporal relationships. METHODOLOGY: We carry out an analysis of our proposed spatio-temporal Bayesian classifier and compare it to a number of classifiers from the machine learning and statistical communities. These are all tested on two datasets of VF and clinical data. We investigate the receiver operating characteristics curves, the resulting network structures and also make use of existing anatomical knowledge of the eye in order to validate the discovered models. RESULTS: Results are very encouraging showing that our classifiers are comparable to existing statistical models whilst also facilitating the understanding of underlying spatial and temporal relationships within VF data. The results reveal the potential of using such models for knowledge discovery within ophthalmic databases, such as networks reflecting the 'nasal step', an early indicator of the onset of glaucoma. CONCLUSION: The results outlined in this paper pave the way for a substantial program of study involving many other spatial and temporal datasets, including retinal image and clinical data.

Algorithms↗

Comparison between neural networks and multiple logistic regression to predict acute coronary syndrome in the emergency room.

OBJECTIVE: Patients with suspicion of acute coronary syndrome (ACS) are difficult to diagnose and they represent a very heterogeneous group. Some require immediate treatment while others, with only minor disorders, may be sent home. Detecting ACS patients using a machine learning approach would be advantageous in many situations. METHODS AND MATERIALS: Artificial neural network (ANN) ensembles and logistic regression models were trained on data from 634 patients presenting an emergency department with chest pain. Only data immediately available at patient presentation were used, including electrocardiogram (ECG) data. The models were analyzed using receiver operating characteristics (ROC) curve analysis, calibration assessments, inter- and intra-method variations. Effective odds ratios for the ANN ensembles were compared with the odds ratios obtained from the logistic model. RESULTS: The ANN ensemble approach together with ECG data preprocessed using principal component analysis resulted in an area under the ROC curve of 80%. At the sensitivity of 95% the specificity was 41%, corresponding to a negative predictive value of 97%, given the ACS prevalence of 21%. Adding clinical data available at presentation did not improve the ANN ensemble performance. Using the area under the ROC curve and model calibration as measures of performance we found an advantage using the ANN ensemble models compared to the logistic regression models. CONCLUSION: Clinically, a prediction model of the present type, combined with the judgment of trained emergency department personnel, could be useful for the early discharge of chest pain patients in populations with a low prevalence of ACS.

Acute Disease↗

Dynamic knowledge validation and verification for CBR teledermatology system.

OBJECTIVE: Case-based reasoning has been of great importance in the development of many decision support applications. However, relatively little effort has gone into investigating how new knowledge can be validated. Knowledge validation is important in dealing with imperfect data collected over time, because inconsistencies in data do occur and adversely affect the performance of a diagnostic system. METHODS: This paper consists of two parts. First, it describes methods that enable the domain expert, who may not be familiar with machine learning, to interactively validate knowledge base of a Web-based teledermatology system. The validation techniques involve decision tree classification and formal concept analysis. Second, it describes techniques to discover unusual relationships hidden in the dataset for building and updating a comprehensive knowledge base, because the diagnostic performance of the system is highly dependent on the content thereof. Therefore, in order to classify different kinds of diseases, it is desirable to have a knowledge base that covers common as well as uncommon diagnoses. RESULTS AND CONCLUSION: Evaluation results show that the knowledge validation techniques are effective in keeping the knowledge base consistent, and that the query refinement techniques are useful in improving the comprehensiveness of the case base.

Dermatology↗

Cardiac surgery risk models: a position article.

Differences in medical outcomes may result from disease severity, treatment effectiveness, or chance. Because most outcome studies are observational rather than randomized, risk adjustment is necessary to account for case mix. This has usually been accomplished through the use of standard logistic regression models, although Bayesian models, hierarchical linear models, and machine-learning techniques such as neural networks have also been used. Many factors are essential to insuring the accuracy and usefulness of such models, including selection of an appropriate clinical database, inclusion of critical core variables, precise definitions for predictor variables and endpoints, proper model development, validation, and audit. Risk models may be used to assess the impact of specific predictors on outcome, to aid in patient counseling and treatment selection, to profile provider quality, and to serve as the basis of continuous quality improvement activities.

Bayes Theorem↗

Functional interaction of nitrogenous organic bases with cytochrome P450: a critical assessment and update of substrate features and predicted key active-site elements steering the access, binding, and orientation of amines.

The widespread use of nitrogenous organic bases as environmental chemicals, food additives, and clinically important drugs necessitates precise knowledge about the molecular principles governing biotransformation of this category of substrates. In this regard, analysis of the topological background of complex formation between amines and P450s, acting as major catalysts in C- and N-oxidative attack, is of paramount importance. Thus, progress in collaborative investigations, combining physico-chemical techniques with chemical-modification as well as genetic engineering experiments, enables substantiation of hypothetical work resulting from the design of pharmacophores or homology modelling of P450s. Based on a general, CYP2D6-related construct, the majority of prospective amine-docking residues was found to cluster near the distal heme face in the six known SRSs, made up by the highly variant helices B', F and G as well as the N-terminal portion of helix C and certain beta-structures. Most of the contact sites examined show a frequency of conservation < 20%, hinting at the requirement of some degree of conformational versatility, while a limited number of amino acids exhibiting a higher level of conservation reside close to the heme core. Some key determinants may have a dual role in amine binding and/or maintenance of protein integrity. Importantly, a series of non-SRS elements are likely to be operative via long-range effects. While hydrophobic mechanisms appear to dominate orientation of the nitrogenous compounds toward the iron-oxene species, polar residues seem to foster binding events through H-bonding or salt-bridge formation. Careful uncovering of structure-function relationships in amine-enzyme association together with recently developed unsupervised machine learning approaches will be helpful in both tailoring of novel amine-type drugs and early elimination of potentially toxic or mutagenic candidates. Also, chimeragenesis might serve in the construction of more efficient P450s for activation of amine drugs and/or bioremediation.

Amines↗

Evolving beyond perfection: an investigation of the effects of long-term evolution on fractal gene regulatory networks.

This paper continues a theme of exploring algorithms based on principles of biological development for tasks such as pattern generation, machine learning and robot control. Previous work has investigated the use of genes expressed as fractal proteins to enable greater evolvability of gene regulatory networks (GRNs). Here, the evolution of such GRNs is investigated further to determine whether evolution exhibits natural tendencies towards efficiency and graceful degradation of developmental programs. Experiments where "perfect" GRNs are evolved for a further thousand generations without the addition of any further selection pressure, confirm this hypothesis. After further evolution, the perfect GRNs operate in a more efficient manner (using fewer proteins) and show an improved ability to function correctly with missing genes. When the algorithm is applied to applications (e.g. robot control) this equates to efficient and fault-tolerant controllers.

Algorithms↗

Generative topographic mapping applied to clustering and visualization of motor unit action potentials.

The identification and visualization of clusters formed by motor unit action potentials (MUAPs) is an essential step in investigations seeking to explain the control of the neuromuscular system. This work introduces the generative topographic mapping (GTM), a novel machine learning tool, for clustering of MUAPs, and also it extends the GTM technique to provide a way of visualizing MUAPs. The performance of GTM was compared to that of three other clustering methods: the self-organizing map (SOM), a Gaussian mixture model (GMM), and the neural-gas network (NGN). The results, based on the study of experimental MUAPs, showed that the rate of success of both GTM and SOM outperformed that of GMM and NGN, and also that GTM may in practice be used as a principled alternative to the SOM in the study of MUAPs. A visualization tool, which we called GTM grid, was devised for visualization of MUAPs lying in a high-dimensional space. The visualization provided by the GTM grid was compared to that obtained from principal component analysis (PCA).

Action Potentials↗

QSAR study of 1,4-dihydropyridine calcium channel antagonists based on gene expression programming.

The gene expression programming, a novel machine learning algorithm, is used to develop quantitative model as a potential screening mechanism for a series of 1,4-dihydropyridine calcium channel antagonists for the first time. The heuristic method was used to search the descriptor space and select the descriptors responsible for activity. A nonlinear, six-descriptor model based on gene expression programming with mean-square errors 0.19 was set up with a predicted correlation coefficient (R2) 0.92. This paper provides a new and effective method for drug design and screening.

Algorithms↗

Proteome-wide structural and interaction analysis using cross-linking mass spectrometry and its applications.

Deciphering the mechanisms of protein-protein interactions (PPIs) and protein structural changes within the native cellular environment is crucial for advancing drug discovery. In vivo chemical cross-linking coupled with mass spectrometry (XL-MS) captures weak, transient, and higher-order interactions that are often dysregulated under altered physiological conditions and remain challenging to detect using conventional methods. Applications of in vivo XL-MS range from targeted mapping of PPIs to large-scale identification of interactome networks within the cells. The integration of quantitative approaches further facilitates comparison across different physiological conditions. The recent incorporation of machine learning (ML) tools into XL-MS workflows is transforming the depth and efficiency of this technology. AI-driven algorithms now enable more accurate identification of cross-linked peptides and the mapping of interaction topologies. Furthermore, the synergistic coupling of in vivo XL-MS data with AI-assisted structural modeling platforms such as AlphaFold allows dynamic and high-throughput prediction of protein networks. This review discusses the broader applications of in vivo XL-MS in complex biological samples, ranging from organelles and cells to whole tissues, and highlights how AI integration is expanding structural biology toward a systems-level understanding of proteome architecture.

Mass Spectrometry↗

Searching for functional sites in protein structures.

An ability to assign protein function from protein structure is important for structural genomics consortia. The complex relationship between protein fold and function highlights the necessity of looking beyond the global fold of a protein to specific functional sites. Many computational methods have been developed that address this issue. These include evolutionary trace methods, methods that involve the calculation and assessment of maximal superpositions, methods based on graph theory, and methods that apply machine learning techniques. Such function prediction techniques have been applied to the identification of enzyme catalytic triads and DNA-binding motifs.

Binding Sites↗

Global survey of organ and organelle protein expression in mouse: combined proteomic and transcriptomic profiling.

Organs and organelles represent core biological systems in mammals, but the diversity in protein composition remains unclear. Here, we combine subcellular fractionation with exhaustive tandem mass spectrometry-based shotgun sequencing to examine the protein content of four major organellar compartments (cytosol, membranes [microsomes], mitochondria, and nuclei) in six organs (brain, heart, kidney, liver, lung, and placenta) of the laboratory mouse, Mus musculus. Using rigorous statistical filtering and machine-learning methods, the subcellular localization of 3274 of the 4768 proteins identified was determined with high confidence, including 1503 previously uncharacterized factors, while tissue selectivity was evaluated by comparison to previously reported mRNA expression patterns. This molecular compendium, fully accessible via a searchable web-browser interface, serves as a reliable reference of the expressed tissue and organelle proteomes of a leading model mammal.

Animals↗

Plasma signals of lung tumor promotion for molecular cancer prevention.

Predicting lung cancer risk would enhance prevention trials. Although the Canakinumab Anti-inflammatory Thrombosis Outcome Study (CANTOS) trial demonstrated reduced lung cancer incidence with interleukin (IL)-1&#x3b2; inhibition, the high number needed to treat (NNT) to prevent lung cancer limits its use in unselected populations. Using machine learning, we identified a 14-protein plasma signature predicting lung cancer more than 5 years before diagnosis. The signature, validated across eight cohorts, was elevated in current smokers and individuals exposed to particulate matter (PM) and linked to lung myeloid and alveolar cells. In epidermal growth factor receptor (EGFR)-driven lung adenocarcinoma, diverse epithelial lineages converged on a keratin8+/claudin4+ alveolar transitional state (KAC), whose transcriptional programs correlated with signature emergence. Components of the signature were induced by PM, oncogenic EGFR, or IL-1&#x3b2;, whereas IL-1&#x3b2; inhibition restrained PM-driven KAC expansion and early tumorigenesis. In CANTOS, the signature identified individuals who seemed to benefit more from anti-IL-1&#x3b2; therapy, lowering the NNT threshold and nominating circulating signals of tumor promotion for prevention.

Humans↗

Detection of antibiotic heteroresistance in clinical microbiology: current and emerging methodologies.

BACKGROUND: Antibiotic heteroresistance (HR) is characterised by the coexistence of susceptible and resistant subpopulations within an apparently isogenic bacterial isolate. Because routine antimicrobial susceptibility testing (AST) primarily assesses the dominant population, HR may escape detection, potentially leading to discrepancies between laboratory susceptibility categorisation and the underlying bacterial population structure. OBJECTIVES: To provide a critical and practice-oriented evaluation of current and emerging methodologies for HR detection and to discuss their strengths, limitations, and potential for clinical implementation. SOURCES: Narrative review based on PubMed searches, complemented by screening of key reference lists and relevant EUCAST and CLSI documents. Peer-reviewed literature was prioritised. CONTENT: Phenotypic approaches, particularly population analysis profiling, remain the reference method for HR definition, but their labour-intensive workflows, long turnaround times, and limited standardisation restrict routine implementation. Alternative strategies, including modified AST assays, metabolic assays, and single-cell platforms, offer gains in speed or throughput but require broader validation. Molecular approaches such as quantitative PCR, droplet digital PCR, targeted deep sequencing, and whole-genome sequencing improve detection of minority resistance determinants. Emerging computational frameworks, including machine learning models integrating phenotypic and genomic data, represent a promising frontier for scalable HR prediction. IMPLICATIONS: Available evidence supports the clinical relevance of HR, although its association with adverse outcomes varies across bacterial species and antibiotic classes. Harmonised methodologies and clinically validated interpretive criteria are needed to support integration of HR assessment into routine diagnostics. Prospective multicentre studies and further standardisation, including engagement with EUCAST and CLSI, will be important to advance clinical implementation.

Antimicrobial resistance↗