Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Simulating soft data to make soft data applicable to simulation.

BACKGROUND: Biomedical processes are often influenced by measures considered "non-crisp", "soft" or "subjective". Despite the growing awareness of the importance of such measures, they are rarely considered in biomedical simulation. This study introduces an input generator for soft data (input generator SD) that makes soft data applicable to simulation. MATERIALS AND METHODS: Machine learning approaches and standard regression techniques were applied to simulate odour intensity ratings. RESULTS: The performance of all the applied methods was satisfactory and the results can be used to modify systems biological mathematical models. CONCLUSION: Soft data should no longer be discounted in systems biological simulations. Exemplarily, it can be demonstrated that the input generator SD produces results that are similar to those that the simulated system can generate. Machine learning and/or appropriate conventional mathematical approaches may be applied to simulate noncrisp processes that can be used to modify mathematical models of any granularity.

Adult↗

Firemaster 550 differentially alters gene expression underlying synaptic function in amygdala of prairie voles after gestational or lactational exposure.

Neurodevelopmental disorders often share similar behavioral diagnostic criteria including socioemotional and cognitive deficits. The prairie vole is a uniquely suitable model to study these deficits because they demonstrate strong social affiliation, bi-parental care, and partner attachment. Previously, we have shown that developmental exposure to the flame-retardant mixture Firemaster 550 (FM 550) impairs socioemotional behavior in the prairie vole and alters underlying neuroanatomy and function. However, the mechanisms for impaired pair bonding in males and increased anxiety in females remain unknown, along with the specific critical window(s) of vulnerability. Herein, we exposed prairie vole dams to FM 550 during gestation or lactation, and performed bulk RNA-seq on the amygdala, a hub of socioemotional processing, in their adult offspring. Two mathematically orthogonal methods were utilized for analysis, a linear statistical method and an ensemble machine learning method, incorporating sex as a biological variable. Gene ontology (GO) pathway analysis was performed following both and results compared to identify potential mechanisms of toxicity. GO results indicated consistent expression changes in the Synapse cellular component in all conditions, and implicated glutamatergic signaling specifically. Additionally, gestational exposure (GE) altered genes underlying modulation of synaptic transmission and neural development, while lactational exposure (LE) impacted genes underlying synaptic plasticity, axon guidance, and mitophagy. Machine learning identified disruption of endocrine system development, regulation of biosynthetic processes in GE animals, and suppression of various neuroinflammatory genes across multiple groups. Finally, we performed RNA expression analysis using Nanostring and demonstrated stronger correlation with the differentially expressed genes (DEG) of interest in females than males. Overall, this study demonstrates both the intersecting and distinct impacts of FM 550 exposure on amygdalar gene expression depending on sex and timing of exposure.

Animals↗

A neural-network based method for prediction of gamma-turns in proteins from multiple sequence alignment.

In the present study, an attempt has been made to develop a method for predicting gamma-turns in proteins. First, we have implemented the commonly used statistical and machine-learning techniques in the field of protein structure prediction, for the prediction of gamma-turns. All the methods have been trained and tested on a set of 320 nonhomologous protein chains by a fivefold cross-validation technique. It has been observed that the performance of all methods is very poor, having a Matthew's Correlation Coefficient (MCC) </= 0.06. Second, predicted secondary structure obtained from PSIPRED is used in gamma-turn prediction. It has been found that machine-learning methods outperform statistical methods and achieve an MCC of 0.11 when secondary structure information is used. The performance of gamma-turn prediction is further improved when multiple sequence alignment is used as the input instead of a single sequence. Based on this study, we have developed a method, GammaPred, for gamma-turn prediction (MCC = 0.17). The GammaPred is a neural-network-based method, which predicts gamma-turns in two steps. In the first step, a sequence-to-structure network is used to predict the gamma-turns from multiple alignment of protein sequence. In the second step, it uses a structure-to-structure network in which input consists of predicted gamma-turns obtained from the first step and predicted secondary structure obtained from PSIPRED.

Databases, Protein↗

Comparison of genetic algorithms and other classification methods in the diagnosis of female urinary incontinence.

Galactica, a newly developed machine-learning system that utilizes a genetic algorithm for learning, was compared with discriminant analysis, logistic regression, k-means cluster analysis, a C4.5 decision-tree generator and a random bit climber hill-climbing algorithm. The methods were evaluated in the diagnosis of female urinary incontinence in terms of prediction accuracy of classifiers, on the basis of patient data. The best methods were discriminant analysis, logistic regression, C4.5 and Galactica. Practically no statistically significant differences existed between the prediction accuracy of these classification methods. We consider that machine-learning systems C4.5 and Galactica are preferable for automatic construction of medical decision aids, because they can cope with missing data values directly and can present a classifier in a comprehensible form. Galactica performed nearly as well as C4.5. The results are in agreement with the results of earlier research, indicating that genetic algorithms are a competitive method for constructing classifiers from medical data.

Algorithms↗

Integrated multi-omics analyses identify an RAS-SLC11A2-associated molecular framework linking iron metabolism with PCOS-related cardiometabolic risk.

INTRODUCTION: PCOS is a common endocrine disorder with elevated cardiometabolic risk, yet the role of the renin-angiotensin system (RAS)-iron metabolism axis in this comorbidity remains unclear. We explored its underlying mechanisms and evaluated the therapeutic potential of gentiopicroside. METHODS: Integrated multi-omics analyses combining transcriptomics, single-cell RNA sequencing, Mendelian randomization, machine learning, molecular docking, and in vitro functional assays were performed to identify shared molecular pathways and therapeutic targets across PCOS, hypertension, NAFLD, and T2DM. RESULTS: SLC11A2 was consistently dysregulated in PCOS transcriptomic datasets, and associated with iron metabolism, inflammatory response and oxidative stress pathways. Genetic analyses validated RAS-related regulation in hypertension susceptibility and revealed shared genetic architecture between PCOS and cardiometabolic traits. Network and single-cell analyses characterized SLC11A2-associated molecular patterns in disease-relevant cell types; machine learning identified disease-classifying molecular signatures. Gentiopicroside alleviated inflammatory and oxidative stress phenotypes, including reduced IL-6 expression and reactive oxygen species accumulation. CONCLUSION: This study defines an RAS-SLC11A2 molecular framework linking iron metabolism dysregulation to PCOS-related cardiometabolic risk, elucidating the mechanisms connecting ovarian dysfunction, inflammation, oxidative stress and hypertension, and supports gentiopicroside as a promising therapeutic candidate.

Humans↗

Evaluation of automatic knowledge acquisition techniques in the diagnosis of acute abdominal pain. Acute Abdominal Pain Study Group.

Clinical diagnosis in acute abdominal pain is still a major problem. Computer-aided diagnosis offers some help; however, existing systems still produce high error rates. We therefore tested machine learning techniques in order to improve standard statistical systems. The investigation was based on a prospective clinical database with 1254 cases, 46 diagnostic parameters and 15 diagnoses. Independence Bayes and the automatic rule induction techniques ID3, NewId, PRISM, CN2, C4.5 and ITRULE were trained with 839 cases and separately tested on 415 cases. No major differences in overall accuracy were observed (43-48%), except for NewId, which was below the average. Between the different techniques some similarities were found, but also considerable differences with respect to specific diagnoses. Machine learning techniques did not improve the results of the standard model Independence Bayes. Problem dimensionality, sample size and model complexity are major factors influencing diagnostic accuracy in computer-aided diagnosis of acute abdominal pain.

Abdominal Pain↗

Comparing syntactic complexity in medical and non-medical corpora.

With the growing use of Natural Language Processing (NLP) techniques as solutions in Medical Informatics, the need to quickly and efficiently create the knowledge structures used by these systems has grown concurrently. Automatic discovery of a lexicon for use by an NLP system through machine learning will require information about the syntax of medical language. Understanding the syntactic differences between medical and non-medical corpora may allow more efficient acquisition of a lexicon. Three experiments designed to quantify the syntactic differences in medical and non-medical corpora were conducted. The results show that the syntax of medical language shows less variation than non-medical language and is likely simpler. The differences were great enough to question the applicability of general language tools on medical language. These differences may reduce the difficulty of some free text machine learning problems by capitalizing on the simpler nature of narrative medical syntax.

Artificial Intelligence↗

A regulatory network underlying idiopathic pulmonary fibrosis.

BACKGROUND: Idiopathic pulmonary fibrosis (IPF) is a progressive interstitial lung disease in which genetic susceptibility interacts with epithelial, immune, and mesenchymal remodeling. Although the chromosome 11p15.5 locus contains established IPF susceptibility signals near MUC5B and TOLLIP, the broader regulatory architecture of this region remains incompletely resolved. METHODS: We integrated IPF genome-wide association study summary statistics with methylation, expression, and protein quantitative trait loci using summary-data-based Mendelian randomization (SMR). SMR-prioritized candidates were evaluated in independent transcriptomic and methylation cohorts and further contextualized using microRNA, transcription-factor, protein-interaction, machine-learning, single-cell, and spatial transcriptomic analyses. Fibrosis-associated expression patterns were assessed in a bleomycin-induced pulmonary fibrosis rat model. RESULTS: The analyses recovered the established MUC5B and TOLLIP signals and prioritized BRSK2 as a comparatively underexplored candidate supported by eQTL-based SMR and independent molecular evidence. The BRSK2 pQTL association did not pass the HEIDI test and was therefore not interpreted as convergent protein-level genetic evidence. Network analyses linked BRSK2 to cell-cycle, metabolic-stress, and senescence-related programs, while cross-cohort machine learning prioritized FOXA2, CDC25B, and NFE2 as informative network features. Single-cell and spatial analyses localized BRSK2 preferentially to fibroblast and myofibroblast compartments and to regions with greater histological fibrosis severity. In fibrotic rat lungs, BRSK2 expression increased, whereas FOXA2 and CDC25B decreased at the transcript and protein levels. CONCLUSIONS: These findings refine the molecular landscape of the chromosome 11p15.5 IPF susceptibility locus and prioritize BRSK2 as a candidate component of an IPF-associated profibrotic fibroblast state. Its causal contribution, direct regulatory relationships, and therapeutic tractability require targeted mechanistic validation.

Idiopathic Pulmonary Fibrosis↗

Optimization of neural network architecture using genetic programming improves detection and modeling of gene-gene interactions in studies of human diseases.

BACKGROUND: Appropriate definition of neural network architecture prior to data analysis is crucial for successful data mining. This can be challenging when the underlying model of the data is unknown. The goal of this study was to determine whether optimizing neural network architecture using genetic programming as a machine learning strategy would improve the ability of neural networks to model and detect nonlinear interactions among genes in studies of common human diseases. RESULTS: Using simulated data, we show that a genetic programming optimized neural network approach is able to model gene-gene interactions as well as a traditional back propagation neural network. Furthermore, the genetic programming optimized neural network is better than the traditional back propagation neural network approach in terms of predictive ability and power to detect gene-gene interactions when non-functional polymorphisms are present. CONCLUSION: This study suggests that a machine learning strategy for optimizing neural network architecture may be preferable to traditional trial-and-error approaches for the identification and characterization of gene-gene interactions in common, complex human diseases.

Algorithms↗

Bayesian analysis, pattern analysis, and data mining in health care.

PURPOSE OF REVIEW: To discuss the current role of data mining and Bayesian methods in biomedicine and heath care, in particular critical care. RECENT FINDINGS: Bayesian networks and other probabilistic graphical models are beginning to emerge as methods for discovering patterns in biomedical data and also as a basis for the representation of the uncertainties underlying clinical decision-making. At the same time, techniques from machine learning are being used to solve biomedical and health-care problems. SUMMARY: With the increasing availability of biomedical and health-care data with a wide range of characteristics there is an increasing need to use methods which allow modeling the uncertainties that come with the problem, are capable of dealing with missing data, allow integrating data from various sources, explicitly indicate statistical dependence and independence, and allow integrating biomedical and clinical background knowledge. These requirements have given rise to an influx of new methods into the field of data analysis in health care, in particular from the fields of machine learning and probabilistic graphical models.

Bayes Theorem↗

Computational identification of residues that modulate voltage sensitivity of voltage-gated potassium channels.

BACKGROUND: Studies of the structure-function relationship in proteins for which no 3D structure is available are often based on inspection of multiple sequence alignments. Many functionally important residues of proteins can be identified because they are conserved during evolution. However, residues that vary can also be critically important if their variation is responsible for diversity of protein function and improved phenotypes. If too few sequences are studied, the support for hypotheses on the role of a given residue will be weak, but analysis of large multiple alignments is too complex for simple inspection. When a large body of sequence and functional data are available for a protein family, mature data mining tools, such as machine learning, can be applied to extract information more easily, sensitively and reliably. We have undertaken such an analysis of voltage-gated potassium channels, a transmembrane protein family whose members play indispensable roles in electrically excitable cells. RESULTS: We applied different learning algorithms, combined in various implementations, to obtain a model that predicts the half activation voltage of a voltage-gated potassium channel based on its amino acid sequence. The best result was obtained with a k-nearest neighbor classifier combined with a wrapper algorithm for feature selection, producing a mean absolute error of prediction of 7.0 mV. The predictor was validated by permutation test and evaluation of independent experimental data. Feature selection identified a number of residues that are predicted to be involved in the voltage sensitive conformation changes; these residues are good target candidates for mutagenesis analysis. CONCLUSION: Machine learning analysis can identify new testable hypotheses about the structure/function relationship in the voltage-gated potassium channel family. This approach should be applicable to any protein family if the number of training examples and the sequence diversity of the training set that are necessary for robust prediction are empirically validated. The predictor and datasets can be found at the VKCDB web site.

Algorithms↗

Predictive models for protein crystallization.

Crystallization of proteins is a nontrivial task, and despite the substantial efforts in robotic automation, crystallization screening is still largely based on trial-and-error sampling of a limited subset of suitable reagents and experimental parameters. Funding of high throughput crystallography pilot projects through the NIH Protein Structure Initiative provides the opportunity to collect crystallization data in a comprehensive and statistically valid form. Data mining and machine learning algorithms thus have the potential to deliver predictive models for protein crystallization. However, the underlying complex physical reality of crystallization, combined with a generally ill-defined and sparsely populated sampling space, and inconsistent scoring and annotation make the development of predictive models non-trivial. We discuss the conceptual problems, and review strengths and limitations of current approaches towards crystallization prediction, emphasizing the importance of comprehensive and valid sampling protocols. In view of limited overlap in techniques and sampling parameters between the publicly funded high throughput crystallography initiatives, exchange of information and standardization should be encouraged, aiming to effectively integrate data mining and machine learning efforts into a comprehensive predictive framework for protein crystallization. Similar experimental design and knowledge discovery strategies should be applied to valid analysis and prediction of protein expression, solubilization, and purification, as well as crystal handling and cryo-protection.

Bayes Theorem↗

Computational intelligence for the detection and classification of malignant lesions in screening mammography.

This report deals with the discussion of the findings obtained from the application of two computational intelligence methodologies for the detection of microcalcifications in screening mammography data. Genetic programming and inductive machine learning have been applied, in order to produce meaningful diagnostic rules for the medical staff. The data used in the experiments correspond to information acquired from two images of each breast of the patient, along with some associated patient information such as the age at time of study. Similar datasets have been previously used in an attempt to facilitate the development of computer algorithms to aid screening. Experienced screening radiologists have double-read the screening mammograms, they have weighted the malignancy ratings and averaged out the levels of suspiciousness assigned to each finding in the screenings. The diagnostic rules which were obtained from both genetic programming and machine learning have been evaluated in detail and then analyzed and discussed by collaborative medical experts, in parallel to findings from related literature. Results seem encouraging for further use and analysis by medical staff specializing in screening mammography.

Artificial Intelligence↗

Computational prediction of the chromosome-damaging potential of chemicals.

We report on the generation of computer-based models for the prediction of the chromosome-damaging potential of chemicals as assessed in the in vitro chromosome aberration (CA) test. On the basis of publicly available CA-test results of more than 650 chemical substances, half of which are drug-like compounds, we generated two different computational models. The first model was realized using the (Q)SAR tool MCASE. Results obtained with this model indicate a limited performance (53%) for the assessment of a chromosome-damaging potential (sensitivity), whereas CA-test negative compounds were correctly predicted with a specificity of 75%. The low sensitivity of this model might be explained by the fact that the underlying 2D-structural descriptors only describe part of the molecular mechanism leading to the induction of chromosome aberrations, that is, direct drug-DNA interactions. The second model was constructed with a more sophisticated machine learning approach and generated a classification model based on 14 molecular descriptors, which were obtained after feature selection. The performance of this model was superior to the MCASE model, primarily because of an improved sensitivity, suggesting that the more complex molecular descriptors in combination with statistical learning approaches are better suited to model the complex nature of mechanisms leading to a positive effect in the CA-test. An analysis of misclassified pharmaceuticals by this model showed that a large part of the false-negative predicted compounds were uniquely positive in the CA-test but lacked a genotoxic potential in other mutagenicity tests of the regulatory testing battery, suggesting that biologically nonsignificant mechanisms could be responsible for the observed positive CA-test result. Since such mechanisms are not amenable to modeling approaches it is suggested that a positive prediction made by the model reflects a biologically significant genotoxic potential. An integration of the machine-learning model as a screening tool in early discovery phases of drug development is proposed.

Chromosomes↗

Detection and analysis of statistical differences in anatomical shape.

We present a computational framework for image-based analysis and interpretation of statistical differences in anatomical shape between populations. Applications of such analysis include understanding developmental and anatomical aspects of disorders when comparing patients versus normal controls, studying morphological changes caused by aging, or even differences in normal anatomy, for example, differences between genders. Once a quantitative description of organ shape is extracted from input images, the problem of identifying differences between the two groups can be reduced to one of the classical questions in machine learning of constructing a classifier function for assigning new examples to one of the two groups while making as few misclassifications as possible. The resulting classifier must be interpreted in terms of shape differences between the two groups back in the image domain. We demonstrate a novel approach to such interpretation that allows us to argue about the identified shape differences in anatomically meaningful terms of organ deformation. Given a classifier function in the feature space, we derive a deformation that corresponds to the differences between the two classes while ignoring shape variability within each class. Based on this approach, we present a system for statistical shape analysis using distance transforms for shape representation and the support vector machines learning algorithm for the optimal classifier estimation and demonstrate it on artificially generated data sets, as well as real medical studies.

Algorithms↗

Molecular hashkeys: a novel method for molecular characterization and its application for predicting important pharmaceutical properties of molecules.

We define a novel numerical molecular representation, called the molecular hashkey, that captures sufficient information about a molecule to predict pharmaceutically interesting properties directly from three-dimensional molecular structure. The molecular hashkey represents molecular surface properties as a linear array of pairwise surface-based comparisons of the target molecule against a common 'basis-set' of molecules. Hashkey-measured molecular similarity correlates well with direct methods of measuring molecular surface similarity. Using a simple machine-learning technique with the molecular hashkeys, we show that it is possible to accurately predict the octanol-water partition coefficient, log P. Using more sophisticated learning techniques, we show that an accurate model of intestinal absorption for a set of drugs can be constructed using the same hashkeys used in the aforementioned experiments. Once a set of molecular hashkeys is calculated, its use in the training and testing of property-based models is very fast. Further, the required amount of data for model construction is very small. Neural network-based hashkey models trained on data sets as small as 30 molecules yield statistically significant prediction of molecular properties. The lack of a requirement for large data sets lends itself well to the prediction of pharmaceutically relevant molecular parameters for which data generation is expensive and slow. Molecular hashkeys coupled with machine-learning techniques can yield models that predict key pharmacological aspects of biologically important molecules and should therefore be important in the design of effective therapeutics.

Drug Design↗

Predicting dire outcomes of patients with community acquired pneumonia.

Community-acquired pneumonia (CAP) is an important clinical condition with regard to patient mortality, patient morbidity, and healthcare resource utilization. The assessment of the likely clinical course of a CAP patient can significantly influence decision making about whether to treat the patient as an inpatient or as an outpatient. That decision can in turn influence resource utilization, as well as patient well being. Predicting dire outcomes, such as mortality or severe clinical complications, is a particularly important component in assessing the clinical course of patients. We used a training set of 1601 CAP patient cases to construct 11 statistical and machine-learning models that predict dire outcomes. We evaluated the resulting models on 686 additional CAP-patient cases. The primary goal was not to compare these learning algorithms as a study end point; rather, it was to develop the best model possible to predict dire outcomes. A special version of an artificial neural network (NN) model predicted dire outcomes the best. Using the 686 test cases, we estimated the expected healthcare quality and cost impact of applying the NN model in practice. The particular, quantitative results of this analysis are based on a number of assumptions that we make explicit; they will require further study and validation. Nonetheless, the general implication of the analysis seems robust, namely, that even small improvements in predictive performance for prevalent and costly diseases, such as CAP, are likely to result in significant improvements in the quality and efficiency of healthcare delivery. Therefore, seeking models with the highest possible level of predictive performance is important. Consequently, seeking ever better machine-learning and statistical modeling methods is of great practical significance.

Community-Acquired Infections↗

Genetically optimized fuzzy decision trees.

In this study, we are concerned with genetically optimized fuzzy decision trees (G-DTs). Decision trees are fundamental architectures of machine learning, pattern recognition, and system modeling. Starting with the generic decision tree with discrete or interval-valued attributes, we develop its fuzzy set-based generalization. In this generalized structure we admit the values of the attributes that are represented by some membership functions. Such fuzzy decision trees are constructed in the setting of genetic optimization. The underlying genetic algorithm optimizes the parameters of the fuzzy sets associated with the individual nodes where they play a role of fuzzy "switches" by distributing a flow of processing completed within the tree. We discuss various forms of the fitness function that help capture the essence of the problem at hand (that could be either of classification nature when dealing with discrete outputs or regression-like when handling a continuous output variable). We quantify a nature of the generalization of the tree by studying an optimally adjusted spreads of the membership functions located at the nodes of the decision tree. A series of experiments exploiting synthetic and machine learning data is used to illustrate the performance of the G-DTs.

Algorithms↗