Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble learning”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

In silico prediction method for plant Nucleotide-binding leucine-rich repeat- and pathogen effector interactions.

Plant Nucleotide-binding leucine-rich repeat (NLR) proteins play a crucial role in effector recognition and activation of Effector triggered immunity following pathogen infection. Genome sequencing advancements have led to the identification of a myriad of NLRs in numerous agriculturally important plant species. However, deciphering which NLRs recognize specific pathogen effectors remains challenging. Predicting NLR-effector interactions in silico will provide a more targeted approach for experimental validation, critical for elucidating function, and advancing our understanding of NLR-triggered immunity. In this study, NLR-effector protein complex structures were predicted using AlphaFold2-Multimer for all experimentally validated NLR-effector interactions reported in literature. Binding affinities- and energies were predicted using 97 machine learning models from Area-Affinity. We show that AlphaFold2-Multimer predicted structures have acceptable accuracy and can be used to investigate NLR-effector interactions in silico. Binding affinities for 58 NLR-effector complexes ranged between -8.5 and -10.6 log(K), and binding energies between -11.8 and -14.4 kcal/mol-1, depending on the Area-Affinity model used. For 2427 "forced" NLR-effector complexes, these estimates showed larger variability, enabling identification of novel NLR-effector interactions with 99% accuracy using an Ensemble machine learning model. The narrow range of binding energies- and affinities for "true" interactions suggest a specific change in Gibbs free energy, and thus conformational change, is required for NLR activation. This is the first study to provide a method for predicting NLR-effector interactions, applicable to all pathosystems. Finally, the NLR-Effector Interaction Classification (NEIC) resource can streamline research efforts by identifying NLRs important for plant-pathogen resistance, advancing our understanding of plant immunity.

Plant Proteins↗

Noninvasive detection and differentiation of gastric malignancy using cell-free DNA biomarkers.

INTRODUCTION: Gastric cancer remains a major global health burden, with high mortality driven by late-stage diagnoses that limit treatment options and reduce survival. Current diagnostic methods such as endoscopy and biopsy are invasive, resource-intensive, and impractical for large-scale early detection. OBJECTIVES: This study aimed to develop and validate an ensemble machine learning model integrating four cell-free DNA (cfDNA) fragmentomic feature classes derived from 5 × whole genome sequencing (WGS) data to non-invasively differentiate malignant gastric cancer from benign gastric lesions in high-risk or symptomatic patients. METHODS: A total of 681 plasma samples were prospectively collected, comprising 329 from patients with gastric cancer or high-grade intraepithelial neoplasia (HGIN) and 352 from individuals with benign gastric conditions. The dataset was divided into a training cohort (n = 333) and a temporally independent validation cohort (n = 348). An external validation cohort of 305 participants was also included. RESULTS: The ensemble model achieved an AUROC of 0.920 in cross-validation testing on the training cohort, 0.912 in the independent validation cohort, and 0.896 (95% CI 0.860-0.932) in the external cohort. At a pre-specified prediction threshold of 0.402, the model demonstrated 93.3% sensitivity and 71.9% specificity in the validation cohort, yielding a PPV of 71.3% and an NPV of 93.5%. In the external cohort, sensitivity and specificity were 91.7% and 69.1%, respectively (PPV 75.7%, NPV 88.8%). Model scores correlated with clinical stage, tumor grade, and histopathological subtype. Approximately 71% of non-cancer patients could have been spared unnecessary endoscopy. CONCLUSIONS: The cfDNA fragmentomics-based ensemble model enables accurate, non-invasive differentiation between gastric cancer and benign gastric lesions in high-risk or symptomatic patients. This approach demonstrates strong potential as a pre-endoscopy triage tool, supporting earlier detection and more efficient use of diagnostic resources.

Humans↗

Metagenomic analyses reveal E. coli-derived siderophores as potential signatures for breast cancer.

BACKGROUND: Breast cancer remains a leading cause of cancer-related mortality in women. Recent evidence implicates the gut microbiome and metabolites in breast cancer pathogenesis. This study explores associations between gut microbial species, their predicted metabolites, and breast cancer to uncover potential mechanistic insights. METHODS: Comprehensive metagenomic analyses were conducted on the gut microbiome of pre- and postmenopausal breast cancer patients, where microbial species were profiled through AMPHORA2 and metabolites were predicted through antiSMASH. Multivariate association analysis was used to identify significant associations between specific microbial species, predicted metabolites, and breast cancer status. A custom ensemble machine learning classifier was developed to classify pre- and postmenopausal breast cancer cases and controls based on microbial and predicted metabolite features. Additionally, a synthetic microbiome dataset was generated through MIDASim to validate the reproducibility of the ML results. Using our results, we explored the underlying dynamics of identified taxa and metabolite in breast cancer through literature and statistical support. RESULTS: Our analysis identified 471 microbial species and predicted 40 key metabolites in the metagenomic data. Multivariate analysis identified significant positive associations (p-value&#x2009;<&#x2009;0.05) of E. coli, siderophore, and thiopeptide with breast cancer. The custom ensemble model achieved accuracy and AUC as high as 78% and 90%, respectively, in classifying pre- and postmenopausal cases and controls. The high-ranking features i.e., E. coli, siderophore, and thiopeptide were consistent with the results of the multivariate association analysis, thereby substantiating their biological significance. Using these findings, we propose a mechanistic model in which E. coli secretes siderophores under iron-limited conditions in breast cancer patients, for iron sequestration from the host, which can potentially promote angiogenesis and tumor progression. CONCLUSION: Our findings suggest that microbial iron acquisition mechanisms may play a critical role in breast cancer pathophysiology. Functional validation of these mechanisms is needed to assess therapeutic potential. This study highlights gut microbiota and their metabolites as promising targets for breast cancer research and intervention.

Breast Neoplasms↗

Firemaster 550 differentially alters gene expression underlying synaptic function in amygdala of prairie voles after gestational or lactational exposure.

Neurodevelopmental disorders often share similar behavioral diagnostic criteria including socioemotional and cognitive deficits. The prairie vole is a uniquely suitable model to study these deficits because they demonstrate strong social affiliation, bi-parental care, and partner attachment. Previously, we have shown that developmental exposure to the flame-retardant mixture Firemaster 550 (FM 550) impairs socioemotional behavior in the prairie vole and alters underlying neuroanatomy and function. However, the mechanisms for impaired pair bonding in males and increased anxiety in females remain unknown, along with the specific critical window(s) of vulnerability. Herein, we exposed prairie vole dams to FM 550 during gestation or lactation, and performed bulk RNA-seq on the amygdala, a hub of socioemotional processing, in their adult offspring. Two mathematically orthogonal methods were utilized for analysis, a linear statistical method and an ensemble machine learning method, incorporating sex as a biological variable. Gene ontology (GO) pathway analysis was performed following both and results compared to identify potential mechanisms of toxicity. GO results indicated consistent expression changes in the Synapse cellular component in all conditions, and implicated glutamatergic signaling specifically. Additionally, gestational exposure (GE) altered genes underlying modulation of synaptic transmission and neural development, while lactational exposure (LE) impacted genes underlying synaptic plasticity, axon guidance, and mitophagy. Machine learning identified disruption of endocrine system development, regulation of biosynthetic processes in GE animals, and suppression of various neuroinflammatory genes across multiple groups. Finally, we performed RNA expression analysis using Nanostring and demonstrated stronger correlation with the differentially expressed genes (DEG) of interest in females than males. Overall, this study demonstrates both the intersecting and distinct impacts of FM 550 exposure on amygdalar gene expression depending on sex and timing of exposure.

Animals↗

Allostatic load is associated with symptoms in chronic fatigue syndrome patients.

OBJECTIVES: To further explore the relationship between chronic fatigue syndrome (CFS) and allostatic load (AL), we conducted a computational analysis involving 43 patients with CFS and 60 nonfatigued, healthy controls (NF) enrolled in a population-based case-control study in Wichita (KS, USA). We used traditional biostatistical methods to measure the association of high AL to standardized measures of physical and mental functioning, disability, fatigue and general symptom severity. We also used nonlinear regression technology embedded in machine learning algorithms to learn equations predicting various CFS symptoms based on the individual components of the allostatic load index (ALI). METHODS: An ALI was computed for all study participants using available laboratory and clinical data on metabolic, cardiovascular and hypothalamic-pituitary-adrenal (HPA) axis factors. Physical and mental functioning/impairment was measured using the Medical Outcomes Study 36-item Short Form Health Survey (SF-36); current fatigue was measured using the 20-item multidimensional fatigue inventory (MFI); frequency and intensity of symptoms was measured using the 19-item symptom inventory (SI). Genetic programming, a nonlinear regression technique, was used to learn an ensemble of different predictive equations rather just than a single one. Statistical analysis was based on the calculation of the percentage of equations in the ensemble that utilized each input variable, producing a measure of the 'utility' of the variable for the predictive problem at hand. Traditional biostatistics methods include the median and Wilcoxon tests for comparing the median levels of subscale scores obtained on the SF-36, the MFI and the SI summary score. RESULTS: Among CFS patients, but not controls, a high level of AL was significantly associated with lower median values (indicating worse health) of bodily pain, physical functioning and general symptom frequency/intensity. Using genetic programming, the ALI was determined to be a better predictor of these three health measures than any subcombination of ALI components among cases, but not controls.

Adult↗

Cortical ensemble activity increasingly predicts behaviour outcomes during learning of a motor task.

When an animal learns to make movements in response to different stimuli, changes in activity in the motor cortex seem to accompany and underlie this learning. The precise nature of modifications in cortical motor areas during the initial stages of motor learning, however, is largely unknown. Here we address this issue by chronically recording from neuronal ensembles located in the rat motor cortex, throughout the period required for rats to learn a reaction-time task. Motor learning was demonstrated by a decrease in the variance of the rats' reaction times and an increase in the time the animals were able to wait for a trigger stimulus. These behavioural changes were correlated with a significant increase in our ability to predict the correct or incorrect outcome of single trials based on three measures of neuronal ensemble activity: average firing rate, temporal patterns of firing, and correlated firing. This increase in prediction indicates that an association between sensory cues and movement emerged in the motor cortex as the task was learned. Such modifications in cortical ensemble activity may be critical for the initial learning of motor tasks.

Analysis of Variance↗

Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory.

Damage to the hippocampal system disrupts recent memory but leaves remote memory intact. The account presented here suggests that memories are first stored via synaptic changes in the hippocampal system, that these changes support reinstatement of recent memories in the neocortex, that neocortical synapses change a little on each reinstatement, and that remote memory is based on accumulated neocortical changes. Models that learn via changes to connections help explain this organization. These models discover the structure in ensembles of items if learning of each item is gradual and interleaved with learning about other items. This suggests that the neocortex learns slowly to discover the structure in ensembles of experiences. The hippocampal system permits rapid learning of new items without disrupting this structure, and reinstatement of new memories interleaves them with others to integrate them into structured neocortical memory systems.

Amnesia, Retrograde↗

Lung cancer cell identification based on artificial neural network ensembles.

An artificial neural network ensemble is a learning paradigm where several artificial neural networks are jointly used to solve a problem. In this paper, an automatic pathological diagnosis procedure named Neural Ensemble-based Detection (NED) is proposed, which utilizes an artificial neural network ensemble to identify lung cancer cells in the images of the specimens of needle biopsies obtained from the bodies of the subjects to be diagnosed. The ensemble is built on a two-level ensemble architecture. The first-level ensemble is used to judge whether a cell is normal with high confidence where each individual network has only two outputs respectively normal cell or cancer cell. The predictions of those individual networks are combined by a novel method presented in this paper, i.e. full voting which judges a cell to be normal only when all the individual networks judge it is normal. The second-level ensemble is used to deal with the cells that are judged as cancer cells by the first-level ensemble, where each individual network has five outputs respectively adenocarcinoma, squamous cell carcinoma, small cell carcinoma, large cell carcinoma, and normal, among which the former four are different types of lung cancer cells. The predictions of those individual networks are combined by a prevailing method, i.e. plurality voting. Through adopting those techniques, NED achieves not only a high rate of overall identification, but also a low rate of false negative identification, i.e. a low rate of judging cancer cells to be normal ones, which is important in saving lives due to reducing missing diagnoses of cancer patients.

Adenocarcinoma↗

[The neuronal ensembles of the brain].

The conclusion about availability of two types of neuron ensembles in the central nervous structures was made on the basis of proper experimental studies and appropriate literature data analysis. The first is presented by neurons functional cooperation (Hebbian ensembles) formed during learning process on the basis of synchronization by high-frequency oscillator activity. The second is presented by morphological-and-functional structures (Kogan's neuron ensembles) formed during neurogenesis process and genetically destined for biologically valuable patterns recognition. The control of functional state of the brain neuron networks and occurred information processes is realized by non-specific systems either restrict the neurons capability of joining in the rate of global synchronization by synchronization and desynchronization actions or form and maintain the common rhythm of activity in considerable neuron populations. The latter block single elements ability to enter into specific interactions.

Animals↗

Explaining the output of ensembles in medical decision support on a case by case basis.

The use of ensembles in machine learning (ML) has had a considerable impact in increasing the accuracy and stability of predictors. This increase in accuracy has come at the cost of comprehensibility as, by definition, an ensemble model is considerably more complex than its component models. This is of significance for decision support systems in medicine because of the reluctance to use models that are essentially black boxes. Work on making ensembles comprehensible has so far focused on global models that mirror the behaviour of the ensemble as closely as possible. With such global models there is a clear tradeoff between comprehensibility and fidelity. In this paper, we pursue another tack, looking at local comprehensibility where the output of the ensemble is explained on a case-by-case basis. We argue that this meets the requirements of medical decision support systems. The approach presented here identifies the ensemble members that best fit the case in question and presents the behaviour of these in explanation.

Anticoagulants↗

Towards an explicit account of implicit learning.

PURPOSE OF REVIEW: The human brain supports acquisition mechanisms that can extract structural regularities implicitly from experience without the induction of an explicit model. Reber defined the process by which an individual comes to respond appropriately to the statistical structure of the input ensemble as implicit learning. He argued that the capacity to generalize to new input is based on the acquisition of abstract representations that reflect underlying structural regularities in the acquisition input. We focus this review of the implicit learning literature on studies published during 2004 and 2005. We will not review studies of repetition priming ('implicit memory'). Instead we focus on two commonly used experimental paradigms: the serial reaction time task and artificial grammar learning. Previous comprehensive reviews can be found in Seger's 1994 article and the Handbook of Implicit Learning. RECENT FINDINGS: Emerging themes include the interaction between implicit and explicit processes, the role of the medial temporal lobe, developmental aspects of implicit learning, age-dependence, the role of sleep and consolidation. SUMMARY: The attempts to characterize the interaction between implicit and explicit learning are promising although not well understood. The same can be said about the role of sleep and consolidation. Despite the fact that lesion studies have relatively consistently suggested that the medial temporal lobe memory system is not necessary for implicit learning, a number of functional magnetic resonance studies have reported medial temporal lobe activation in implicit learning. This issue merits further research. Finally, the clinical relevance of implicit learning remains to be determined.

Brain↗

An experimental bias-variance analysis of SVM ensembles based on resampling techniques.

Recently, bias-variance decomposition of error has been used as a tool to study the behavior of learning algorithms and to develop new ensemble methods well suited to the bias-variance characteristics of base learners. We propose methods and procedures, based on Domingo's unified bias-variance theory, to evaluate and quantitatively measure the bias-variance decomposition of error in ensembles of learning machines. We apply these methods to study and compare the bias-variance characteristics of single support vector machines (SVMs) and ensembles of SVMs based on resampling techniques, and their relationships with the cardinality of the training samples. In particular, we present an experimental bias-variance analysis of bagged and random aggregated ensembles of SVMs in order to verify their theoretical variance reduction properties. The experimental bias-variance analysis quantitatively characterizes the relationships between bagging and random aggregating, and explains the reasons why ensembles built on small subsamples of the data work with large databases. Our analysis also suggests new directions for research to improve on classical bagging.

Algorithms↗

Dynamics of population code for working memory in the prefrontal cortex.

Some neurons (delay cells) in the prefrontal cortex elevate their activities throughout the time period during which the animal is required to remember past events and prepare future behavior, suggesting that working memory is mediated by continuous neural activity. It is unknown, however, how working memory is represented within a population of prefrontal cortical neurons. We recorded from neuronal ensembles in the prefrontal cortex as rats learned a new delayed alternation task. Ensemble activities changed in parallel with behavioral learning so that they increasingly allowed correct decoding of previous and future goal choices. In well-trained rats, considerable decoding was possible based on only a few neurons and after removing continuously active delay cells. These results show that neural activity in the prefrontal cortex changes dynamically during new task learning so that working memory is robustly represented and that working memory can be mediated by sequential activation of different neural populations.

Action Potentials↗

Are spatial memories strengthened in the human hippocampus during slow wave sleep?

In rats, the firing sequences observed in hippocampal ensembles during spatial learning are replayed during subsequent sleep, suggesting a role for posttraining sleep periods in the offline processing of spatial memories. Here, using regional cerebral blood flow measurements, we show that, in humans, hippocampal areas that are activated during route learning in a virtual town are likewise activated during subsequent slow wave sleep. Most importantly, we found that the amount of hippocampal activity expressed during slow wave sleep positively correlates with the improvement of performance in route retrieval on the next day. These findings suggest that learning-dependent modulation in hippocampal activity during human sleep reflects the offline processing of recent episodic and spatial memory traces, which eventually leads to the plastic changes underlying the subsequent improvement in performance.

Adult↗

Examples and experience: on the uncertainty of medicine.

After a brief account of the uncertainty of medicine in early modern thought, this paper focuses on two supple, sophisticated accounts of medicine by 'non-medical' writers--Michel de Montaigne's views of medical theory and medical practice and Francis Bacon's proposals for renovating both--in which the claims of individual sufferers are set against the normativity of medicine as a whole. From around 1500 to around 1680, in the common ensemble of both learned and popular invective, medicine was disparaged as poor philosophy and worse practice, even as the 'lowest of professions'. In remarkably broad, elegant interventions, Montaigne argues that medicine is based on 'examples and experience' (and 'so is my opinion', he adds), impugning its universalizing claims with the tractable experience of his own embodiment, with his own historia and consilium, while Francis Bacon enlists dietetics, Hippocratic case-taking and medical history in his broad programme for the reform of medicine. He more or less accepts Montaigne's argument for particularity in medical theory and practice, but presses the particular into service in his reformist programme. Like many sixteenth- and early seventeenth-century scholars and physicians frustrated with Galenic methods and models, both turn to Hippocratic practice and to hygiene and dietetics as salves for an ailing discipline. Finally, I argue that both writers enquire into viable means for inflecting learned medicine with particular experience, and both settle on rhetorical tools - analogy and exemplarity - as the means by which universalized medical models might be particularized or reformed.

History of Medicine↗

A SuperLearner-based pipeline for the development of DNA methylation-derived predictors of phenotypic traits.

BACKGROUND: DNA methylation (DNAm) provides a window to characterize the impacts of environmental exposures and the biological aging process. Epigenetic clocks are often trained on DNAm using penalized regression of CpG sites, but recent evidence suggests potential benefits of training epigenetic predictors on principal components. METHODOLOGY/FINDINGS: We developed a pipeline to simultaneously train three epigenetic predictors; a traditional CpG Clock, a PCA Clock, and a SuperLearner PCA Clock (SL PCA). We gathered publicly available DNAm datasets to generate i) a novel childhood epigenetic clock, ii) a reconstructed Hannum adult blood clock, and iii) as a proof of concept, a predictor of polybrominated biphenyl exposure using the three developmental methodologies. We used correlation coefficients and median absolute error to assess fit between predicted and observed measures, as well as agreement between duplicates. The SL PCA clocks improved fit with observed phenotypes relative to the PCA clocks or CpG clocks across several datasets. We found evidence for higher agreement between duplicate samples run on alternate DNAm arrays when using SL PCA clocks relative to traditional methods. Analyses examining associations between relevant exposures and epigenetic age acceleration (EAA) produced more precise effect estimates when using predictions derived from SL PCA clocks. CONCLUSIONS: We introduce a novel method for the development of DNAm-based predictors that combines the improved reliability conferred by training on principal components with advanced ensemble-based machine learning. Coupling SuperLearner with PCA in the predictor development process may be especially relevant for studies with longitudinal designs utilizing multiple array types, as well as for the development of predictors of more complex phenotypic traits.

DNA Methylation↗

Efficient Partition of Learning Data Sets for Neural Network Training.

This study investigates the emerging possibilities of combining unsupervised and supervised learning in neural network ensembles. Such strategy is used to get an efficient partition of a noisy input data set in order to focus the training of neural networks on the most complex and informative domains of the data set and accelerate the learning phase. The proposed algorithm provides a good prediction accuracy using fewer cases from non-informative domains according to a correlative measure of dependency between cases of the training set. This measure takes into account internal relationships amid analyzed data and can be used to cluster neighbor cases in a multidimensional space and to filter out the outliers. The possible relation of the proposed algorithm to brain processing occurring in the thalamo-cortical pathway is discussed.

Journal Article↗

Medical diagnosis with C4.5 Rule preceded by artificial neural network ensemble.

Comprehensibility is very important for any machine learning technique to be used in computer-aided medical diagnosis. Since an artificial neural network ensemble is composed of multiple artificial neural networks, its comprehensibility is worse than that of a single artificial neural network. In this paper, C4.5 Rule-PANE which combines artificial neural network ensemble with rule induction by regarding the former as a preprocess of the latter, is proposed. At first, an artificial neural network ensemble is trained. Then, a new training data set is generated by feeding the feature vectors of the original training instances to the trained ensemble and replacing the expected class labels of the original training instances with the class labels output from the ensemble. Additional training data may also be appended by randomly generating feature vectors and combining them with their corresponding class labels output from the ensemble. Finally, a specific rule induction approach, i.e., C4.5 Rule, is used to learn rules from the new training data set. Case studies on diabetes, hepatitis, and breast cancer show that C4.5 Rule-PANE could generate rules with strong generalization ability, which profits from artificial neural network ensemble, and strong comprehensibility, which profits from rule induction.

Diagnosis↗