Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble learning”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Computation-directed identification of OxyR DNA binding sites in Escherichia coli.

A computational search was carried out to identify additional targets for the Escherichia coli OxyR transcription factor. This approach predicted OxyR binding sites upstream of dsbG, encoding a periplasmic disulfide bond chaperone-isomerase; upstream of fhuF, encoding a protein required for iron uptake; and within yfdI. DNase I footprinting assays confirmed that oxidized OxyR bound to the predicted site centered 54 bp upstream of the dsbG gene and 238 bp upstream of a known OxyR binding site in the promoter region of the divergently transcribed ahpC gene. Although the new binding site was near dsbG, Northern blotting and primer extension assays showed that OxyR binding to the dsbG-proximal site led to the induction of a second ahpCF transcript, while OxyR binding to the ahpCF-proximal site leads to the induction of both dsbG and ahpC transcripts. Oxidized OxyR binding to the predicted site centered 40 bp upstream of the fhuF gene was confirmed by DNase I footprinting, but these assays further revealed a second higher-affinity site in the fhuF promoter. Interestingly, the two OxyR sites in the fhuF promoter overlapped with two regions bound by the Fur repressor. Expression analysis revealed that fhuF was repressed by hydrogen peroxide in an OxyR-dependent manner. Finally, DNase I footprinting experiments showed OxyR binding to the site predicted to be within the coding sequence of yfdI. These results demonstrate the versatile modes of regulation by OxyR and illustrate the need to learn more about the ensembles of binding sites and transcripts in the E. coli genome.

Bacterial Outer Membrane Proteins↗

An incremental learning algorithm with confidence estimation for automated identification of NDE signals.

An incremental learning algorithm is introduced for learning new information from additional data that may later become available, after a classifier has already been trained using a previously available database. The proposed algorithm is capable of incrementally learning new information without forgetting previously acquired knowledge and without requiring access to the original database, even when new data include examples of previously unseen classes. Scenarios requiring such a learning algorithm are encountered often in nondestructive evaluation (NDE) in which large volumes of data are collected in batches over a period of time, and new defect types may become available in subsequent databases. The algorithm, named Learn++, takes advantage of synergistic generalization performance of an ensemble of classifiers in which each classifier is trained with a strategically chosen subset of the training databases that subsequently become available. The ensemble of classifiers then is combined through a weighted majority voting procedure. Learn++ is independent of the specific classifier(s) comprising the ensemble, and hence may be used with any supervised learning algorithm. The voting procedure also allows Learn++ to estimate the confidence in its own decision. We present the algorithm and its promising results on two separate ultrasonic weld inspection applications.

Journal Article↗

Automated protein classification using consensus decision.

We propose a novel technique for automatically generating the SCOP classification of a protein structure with high accuracy. High accuracy is achieved by combining the decisions of multiple methods using the consensus of a committee (or an ensemble) classifier. Our technique is rooted in machine learning which shows that by judicially employing component classifiers, an ensemble classifier can be constructed to outperform its components. We use two sequence- and three structure-comparison tools as component classifiers. Given a protein structure, using the joint hypothesis, we first determine if the protein belongs to an existing category (family, superfamily, fold) in the SCOP hierarchy. For the proteins that are predicted as members of the existing categories, we compute their family-, superfamily-, and fold-level classifications using the consensus classifier. We show that we can significantly improve the classification accuracy compared to the individual component classifiers. In particular, we achieve error rates that are 3-12 times less than the individual classifiers' error rates at the family level, 1.5-4.5 times less at the superfamily level, and 1.1-2.4 times less at the fold level.

Algorithms↗

Learning to control a brain-machine interface for reaching and grasping by primates.

Reaching and grasping in primates depend on the coordination of neural activity in large frontoparietal ensembles. Here we demonstrate that primates can learn to reach and grasp virtual objects by controlling a robot arm through a closed-loop brain-machine interface (BMIc) that uses multiple mathematical models to extract several motor parameters (i.e., hand position, velocity, gripping force, and the EMGs of multiple arm muscles) from the electrical activity of frontoparietal neuronal ensembles. As single neurons typically contribute to the encoding of several motor parameters, we observed that high BMIc accuracy required recording from large neuronal ensembles. Continuous BMIc operation by monkeys led to significant improvements in both model predictions and behavioral performance. Using visual feedback, monkeys succeeded in producing robot reach-and-grasp movements even when their arms did not move. Learning to operate the BMIc was paralleled by functional reorganization in multiple cortical areas, suggesting that the dynamic properties of the BMIc were incorporated into motor and sensory cortical representations.

Animals↗

The Bioperl toolkit: Perl modules for the life sciences.

The Bioperl project is an international open-source collaboration of biologists, bioinformaticians, and computer scientists that has evolved over the past 7 yr into the most comprehensive library of Perl modules available for managing and manipulating life-science information. Bioperl provides an easy-to-use, stable, and consistent programming interface for bioinformatics application programmers. The Bioperl modules have been successfully and repeatedly used to reduce otherwise complex tasks to only a few lines of code. The Bioperl object model has been proven to be flexible enough to support enterprise-level applications such as EnsEMBL, while maintaining an easy learning curve for novice Perl programmers. Bioperl is capable of executing analyses and processing results from programs such as BLAST, ClustalW, or the EMBOSS suite. Interoperation with modules written in Python and Java is supported through the evolving BioCORBA bridge. Bioperl provides access to data stores such as GenBank and SwissProt via a flexible series of sequence input/output modules, and to the emerging common sequence data storage format of the Open Bioinformatics Database Access project. This study describes the overall architecture of the toolkit, the problem domains that it addresses, and gives specific examples of how the toolkit can be used to solve common life-sciences problems. We conclude with a discussion of how the open-source nature of the project has contributed to the development effort.

Algorithms↗

Decision forest analysis of large-scale sib-pair identical-by-decent profiles for locating the underlying disease genes for alcoholism in human.

OBJECTIVE: To extract the relevant SNPs for alcoholism using sib-pair IBD profiles of pedigrees. METHODS: We used the ensemble decision approach, a supervised learning approach based on decision forests, to locate alcoholism relevant SNPs using genome-wide SNP data. RESULTS: Application to a publicly available large dataset of 100 simulated replicates for three American populations (http://www.gaworkshop.org/) demonstrates that the proposed approach has successfully located all of the simulated true loci. CONCLUSION: The numerical results establish the proposed decision forest analysis to be a powerful and practical alternative for large-scale family-based association study.

Alcoholism↗

Dramatic interactions: theater work and the formation of learning communities.

This article examines the relationship of theater and dramatic study to models of learning communities in promoting identity, diversity, and culture. Theater is an example of how learning community can be achieved and levels of theater use in education are presented as ways in which educators can create ensemble and foster community. Strategies for developing learning communities using the performing arts are provided.

Curriculum↗

Decision tree based information integration for automated protein classification.

We propose a novel technique for automatically generating the SCOP classification of a protein structure with high accuracy. We achieve accurate classification by combining the decisions of multiple methods using the consensus of a committee (or an ensemble) classifier. Our technique, based on decision trees, is rooted in machine learning which shows that by judicially employing component classifiers, an ensemble classifier can be constructed to outperform its components. We use two sequence- and three structure-comparison tools as component classifiers. Given a protein structure and using the joint hypothesis, we first determine if the protein belongs to an existing category (family, superfamily, fold) in the SCOP hierarchy. For the proteins that are predicted as members of the existing categories, we compute their family-, superfamily-, and fold-level classifications using the consensus classifier. We show that we can significantly improve the classification accuracy compared to the individual component classifiers. In particular, we achieve error rates that are 3-12 times less than the individual classifiers' error rates at the family level, 1.5-4.5 times less at the superfamily level, and 1.1-2.4 times less at the fold level.

Algorithms↗

N-Terminal myristoylation predictions by ensembles of neural networks.

N-terminal myristoylation is a post-translational modification that causes the addition of a myristate to a glycine in the N-terminal end of the amino acid chain. This work presents neural network (NN) models that learn to discriminate myristoylated and nonmyristoylated proteins. Ensembles of 25 NNs and decision trees were trained on 390 positive sequences and 327 negative sequences. Experiments showed that NN ensembles were more accurate than decision tree ensembles. Our NN predictor evaluated by the leave-one-out procedure, obtained a false positive error rate equal to 2.1%. That was better than the PROSITE pattern for myristoylation for which the false positive error rate was 22.3%. On a recent version of Swiss-Prot (41.2), the NN ensemble predicted 876 myristoylated proteins, while 1150 proteins were predicted by the PROSITE pattern for myristoylation. Finally, compared to the well-known NMT predictor, the NN predictor gave similar results. Our tool is available under http://www.expasy.org/tools/myristoylator/myristoylator.html.

Amino Acid Sequence↗

Representations of odors in the rat orbitofrontal cortex change during and after learning.

Cells in the orbitofrontal cortex (OF) respond to odors and their associated rewards. To determine how these responses are acquired and maintained, the authors recorded single OF units in rats performing an odor discrimination task. Approximately 64% of all cells differentiated between rewarded and nonrewarded odors. These odor valence responses changed during learning in 26% of all cells, and these changes were positively correlated with improving performance, supporting the idea that the information provided by these cells is used in learning the task. However, changes in odor valence responses were also observed after learning, and included not only increases in odor discrimination, but also decreases or mixed increases and decreases. Thus, only some of the changes in firing reflected acquisition of the task. The results suggest that learning triggers a continuing reorganization of OF neural ensembles representing odors and their rewards.

Action Potentials↗

Changes in electrical activity of rabbit olfactory bulb and cortex to conditioned odor stimulation.

Rabbits with chronically implanted electrodes in olfactory bulb and cortex were classically conditioned to give an increase in relative frequency of sniffing to odor stimuli (CS+) reinforced with mild electric shock. Electroencephalographic high-frequency (35-85 Hz) bursts were recorded from an ensemble of nine bulbar depth electrodes and a second ensemble of 50 cortical surface electrodes. The olfactory cortex responded to the CS+ with sustained elevation of burst amplitude even though the olfactory bulb, from which it receives its primary centripetal input, underwent a marked decline in burst amplitude during the same time period. The amplitude reduction was not spatially uniform: The burst of the bulbar region that declined most in amplitude had the greatest phase lag with respect to the bulbar ensemble average burst. These effects were learning related because they did not occur for CS+ trials at the beginning of conditioning or for unreinforced control trials at any time.

Action Potentials↗

Social disconnection integrates genetic and proteomic risks in suicidal ideation and depression.

Suicidal ideation (SI) and major depressive disorder (MDD) are complex psychiatric conditions arising from the interplay of genetic liability, molecular processes, and psychosocial factors. While these dimensions have been extensively studied in isolation, their joint contribution to SI and MDD remains unclear. This study integrates multi-modal data to elucidate these synergistic effects and develop robust models for individual-level risk stratification. Leveraging longitudinal multi-modal data from 13,085 UK Biobank participants, we integrated genomic, proteomic, and social connection profiles. We developed interpretable risk scores using a rigorous supervised machine learning framework encompassing diverse linear and ensemble classifiers. Permutation importance was employed to quantify feature contributions and derive transparent, weighted risk metrics across diverse classifiers. These scores were validated through association, interaction, and mediation analyses. Social connection-based risk scores significantly differentiated cases and controls across the two suicidal ideation phenotypes at 2017 and 2023 with cross-sectional analyses (AUCs: 0.70 - 0.73), outperforming proteomic-only models. Functional dimensions of social connection emerged as the most informative predictors. Longitudinal analyses revealed that social risk scores at baseline predicted suicidal ideation onset six years later, independent of demographic covariates. Interaction analyses demonstrated that polygenic risk for suicide attempt significantly interacted with both social and proteomic risk features in relation to depression. Structural equation models further confirmed that social disconnection acts as a key mediator linking genetic predisposition to MDD and SI. Social disconnection is a critical risk factor mediating the impact of genetic vulnerability on psychiatric outcomes. Integrating social, genetic, and molecular data supports a multilevel framework for risk stratification and highlights the potential of socially oriented interventions to mitigate biological risk.

Humans↗

Teaching computers to fold proteins.

A new general algorithm for optimization of potential functions for protein folding is introduced. It is based upon gradient optimization of the thermodynamic stability of native folds of a training set of proteins with known structure. The iterative update rule contains two thermodynamic averages which are estimated by (generalized ensemble) Monte Carlo. We test the learning algorithm on a Lennard-Jones (LJ) force field with a torsional angle degrees-of-freedom and a single-atom side-chain. In a test with 24 peptides of known structure, none folded correctly with the initial potential functions, but two-thirds came within 3 A to their native fold after optimizing the potential functions.

Algorithms↗

Topologically distinct intratumoral heterogeneity scores for predicting high-risk pathological grades in invasive lung adenocarcinoma: A multicenter study across four institutions.

High-risk subtypes of invasive lung adenocarcinoma (IAC), particularly micropapillary- or solid-predominant patterns, are closely associated with poor prognosis. This multicenter retrospective study developed and validated a predictive model for the preoperative identification of these high-risk subtypes using topologically distinct intratumoral heterogeneity (ITH) scores derived from CT images. The study included 1,051 patients with IAC. Two complementary ITH scores were developed: a two-dimensional ITH score, which integrated local radiomics features with global pixel distribution patterns on the largest cross-sectional CT slice, and a three-dimensional ITH score, which extended this quantification across the entire tumor volume. Clinicoradiological features and ITH scores were incorporated as model inputs to construct six base machine learning classifiers and a final stacking ensemble classifier. Model interpretability and robustness were evaluated using SHapley Additive exPlanations (SHAP)-based ablation analyses. An independent dataset from The Cancer Imaging Archive (TCIA) was used for external validation to investigate associations between ITH scores and pathological characteristics, genomic features, recurrence-free survival, and overall survival. The stacking ensemble classifier achieved the best predictive performance, with an area under the receiver operating characteristic curve of 0.875, outperforming models based solely on radiomics features (0.834) or clinicoradiological features (0.792). SHAP analysis identified the 3D ITH score as the most influential contributor to model output, and TCIA validation showed that higher 3D ITH scores were associated with more aggressive tumor biology and poorer survival outcomes. The topologically distinct 3D ITH score may provide a clinically meaningful imaging biomarker for preoperative risk stratification in IAC.

Journal Article↗

Hippocampal encoding of non-spatial trace conditioning.

Trace eyeblink classical conditioning is a non-spatial learning paradigm that requires an intact hippocampus. This task is hippocampus-dependent because the auditory tone conditioned stimulus (CS) is temporally separated from the corneal airpuff unconditioned stimulus (US) by a 500-ms trace interval. Our laboratory has performed a series of neurophysiological experiments that have examined the activity of pyramidal cells in the CA1 area of the hippocampus during trace eyeblink conditioning. We have found that the non-spatial stimuli involved in this paradigm are encoded in the hippocampus in a logical order that is necessary for their association and the subsequent expression of behavioral learning. Although there were many profiles of single neurons responding to the CS-US trial during training, the majority of the neurons showed an increase in activity to the airpuff-US. Prior to learning, it appears that hippocampal cells and ensembles of cells were preferentially attending to the stimulus with immediate behavioral importance, the US. Hippocampal cells then began to respond to the associated neutral stimulus, the CS. Shortly thereafter, animals began to show increases in the behavioral expression of CRs. In some experiments, hippocampal neurons from aged animals exhibited impairments in the encoding of CS and US information. These aged animals were not able to associate these stimuli and acquire trace eyeblink CRs. Our findings along with the findings of other spatial learning studies, suggest that the hippocampus is involved in encoding information about discontiguous sets of stimuli, either spatial or nonspatial, especially early in the learning process.

Animals↗

Differential corticostriatal plasticity during fast and slow motor skill learning in mice.

BACKGROUND: Motor skill learning usually comprises "fast" improvement in performance within the initial training session and "slow" improvement that develops across sessions. Previous studies have revealed changes in activity and connectivity in motor cortex and striatum during motor skill learning. However, the nature and dynamics of the plastic changes in each of these brain structures during the different phases of motor learning remain unclear. RESULTS: By using multielectrode arrays, we recorded the simultaneous activity of neuronal ensembles in motor cortex and dorsal striatum of mice during the different phases of skill learning on an accelerating rotarod. Mice exhibited fast improvement in the task during the initial session and also slow improvement across days. Throughout training, a high percentage of striatal (57%) and motor cortex (55%) neurons were task related; i.e., changed their firing rate while mice were running on the rotarod. Improvement in performance was accompanied by substantial plastic changes in both striatum and motor cortex. We observed parallel recruitment of task-related neurons in both structures specifically during the first session. Conversely, during slow learning across sessions we observed differential refinement of the firing patterns in each structure. At the neuronal ensemble level, we observed considerable changes in activity within the first session that became less evident during subsequent sessions. CONCLUSIONS: These data indicate that cortical and striatal circuits exhibit remarkable but dissociable plasticity during fast and slow motor skill learning and suggest that distinct neural processes mediate the different phases of motor skill learning.

Analysis of Variance↗

Applications of the Vitamin D sterol-Vitamin D receptor (VDR) conformational ensemble model.

Over the past 20 years much has been learned about the cellular actions of the steroid hormone 1alpha,25(OH)2-Vitamin D3 (1,25D). Perhaps most importantly structure-function studies led to the discovery that different chemical and physical features of 1,25D are preferred to initiate either exonuclear, non-genomic or endonuclear, genomic cellular signaling. It is well documented that both a 1alpha-OH and 25-OH, and a 6-s-trans, bowl-shaped, sterol conformation are absolutely required for efficient gene transcription, while 6-s-cis locked analogs and 1-deoxy, 25(OH)D3 metabolites activate a variety of non-genomic, rapid responses. These results and the observation that S237 (helix-3; H3) and R274 (H5) are the most static residues in the human 1,25D-Vitamin D receptor (VDR) X-ray construct (see B-values in pdb: 1DB1) and form H-bonds with the 1alpha-OH of 1,25D in the X-ray, genomic pocket (G-pocket), provided the basis for the molecular modeling experiments that led to the discovery of a putative VDR alternative ligand binding pocket (A-pocket). The conformational ensemble model generated from the in silico results provides an explanation for how the VDR can function as a receptor propagating both genomic and non-genomic signaling events. In this report the theoretical gating properties controlling ligand access to the A- and G-pockets will be compared and the model will be used to provide a molecular explanation for the confusing structure-function results pertaining to 1,25D, its side-chain metabolite, 23S,25R-1alpha,25(OH)2-D3-26,23-lactone (BS), and its synthetic two side-chain analog, 21-(3'-hydroxy-3'-methylbutyl)-1alpha,25(OH)2-D3 (KH or Gemini). In addition, evidence that the model is consistent with the pH requirement for Vitamin D sterol-VDR crystallization will be presented.

Arginine↗