Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Generating uniformly distributed random networks.

The analysis of real networks taken from the biological, social, and physical sciences often requires a carefully posed statistical null-hypothesis approach. One common method requires comparing real networks to an ensemble of random matrices that satisfy realistic constraints in which each different matrix member is equiprobable. We discuss existing methods for generating uniformly distributed (constrained) random matrices, describe their shortcomings, and present an efficient technique that should have many practical applications.

Journal Article↗

Comparison between neural networks and multiple logistic regression to predict acute coronary syndrome in the emergency room.

OBJECTIVE: Patients with suspicion of acute coronary syndrome (ACS) are difficult to diagnose and they represent a very heterogeneous group. Some require immediate treatment while others, with only minor disorders, may be sent home. Detecting ACS patients using a machine learning approach would be advantageous in many situations. METHODS AND MATERIALS: Artificial neural network (ANN) ensembles and logistic regression models were trained on data from 634 patients presenting an emergency department with chest pain. Only data immediately available at patient presentation were used, including electrocardiogram (ECG) data. The models were analyzed using receiver operating characteristics (ROC) curve analysis, calibration assessments, inter- and intra-method variations. Effective odds ratios for the ANN ensembles were compared with the odds ratios obtained from the logistic model. RESULTS: The ANN ensemble approach together with ECG data preprocessed using principal component analysis resulted in an area under the ROC curve of 80%. At the sensitivity of 95% the specificity was 41%, corresponding to a negative predictive value of 97%, given the ACS prevalence of 21%. Adding clinical data available at presentation did not improve the ANN ensemble performance. Using the area under the ROC curve and model calibration as measures of performance we found an advantage using the ANN ensemble models compared to the logistic regression models. CONCLUSION: Clinically, a prediction model of the present type, combined with the judgment of trained emergency department personnel, could be useful for the early discharge of chest pain patients in populations with a low prevalence of ACS.

Acute Disease↗

Real-time imaging of evoked activity in local circuits of the salamander olfactory bulb.

The encoding of olfactory information in the central nervous system (CNS) depends on spatially distributed patterns of activity generated simultaneously in many neuronal circuits. Optical neurophysiological recording permits analysis of neural activity non-invasively and with high spatial and temporal resolution. Here, a video method for imaging voltage-sensitive dye fluorescence in vivo is used to map neuronal activity in local circuits of the salamander olfactory bulb. The method permits the imaging of simultaneous ensemble transmembrane activity in real time. After electrical stimulation of the olfactory nerve, activity spreads centripetally from the sites of synaptic input to generate nonhomogeneous response patterns that are presumably mediated by local circuits within the bulbar layers. The results also show the overlapping temporal sequences of activation of cell groups in each layer. The method thus provides high resolution, sequential video images of the spatial and temporal progression of transmembrane events in neuronal circuits after afferent stimulation and offers the opportunity for studying ensemble events in other brain regions.

Animals↗

JIGSAW: integration of multiple sources of evidence for gene prediction.

MOTIVATION: Computational gene finding systems play an important role in finding new human genes, although no systems are yet accurate enough to predict all or even most protein-coding regions perfectly. Ab initio programs can be augmented by evidence such as expression data or protein sequence homology, which improves their performance. The amount of such evidence continues to grow, but computational methods continue to have difficulty predicting genes when the evidence is conflicting or incomplete. Genome annotation pipelines collect a variety of types of evidence about gene structure and synthesize the results, which can then be refined further through manual, expert curation of gene models. RESULTS: JIGSAW is a new gene finding system designed to automate the process of predicting gene structure from multiple sources of evidence, with results that often match the performance of human curators. JIGSAW computes the relative weight of different lines of evidence using statistics generated from a training set, and then combines the evidence using dynamic programming. Our results show that JIGSAW's performance is superior to ab initio gene finding methods and to other pipelines such as Ensembl. Even without evidence from alignment to known genes, JIGSAW can substantially improve gene prediction accuracy as compared with existing methods. AVAILABILITY: JIGSAW is available as an open source software package at http://cbcb.umd.edu/software/jigsaw.

Algorithms↗

Statistical ensemble of scale-free random graphs.

A thorough discussion of the statistical ensemble of scale-free connected random tree graphs is presented. Methods borrowed from field theory are used to define the ensemble and to study analytically its properties. The ensemble is characterized by two global parameters, the fractal and the spectral dimensions, which are explicitly calculated. It is discussed in detail how the geometry of the graphs varies when the weights of the nodes are modified. The stability of the scale-free regime is also considered: when it breaks down, either a scale is spontaneously generated or else, a "singular" node appears and the graphs become crumpled. A new computer algorithm to generate these random graphs is proposed. Possible generalizations are also discussed. In particular, more general ensembles are defined along the same lines and the computer algorithm is extended to arbitrary (degenerate) scale-free random graphs.

Journal Article↗

Efficient sampling in collective coordinate space.

Collective motions in biological macromolecules have been shown to be important for function. The most important collective motions occur on slow time scales, which poses a sampling problem in dynamic simulation of biomolecules. We present a novel method for efficient conformational sampling. The method combines the simulation of an ensemble of concurrent trajectories with restraints acting on the ensemble of structures as a whole. Two properties of the ensemble may be restrained: (i) the variance of the ensemble and (ii) the average position of the ensemble. Both properties are defined in a subspace of collective coordinate space spanned by an arbitrary number of modes. We show that weak restraints on the ensemble variance suffice for an increase in sampling efficiency along soft modes by two orders of magnitudes. The resulting trajectories exhibit virtually the same structural quality as trajectories generated by restraint-free-molecular dynamics simulation, as judged by standard structure validation tools. The method is used to probe the resistance of a structure against conformational changes along collective modes and clearly distinguishes soft from stiff modes. Further applications are discussed. Proteins 2000;39:82-88.

Computer Simulation↗

A novel ensemble machine learning for robust microarray data classification.

Microarray data analysis and classification has demonstrated convincingly that it provides an effective methodology for the effective diagnosis of diseases and cancers. Although much research has been performed on applying machine learning techniques for microarray data classification during the past years, it has been shown that conventional machine learning techniques have intrinsic drawbacks in achieving accurate and robust classifications. This paper presents a novel ensemble machine learning approach for the development of robust microarray data classification. Different from the conventional ensemble learning techniques, the approach presented begins with generating a pool of candidate base classifiers based on the gene sub-sampling and then the selection of a sub-set of appropriate base classifiers to construct the classification committee based on classifier clustering. Experimental results have demonstrated that the classifiers constructed by the proposed method outperforms not only the classifiers generated by the conventional machine learning but also the classifiers generated by two widely used conventional ensemble learning methods (bagging and boosting).

Algorithms↗

Periodic orbits of the ensemble of Sinai-Arnold cat maps and pseudorandom number generation.

We propose methods for constructing high-quality pseudorandom number generators (RNGs) based on an ensemble of hyperbolic automorphisms of the unit two-dimensional torus (Sinai-Arnold map or cat map) while keeping a part of the information hidden. The single cat map provides the random properties expected from a good RNG and is hence an appropriate building block for an RNG, although unnecessary correlations are always present in practice. We show that introducing hidden variables and introducing rotation in the RNG output, accompanied with the proper initialization, dramatically suppress these correlations. We analyze the mechanisms of the single-cat-map correlations analytically and show how to diminish them. We generalize the Percival-Vivaldi theory in the case of the ensemble of maps, find the period of the proposed RNG analytically, and also analyze its properties. We present efficient practical realizations for the RNGs and check our predictions numerically. We also test our RNGs using the known stringent batteries of statistical tests and find that the statistical properties of our best generators are not worse than those of other best modern generators.

Journal Article↗

Polarization structure of lidar signals reflected from ice crystal clouds.

Polarization characteristics of signals of a monostatic lidar intended for sensing of homogeneous ice crystal clouds are calculated by the Monte Carlo method. Clouds are modeled as monodisperse ensembles of randomly oriented hexagonal ice crystals. The polarization state of multiply scattered lidar signal components is analyzed for different scattering orders depending on the crystal shapes and sizes as well as on the optical and geometrical conditions of observation. Light-scattering phase matrices (SPMs), calculated by the beam splitting method (BSM), are used as input data for solving the vector radiative transfer equation. The principles of the BSM method are briefly described, and the SPM components are given for hexagonal ice plates and columns of different sizes and linearly polarized incident radiation with the wavelength lambda = 0.55 microm.

Journal Article↗

Property-based design of GPCR-targeted library.

The design of a GPCR-targeted library, based on a scoring scheme for the classification of molecules into "GPCR-ligand-like" and "non-GPCR-ligand-like", is outlined. The methodology is a valuable tool that can aid in the selection and prioritization of potential GPCR ligands for bioscreening from large collections of compounds. It is based on the distillation of knowledge from large databases of GPCR and non-GPCR active agents. The method employed a set of descriptors for encoding the molecular structures and by training of a neural network for classifying the molecules. The molecular requirements were profiled and validated by using available databases of GPCR- and non-GPCR-active agents [5736 diverse GPCR-active molecules and 7506 diverse non-GPCR-active molecules from the Ensemble Database (Prous Science, 2002)]. The method enables efficient qualification or disqualification of a molecule as a potential GPCR ligand and represents a useful tool for constraining the size of GPCR-targeted libraries that will help speed up the development of new GPCR-active drugs.

Databases, Factual↗

Analysis methods for comparison of multiple molecular dynamics trajectories: applications to protein unfolding pathways and denatured ensembles.

In molecular dynamics simulations of protein unfolding, the pathway of one protein molecule is studied at a time. In contrast, experimental denaturation studies sample from large ensembles of molecules passing from the native to unfolded state. If reasonable comparisons with experiment are to be made, then the generality of the simulations needs to be confirmed by performing multiple unfolding simulations. Given that protein unfolding trajectories are very complicated functions of the proteins and the environment, comparing different trajectories, even under the same conditions, is not straightforward. Several methods are presented here that attempt to accomplish this task at different levels of complexity. The simpler methods are geometry based and make use of the root-mean-squared deviations between structures, while the more complicated methods are based on the time variation of the various properties of the system during the unfolding process. These methods are applied to multiple simulations of three different proteins, bovine pancreatic trypsin inhibitor, chymotrypsin inhibitor 2, and barnase. In general, for these three proteins protein unfolding proceeded via expansion of the core and fraying of secondary structure to yield the major transition state. Once past the transition state, the trajectories for a given protein diverged as the protein lost further secondary and tertiary structure by a variety of mechanisms. Although the unfolding pathways diverged, similar conformations were populated in the denatured state even when the unfolding occurred via different pathways. The multitude of different pathways leading to the denatured state agrees with the funnel description of protein folding. Although the pathways differed in conformational space, the physical properties of the conformations were often similar, highlighting the danger of assuming that similar observed properties imply similar conformations. In fact, there may be many different "conformational pathways" of unfolding that fit within a preferred "property space pathway".

Aprotinin↗

Exploring protein energy landscapes with hierarchical clustering.

In this work we present a new method for investigating local energy minima on a protein energy landscape. The CABS (CAlpha, CBeta and the center of mass of the Side chain) method was employed for generating protein models, but any other method could be used instead. Cα traces from an ensemble of models are hierarchical clustered with the HCPM (Hierarchical Clustering of Protein Models) method. The efficiency of this method for sampling and analyzing energy landscapes is shown.

Journal Article↗

Microcanonical calculations of excess thermodynamic properties of dense binary systems.

We derive a formulation to calculate the excess chemical potential of a fraction of N1 particles interacting with N2 particles of a different species. The excess chemical potential is calculated numerically from first principles by coupling molecular dynamics and Thomas-Fermi density functional theory to take into account the contribution arising from the quantum electrons on the forces acting on the ions. The choice of this simple functional is motivated by the fact that the present paper is devoted to the derivation and the validation of the method but more complicated functionals can and will be implemented in the future. This method is applied in the microcanonical ensemble, the most natural ensemble for molecular dynamics simulations. This avoids the introduction of a thermostat in the simulation and thus uncontrolled modifications of the trajectories calculated from the forces between particles. The calculations are conducted for three values of the input thermodynamic quantities, energy and density, and for different total numbers of particles in order to examine the uncertainties due to finite-size effects. This method and these calculations lie the basic foundation to study the thermodynamic stability of dense mixtures, without any a priori assumption on the degree of ionization of the different species.

Journal Article↗

Generalized-ensemble algorithms for molecular simulations of biopolymers.

In complex systems with many degrees of freedom such as peptides and proteins, there exists a huge number of local-minimum-energy states. Conventional simulations in the canonical ensemble are of little use, because they tend to get trapped in states of these energy local minima. A simulation in generalized ensemble performs a random walk in potential energy space and can overcome this difficulty. From only one simulation run, one can obtain canonical-ensemble averages of physical quantities as functions of temperature by the single-histogram and/or multiple-histogram reweighting techniques. In this article we review uses of the generalized-ensemble algorithms in biomolecular systems. Three well-known methods, namely, multicanonical algorithm, simulated tempering, and replica-exchange method, are described first. Both Monte Carlo and molecular dynamics versions of the algorithms are given. We then present three new generalized-ensemble algorithms that combine the merits of the above methods. The effectiveness of the methods for molecular simulations in the protein folding problem is tested with short peptide systems.

Algorithms↗

Modeling Alternative Conformational States in CASP16.

The CASP16 Ensemble Prediction experiment assessed advances in methods for modeling proteins, nucleic acids, and their complexes in multiple conformational states. Targets included systems with experimental structures determined in two or three states, evaluated by direct comparison to experimental coordinates, as well as domain-linker-domain (D-L-D) targets assessed against statistical models from NMR and SAXS data. This paper focuses on the former class of multi-state targets. Ten ensembles were released as community challenges, including ligand-induced conformational changes, protein-DNA complexes, a trimeric protein, a stem-loop RNA, and multiple oligomeric states of a single RNA. For five targets, some groups produced reasonably accurate models of both reference states (best TM-score >0.75). However, with the exception of one protein-ligand complex (T1214), where an apo structure was available as a template, predictors generally failed to capture key structural details distinguishing the states. Overall, accuracy was significantly lower than for single-state targets in other CASP experiments. The most successful approaches generated multiple AlphaFold2 models using enhanced multiple sequence alignments and sampling protocols, followed by model quality based selection. While the AlphaFold3 server performed well on several targets, individual groups outperformed it in specific cases. By contrast, predictions for one protein-DNA complex, three RNA targets, and multiple oligomeric RNA states consistently fell short (TM-score <0.75). These results highlight both progress and persistent challenges in multi-state prediction. Despite recent advances, accurate modeling of conformational ensembles, particularly RNA and large multimeric assemblies, remains a critical frontier for structural biology.

AlphaFold2↗

An evaluation method for the mechanical performance of guide-wires and catheters in accessing the upper urinary tract.

The placement of guide-wires and catheters to gain access to the upper urinary tract can induce undesirable stresses on tissues. Previous studies have characterized the performance of wires and catheters by evaluating their physical properties such as stiffness and friction coefficient. However, the results of these studies do not directly quantify the wire's effects on tissues. Furthermore, the individual physical properties of wires and catheters investigated in previous studies cannot be simply summed up to characterize the behavior of an entire wire/catheter ensemble. This paper presents an objective method for testing guide-wires and catheters that estimates the forces applied by these instruments to anatomical structures during urological procedures. Our model utilizes a computer-controlled test stand that simulates a urological environment by including a tortuous path and a stone obstruction. Experimental results using this model show significant promise in reflecting the performance of guide-wires and catheters measuring the stress exerted upon relevant anatomical structures. Furthermore, due to the modularity of the approach, the model can be easily reconfigured to simulate environments relevant to other medical fields.

Catheterization↗

Determination of a transition state at atomic resolution from protein engineering data.

We present a method for determining the structure of the transition state ensemble (TSE) of a protein by using phi values derived from protein engineering experiments as restraints in molecular dynamics simulations employing a realistic all-atom molecular mechanics energy function. The method uses a biasing potential to select an ensemble of structures having phi values in agreement with the experimental data set. An application to acylphosphatase (AcP), a protein for which phi values have been measured for 24 out of 98 residues, illustrates the approach. The properties of the TSE determined in this way are compared with those of a coarse-grained model obtained using a Monte Carlo (MC) sampling method based on a C(alpha) representation of the structure. The two TSEs determined at different structural resolution are consistent and complementary. While the C(alpha) model allows better sampling of the conformation space occupied by the transition state, the all-atom model offers a more detailed description of the structural and energetic properties of the conformations included in the TSE. The combination of low-resolution C(alpha) results with all-atom molecular dynamics simulations provides a powerful and general method for determining the nature of TSEs from protein engineering data.

Acid Anhydride Hydrolases↗

An efficient protein complex purification method for functional proteomics in higher eukaryotes.

The ensemble of expressed proteins in a given cell is organized in multiprotein complexes. The identification of the individual components of these complexes is essential for their functional characterization. The introduction of the 'tandem affinity purification' (TAP) methodology substantially improved the purification and systematic genome-wide characterization of protein complexes in yeast. The use of this approach in higher eukaryotic cells has lagged behind its use in yeast because the tagged proteins are normally expressed in the presence of the untagged endogenous version, which may compete for incorporation into multiprotein complexes. Here we describe a strategy in which the TAP approach is combined with double-stranded RNA interference (RNAi) to avoid competition from corresponding endogenous proteins while isolating and characterizing protein complexes from higher eukaryotic cells. This strategy allows the determination of the functionality of the tagged protein and increases the specificity and the efficiency of the purification.

Animals↗