Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Mesoscale simulation of polymer reaction equilibrium: combining dissipative particle dynamics with reaction ensemble Monte Carlo. I. Polydispersed polymer systems.

We present a mesoscale simulation technique, called the reaction ensemble dissipative particle dynamics (RxDPD) method, for studying reaction equilibrium of polymer systems. The RxDPD method combines elements of dissipative particle dynamics (DPD) and reaction ensemble Monte Carlo (RxMC), allowing for the determination of both static and dynamical properties of a polymer system. The RxDPD method is demonstrated by considering several simple polydispersed homopolymer systems. RxDPD can be used to predict the polydispersity due to various effects, including solvents, additives, temperature, pressure, shear, and confinement. Extensions of the method to other polymer systems are straightforward, including grafted, cross-linked polymers, and block copolymers. To simulate polydispersity, the system contains full polymer chains and a single fractional polymer chain, i.e., a polymer chain with a single fractional DPD particle. The fractional particle is coupled to the system via a coupling parameter that varies between zero (no interaction between the fractional particle and the other particles in the system) and one (full interaction between the fractional particle and the other particles in the system). The time evolution of the system is governed by the DPD equations of motion, accompanied by changes in the coupling parameter. The coupling-parameter changes are either accepted with a probability derived from the grand canonical partition function or governed by an equation of motion derived from the extended Lagrangian. The coupling-parameter changes mimic forward and reverse reaction steps, as in RxMC simulations.

Journal Article↗

Dissipative particle dynamics simulations in the grand canonical ensemble: applications to polymer brushes.

We have used the dissipative particle dynamics (DPD) method in the grand canonical ensemble to study the compression of grafted polymer brushes in good solvent conditions. The force-distance profiles calculated from DPD simulations in the grand canonical ensemble are in very good agreement with the self-consistent field (SCF) theoretical models and with experimental results for two polystyrene brush layers grafted onto mica surfaces in toluene.

Journal Article↗

Support vector machines committee classification method for computer-aided polyp detection in CT colonography.

RATIONALE AND OBJECTIVES: A new classification scheme for the computer-aided detection of colonic polyps in computed tomographic colonography is proposed. MATERIALS AND METHODS: The scheme involves an ensemble of support vector machines (SVMs) for classification, a smoothed leave-one-out (SLOO) cross-validation method for obtaining error estimates, and use of a bootstrap aggregation method for training and model selection. Our use of an ensemble of SVM classifiers with bagging (bootstrap aggregation), built on different feature subsets, is intended to improve classification performance compared with single SVMs and reduce the number of false-positive detections. The bootstrap-based model-selection technique is used for tuning SVM parameters. In our first experiment, two independent data sets were used: the first, for feature and model selection, and the second, for testing to evaluate the generalizability of our model. In the second experiment, the test set that contained higher resolution data was used for training and testing (using the SLOO method) to compare SVM committee and single SVM performance. RESULTS: The overall sensitivity on independent test set was 75%, with 1.5 false-positive detections/study, compared with 76%-78% sensitivity and 4.5 false-positive detections/study estimated using the SLOO method on the training set. The sensitivity of the SVM ensemble retrained on the former test set estimated using the SLOO method was 81%, which is 7%-10% greater than the sensitivity of a single SVM. The number of false-positive detections per study was 2.6, a 1.5 times reduction compared with a single SVM. CONCLUSION: Training an SVM ensemble on one data set and testing it on the independent data has shown that the SVM committee classification method has good generalizability and achieves high sensitivity and a low false-positive rate. The model selection and improved error estimation method are effective for computer-aided polyp detection.

Algorithms↗

Foldamer simulations: novel computational methods and applications to poly-phenylacetylene oligomers.

We apply several methods to probe the ensemble kinetic and structural properties of a model system of poly-phenylacetylene (pPA) oligomer folding trajectories. The kinetic methods employed included a brute force accounting of conformations, a Markovian state matrix method, and a nonlinear least squares fit to a minimalist kinetic model used to extract the folding time. Each method gave similar measures for the folding time of the 12-mer chain, calculated to be on the order of 7 ns for the complete folding of the chain from an extended conformation. Utilizing both a linear and a nonlinear scaling relationship between the viscosity and the folding time to correct for a low simulation viscosity, we obtain an upper and a lower bound for the approximate folding time within the range 70 ns<tau<350 ns. This is in agreement with the experimentally measured folding time on the order of 160 ns. The kinetic model used to fit the kinetic behavior of the ensemble of trajectories provides a framework to describe the bulk folding mechanism. We were able to identify two unique clusters of conformations that provide a structural basis to account for the appearance of a kinetic intermediate in the mechanism. We discuss the implications of these findings in the context of helix-coil theory.

Journal Article↗

Systemic properties of ensembles of metabolic networks: application of graphical and statistical methods to simple unbranched pathways.

MOTIVATION: Mathematical models are the only realistic method for representing the integrated dynamic behavior of complex biochemical networks. However, it is difficult to obtain a consistent set of values for the parameters that characterize such a model. Even when a set of parameter values exists, the accuracy of the individual values is questionable. Therefore, we were motivated to explore statistical techniques for analyzing the properties of a given model when knowledge of the actual parameter values is lacking. RESULTS: The graphical and statistical methods presented in the previous paper are applied here to simple unbranched biosynthetic pathways subject to control by feedback inhibition. We represent these pathways within a canonical nonlinear formalism that provides a regular structure that is convenient for randomly sampling the parameter space. After constructing a large ensemble of randomly generated sets of parameter values, the structural and behavioral properties of the model with these parameter sets are examined statistically and classified. The results of our analysis demonstrate that certain properties of these systems are strongly correlated, thereby revealing aspects of organization that are highly probable independent of selection. Finally, we show how specification of a given behavior affects the distribution of acceptable parameter values.

Amino Acids↗

Isomolar semigrand ensemble molecular dynamics: development and application to liquid-liquid equilibria.

An extended system molecular dynamics method for the isomolar semigrand ensemble (fixed number of particles, pressure, temperature, and fugacity fraction) is developed and applied to the calculation of liquid-liquid equilibria (LLE) for two Lennard-Jones mixtures. The method utilizes an extended system variable to dynamically control the fugacity fraction xi of the mixture by gradually transforming the identity of particles in the system. Two approaches are used to compute coexistence points. The first approach uses multiple-histogram reweighting techniques to determine the coexistence xi and compositions of each phase at temperatures near the upper critical solution temperature. The second approach, useful for cases in which there is no critical solution temperature, is based on principles of small system thermodynamics. In this case a coexistence point is found by running N-P-T-xi simulations at a common temperature and pressure and varying the fugacity fraction to map out the difference in chemical potential between the two species A and B (mu(A)-mu(B)) as a function of composition. Once this curve is known the equal-distance/equal-area criterion is used to determine the coexistence point. Both approaches give results that are comparable to those of previous Monte Carlo (MC) simulations. By formulating this approach in a molecular dynamics framework, it should be easier to compute the LLE of complex molecules whose intramolecular degrees of freedom are often difficult to properly sample with MC techniques.

Journal Article↗

An ensemble of K-local hyperplanes for predicting protein-protein interactions.

Prediction of protein-protein interaction is a difficult and important problem in biology. In this paper, we propose a new method based on an ensemble of K-local hyperplane distance nearest neighbor (HKNN) classifiers, where each HKNN is trained using a different physicochemical property of the amino acids. Moreover, we propose a new encoding technique that combines the amino acid indices together with the 2-Grams amino acid composition. A fusion of HKNN classifiers combined with the 'Sum rule' enables us to obtain an improvement over other state-of-the-art methods. The approach is demonstrated by building a learning system based on experimentally validated protein-protein interactions in human gastric bacterium Helicobacter pylori and in Human dataset.

Algorithms↗

Generation of initial trajectories for transition path sampling of chemical reactions with ab initio molecular dynamics.

Transition path sampling is an innovative method for focusing a molecular dynamics simulation on a reactive event. Although transition path sampling methods can generate an ensemble of reactive trajectories, an initial reactive trajectory must be generated by some other means. In this paper, the authors have evaluated three methods for generating initial reactive trajectories for transition path sampling with ab initio molecular dynamics. The authors have tested each of these methods on a set of chemical reactions involving the breaking and making of covalent bonds: the 1,2-hydrogen elimination in the borane-ammonia adduct, a tautomerization, and the Claisen rearrangement. The first method is to initiate trajectories from the potential energy transition state, which was effective for all reactions in the test set. Assigning atomic velocities found using normal mode analysis greatly improved the success of this method. The second method uses a high temperature molecular dynamics simulation and then iteratively reduces the total energy of the simulation until a low temperature reactive trajectory is found. This was effective in generating a low temperature trajectory from an initial trajectory run at 3000 K of the tautomerization reaction, although it failed for the other two. The third uses an orbital based bias potential to find a reactive trajectory and uses this trajectory to initiate an unbiased trajectory. The authors found that a highest occupied molecular orbital-lowest unoccupied molecular orbital bias could be used to find a reactive trajectory for the Claisen rearrangement, although it failed for the other two reactions. These techniques will help make it practical to use transition path sampling to study chemical reaction mechanisms that involve bond breaking and forming.

Journal Article↗

A late-stopping method for optimal aggregation of neural networks.

Ensembles of artificial neural networks have been used in the last years as classification/regression machines, showing improved generalization capabilities that outperform those of single networks. However, it has been recognized that for aggregation to be effective the individual networks must be as accurate and diverse as possible. An important problem is, then, how to tune the aggregate members in order to have an optimal compromise between these two conflicting conditions. We propose here a simple method for constructing regression/classification ensembles of neural networks that leads to overtrained aggregate members with an adequate balance between accuracy and diversity. The algorithm is favorably tested against other methods recently proposed in the literature, producing an improvement in performance on the standard statistical databases used as benchmarks. In addition, and as a concrete application, we apply our method to the sunspot time series and predict the remainder of the current cycle 23 of solar activity.

Data Collection↗

Monte Carlo sampling of near-native structures of proteins with applications.

Since a protein's dynamic fluctuation inside cells affects the protein's biological properties, we present a novel method to study the ensemble of near-native structures (NNS) of proteins, namely, the conformations that are very similar to the experimentally determined native structure. We show that this method enables us to (i) quantify the difficulty of predicting a protein's structure, (ii) choose appropriate simplified representations of protein structures, and (iii) assess the effectiveness of knowledge-based potential functions. We found that well-designed simple representations of protein structures are likely as accurate as those more complex ones for certain potential functions. We also found that the widely used contact potential functions stabilize NNS poorly, whereas potential functions incorporating local structure information significantly increase the stability of NNS.

Computer Simulation↗

Evaluation of Interaction Forces between Macroparticles in Simple Fluids by Molecular Dynamics Simulation.

The present article provides the description of the solvation forces between large spheres in a fluid. The molecular dynamics (MD) method was applied to the relatively simple systems in which a pair of structureless macroparticles, either solvophobic or solvophilic, is immersed in a simple fluid of two types, either a soft-sphere or a Lennard-Jones fluid. When a pair of solvophobic macroparticles was in the attractive Lennard-Jones fluid, no dense layer of the solvent particles formed near the surface of the macroparticles and the strong attractive forces were induced between them. In the other combinations of macroparticles and fluids, the dense layers formed and the solvation forces oscillated, exhibiting the attraction and repulsion, whose periodic distance was about the diameter of solvent particles. Our results agreed well with those of the other simulation and theoretical studies with respect to the solvent density profile near a macroparticle and the force-distance profile between macroparticles. The benefit of our approach would be the simplicity in specifying or finding the bulk condition that is in equilibrium with the thin film of molecules between large surfaces. The present method can be applied straightforward to macroparticles immersed in mixtures and complex fluids described by the bead-spring model, to which the conventional grand canonical ensemble Monte Carlo (GCEMC) method is hardly accessible. Copyright 1999 Academic Press.

Journal Article↗

Inequivalence of pure state ensembles for open quantum systems: the preferred ensembles are those that are physically realizable.

An open quantum system in steady state rho(ss) can be represented by a weighted ensemble of pure states rho(ss) = [equation: see text] in infinitely many ways. A physically realizable (PR) ensemble is one for which some continuous measurement of the environment will collapse the system into a pure state /psi(t)>, stochastically evolving such that the proportion of time for which /psi(t)> = /psi(k)> equals Weierstrass p(k). Some, but not all, ensembles are PR. This constitutes the preferred ensemble fact. We present the necessary and sufficient conditions for a given ensemble to be PR, and illustrate the method by showing that the coherent state ensemble is not PR for an atom laser.

Journal Article↗

Noninvasive optical imaging by speckle ensemble.

We propose a new method imaging through scattering media. An object hidden between two biological tissues (chicken breast) is reconstructed from any speckled images obtained from the output of a multichannel optical imaging system. The effect of multiple imaging is achieved with a microlens array. Each lens is the array projects a different speckled image onto a digital camera. The set of speckled images from the entire array is first shifted to a common center and then accumulated into a single average picture.

Animals↗

Searching sequence space to engineer proteins: exponential ensemble mutagenesis.

We describe an efficient method for generating combinatorial libraries with a high percentage of unique and functional mutants. Combinatorial libraries have been successfully used in the past to express ensembles of mutant proteins in which all possible amino acids are encoded at a few positions in the sequence. However, as more positions are mutagenized the proportion of functional mutants is expected to decrease exponentially. Small groups of residues were randomized in parallel to identify, at each altered position, amino acids which lead to functional proteins. By using optimized nucleotide mixtures deduced from the sequences selected from the random libraries, we have simultaneously altered 16 sites in a model pigment binding protein: approximately one percent of the observed mutants were functional. Mathematical formalization and extrapolation of our experimental data suggests that a 10(7)-fold increase in the throughput of functional mutants has been obtained relative to the expected frequency from a random combinatorial library. Exponential ensemble mutagenesis should be advantageous in cases where many residues must be changed simultaneously to achieve a specific engineering goal, as in the combinatorial mutagenesis of phage displayed antibodies. With the enhanced functional mutant frequencies obtained by this method, entire proteins could be mutagenized combinatorially.

Amino Acid Sequence↗

Prediction of solvent accessibility and sites of deleterious mutations from protein sequence.

Residues that form the hydrophobic core of a protein are critical for its stability. A number of approaches have been developed to classify residues as buried or exposed. In order to optimize the classification, we have refined a suite of five methods over a large dataset and proposed a metamethod based on an ensemble average of the individual methods, leading to a two-state classification accuracy of 80%. Many studies have suggested that hydrophobic core residues are likely sites of deleterious mutations, so we wanted to see to what extent these sites can be predicted from the putative buried residues. Residues that were most confidently classified as buried were proposed as sites of deleterious mutations. This proposition was tested on six proteins for which sites of deleterious mutations have previously been identified by stability measurement or functional assay. Of the total of 130 residues predicted as sites of deleterious mutations, 104 (or 80%) were correct.

Amino Acid Sequence↗

Equation-free dynamic renormalization of a Kardar-Parisi-Zhang-type equation.

In the context of equation-free computation, we devise and implement a procedure for using short-time direct simulations of a Kardar-Parisi-Zhang-(KPZ-) type equation to calculate the self-similar solution for its ensemble averaged correlation function. The method involves "lifting" from candidate pair-correlation functions to consistent realization ensembles, short bursts of KPZ-type evolution, and appropriate rescaling of the resulting averaged pair correlation functions. Both the self-similar shapes and their similarity exponents are obtained at a computational cost significantly reduced to that required to reach saturation in such systems.

Journal Article↗

An ENSEMBLE machine learning approach for the prediction of all-alpha membrane proteins.

MOTIVATION: All-alpha membrane proteins constitute a functionally relevant subset of the whole proteome. Their content ranges from about 10 to 30% of the cell proteins, based on sequence comparison and specific predictive methods. Due to the paucity of membrane proteins solved with atomic resolution, the training/testing sets of predictive methods for protein topography and topology routinely include very few well-solved structures mixed with a hundred proteins known with low resolution. Moreover, available predictors fail in predicting recently crystallised membrane proteins (Chen et al., 2002). Presently the number of well-solved membrane proteins comprises some 59 chains of low sequence homology. It is therefore possible to train/test predictors only with the set of proteins known with atomic resolution and evaluate more thoroughly the performance of different methods. RESULTS: We implement a cascade-neural network (NN), two different hidden Markov models (HMM), and their ensemble (ENSEMBLE) as a new method. We train and test in cross validation the three methods and ENSEMBLE on the 59 well resolved membrane proteins. ENSEMBLE scores with a per-protein accuracy of 90% for topography and 71% for topology, outperforming the best single method of 7 and 5 percentage points, respectively. When tested on a low resolution set of 151 proteins, with no homology with the 59 proteins, the per-protein accuracy of ENSEMBLE is 76% for topography and 68% for topology. Our results also indicate that the performance of ENSEMBLE is higher than that of the best predictors presently available on the Web.

Algorithms↗

Grand canonical ensemble Monte Carlo simulation of the dCpG/proflavine crystal hydrate.

The grand canonical ensemble Monte Carlo molecular simulation method is used to investigate hydration patterns in the crystal hydrate structure of the dCpG/proflavine intercalated complex. The objective of this study is to show by example that the recently advocated grand canonical ensemble simulation is a computationally efficient method for determining the positions of the hydrating water molecules in protein and nucleic acid structures. A detailed molecular simulation convergence analysis and an analogous comparison of the theoretical results with experiments clearly show that the grand ensemble simulations can be far more advantageous than the comparable canonical ensemble simulations.

Binding Sites↗