Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Using respirators and goggles to control exposure to air pollutants in an anatomy laboratory.

BACKGROUND: Engineering or administrative methods are often insufficient or impractical to control exposure to chemicals in anatomy laboratories. This study explored the feasibility of wearing one or a combination of respirators and goggles used as personal protective equipment (PPE) to control exposure in one such laboratory. METHODS: A group of 28 subjects were briefly trained in wearing PPE, fit-tested, and asked to complete a questionnaire regarding their subjective reaction after wearing the assigned PPE ensemble while working in the laboratory. The subjects' exposure to formaldehyde was also measured and generally exceeded the recommended limits. RESULTS: When a full-face respirator or the combination of a half-mask respirator and goggles was worn, a majority of subjects reported no odor problem and no irritation to eyes or upper respiratory system. Subjects accepted the PPE to certain degrees, but those using respirators encountered difficulties communicating with others. CONCLUSIONS: The combination of a half-mask respirator and goggles was the most feasible ensemble to control exposure to air pollutants in an anatomy laboratory.

Adult↗

Random forest: a classification and regression tool for compound classification and QSAR modeling.

A new classification and regression tool, Random Forest, is introduced and investigated for predicting a compound's quantitative or categorical biological activity based on a quantitative description of the compound's molecular structure. Random Forest is an ensemble of unpruned classification or regression trees created by using bootstrap samples of the training data and random feature selection in tree induction. Prediction is made by aggregating (majority vote or averaging) the predictions of the ensemble. We built predictive models for six cheminformatics data sets. Our analysis demonstrates that Random Forest is a powerful tool capable of delivering performance that is among the most accurate methods to date. We also present three additional features of Random Forest: built-in performance assessment, a measure of relative importance of descriptors, and a measure of compound similarity that is weighted by the relative importance of descriptors. It is the combination of relatively high prediction accuracy and its collection of desired features that makes Random Forest uniquely suited for modeling in cheminformatics.

Journal Article↗

How enzyme dynamics helps catalyze a reaction in atomic detail: a transition path sampling study.

We have applied the Transition Path Sampling algorithm to the reaction catalyzed by the enzyme Lactate Dehydrogenase. This study demonstrates the ease of scaling Transition Path Sampling for applications on many degree of freedom systems, whose energy surface is a complex terrain of valleys and saddle points. As a Monte Carlo importance sampling method, transition path sampling is capable of surmounting barriers in path phase space and focuses simulation on the rare event of enzyme catalyzed atom transfers. Generation of the transition path ensemble, for this reaction, resolves a paradox in the literature in which some studies exposed the catalytic mechanism of hydride and proton transfer by lactate dehydrogenase to be concerted and others stepwise. Transition path sampling has confirmed both mechanisms as possible paths from reactants to products. With the objective to identify a generalized, reduced reaction coordinate, time series of both donor-acceptor distances and residue distances from the active site have been examined. During the transition from pyruvate to lactate, residues located behind the transferring hydride collectively compress toward the active site causing residues located behind the hydride acceptor to relax away. It is demonstrated that an incomplete compression/relaxation transition across the donor-acceptor axis compromises the reaction.

Catalysis↗

Annotation matters: the effect of structural gene annotation on orthology inference.

MOTIVATION: In silico gene annotation, the process of identifying the genes present in a genome, remains a challenging task. As genome assemblies rapidly increase, the corresponding gene models and repertoires often fall short in quality. Despite advances in annotation methods, a lack of community standards means that most published gene annotations result from ad hoc pipelines. As a result, only a few species have nearly complete and accurate gene models. This annotation quality is thought to affect downstream analyses, including orthology inference, often the first step of comparative genomics studies. RESULTS: We show that different annotation methods yield markedly distinct orthology inferences. We compared orthology assignments of gene models obtained by four prominent protein-coding gene model sources: the NCBI Eukaryotic Genome Annotation Pipeline, the Ensembl Gene Annotation System, the UniProt Reference Proteomes, and Augustus 3.4 (an ab initio pipeline). We observe significant discrepancies between sources, namely in the proportion of orthologous genes per genome, the completeness of Hierarchical Orthologous Groups, and the accuracy and recall of the predicted orthologs on a standard orthology benchmark.

Molecular Sequence Annotation↗

Geometric approach to the pressure tensor and the elastic constants.

Expressions are obtained for the pressure tensor in the canonical and the microcanonical ensemble for both isolated and periodic systems, using the same geometric approach to thermodynamic derivatives as has been used previously to define the configurational temperature. The inherent freedom of the method leads to a straightforward proof of the equivalence of atomic and molecular pressures, for short molecules and for molecules exceeding the dimensions of a periodic simulation box. The effect of holonomic constraints on the pressure is discussed. Expressions for the elastic constants are derived in the same manner.

Journal Article↗

Influence of polymer architecture and polymer-wall interaction on the adsorption of polymers into a slit-pore.

The effects of molecular topology and polymer-surface interaction on the properties of isolated polymer chains trapped in a slit were investigated using off-lattice Monte Carlo simulations. Various methods were implemented to allow efficient simulation of molecular structure, confinement force, and free energy for a chain interacting with such "sticky" surfaces. The simulations were performed in the canonical ensemble, and the free energy was sampled via virtual slit-separation moves. Six different chain architectures were studied: linear, star-branched, dendritic, cyclic, two-node (i.e., containing two tetrafunctional intramolecular crosslinks), and six-node molecules. The first three topologies entail increasing degrees of branching, and the last three topologies entail increasing degrees of intramolecular bonding. The confinement force, monomer density profile, and conformational properties for all these systems were compared (for identical molecular weight N) and analyzed as a function of adsorption strength. The compensation point where the wall attraction counterbalances the polymer-slit exclusion effects was the focus of our study. It was found that the attractive energy at the compensation point, epsilon(c), is a weak increasing function of the chain length for excluded-volume chains. The value of epsilon(c) differs significantly for different topologies, and smaller values are associated with better-adsorbing molecules. Due to their globular shape and numerous chain ends, branched molecules (e.g., stars and dendrimers) experience a relatively small entropic penalty for adsorption at low adsorption force and moderate confinement. However, as the adsorption force increases, the more flexible linear chains reach the compensation point at a weaker attractive energy because of the ease with which monomers can be packed near the walls. In moderate to weak confinement, molecules with intramolecular cross-links, such as cyclic, two-node, and six-node molecules, always adsorb better than the other chains (with the same N). Especially at strong adsorption, two-node and six node molecules are highly localized in the region near the walls. Under strong confinement conditions, chain rigidity becomes the dominating factor and the more flexible linear chain adsorbs the best at all adsorption strengths. These results provide useful insights for controlling confinement and depletion forces of polymers with different molecular architectures in the presence of attractive polymer-surface interactions.

Adsorption↗

Temporal correlations and neural spike train entropy.

Sampling considerations limit the experimental conditions under which information theoretic analyses of neurophysiological data yield reliable results. We develop a procedure for computing the full temporal entropy and information of ensembles of neural spike trains, which performs reliably for limited samples of data. This approach also yields insight to the role of correlations between spikes in temporal coding mechanisms. The method, when applied to recordings from complex cells of the monkey primary visual cortex, results in lower rms error information estimates in comparison to a "brute force" approach.

Action Potentials↗

Limited flexibility of lactose detected from residual dipolar couplings using molecular dynamics simulations and steric alignment methods.

The conformational flexibility of lactose in solution has been investigated by residual dipolar couplings (RDCs). One-bond carbon-proton and proton-proton coupling constants have been measured in two oriented media and interpreted in combination with molecular dynamics simulations (MD). Two different approaches, known as PALES (Zweckstetter et al., J. Am. Chem. Soc. 2000, 122, 3791-3792) and TRAMITE (Azurmendi et al., J. Am. Chem. Soc. 2002, 124, 2426-2427), have been used to determine the alignment tensor from a shape-induced alignment model with the oriented medium. The steric alignment of the structures from several MD trajectories has provided ensemble averaged RDCs that have been compared with the experimental ones. The obtained results reveal the almost exclusive presence of a major low energy region defined as syn-phi/syn-psi (> 97%), for which sampling occurs in a dynamic manner. This result satisfactorily agrees with that determined by standard NOE-based methods.

Carbohydrate Conformation↗

A method for the rapid exchange of solutions bathing excised membrane patches.

In this communication we describe a technique for rapidly exchanging solutions bathing excised membrane patches, and present examples of its implementation using both outside-out and inside-out patches. The ability to make step changes in the concentration of channel-activating ligands (e.g., acetylcholine, calcium) offers a novel and direct means of measuring kinetic processes in the 10-100-ms range. The responses to step ligand concentration changes are well suited to ensemble variance analysis, yielding estimates of the number of channels in a patch, and testing assumptions of channel independence and homogeneity. Kinetic analysis of the pseudomacroscopic currents obtained by averaging large numbers of responses can be compared and correlated with analysis of the microscopic behavior of single channels, using the same membrane patch for both approaches. Practical and theoretical limitations associated with the method are briefly discussed.

Animals↗

Ensemble-based convergence analysis of biomolecular trajectories.

Assessing the convergence of a biomolecular simulation is an essential part of any careful computational investigation, because many fundamental aspects of molecular behavior depend on the relative populations of different conformers. Here we present a physically intuitive method to self-consistently assess the convergence of trajectories generated by molecular dynamics and related methods. Our approach reports directly and systematically on the structural diversity of a simulation trajectory. Straightforward clustering and classification steps are the key ingredients, allowing the approach to be trivially applied to systems of any size. Our initial study on met-enkephalin strongly suggests that even fairly long trajectories (approximately 50 ns) may not be converged for this small--but highly flexible--system.

Biopolymers↗

Amplitudes and directions of internal protein motions from a JAM analysis of 15N relaxation data.

A method has been developed for characterizing dynamic structures of proteins in solution by using nuclear magnetic resonance (NMR) restraints and 15N relaxation data. This method is based on the concept of the jumping-among-minima (JAM) model. In this model we assume that protein dynamics can be described on the basis of conformational substates, and involves intra- and inter-substate motion. A set of substates is created by picking energy-minimized conformations from the conformational space consistent with the geometric NMR restraints. Intra-substate motions, which occur on the timescale of approximately 10 ps, are simulated with molecular dynamics (MD) calculations with force-field energy terms. Statistical weights of the conformational substates are determined to reproduce the NMR relaxation parameters. The refinement procedure consists of four stages: (i) determination of the ensemble of structures that satisfy NMR restraints, (ii) determination of intra-substate fluctuation, (iii) determination of statistical weights of conformational substates to reproduce model-free relaxation parameters, and (iv) analysis of the resulting dynamic structure to determine amplitudes and directions of internal protein motions. This method was employed to investigate structure and dynamics of the adhesion domain of human CD2 (hCD2) in solution. Two major collective modes, whose contributions to atomic mean-square fluctuations are 77.1% in total, are identified by the refinement. The first mode is interpreted as a rigid-body motion of a protein segment consisting of a part of the B--C loop, a part of the F strand, and the F--G loop. Another type of smaller-amplitude mode is indicated for the C'--C'' loop. The motions affect primarily the curvature of the slightly concave counterreceptor-binding site and represent transitions between a concave (closed) and flat (open) binding face. By comparing the ensemble of structures in solution to the complex structure with counterreceptor CD58, we found that these two types of motions resemble the change upon counterreceptor binding.

Models, Molecular↗

Folding of a model three-helix bundle protein: a thermodynamic and kinetic analysis.

The kinetics and thermodynamics of an off-lattice model for a three-helix bundle protein are investigated as a function of a bias gap parameter that determines the energy difference between native and non-native contacts. A simple dihedral potential is used to introduce the tendency to form right-handed helices. For each value of the bias parameter, 100 trajectories of up to one microsecond are performed. Such statistically valid sampling of the kinetics is made possible by the use of the discrete molecular dynamics method with square-well interactions. This permits much faster simulations for off-lattice models than do continuous potentials. It is found that major folding pathways can be defined, although ensembles with considerable structural variation are involved. The large gap models generally fold faster than those with a smaller gap. For the large gap models, the kinetic intermediates are non-obligatory, while both obligatory and non-obligatory intermediates are present for small gap models. Certain large gap intermediates have a two-helix microdomain with one helix extended outward (as in domain-swapped dimers); the small gap intermediates have more diverse structures. The importance of studying the kinetic, as well as the thermodynamics, of folding for an understanding of the mechanism is discussed and the relation between kinetic and equilibrium intermediates is examined. It is found that the behavior of this model system has aspects that encompass both the "new" view and the "old" view of protein folding.

Algorithms↗

Diffusional and compartmental models for tracer washout records of Na in dog carotid.

This paper re-examines studies of Na kinetics in canine carotid arterial wall previously reported by three groups of investigators. The similarities and differences in experimental and analytical approaches for Na arterial wall washout therein are reviewed. Major conclusions are: (1) A three-compartment model in series is adequate for all short and long (after subtracting the small slow exponential) records. (2) Values for the diffusion coefficient, however, varied markedly among the different groups. (3) A continuous sampling method as opposed to an intermittent one measures better the kinetics of a rapidly exchanging ion as Na+. (4) Models based on individual washout records are better than those based on ensemble averaging. Characterization of sets of results in terms of mean values for comparing parameters provides significant potential for distinguishing quantitative features between control and treated sets despite inadequacies in data sampling or model specifications.

Animals↗

Multicanonical schemes for mapping out free-energy landscapes of single-component and multicomponent systems.

Multicanonical (MUCA) sampling is a powerful approach for simulating large domains of thermodynamic macrostate space that relies on mapping out either the density of states or a free energy of the system as a function of a suitable "order parameter." The purpose of this study is to extend and apply to more complex systems the method introduced in a previous paper [M. K. Fenwick and F. A. Escobedo, J. Chem. Phys. 120, 3066 (2004)] that uses Bennett's acceptance ratio method for estimating MUCA free energies. Four types of MUCA schemes are considered according to what order parameter is adopted and how the macrostate space is traversed: a la grand canonical ensemble, a la semigrand canonical ensemble, a la semigrand isothermal-isobaric ensemble, and a la isothermal-isobaric ensemble. Two types of systems are studied, the first is a two-component Lennard-Jones mixture that exhibits a vapor-liquid transition, and the second is a hard-cuboid containing system that exhibits an isotropic-liquid crystalline transition. These systems are simulated with different MUCA schemes and the resulting free-energy profiles are used to determine phase-coexistence conditions. For the Lennard-Jones systems, it is also demonstrated that different types of MUCA simulations can be conveniently performed over different macrostate regions and the results can be subsequently pieced together into a continuous weighting function.

Journal Article↗

Unspecific hydrophobic stabilization of folding transition states.

Here we present a method for determining the inference of non-native conformations in the folding of a small domain, alpha-spectrin Src homology 3 domain. This method relies on the preservation of all native interactions after Tyr/Phe exchanges in solvent-exposed, contact-free positions. Minor changes in solvent exposure and free energy of the denatured ensemble are in agreement with the reverse hydrophobic effect, as the Tyr/Phe mutations slightly change the polypeptide hydrophilic/hydrophobic balance. Interestingly, more important Gibbs energy variations are observed in the transition state ensemble (TSE). Considering the small changes induced by the H/OH replacements, the observed energy variations in the TSE are rather notable, but of a magnitude that would remain undetected under regular mutations that alter the folded structure free energy. Hydrophobic residues outside of the folding nucleus contribute to the stability of the TSE in an unspecific nonlinear manner, producing a significant acceleration of both unfolding and refolding rates, with little effect on stability. These results suggest that sectors of the protein transiently reside in non-native areas of the landscape during folding, with implications in the reading of phi values from protein engineering experiments. Contrary to previous proposals, the principle that emerges is that non-native contacts, or conformations, could be beneficial in evolution and design of some fast folding proteins.

Biophysics↗

SSAHA: a fast search method for large DNA databases.

We describe an algorithm, SSAHA (Sequence Search and Alignment by Hashing Algorithm), for performing fast searches on databases containing multiple gigabases of DNA. Sequences in the database are preprocessed by breaking them into consecutive k-tuples of k contiguous bases and then using a hash table to store the position of each occurrence of each k-tuple. Searching for a query sequence in the database is done by obtaining from the hash table the "hits" for each k-tuple in the query sequence and then performing a sort on the results. We discuss the effect of the tuple length k on the search speed, memory usage, and sensitivity of the algorithm and present the results of computational experiments which show that SSAHA can be three to four orders of magnitude faster than BLAST or FASTA, while requiring less memory than suffix tree methods. The SSAHA algorithm is used for high-throughput single nucleotide polymorphism (SNP) detection and very large scale sequence assembly. Also, it provides Web-based sequence search facilities for Ensembl projects.

Algorithms↗

Designing protein beta-sheet surfaces by Z-score optimization.

Studies of lattice models of proteins have suggested that the appropriate energy expression for protein design may include nonthermodynamic terms to accommodate negative design concerns. One method, developed in lattice model studies, maximizes a quantity known as the " Z-score," which compares the lowest energy sequence whose ground state structure is the target structure to an ensemble of random sequences. Here we show that, in certain circumstances, the technique can be applied to real proteins. The resulting energy expression is used to design the beta-sheet surfaces of two real proteins. We find experimentally that the designed proteins are stable and well folded, and in one case is even more thermostable than the wild type.

Models, Chemical↗

Simulation of phase transitions in highly asymmetric fluid mixtures.

We present a novel method for the accurate numerical determination of the phase behavior of fluid mixtures having large particle-size asymmetries. By incorporating the recently developed geometric cluster algorithm within a restricted Gibbs ensemble, we are able to probe directly the density and concentration fluctuations that drive phase transitions, but that are inaccessible to conventional simulation algorithms. We develop a finite-size scaling theory that relates these density fluctuations to those of the grand-canonical ensemble, thereby enabling accurate location of critical points and coexistence curves of multicomponent fluids. Several illustrative examples are presented.

Journal Article↗