Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Entanglement assisted metrology.

We propose a new approach to the measurement of a single spin state, based on nuclear magnetic resonance (NMR) techniques and inspired by the coherent control over many-body systems envisaged by quantum information processing. A single target spin is coupled via the magnetic dipolar interaction to a large ensemble of spins. Applying radio frequency pulses, we can control the evolution so that the spin ensemble reaches one of two orthogonal states whose collective properties differ depending on the state of the target spin and are easily measured. We first describe this measurement process using quantum gates; then we show how equivalent schemes can be defined in terms of the Hamiltonian and thus implemented under conditions of real control, using well established NMR techniques. We demonstrate this method with a proof of principle experiment in ensemble liquid state NMR and simulations for small spin systems.

Journal Article↗

Simulating rare events in equilibrium or nonequilibrium stochastic systems.

We present three algorithms for calculating rate constants and sampling transition paths for rare events in simulations with stochastic dynamics. The methods do not require a priori knowledge of the phase-space density and are suitable for equilibrium or nonequilibrium systems in stationary state. All the methods use a series of interfaces in phase space, between the initial and final states, to generate transition paths as chains of connected partial paths, in a ratchetlike manner. No assumptions are made about the distribution of paths at the interfaces. The three methods differ in the way that the transition path ensemble is generated. We apply the algorithms to kinetic Monte Carlo simulations of a genetic switch and to Langevin dynamics simulations of intermittently driven polymer translocation through a pore. We find that the three methods are all of comparable efficiency, and that all the methods are much more efficient than brute-force simulation.

Journal Article↗

Exploring reaction pathways with transition path and umbrella sampling: application to methyl maltoside.

The transition path sampling (TPS) method is a powerful approach to study chemical reactions or transitional properties on complex potential energy landscapes. One of the main advantages of the method over potential of mean force methods is that reaction rates can be directly accessed without knowledge of the exact reaction coordinate. We have investigated the complementary nature of these two differing approaches, comparing transition path sampling with the weighted histogram analysis method to study a conformational change in a small model system. In this case study, the transition paths for a transition between two rotational conformers of a model disaccharide molecule, methyl beta-D-maltoside, were compared with a free energy surface constrained by the two commonly used glycosidic (phi,psi) torsional angles. The TPS method revealed a reaction channel that was not apparent from the potential of mean force method, and the suitability of phi and psi as reaction coordinates to describe the isomerization in vacuo was confirmed by examination of the transition path ensemble. Using both transition state theory and transition path sampling methods, the transition rate was estimated. We have estimated a characteristic time between transitions of approximately 160 ns for this rare isomerization event between the two conformations of the carbohydrate. We conclude that transition path sampling can extract subtle information about the dynamics not apparent from the potential of mean force method. However, in calculating the reaction rate, the transition path sampling method required 27.5 times the computational effort than was needed by the potential of mean force method.

Algorithms↗

[Prediction of the ensembles of RNA secondary structures. Kinetic analysis of self-organization].

A new approach to the problem of prediction of secondary structures of RNA, which is based on the kinetic analysis of self-organising molecules is proposed. Structural reconstructions that take place during formation of secondary structures are described in terms of Markov process. A set of states and probability transition were defined. Monte-Carlo methods were used to describe this process. Probability distributions of various secondary structures depending on time are given. Examples of calculations for ensembles of secondary structures of some tRNAs are described. An effective method of steady-state ensemble research, which is based on a quick RESETTING of all possible variance of the secondary structures of RNAs is given. By ascribing to each of these structures the value of probabilities as a function of free energy it was possible to obtain the Boltzmann ensemble of secondary structures.

Escherichia coli↗

RNA secondary structure prediction by centroids in a Boltzmann weighted ensemble.

Prediction of RNA secondary structure by free energy minimization has been the standard for over two decades. Here we describe a novel method that forsakes this paradigm for predictions based on Boltzmann-weighted structure ensemble. We introduce the notion of a centroid structure as a representative for a set of structures and describe a procedure for its identification. In comparison with the minimum free energy (MFE) structure using diverse types of structural RNAs, the centroid of the ensemble makes 30.0% fewer prediction errors as measured by the positive predictive value (PPV) with marginally improved sensitivity. The Boltzmann ensemble can be separated into a small number (3.2 on average) of clusters. Among the centroids of these clusters, the "best cluster centroid" as determined by comparison to the known structure simultaneously improves PPV by 46.5% and sensitivity by 21.7%. For 58% of the studied sequences for which the MFE structure is outside the cluster containing the best centroid, the improvements by the best centroid are 62.5% for PPV and 31.4% for sensitivity. These results suggest that the energy well containing the MFE structure under the current incomplete energy model is often different from the one for the unavailable complete model that presumably contains the unique native structure. Centroids are available on the Sfold server at http://sfold.wadsworth.org.

Base Pairing↗

Multi-class protein fold classification using a new ensemble machine learning approach.

Protein structure classification represents an important process in understanding the associations between sequence and structure as well as possible functional and evolutionary relationships. Recent structural genomics initiatives and other high-throughput experiments have populated the biological databases at a rapid pace. The amount of structural data has made traditional methods such as manual inspection of the protein structure become impossible. Machine learning has been widely applied to bioinformatics and has gained a lot of success in this research area. This work proposes a novel ensemble machine learning method that improves the coverage of the classifiers under the multi-class imbalanced sample sets by integrating knowledge induced from different base classifiers, and we illustrate this idea in classifying multi-class SCOP protein fold data. We have compared our approach with PART and show that our method improves the sensitivity of the classifier in protein fold classification. Furthermore, we have extended this method to learning over multiple data types, preserving the independence of their corresponding data sources, and show that our new approach performs at least as well as the traditional technique over a single joined data source. These experimental results are encouraging, and can be applied to other bioinformatics problems similarly characterised by multi-class imbalanced data sets held in multiple data sources.

Amino Acid Sequence↗

Highly potent side-chain to side-chain cyclized enkephalin analogues containing a carbonyl bridge: synthesis, biology and conformation.

Six novel cyclic enkephalin analogues have been synthesized. Cyclization of the linear peptides containing basic amino acid residues in position 2 and 5 was achieved by treatment with bis(4-nitrophenyl)carbonate. It was found that some of the compounds exibit unusually high mu-opioid activity in the guinea pig ileum (GPI) assay. The 18-membered analogue cyclo(N(epsilon),N(beta)-carbonyl-D-Lys2,Dap5)-enkephalinamide turned out to be one of the most potent mu-agonists reported so far. NMR spectra of the peptides were recorded and structural parameters were determined. The conformational space was exhaustively examined for each of them using the electrostatically driven Monte Carlo method. Each peptide was finally described as an ensemble of conformations. A model of the bioactive conformation of this class of opioid peptides was proposed.

Animals↗

A method for studying optical anisotropy of polymers as a function of molar mass.

Optical properties of polymers are extremely important in many end-use applications, as is the ability of anisotropic polymers to depolarize incident radiation. To date, most light-scattering studies of the optical anisotropy of macromolecules have dealt with the bulk state or measured ensemble properties of dilute solutions. Here, we introduce a method to determine the optical anisotropy as a continuous function of molar mass. By direct, on-line coupling of size-exclusion chromatography and depolarization multiangle light scattering (SEC/D-MALS), molar mass averages, polydispersities, molar mass distributions, and the distribution of the optical anisotropy as a function of molar mass may all be determined. To quantify the anisotropy, it has been expressed in terms of the Cabannes factor, thus permitting the Rayleigh ratio necessary for light-scattering calculations to be corrected for anisotropy. The effects of tacticity, heavy atom substitution on the main chain, and chain helicity on the depolarization behavior of polymers have been studied using atactic and isotactic PMMA; atactic and brominated PS; and the semiflexible polypeptide PBLG, which maintains an extended structure in solution. An introduction to the theory of SEC/D-MALS is given.

Journal Article↗

Conformational variability of solution nuclear magnetic resonance structures.

In structure determination by X-ray crystallography and solution NMR spectroscopy, experimental data are collected as time and ensemble-averages. Thus, in principle, appropriate time and ensemble-averaged models should be used. Refinement of an ensemble of conformers rather than one single structure against the experimental NMR data could, however, result in overfitting the data because of the significantly increased number of parameters. To avoid overfitting, complete cross-validation, which provides an unbiased measure of the fit, has been applied to nuclear Overhauser effect derived distance refinement. Using two synthetic test cases, a correlation was demonstrated between the cross-validated measure to the fit (defined in terms of root-mean-square deviations from the distance restraints and number of violations) and the number of models that best reproduce the conformational variability in solution. A new method, based on a probability map, has been used to generate good representations of the resulting ensembles of structures. The method has also been applied to observed NMR data for two proteins, interleukin 4 and interleukin 8. For interleukin 4, cross-validation indicates that a single-conformer model gives the most accurate representation of the structure, whereas conventional measures of fit between the experimental data and those calculated from the model decrease when increasing the number of conformers, indicating overfitting. For interleukin 8, complete cross-validation predicts a twin-conformer model to be the most faithful representation of the experimental data. Two distinct conformations for the loop formed by residues 16 to 22 emerge from the family of twin-conformer structures. The putative alternate conformation of the loop is not observed in the crystal structure of interleukin 8. However, because of crystal packing contacts in this region this does not necessarily exclude the presence of the alternate conformation in solution. The twin-conformer model is supported by observed chemical exchange line broadening for the amide of His18 obtained by 15N relaxation studies. This region has also been implied to be involved in receptor binding.

Allergens↗

Estimation of exposure from spilled glutaraldehyde solutions in a hospital setting.

Glutaraldehyde is commonly used in hospitals for cold disinfection of instruments which may be damaged by autoclaving. The increased use of automatic washer/disinfection machines has resulted in a greater risk of spills than with manual methods. A series of experiments was conducted to answer two related research questions: what was the likely range of airborne concentrations when glutaraldehyde is spilled, and are commonly used personal protective equipment ensembles effective and practicable in use? Objective measurements using three sampling methods (two pumped methods based on OSHA 64, one using treated filters and the other based on adsorbent tubes, and a Glutaraldemeter direct reading instrument) were conducted with spills of various surface areas of both 2 and 50% solutions of glutaraldehyde. Results ranged between < 0.01 and 1.4 ppm. Two personal protective equipment ensembles were tested. One was based on a half-facepiece respirator with gas-tight goggles, while the other comprised a full-facepiece cartridge respirator. Both ensembles gave adequate protection against irritation, although in use the half-facepiece respirator and goggles tended to interfere with each other. The direct reading instrument generally underestimated the glutaraldehyde concentrations, although there was a significant association with the results obtained using the method based on adsorbent tubes.

Accidents, Occupational↗

Dopamine and the mechanisms of cognition: Part I. A neural network model predicting dopamine effects on selective attention.

BACKGROUND: Dopamine affects neural information processing, cognition, and behavior; however, the mechanisms through which these three levels of function are affected have remained unspecified. We present a parallel-distributed processing model of dopamine effects on neural ensembles that accounts for effects on human performance in a selective attention task. METHODS: Task performance is stimulated using principles and mechanisms that capture salient aspects of information processing in neural ensembles. Dopamine effects are simulated as a change in gain of neural assemblies in the area of release. RESULTS: The model leads to different predictions as a function of the hypothesized location of dopamine effects. Motor system effects are simulated as a change in gain over the response layer of the model. This induces speeding of reaction times but an impairment of accuracy. Cognitive attentional effects are simulated as a change in gain over the attention layer. This induces a speeding of reaction times and an improvement of accuracy, especially at very fast reaction times and when processing of the stimulus requires selective attention. CONCLUSIONS: A computer simulation using widely accepted principles of processing in neural ensembles can account for reaction time distributions and time-accuracy curves in a selective attention task. The simulation can be used to generate predictions about the effects of dopamine agonists on performance. An empirical study evaluating these predictions is described in a companion paper.

Attention↗

Imaging, image processing and pattern analysis of skin capillary ensembles.

BACKGROUND/AIMS: The capillary bed is recognized as the site where metabolic and nutrient processes occur for living tissues at all levels. The evaluation of this vital process is a major concern in microcirculation. Unlike traditional approaches that concentrated on the extreme local properties of this process, a more global analysis toward capillary ensembles is employed here, since capillaries work as a cooperative entirety. As a first step toward ensemble analysis, the static and planar geometric parameters are investigated. Parameters such as the capillary adjacency and size information are very important in predicting and analysing certain malfunctions in the microvascular bed. METHODS/RESULTS: In order to achieve an objective and accurate analysis of these vital parameters, a computerized imaging system is proposed. Not only the number of capillaries and the capillary cross-sectional areas are important in describing the microvascular bed but the planar distribution pattern of the capillaries also carries valid information. This information, unique to the ensemble analysis, can be used to reveal, visualise and quantify the clustering of capillaries; and this information, according to the Krogh model, is fundamental in estimating the tissue oxygen supply. Two spatial models, the closest neighbor and triangulation methods, have been applied to the captured images of capillary ensembles. The closest neighbor technique generates a minimal distance map or displays a distribution, which depicts the local clustering of capillaries. The triangulation technique, on the other hand, generates a mutual distance map, which is a global description of the capillary positions. Triangulation methods have been evaluated but all except the Greedy triangulation method have been rejected due to lack of robustness and model weakness. Therefore, the capillaries are triangulated by the Greedy triangulation method, and the capillary distribution uniformity is defined as one minus the coefficient of variance of the edge lengths of the mutual distance map. CONCLUSIONS: A series of advanced image processing methods have been developed that efficiently extract the capillary position, size and distribution information from the images. These results facilitate the automatic counting of capillaries and the capillary size-related pathological analysis.

Journal Article↗

Hydrogen exchange in a large 29 kD protein and characterization of molten globule aggregation by NMR.

The nature of denatured ensembles of the enzyme human carbonic anhydrase (HCA) has been extensively studied by various methods in the past. The protein constitutes an interesting model for folding studies that does not unfold by a simple two-state transition, instead a molten globule intermediate is highly populated at 1.5 M GuHCl. In this work, NMR and H/D exchange studies have been conducted on one of the isozymes, HCA I. The H/D exchange studies, which were enabled by the previously obtained resonance assignment of HCA I, have been used to identify unfolded forms that are accessible from the native state. In addition, the GuHCl-induced unfolded states of HCA I have also been characterized by NMR at GuHCl concentrations in the 0-5 M range. The most important findings in this work are as follows: (1) Amide protons located in the center of the beta-sheet require global unfolding events for efficient H/D exchange. (2) The molten globule and the native state give similar protection against H/D exchange for all of the observable amide protons (i.e., water seems not to efficiently penetrate the interior of the molten globule). (3) At high protein concentrations, the molten globule can form large aggregates, which are not detectable by solution-state NMR methods. (4) The unfolded state (U), present at GuHCl concentrations above 2 M, is composed of an ensemble of conformations having residual structures with different stabilities.

Amides↗

FCS cell surface measurements--photophysical limitations and consequences on molecular ensembles with heterogenic mobilities.

BACKGROUND: Fluorescence Correlation Spectroscopy is a powerful method to analyze densities and diffusive behavior of molecules in membranes, but effects of photodegradation can easily be overlooked. METHOD: Based on experimental photophysical parameters, calculations were performed to analyze the consequences of photobleaching in fluorescence correlation spectroscopy (FCS) cell surface experiments, covering a range of standard measurement conditions. RESULTS: Cumulative effects of photobleaching can be prominent, although an absolute majority of the fluorescent molecules would pass the laser excitation beam without being photo-bleached. Given a distribution of molecules on a cell surface with different diffusive properties, the fraction of molecules that is actually analyzed depends strongly on the excitation intensities and measurement times, as well as on the size of the reservoir of freely diffusing molecules. Both the slower and the faster diffusing molecules can be disfavored. CONCLUSIONS: Apart from quantifying photobleaching effects, the calculations suggest that the effects can be used to extract additional information, for instance about the size of the reservoirs of free diffusion. By certain choices of measurement conditions, it may be possible to more specifically analyze certain species within a population, based on their different diffusive properties, different areas of free diffusion, or different kinetics of possible transient binding.

Animals↗

Using partial directed coherence to describe neuronal ensemble interactions.

This paper illustrates the use of the recently introduced method of partial directed coherence in approaching how interactions among neural structures change over short time spans that characterize well defined behavioral states. Central to the method is its use of multivariate time series modelling in conjunction with the concept of Granger causality. Simulated neural network models were used to illustrate the technique's power and limitations when dealing with neural spiking data. This was followed by the analysis of multi-unit activity data illustrating dynamical change in the interaction of thalamo-cortical structures in a behaving rat.

Action Potentials↗

Ensemble docking of multiple protein structures: considering protein structural variations in molecular docking.

One approach to incorporate protein flexibility in molecular docking is the use of an ensemble consisting of multiple protein structures. Sequentially docking each ligand into a large number of protein structures is computationally too expensive to allow large-scale database screening. It is challenging to achieve a good balance between docking accuracy and computational efficiency. In this work, we have developed a fast, novel docking algorithm utilizing multiple protein structures, referred to as ensemble docking, to account for protein structural variations. The algorithm can simultaneously dock a ligand into an ensemble of protein structures and automatically select an optimal protein structure that best fits the ligand by optimizing both ligand coordinates and the conformational variable m, where m represents the m-th structure in the protein ensemble. The docking algorithm was validated on 10 protein ensembles containing 105 crystal structures and 87 ligands in terms of binding mode and energy score predictions. A success rate of 93% was obtained with the criterion of root-mean-square deviation <2.5 A if the top five orientations for each ligand were considered, comparable to that of sequential docking in which scores for individual docking are merged into one list by re-ranking, and significantly better than that of single rigid-receptor docking (75% on average). Similar trends were also observed in binding score predictions and enrichment tests of virtual database screening. The ensemble docking algorithm is computationally efficient, with a computational time comparable to that for docking a ligand into a single protein structure. In contrast, the computational time for the sequential docking method increases linearly with the number of protein structures in the ensemble. The algorithm was further evaluated using a more realistic ensemble in which the corresponding bound protein structures of inhibitors were excluded. The results show that ensemble docking successfully predicts the binding modes of the inhibitors, and discriminates the inhibitors from a set of noninhibitors with similar chemical properties. Although multiple experimental structures were used in the present work, our algorithm can be easily applied to multiple protein conformations generated by computational methods, and helps improve the efficiency of other existing multiple protein structure(MPS)-based methods to accommodate protein flexibility.

Algorithms↗

Optimization of dynamic measurement of receptor kinetics by wavelet denoising.

The most important technical limitation affecting dynamic measurements with PET is low signal-to-noise ratio (SNR). Several reports have suggested that wavelet processing of receptor kinetic data in the human brain can improve the SNR of parametric images of binding potential (BP). However, it is difficult to fully assess these reports because objective standards have not been developed to measure the tradeoff between accuracy (e.g. degradation of resolution) and precision. This paper employs a realistic simulation method that includes all major elements affecting image formation. The simulation was used to derive an ensemble of dynamic PET ligand (11C-raclopride) experiments that was subjected to wavelet processing. A method for optimizing wavelet denoising is presented and used to analyze the simulated experiments. Using optimized wavelet denoising, SNR of the four-dimensional PET data increased by about a factor of two and SNR of three-dimensional BP maps increased by about a factor of 1.5. Analysis of the difference between the processed and unprocessed means for the 4D concentration data showed that more than 80% of voxels in the ensemble mean of the wavelet processed data deviated by less than 3%. These results show that a 1.5x increase in SNR can be achieved with little degradation of resolution. This corresponds to injecting about twice the radioactivity, a maneuver that is not possible in human studies without saturating the PET camera and/or exposing the subject to more than permitted radioactivity.

Algorithms↗

Classification scheme for the design of serine protease targeted compound libraries.

The development of a scoring scheme for the classification of molecules into serine protease (SP) actives and inactives is described. The method employed a set of pre-selected descriptors for encoding the molecular structures, and a trained neural network for classifying the molecules. The molecular requirements were profiled and validated by using available databases of SP- and non-SP-active agents [1,439 diverse SP-active molecules, and 5,131 diverse non-SP-active molecules from the Ensemble Database (Prous Science, 2002)] and Sensitivity Analysis. The method enables an efficient qualification or disqualification of a molecule as a potential serine protease ligand. It represents a useful tool for constraining the size of virtual libraries that will help accelerate the development of new serine protease active drugs.

Computer Simulation↗