Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

The performance of information-theoretic criteria in detecting the number of independent signals in multi-lead ECGs.

Three different methods to detect the number of independent signals in multilead ECGs were evaluated by using different ECG realizations with known specifications for signal and noise: a threshold method (TM), the minimum description length (MDL) and Akaike's information criterion (AIC). The fundamental assumption in this kind of signal processing is that both the signal and the noise stem from independent stochastic generators. The consequence is that the detection of the number of signals is only possible if the noise is white, or if the noise properties have been specified. The evaluation was performed with respect to the QRS complex of individual multilead ECG simulations and to the entire ensemble. In the simulated ECGs the number of independent signals was fixed (eight). It was found that, out of the three methods studied, the performance of MDL was the best, especially when the number of available observations in the noise (used to estimate the noise specifications) was moderate.

Data Collection↗

Application of Lempel-Ziv complexity to the analysis of neural discharges.

Pattern matching is a simple method for studying the properties of information sources based on individual sequences (Wyner et al 1998 IEEE Trans. Inf. Theory 44 2045-56). In particular, the normalized Lempel-Ziv complexity (Lempel and Ziv 1976 IEEE Trans. Inf. Theory 22 75-88), which measures the rate of generation of new patterns along a sequence, is closely related to such important source properties as entropy and information compression ratio. We make use of this concept to characterize the responses of neurons of the primary visual cortex to different kinds of stimulus, including visual stimulation (sinusoidal drifting gratings) and intracellular current injections (sinusoidal and random currents), under two conditions (in vivo and in vitro preparations). Specifically, we digitize the neuronal discharges with several encoding techniques and employ the complexity curves of the resulting discrete signals as fingerprints of the stimuli ensembles. Our results show, for example, that if the neural discharges are encoded with a particular one-parameter method ('interspike time coding'), the normalized complexity remains constant within some classes of stimuli for a wide range of the parameter. Such constant values of the normalized complexity allow then the differentiation of the stimuli classes. With other encodings (e.g. 'bin coding'), the whole complexity curve is needed to achieve this goal. In any case, it turns out that the normalized complexity of the neural discharges in vivo are higher (and hence carry more information in the sense of Shannon) than in vitro for the same kind of stimulus.

Action Potentials↗

Search for folding nuclei in native protein structures.

UNLABELLED: The problem of finding folding nuclei (a set of native contacts that play an important role in folding) along with identifying folding pathways (a time-ordered sequence of folding events) of proteins is one of the most important problems in protein chemistry. Here we propose a novel and simple approach to address this problem as follows: given the topology of the native state, identify native contacts that form folding nuclei based on a graph-theoretical approach that considers effective contact order (effective loop closure) as its objective function. MOTIVATION: A number of computational methods for the prediction of folding nuclei already exists in the literature, but most of them rely on restrictive assumptions about the nature of nuclei or the process of folding. Our motivation is to develop a simple, efficient and robust algorithm to find an ensemble of pathways with the lowest effective contact order and to identify contacts that are crucial for folding. RESULTS: Our approach is different from the previously used methods in that it uses efficient graph algorithms and does not formulate restrictive assumptions about folding nuclei. Our predictions provide more details concerning the protein folding pathway than most other methods in the literature. We demonstrate the success of our approach by predicting folding nuclei for a dataset of proteins for which experimental kinetic data is available. We show that our method compares favourably with other methods in the literature and that its results agree with experimental results. AVAILABILITY: The executable for the proposed algorithm is available at http://www.cs.ubc.ca/~/foldingnuclei.html

Algorithms↗

A proposed architecture and method of operation for improving the protection of privacy and confidentiality in disease registers.

BACKGROUND: Disease registers aim to collect information about all instances of a disease or condition in a defined population of individuals. Traditionally methods of operating disease registers have required that notifications of cases be identified by unique identifiers such as social security number or national identification number, or by ensembles of non-unique identifying data items, such as name, sex and date of birth. However, growing concern over the privacy and confidentiality aspects of disease registers may hinder their future operation. Technical solutions to these legitimate concerns are needed. DISCUSSION: An alternative method of operation is proposed which involves splitting the personal identifiers from the medical details at the source of notification, and separately encrypting each part using asymmetrical (public key) cryptographic methods. The identifying information is sent to a single Population Register, and the medical details to the relevant disease register. The Population Register uses probabilistic record linkage to assign a unique personal identification (UPI) number to each person notified to it, although not necessarily everyone in the entire population. This UPI is shared only with a single trusted third party whose sole function is to translate between this UPI and separate series of personal identification numbers which are specific to each disease register. SUMMARY: The system proposed would significantly improve the protection of privacy and confidentiality, while still allowing the efficient linkage of records between disease registers, under the control and supervision of the trusted third party and independent ethics committees. The proposed architecture could accommodate genetic databases and tissue banks as well as a wide range of other health and social data collections. It is important that proposals such as this are subject to widespread scrutiny by information security experts, researchers and interested members of the general public, alike.

Computer Security↗

Testing a flexible-receptor docking algorithm in a model binding site.

Sampling receptor flexibility is challenging for database docking. We consider a method that treats multiple flexible regions of the binding site independently, recombining them to generate different discrete conformations. This algorithm scales linearly rather than exponentially with the receptor's degrees of freedom. The method was first evaluated for its ability to identify known ligands of a hydrophobic cavity mutant of T4 lysozyme (L99A). Some 200000 molecules of the Available Chemical Directory (ACD) were docked against an ensemble of cavity conformations. Surprisingly, the enrichment of known ligands from among a much larger number of decoys in the ACD was worse than simply docking to the apo conformation alone. Large decoys, accommodated in the larger cavity conformations sampled in the ensemble, were ranked better than known small ligands. The calculation was redone with an energy correction term that considered the cost of forming the larger cavity conformations. Enrichment improved, as did the balance between high-ranking large and small ligands. In a second retrospective test, the ACD was docked against a conformational ensemble of thymidylate synthase. Compared to docking against individual enzyme conformations, the flexible receptor docking approach improved enrichment of known ligands. Including a receptor conformational energy weighting term improved enrichment further. To test the method prospectively, the ACD database was docked against another cavity mutant of lysozyme (L99A/M102Q). A total of 18 new compounds predicted to bind this polar cavity and to change its conformation were tested experimentally; 14 were found to bind. The bound structures for seven ligands were determined by X-ray crystallography. The predicted geometries of these ligands all corresponded to the observed geometries to within 0.7A RMSD or better. Significant conformational changes of the cavity were observed in all seven complexes. In five structures, part of the observed accommodations were correctly predicted; in two structures, the receptor conformational changes were unanticipated and thus never sampled. These results suggest that although sampling receptor flexibility can lead to novel ligands that would have been missed when docking a rigid structure, it is also important to consider receptor conformational energy.

Algorithms↗

Combined use of ESI-MS and UV diode-array detection for localization of disulfide bonds in proteins: application to an alpha-L-fucosidase of pea.

A simplified strategy is described for the assignment of disulfide bonds in proteins of medium to high molecular mass (10-30 kDa). The method combines the use of high-performance liquid chromatography coupled to electrospray ionization mass spectrometry (HPLC-ESI-MS) and HPLC with UV diode-array detection (HPLC diode array). The denatured protein is subjected to proteolysis and the peptide mixture is divided into three fractions: (i) underivatized peptides, (ii) ethylpyridylated peptides, and (iii) reduced and ethylpyridylated peptides. The three peptide ensembles are then subjected to chromatographic and spectroscopic analysis. A systematic methodology is described to analyze the large amount of data obtained. The method was applied to the localization of disulfide bonds in alpha-L-fucosidase from pea. The two disulfide bonds were located between residues Cys64 and Cys109 and between Cys162 and Cys169, while Cys127 was free.

Algorithms↗

Principal components analysis of protein structure ensembles calculated using NMR data.

One important problem when calculating structures of biomolecules from NMR data is distinguishing converged structures from outlier structures. This paper describes how Principal Components Analysis (PCA) has the potential to classify calculated structures automatically, according to correlated structural variation across the population. PCA analysis has the additional advantage that it highlights regions of proteins which are varying across the population. To apply PCA, protein structures have to be reduced in complexity and this paper describes two different representations of protein structures which achieve this. The calculated structures of a 28 amino acid peptide are used to demonstrate the methods. The two different representations of protein structure are shown to give equivalent results, and correct results are obtained even though the ensemble of structures used as an example contains two different protein conformations. The PCA analysis also correctly identifies the structural differences between the two conformations.

Macromolecular Substances↗

An algorithm for protein engineering: simulations of recursive ensemble mutagenesis.

An algorithm for protein engineering, termed recursive ensemble mutagenesis, has been developed to produce diverse populations of phenotypically related mutants whose members differ in amino acid sequence. This method uses a feedback mechanism to control successive rounds of combinatorial cassette mutagenesis. Starting from partially randomized "wild-type" DNA sequences, a highly parallel search of sequence space for peptides fitting an experimenter's criteria is performed. Each iteration uses information gained from the previous rounds to search the space more efficiently. Simulations of the technique indicate that, under a variety of conditions, the algorithm can rapidly produce a diverse population of proteins fitting specific criteria. In the experimental analog, genetic selection or screening applied during recursive ensemble mutagenesis should force the evolution of an ensemble of mutants to a targeted cluster of related phenotypes.

Algorithms↗

Structural kinetic modeling of metabolic networks.

To develop and investigate detailed mathematical models of metabolic processes is one of the primary challenges in systems biology. However, despite considerable advance in the topological analysis of metabolic networks, kinetic modeling is still often severely hampered by inadequate knowledge of the enzyme-kinetic rate laws and their associated parameter values. Here we propose a method that aims to give a quantitative account of the dynamical capabilities of a metabolic system, without requiring any explicit information about the functional form of the rate equations. Our approach is based on constructing a local linear model at each point in parameter space, such that each element of the model is either directly experimentally accessible or amenable to a straightforward biochemical interpretation. This ensemble of local linear models, encompassing all possible explicit kinetic models, then allows for a statistical exploration of the comprehensive parameter space. The method is exemplified on two paradigmatic metabolic systems: the glycolytic pathway of yeast and a realistic-scale representation of the photosynthetic Calvin cycle.

Computer Simulation↗

Structure refinement with molecular dynamics and a Boltzmann-weighted ensemble.

Time-averaging restraints in molecular dynamics simulations were introduced to account for the averaging implicit in spectroscopic data. Space- or molecule-averaging restraints have been used to overcome the fact that not all molecular conformations can be visited during the finite time of a simulation of a single molecule. In this work we address the issue of using the correct Boltzmann weighting for each member of an ensemble, both in time and in space. It is shown that the molecular- or space-averaging method is simple in theory, but requires a priori knowledge of the behaviour of a system. This is illustrated using a five-atom model system and the small cycle peptide analogue somatostatin. When different molecular conformers that are separated by energy barriers insurmountable on the time scale of a simulation contribute significantly to a measured NOE intensity, the use of space- or molecule-averaged distance restraints yields a more appropriate description of the measured data than conventional single-molecule refinement with or without application of time averaging.

Algorithms↗

Effects of guanidine hydrochloride on the proton inventory of proteins: implications on interpretations of protein stability.

The DeltaG degrees (N)(-)(D) value obtained from extrapolation to zero denaturant concentration by the linear extrapolation method (LEM) is commonly interpreted to represent the Gibbs energy difference between native (N) and denatured (D) ensembles at the limit of zero denaturant concentration. For DeltaG degrees (N)(-)(D) to be interpreted solely in terms of N and D, as is common practice, it must be shown to be independent of denaturant concentration. Because DeltaG degrees (N)(-)(D) is often observed to be dependent on the nature of the denaturant, it is necessary to determine the circumstances under which DeltaG degrees (N)(-)(D) can be interpreted as a property solely of the protein. Here, we use proton inventory, a thermodynamic property of both the native and denatured ensembles, to monitor the thermodynamic character of denaturant-dependent aspects of N and D ensembles and the N right arrow over left arrow D transition. Use of a thermodynamic rather than a spectral parameter to monitor denaturation provides insight into the manner in which denaturant affects the meaning of DeltaG degrees (N)(-)(D) and the nature of the N right arrow over left arrow D transition. Three classes of proteins are defined in terms of the thermodynamic behaviors of their N right arrow over left arrow D transition and N and D ensembles. With guanidine hydrochloride as a denaturant, the classification of protein denaturations by these procedures determines when the LEM gives readily interpretable DeltaG degrees (N)(-)(D) values with this denaturant and when it does not.

Chymotrypsin↗

Biased sampling of nonequilibrium trajectories: can fast switching simulations outperform conventional free energy calculation methods?

We have investigated the maximum computational efficiency of reversible work calculations that change control parameters in a finite amount of time. Because relevant nonequilibrium averages are slow to converge, a bias on the sampling of trajectories can be beneficial. Such a bias, however, can also be employed in conventional methods for computing reversible work, such as thermodynamic integration or umbrella sampling. We present numerical results for a simple one-dimensional model and for a Widom insertion in a soft sphere liquid, indicating that, with an appropriately chosen bias, conventional methods are in fact more efficient. We describe an analogy between nonequilibrium dynamics and mappings between equilibrium ensembles, which suggests that the practical inferiority of fast switching is quite general. Finally, we discuss the relevance of adiabatic invariants in slowly driven Hamiltonian systems for the application of Jarzynski's theorem.

Chemistry, Physical↗

Biomolecular free energy profiles by a shooting/umbrella sampling protocol, "BOLAS".

We develop an efficient technique for computing free energies corresponding to conformational transitions in complex systems by combining a Monte Carlo ensemble of trajectories generated by the shooting algorithm with umbrella sampling. Motivated by the transition path sampling method, our scheme "BOLAS" (named after a cowboy's lasso) preserves microscopic reversibility and leads to the correct equilibrium distribution. This makes possible computation of free energy profiles along complex reaction coordinates for biomolecular systems with a lower systematic error compared to traditional, force-biased umbrella sampling protocols. We demonstrate the validity of BOLAS for a bistable potential, and illustrate the method's scope with an application to the sugar repuckering transition in a solvated deoxyadenosine molecule.

Algorithms↗

Equilibrium free energy estimates based on nonequilibrium work relations and extended dynamics.

Jarzynski's relation and the fluctuation theorem have established important connections between nonequilibrium statistical mechanics and equilibrium thermodynamics. In particular, an exact relationship between the equilibrium free energy and the nonequilibrium work is useful for computer simulations. In this paper, we exploit the fact that the free energy is a state function, independent of the pathway taken to change the equilibrium ensemble. We show that a generalized expression is advantageous for computer simulations of free energy differences. Several methods based on this idea are proposed. The accuracy and efficiency of the proposed methods are evaluated with a model problem.

Algorithms↗

A coarse graining method for the identification of transition rates between molecular conformations.

The coarse graining method to be advocated in this paper consists of two main steps. First, the propagation of an ensemble of molecular states is described as a Markov chain by a transition probability matrix in a finite state space. Second, we obtain metastable conformations by an aggregation of variables via Robust Perron Cluster Analysis (PCCA+). Up to now, it has been an open question as to how this coarse graining in space can be transformed to a coarse graining of the Markov chain while preserving the essential dynamic information. In this article, we construct a coarse matrix that is the correct propagator in the space of conformations. This coarse graining procedure carries over to rate matrices and allows to extract transition rates between molecular conformations. This approach is based on the fact that PCCA+ computes molecular conformations as linear combinations of the dominant eigenvectors of the transition matrix.

Journal Article↗

The role of the fast motion of the spin label in the interpretation of EPR spectra for spin-labeled macromolecules.

The spin label method was used to observe the nature of the fast motions of side chains in protein monocrystals. The EPR spectra of spin-labeled lysozyme monocrystals (with different orientations of the tetragonal protein crystal in relation to the direction of the magnetic field) were interpreted using the method of molecular dynamics (MD). Within the proposed simple model, MD calculations of the spin label motion trajectories are performed in a reasonable real time. The model regards the protein molecule as frozen as a whole and the spin-labeled amino acid residue as unfrozen. To calculate the trajectories in vacuum, a model of spin-labeled lysozyme was assembled, and the parameters of the force fields were specified for atoms of the protein molecule, including the spin label. The calculations show that the protein environment sterically limits the area of the possible angular reorientations for the NO reporter group of the nitroxide (within the spin label), and this, in turn, affects the shape of the EPR spectrum. However, it turned out that the spread in the positions of the reporter group in the angle space strictly adheres to the Gaussian distribution. Using the coordinates of the spin label atoms obtained by the MD method within a selected time range and considering the distribution of the spin label states over the ensemble of spin-labeled macromolecules in a crystal, the EPR spectra of spin-labeled lysozyme monocrystals were simulated. The resultant theoretical EPR spectra appeared to be similar to experimental ones.

Animals↗

Particle size distributions of inert spheres and pelletized pharmaceutical products by image analysis.

Image analysis was used to measure particle size distributions (PSDs) of ensembles of 425 to 1400 microm-size materials. Repeatability of a measurement, suitable sample sizes, and methods of sampling were assessed. Two lots of inert spheres were compared prior to drug layering in a Glatt GPCG-5 rotor. The differences in PSD in the starting materials were reflected in the rotor-granulated products. Such detailed information was not available from sieving with U.S. standard wire mesh sieves. The products from the rotor process were polymer-coated in a Wurster process in a Glatt GPCG-3, 4-in. Wurster. The resolution of the technique was sufficient to measure differences in diameter equating to 4-microm coat thickness, which resulted from applying 2% polymer coat weight. The utility of the technique for monitoring commercial scale processes was demonstrated by measuring diameter after layering drug onto nonpareils in a Glatt RG-150 rotor, and by measuring the diameter after application of a polymer solution in a Glatt 46-in. Wurster coating process. The similarity of samples removed from the sample port in situ and samples from the batch suggested that processes in the fluid bed are intensively mixed and inherently random.

Chemistry, Pharmaceutical↗

A hint to search for metalloproteins in gene banks.

MOTIVATION: With the advent of genome sequencing, a huge database of protein primary sequences has been accumulating. In parallel, a number of tools to investigate and expand upon this information, e.g. reconstructing and building relationships between protein families and superfamilies, have been developed. Metalloproteins are proteins capable of binding one or more metal ions, which are required for their biological function or for regulation of their activities or for structural purposes. Sometimes, metal binding can be observed in vitro but not be physiologically relevant. At present, there is a lack of specific tools to address the matter of the identification of metalloproteins in databases of gene sequences. RESULTS: In the present work, an approach exploiting metal-binding patterns (MBPs) of metalloproteins present in the Protein Data Bank to search gene banks for new metalloproteins is presented and applied to copper proteins. Nearly 100 different MBPs have been identified and then used for subsequent applications. The ensemble of sequences of the whole PDB is used to assess the potentiality and limits of the method and to identify levels of confidence for the predictions output by the search. It appears that copper-binding capabilities are identified with a confidence >90% when the percentage of identical amino acids aligned around the MBP by PHI-BLAST is at least 20% with respect to the entire protein domain length. If this percentage is between 10% and 20%, the level of confidence is approximately 50%. Application of the methodology to the entire genome sequences of Pyrococcus furiosus, Escherichia coli, Drosophila melanogaster and Homo sapiens suggests some differentiation between prokaryotes and eukaryotes. SUPPLEMENTARY INFORMATION: A table reporting statistics on the MBP identified; a list of all hits retrieved for the four organisms considered; a figure showing the number of hits for the four organisms as a function of I(d)(Global).

Algorithms↗