Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Random-energy model in random fields.

The random-energy model is studied in the presence of random fields. The problem is solved exactly both in the microcanonical ensemble, without recourse to the replica method, and in the canonical ensemble using the replica formalism. The phase diagrams for bimodal and Gaussian random fields are investigated in detail. In contrast to the Gaussian case, the bimodal random field may lead to a tricritical point and a first-order transition. An interesting feature of the phase diagram is the possibility of a first-order transition from paramagnetic to mixed phase.

Journal Article↗

De novo ligand design to an ensemble of protein structures.

We describe a combinatorial method for de novo ligand design to an ensemble of receptor structures. Receptor conformations, protonation states, and structural water molecules are considered consistently within the framework of de novo ligand design. The method relies on Monte Carlo optimization to search the space of ligand structures, conformations, and rigid-body movements as well as receptor models. The method is applied to an ensemble of HIV protease and human collagenase receptor models. Ligand structures generated de novo exhibit the correct hydrogen-bonding pattern in the core of the active site, with hydrophobic groups extending into the receptor S1 and S1' pocket space. Furthermore, it is shown that known ligands are recovered in the correct binding mode and in the native, most tightly binding receptor model.

Binding Sites↗

Clustering ensembles of neural network models.

We show that large ensembles of (neural network) models, obtained e.g. in bootstrapping or sampling from (Bayesian) probability distributions, can be effectively summarized by a relatively small number of representative models. In some cases this summary may even yield better function estimates. We present a method to find representative models through clustering based on the models' outputs on a data set. We apply the method on an ensemble of neural network models obtained from bootstrapping on the Boston housing data, and use the results to discuss bootstrapping in terms of bias and variance. A parallel application is the prediction of newspaper sales, where we learn a series of parallel tasks. The results indicate that it is not necessary to store all samples in the ensembles: a small number of representative models generally matches, or even surpasses, the performance of the full ensemble. The clustered representation of the ensemble obtained thus is much better suitable for qualitative analysis, and will be shown to yield new insights into the data.

Algorithms↗

Discrimination between modes of toxic action of phenols using rule based methods.

Rule-based ensemble modelling has been used to develop a model with high accuracy and predictive capabilities for distinguishing between four different modes of toxic action for a set of 220 phenols. The model not only predicts the majority class (polar narcotics) well but also the other three classes (weak acid respiratory uncouplers, pro-electrophiles and soft electrophiles) of toxic action despite the severely skewed distribution among the four investigated classes. Furthermore, the investigation also highlights the merits of using ensemble (or consensus) modelling as an alternative to the more traditional development of a single model in order to promote robustness and accuracy with respect to the predictive capability for the derived model.

Databases, Factual↗

Theoretical analysis of drug release into a finite medium from sphere ensembles with various size and concentration distributions.

Release kinetics for heterogeneous sphere ensembles with a dissolved drug, i.e., initial drug loading below or equal to the drug solubility in the matrix, in a finite external medium was modeled with consideration of heterogeneity among and within spheres. Numerical solutions were obtained using the finite element method for sphere ensemble with normal or log-normal distribution of particle size or initial drug loading among spheres. Exact series solutions were derived for ensembles with various initial loading distributions within spheres, namely linear, quadratic, sigmoidal and uniform distribution, using their mean or average radii. Simplified solutions retaining only one term of the series for non-uniform distributions and three terms for uniform distribution were suggested because of their good approximation to the exact solution. The results of finite element analysis showed that the release rate of an ensemble decreased with increasing standard deviation of particle size. Using weight-average radii in the exact solution gave a prediction of release profile closer to that from the actual size distribution than using mean radii. The three non-uniform loading patterns within spheres all showed reduced initial burst and release rate, leading to more steady release rates than uniform loading, among which the sigmoidal distribution offered the best near-zero order release. Non-uniform initial loading among spheres seemed to have insignificant influence on the release profiles. The volume ratio of liquid to a sphere ensemble played an important role in release kinetics. The derived analytical solutions are applicable to multiple spheres or a single sphere in a finite medium or in a perfect sink.

Algorithms↗

Field squeeze operators in optical cavities with atomic ensembles.

We propose a method of generating unitarily single and two-mode field squeezing in an optical cavity with an atomic cloud. Through a suitable laser system, we are able to engineer a squeeze field operator decoupled from the atomic degrees of freedom, yielding a large squeeze parameter that is scaled up by the number of atoms, and realizing degenerate and nondegenerate parametric amplification. By means of the input-output theory we show that ideal squeezed states and perfect squeezing could be approached at the output. The scheme is robust to decoherence processes.

Journal Article↗

Protein design based on the relative entropy.

An approach to protein design is proposed based on the relative entropy and a reduced amino acid alphabet. In this approach, the relative entropy is used as a minimization object function. The method has been tested on a real protein's off-lattice model successfully, and the results are similar to those obtained from other design studies. It can be applied as a uniform frame for both folding and inverse folding of protein. An iterative calculation method of the ensemble average of the contact strength is proposed at the same time.

Algorithms↗

Mean dynamic topography: inter-comparisons and errors.

Knowledge of the ocean dynamic topography, defined as the height of the sea surface above its rest-state (the geoid), would allow oceanographers to study the absolute circulation of the ocean and determine the associated geostrophic surface currents that help to regulate the Earth's climate. Here a novel approach to computing a mean dynamic topography (MDT), together with an error field, is presented for the northern North Atlantic. The method uses an ensemble of MDTs, each of which has been produced by the assimilation of hydrographic data into a numerical ocean model, to form a composite MDT, and uses the spread within the ensemble as a measure of the error on this MDT. The r.m.s. error for the composite MDT is 3.2 cm, and for the associated geostrophic currents the r.m.s. error is 2.5 cms(-1). Taylor diagrams are used to compare the composite MDT with several MDTs produced by a variety of alternative methods. Of these, the composite MDT is found to agree remarkably well with an MDT based on the GRACE geoid GGM01C. It is shown how the composite MDT and its error field are useful validation products against which other MDTs and their error fields can be compared.

Image Processing, Computer-Assisted↗

Conformation spaces of proteins.

We report a simple method for measuring the accessible conformational space explored by an ensemble of protein structures. The method is useful for diverse ensembles derived from molecular dynamics trajectories, molecular modeling, and molecular structure determinations. It can be used to examine a wide range of time scales. The central tactic we use, which has been previously employed, is to replace the true mechanical degrees of freedom of a molecular system with the conformationally effective degrees of freedom as measured by the root-mean squared cartesian distances among all pairs of conformations. Each protein conformation is treated as a point in a high dimensional euclidean space. In this article, we model this space in a novel way by representing it as an N-dimensional hypercube, describable with only two parameters: the number of dimensions and the edge length. To validate this approach, we provide a number of elementary test cases and then use the N-cube method for measuring the size and shape of conformational space covered by molecular dynamics trajectories spanning 10 orders of magnitude in time. These calculations were performed on a small protein, the villin headpiece subdomain, exploring both the native state and the misfolded/folding regime. Distinct features include single, vibrationally averaged, substate minima on the 0.1-1-ps time scale, thermally averaged conformational states that persist for 1-100 ps and transitions between these local minima on nanosecond time scales. Large-scale refolding modes appear to become uncorrelated on the microsecond time scale. Associated length scales for these events are 0.2 A for the vibrational minima; 0.5 A for the conformational minima; and 1-2 A for the nanosecond events. We find that the conformational space that is dynamically accessible during folding of villin has enough volume for approximately 10(9) minima of the variety that persist for picoseconds. Molecular dynamics trajectories of the native protein and experimentally derived solution ensembles suggest the native state to be composed of approximately 10(2) of these thermally accessible minima. Thus, based on random exploration of accessible folding space alone, protein folding for a small protein is predicted to be a milliseconds time scale event. This time can be compared with the experimental folding time for villin of 10-100 micros. One possible explanation for the 10-100-fold discrepancy is that the slope of the "folding funnel" increases the rate 1-2 orders of magnitude above random exploration of substates.

Algorithms↗

Efficient conformational sampling of local side-chain flexibility.

Side-chain flexibility of ligand-binding sites needs to be considered in the rational design of novel inhibitors. We have developed a method to generate conformational ensembles that efficiently sample local side-chain flexibility from a single crystal structure. The rotamer-based approach is tested here for the S1' pocket of human collagenase-1 (MMP-1), which is known to undergo conformational changes in multiple side-chains upon binding of certain inhibitors. First, a raw ensemble consisting of a large number of conformers of the S1' pocket was generated using an exhaustive search of rotamer combinations on a template crystal structure. A combination of principal component analysis and fuzzy clustering was then employed to successfully identify a core ensemble consisting of a low number of representatives from the raw ensemble. The core ensemble contained geometrically diverse conformers of stable nature, as indicated in several cases by a relative energy lower than that of the minimised template crystal structure. Through comparisons with X-ray crystallography and NMR structural data we show that the core ensemble occupied a conformational space similar to that observed under experimental conditions. The synthetic inhibitor RS-104966 is known to induce a conformational change in the side-chains of the S1' pocket of MMP-1 and could not be docked in the template crystal structure. However, the experimental binding mode was reproduced successfully using members of the core ensemble as the docking target, establishing the usefulness of the method in drug design.

Arginine↗

Predicting protein mutant energetics by self-consistent ensemble optimization.

In this paper we present a self-consistent ensemble optimization (SCEO) theory for efficient conformational search, which we have applied to predicting the effects of mutations on protein thermostability. This approach takes advantage of a statistical mechanical self-consistency condition to home in iteratively on the global minimum structure. We employ a fast potential of mean-force approximation to cut computation time to a few minutes for a typical protein mutation, with only linear time-dependence on the size of the prediction problem. Rather than seeking a single, static structure of minimum energy, the new method optimizes an ensemble of many conformations, seeking to predict the most likely ensemble for the native state at a desired temperature. Testing this approach with a simple physical model focusing entirely on steric interactions and side-chain rearrangement, we obtain robustly convergent prediction of core side-chain conformation, and of hydrophobic core mutations' effects on protein stability. Self-consistent ensemble optimization is superior to simulated annealing in its speed and convergence to the global minimum, and insensitive to starting conformation. In calculations on lambda repressor protein, structural predictions for an eight-residue molten-zone had side-chain r.m.s. error of 0.49 A for the wild-type protein. Evaluation of the method's mutant structure predictions should become possible, as structures of these mutant repressors are solved. Predicted energies for a series of nine hydrophobic core mutants correlated with measured free energies of unfolding with a coefficient of 0.82.

Amino Acid Sequence↗

Source density analysis of scalp potentials during linguistic and non-linguistic processing of visual stimuli.

Event-related potentials (ERPs) were recorded from 40 locations, covering most of the scalp, during repeated tasks in which the observer (O) had to judge either the tense of a printed verb (V) or the symmetry of a spatial pattern (S). Stimuli were drawn at random from large ensembles. A simplified method of Laplacean analysis (MacKay 1983, 1984) allowed the corresponding source densities to be mapped at up to 28 locations, relatively free of artefacts due to eye movements or tongue movements. O signalled his judgement in each case by pressing one of two buttons on a given cue. The decision time allowed was kept short (about 1 s) but long enough for the task to be handled successfully. When stimuli 'V' and 'S' were drawn from geometrically different ensembles, the source-density distributions for the two tasks differed significantly at a number of locations. When 'V' and 'S' were drawn from a common ensemble, however, and O was instructed on each trial (in random order) to assess each stimulus as a word or as a geometrical pattern, the similarities in the source-density maps were more striking than the differences. It would seem that during sufficiently rapid verbal and spatial judgments, little sign of hemispheric specialization or task-specific differences may appear in the spatiotemporal profile of ERP source densities. More salient differences, some lateralized, appeared during the preparation interval prior to verbal and spatial tasks; but their pattern varied widely from subject to subject.

Cerebral Cortex↗

Defining the precision with which a protein structure is determined by NMR. Application to motilin.

A simple procedure is introduced for accurately defining the precision with which the Cartesian coordinates of any macromolecular structure are determined by nuclear Overhauser data. The method utilizes an ensemble of structures obtained from an array of independent simulated data sets derived from a final structure. Using the noise-free, back-calculated NOE spectrum as the "true" NOE spectrum, simulated Monte Carlo data sets are created by superimposing onto the "true" spectrum Gaussian distributed noise with a standard deviation equal to that of the residuals. Full relaxation matrix refinements of the simulated data sets provide probability distributions of the Cartesian coordinates for each atom in the model. Molecular dynamics simulations are included to estimate the effect of sparse information on the precision. The procedure is applied here to the 22-residue peptide hormone motilin, and the results are compared to those obtained using the conventional method of analyzing multiple refinements using a single distance constraint set. The average root mean square deviation for alpha-carbon atoms in the central portion (Arg12-Arg18) of the single helix of motilin was determined to be 0.72 A by the Monte Carlo method, compared to 1.3 A determined by an analysis of the 10 best DIANA structures using the same number of constraints between the same atoms. The origin of the bias of the conventional method is discussed.

Amino Acid Sequence↗

Sequence-specific solvent accessibilities of protein residues in unfolded protein ensembles.

Protein stability cannot be understood without the correct description of the unfolded state. We present here an efficient method for accurate calculation of atomic solvent exposures for denatured protein ensembles. The method used to generate the ensembles has been shown to reproduce diverse biophysical experimental data corresponding to natively and chemically unfolded proteins. Using a data set of 19 nonhomologous proteins containing from 98 to 579 residues, we report average accessibilities for all residue types. These averaged accessibilities are considerably lower than those previously reported for tripeptides and close to the lower limit reported by Creamer and co-workers. Of importance, we observe remarkable sequence dependence for the exposure to solvent of all residue types, which indicates that average residue solvent exposures can be inappropriate to interpret mutational studies. In addition, we observe smaller influences of both protein size and protein amino acid composition in the averaged residue solvent exposures for individual proteins. Calculating residue-specific solvent accessibilities within the context of real sequences is thus necessary and feasible. The approach presented here may allow a more precise parameterization of protein energetics as a function of polar- and apolar-area burial and opens new ways to investigate the energetics of the unfolded state of proteins.

Computer Simulation↗

Design of synthetic gene libraries encoding random sequence proteins with desired ensemble characteristics.

Libraries of random sequence polypeptides are useful as sources of unevolved proteins, novel ligands, and potential lead compounds for the development of vaccines and therapeutics. The expression of small random peptides has been achieved previously using DNA synthesized with equimolar mixtures of nucleotides. For many potential uses of random polypeptide libraries, concerns such as avoiding termination codons and matching target amino acid compositions make more complex designs necessary. In this study, three mixtures of nucleotides, corresponding to the three positions in the codon, were designed such that semirandom DNA synthesized by repeated cycles of the three mixtures created an open reading frame encoding random sequence polypeptides with desired ensemble characteristics. Two methods were used to design the nucleotide mixtures: the manual use of a spreadsheet and a refining grid search algorithm. Using design targets of less than or equal to 1% stop codons and an amino acid composition based on the average ratios observed in natural, globular proteins, the search methods yielded similar nucleotide ratios, Semirandom DNA, synthesized with a designed, three-residue repeat pattern, can encode libraries of very high diversity and represents an important tool for the construction of random polypeptide libraries.

Amino Acid Sequence↗

The effects of cryogenic blockade of the centrifugal, bulbopetal pathways on the dynamic and static response characteristics of goldfish olfactory bulb mitral cells.

The responses of single goldfish olfactory bulb mitral cells were studied by extracellular recordings before and during cryogenic blockade of the efferent, centrifugal pathways in the ipsilateral olfactory tract. In each experiment the same odour was presented 40 times before and then 40 times during cooling. Each stimulus period (at least 30 s) was preceded by a stimulus-free interval (at least 30 s), during which a steady stream of tap water was applied. These procedures allow the investigation of activity changes of single neurons and of cell ensembles using statistical methods. i) In comparison with the pre-cooling activity, cooling of the efferent pathways did not cause a generalized disinhibition in mitral cell responses. Significant disinhibitory, significant inhibitory and indifferent effects occurred in about the same proportion during repetitive water and odour applications. ii) Abrupt or slow changes of single mitral cell discharge patterns during the 40 water and odour applications were observed before and during blocking of the efferent fibre systems: These pattern changes are therefore not necessarily a consequence of the efferent signals, and may thus have been a result of intrabulbar plasticity. iii) The most notable effect of efferent fibre blockade across all experiments was a significant (Wilcoxon-rank-test, P = 0.01) decrease of the signal to noise ratio i.e., the ratio between the activity during the "spontaneous" (water) and the stimulus (odour) phase, which could be demonstrated for both the phasic (immediately after stimulus onset) and tonic (during long term stimulation) components of the mitral cell responses.

Action Potentials↗

A stochastic model in liquid penetration through fibrous media.

The statistical genesis of the process of liquid penetration through fibrous media can be regarded as the interaction and the resulting balance among media and liquid cells that comprise the ensemble. A stochastic method, Ising's model, combined with Monte Carlo simulation, can therefore be employed in the study of liquid penetration through fibrous media. This process is driven by the difference of energy of the system after and before a liquid moves from one cell to the other. The energy of the system comprises the internal energy, work done by external force to the system, and the mechanical energy. For experimental verification, the process of water penetration through isotropic fiber mats, both spontaneously and under pressure, is examined. Simulation results are in good agreement with the experiments, indicating a good prospect of the method to be applied in this area.

Journal Article↗

Structure of a beta-alanine-linked polyamide bound to a full helical turn of purine tract DNA in the 1:1 motif.

Polyamides composed of N-methylpyrrole (Py), N-methylimidazole (Im) and N-methylhydroxypyrrole (Hp) amino acids linked by beta-alanine (beta) bind the minor groove of DNA in 1:1 and 2:1 ligand to DNA stoichiometries. Although the energetics and structure of the 2:1 complex has been explored extensively, there is remarkably less understood about 1:1 recognition beyond the initial studies on netropsin and distamycin. We present here the 1:1 solution structure of ImPy-beta-Im-beta-ImPy-beta-Dp bound in a single orientation to its match site within the DNA duplex 5'-CCAAAGAGAAGCG-3'.5'-CGCTTCTCTTTGG-3' (match site in bold), as determined by 2D (1)H NMR methods. The representative ensemble of 12 conformers has no distance constraint violations greater than 0.13 A and a pairwise RMSD over the binding site of 0.80 A. Intermolecular NOEs place the polyamide deep inside the minor groove, and oriented N-C with the 3'-5' direction of the purine-rich strand. Analysis of the high-resolution structure reveals the ligand bound 1:1 completely within the minor groove for a full turn of the DNA helix. The DNA is B-form (average rise=3.3 A, twist=38 degrees ) with a narrow minor groove closing down to 3.0-4.5 A in the binding site. The ligand and DNA are aligned in register, with each polyamide NH group forming bifurcated hydrogen bonds of similar length to purine N3 and pyrimidine O2 atoms on the floor of the minor groove. Each imidazole group is hydrogen bonded via its N3 atom to its proximal guanine's exocyclic amino group. The important roles of beta-alanine and imidazole for 1:1 binding are discussed.

Base Pairing↗