Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

A methodology for testing for statistically significant differences between fully 3D PET reconstruction algorithms.

We present a practical methodology for evaluating 3D PET reconstruction methods. It includes generation of random samples from a statistically described ensemble of 3D images resembling those to which PET would be applied in a medical situation, generation of corresponding projection data with noise and detector point spread function simulating those of a 3D PET scanner, assignment of figures of merit appropriate for the intended medical applications, optimization of the reconstruction algorithms on a training set of data, and statistical testing of the validity of hypotheses that say that two reconstruction algorithms perform equally well (from the point of view of a particular figure of merit) as compared to the alternative hypotheses that say that one of the algorithms outperforms the other. Although the methodology was developed with the 3D PET in mind, it can be used, with minor changes, for other 3D data collection methods, such as fully 3D cr or SPECT.

Algorithms↗

Simulation of phase transitions in fluids.

This review provides a discussion of recent techniques for simulation of phase equilibria of complex fluids. Monte Carlo methods are emphasized over molecular dynamics methods. We describe recent developments, such as the use of expanded-ensemble, tempering, or histogram reweighting techniques. Our discussion of such developments is aimed at a general audience and is intended to provide an overview of the main advantages and limitations of each particular technique. References are provided to allow interested readers to identify and trace back most recent applications of a particular simulation technique. We conclude with general guidelines regarding selection of suitable simulation methods for particular problems and systems of interest.

Journal Article↗

Gene finding in the chicken genome.

BACKGROUND: Despite the continuous production of genome sequence for a number of organisms, reliable, comprehensive, and cost effective gene prediction remains problematic. This is particularly true for genomes for which there is not a large collection of known gene sequences, such as the recently published chicken genome. We used the chicken sequence to test comparative and homology-based gene-finding methods followed by experimental validation as an effective genome annotation method. RESULTS: We performed experimental evaluation by RT-PCR of three different computational gene finders, Ensembl, SGP2 and TWINSCAN, applied to the chicken genome. A Venn diagram was computed and each component of it was evaluated. The results showed that de novo comparative methods can identify up to about 700 chicken genes with no previous evidence of expression, and can correctly extend about 40% of homology-based predictions at the 5' end. CONCLUSIONS: De novo comparative gene prediction followed by experimental verification is effective at enhancing the annotation of the newly sequenced genomes provided by standard homology-based methods.

Animals↗

A method for determining neural connectivity and inferring the underlying network dynamics using extracellular spike recordings.

In the present paper we propose a novel method for the identification and modeling of neural networks using extracellular spike recordings. We create a deterministic model of the effective network, whose dynamic behavior fits experimental data. The network obtained by our method includes explicit mathematical models of each of the spiking neurons and a description of the effective connectivity between them. Such a model allows us to study the properties of the neuron ensemble independently from the original data. It also permits to infer properties of the ensemble that cannot be directly obtained from the observed spike trains. The performance of the method is tested with spike trains artificially generated by a number of different neural networks.

Action Potentials↗

Writing brains: tracing the psyche with the graphical method.

At the end of the 19th century, the graphic method kindled attempts to use it for investigating psychic processes. In Germany, Hans Berger took up this line of research, later to become the pioneer of electroencephalography (EEG). This trajectory of Berger's work is analyzed as an "enabling constraint" guiding him toward the EEG at a time when nobody else was pursuing this line of research and also causing serious methodological problems. In the epistemological perspective of this analysis, many of his problems extend beyond the local context of his work and point toward ambiguities surrounding the project to trace the psyche with the graphic method. From the mid-1930s, the EEG inspired ongoing attempts to decipher the specific meanings of these recordings, and large ensembles of machinery were mobilized, molding concepts of the psyche according to the results and the specifications of the graphic method.

Electroencephalography↗

Development of an extended simulated annealing method: application to the modeling of complementary determining regions of immunoglobulins.

An extended simulated annealing process (ESAP) has been developed in order to obtain an ensemble of conformations of a peptide segment from a protein fluctuating at a given temperature. The annealing process was performed with a fast Monte Carlo method using the scaled collective variables developed by Noguti and Go. The system was divided into two parts: one consists of one or more peptide segments and is flexible around the main-chain and side-chain torsional angles; the other represents the rest of the molecule and was maintained fixed at the atomic positions determined by x-ray experiments. The target function included the nonbonding atomic interactions and a distance function to anchor the N and C terminal ends of each segment to the fixed part. Three systems of complementary determining regions (CDR) of antibodies were tested and compared to x-ray data: L2 loop (7 residues) of the light chain of lambda-type Bence-Jones protein, H1 and the H2 loops (14 residues) of McPC603, and H1 and H2 loops (12 residues) of HyHEL-5. Each state of CDR conformations was characterized at room temperature by the average of their coordinates (average conformation) and the internal energy. With a limited number of annealing processes (10), starting from the extended conformation, we have obtained states with conformations close to the observed x-ray structures, from 1.1 to 1.7 A root mean square deviation (rmsd) of main-chain atoms depending on the system. These states were identical or within 0.25 A rmsd of those of lowest internal energy. For unknown CDR structures the criteria of lowest internal energies from ESAP can be used to predict hypervariable loop structures in antibodies with an accuracy comparable to other methods.

Amino Acid Sequence↗

Exploration of compact protein conformations using the guided replication Monte Carlo method.

We have studied the use of a new Monte Carlo (MC) chain generation algorithm, introduced by T. Garel and H. Orland [(1990) Journal of Physics A, Vol. 23, pp. L621-L626], for examining the thermodynamics of protein folding transitions and for generating candidate C(alpha) backbone structures as starting points for a de novo protein structure paradigm. This algorithm, termed the guided replication Monte Carlo method, allows a rational approach to the introduction of known "native" folded characteristics as constraints in the chain generation process . We have shown this algorithm to be computationally very efficient in generating large ensembles of candidate C(alpha) chains on the face centered cubic lattice, and illustrate its use by calculating a number of thermodynamic quantities related to protein folding characteristics. In particular, we have used this static MC algorithm to compare such temperature-dependent quantities as the ensemble mean energy, ensemble mean free energy, the heat capacity, and the mean-square radius of gyration. We also demonstrate the use of several simple "guide fields" for introducing protein-specific constraints into the ensemble generation process. Several extensions to our current model are suggested, and applications of the method to other folding related problems are discussed.

Algorithms↗

The information transmitted by ensembles of primary spindle afferents is diminished when ketamine is used as a pre-anaesthetic.

The effect of pre-anaesthetic ketamine on ensemble coding of different stimuli consisting of muscle stretches of various amplitudes was studied for ensembles of simultaneously recorded primary muscle spindle afferents (MSAs). The experiments were conducted on 8 alpha-chloralose anaesthetised cats. Three of the cats received a pre-anaesthetic dose of ketamine (25 mg/kg) injected subcutaneously (ketamine group), while the remaining five animals did not (non-ketamine group). Data for ensemble coding were collected both before and after cutting the ventral root. A method based on principal component analysis and algorithms was used to quantify stimulus discrimination and an ANOVA tested differences between groups as well as differences due to ventral root cutting. When the fusimotor supply was intact, a general trend of an increase in the ability to discriminate stimuli with increasing ensemble size was observed for both groups, however, this trend was significantly greater for the non-ketamine group as compared to the ketamine group. When the ventral root was cut, the discrimination pattern for the non-ketamine group decreased significantly (as compared to before ventral root cutting), however, no change occurred for the ketamine group. Consequently, no difference in discrimination pattern was detected between groups after ventral root cutting. The reduction in information transmitted by ensembles of primary MSAs when ketamine is used as a pre-anaesthetic may suggest that ketamine elicits an adverse affect on the fusimotor system.

Afferent Pathways↗

Robust sound classification through the representation of similarity using response fields derived from stimuli during early experience.

Models of auditory processing, particularly of speech, face many difficulties. Included in these are variability among speakers, variability in speech rate, and robustness to moderate distortions such as time compression. We constructed a system based on ensembles of feature detectors derived from fragments of an onset-sensitive sound representation. This method is based on the idea of 'spectro-temporal response fields' and uses convolution to measure the degree of similarity through time between the feature detectors and the stimulus. The output from the ensemble was used to derive segmentation cues and patterns of response, which were used to train an artificial neural network (ANN) classifier. This allowed us to estimate a lower bound for the mutual information between the class of the input and the class of the output. Our results suggest that there is significant information in the output of our system, and that this is robust with respect to the exact choice of feature set, time compression in the stimulus, and speaker variation. In addition, the robustness to time compression in the stimulus has features in common with human psychophysics. Similar experiments using feature detectors derived from fragments of non-speech sounds performed less well. This result is interesting in the light of results showing aberrant cortical development in animals exposed to impoverished auditory environments during the developmental phase.

Acoustic Stimulation↗

Conformation and hydrogen ion titration of proteins: a continuum electrostatic model with conformational flexibility.

A new method for including local conformational flexibility in calculations of the hydrogen ion titration of proteins using macroscopic electrostatic models is presented. Intrinsic pKa values and electrostatic interactions between titrating sites are calculated from an ensemble of conformers in which the positions of titrating side chains are systematically varied. The method is applied to the Asp, Glu, and Tyr residues of hen lysozyme. The effects of different minimization and/or sampling protocols for both single-conformer and multi-conformer calculations are studied. For single-conformer calculations it is found that the results are sensitive to the choice of all-hydrogen versus polar-hydrogen-only atomic models and to the minimization protocol chosen. The best overall agreement of single-conformer calculations with experiment is obtained with an all-hydrogen model and either a two-step minimization process or minimization using a high dielectric constant. Multi-conformational calculations give significantly improved agreement with experiment, slightly smaller shifts between model compound pKa values and calculated intrinsic pKa values, and reduced sensitivity of the intrinsic pKa calculations to the initial details of the structure compared to single-conformer calculations. The extent of these improvements depends on the type of minimization used during the generation of conformers, with more extensive minimization giving greater improvements. The ordering of the titrations of the active-site residues, Glu-35 and Asp-52, is particularly sensitive to the minimization and sampling protocols used. The balance of strong site-site interactions in the active site suggests a need for including site-site conformational correlations.

Animals↗

Improved time-frequency filtering of signal-averaged ECGs.

A recently proposed time-frequency filtering technique has shown promising results for the enhancement of signal-averaged electrocardiograms. This method weights the short-time Fourier transform of the ensemble-averaged signal, analogous to the spectral domain Wiener filtering of stationary signals. In effect, it is a self-designing, time-varying Wiener filter applied to the high-resolution electrocardiogram (HRECG). In this study, the authors empirically show that the performance of the proposed technique is about 2-3 dB lower over the critical late-potential portion of the HRECG than the optimal fixed-window, time-frequency filter based on ideal a priori knowledge of statistics. Although this ideal knowledge and performance is unattainable in practice, these results suggest that there remains potential for modest improvement. To narrow this gap in performance, improvements based on alternative structures for the time-frequency filter, including time-varying short-time Fourier transform windows, are proposed. Simulation results show that an improved fixed-window technique can potentially yield an improvement of about 1-1.5 dB. By using properly chosen time-varying windows, the performance could potentially be improved of about 1-1.5 dB. By using properly chosen time-varying windows, the performance could potentially be improved even further. Thus, the improved techniques could produce an HRECG using fewer averages than the existing method, or that could tolerate a lower initial signal-to-noise ratio.

Algorithms↗

Identification of physiological systems: estimation of linear time-varying dynamics with non-white inputs and noisy outputs.

A new technique to identify linear time-varying systems from ensembles of input-output realisations is presented. First, a correlation-based least-squares method is derived. This method consists of solving, for each sampling time, a matrix equation involving estimates of the input autocorrelation and input-output cross-correlation functions computed from data across the ensemble. Then, the matrix inverse needed to solve this matrix equation is replaced with a pseudo-inverse. The model is thus constrained to describe only those components of the dynamics that can be reliably identified. Ignoring 'unidentifiable' components has virtually no adverse effect on the predicted outputs. Simulation results demonstrate that the pseudoinverse technique yields more reliable estimates of the dynamics than a previously proposed least-squares technique when the inputs are coloured and the output signal-to-noise ratio (SNR) is low. With the input spectrum flat up to approximately 10% of the sampling rate and an output SNR of 5dB, the mean variance accounted for (VAF) between the true instantaneous impulse response functions (IRFs) and the instantaneous IRFs estimated with the least-squares technique was 0.2%. In contrast, the mean VAF between the true instantaneous IRFs and the instantaneous IRFs estimated with the pseudoinverse technique was 89.0%.

Ankle Joint↗

Resolving clusters in chaotic ensembles of globally coupled identical oscillators.

Clustering in ensembles of globally coupled identical chaotic oscillators is reconsidered using a twofold approach. Stability of clusters towards "emanation" of the elements is described with the evaporation Lyapunov exponents. It appears that direct numerical simulations of ensembles often lead to spurious clusters that have positive evaporation exponents, due to a numerical trap. We propose a numerical method that surmounts the spurious clustering. We also demonstrate that clustering can be very sensitive to the number of elements in the ensemble.

Journal Article↗

Conformational studies of an amphipathic peptide corresponding to human apolipoprotein A-II residues 18-30 with a C-terminal lipid binding motif EWLNS.

A peptide was designed and synthesized to enhance the lipid binding properties of a 13-residue fragment of apolipoprotein A-II. The peptide, VTDYGKDLMEKVKEWLNS [apoA-II(18-30)+], contains a five-residue amphipathic motif, EWLNS, at the C-terminus of apolipoprotein A-II residues 18-30. The lipid binding properties of apoA-II(18-30)+ were assessed using optical spectroscopy in the presence of sodium dodecyl sulfate (SDS), dodecylphosphocholine (DPC), tetradecyltrimethyl ammonium chloride (TMA) and dimyristoylphosphatidylcholine (DMPC). The fluorescence emission spectra and the circular dichroism data suggested that apoA-II(18-30)+ interacted most strongly with SDS and most weakly with DMPC. An ensemble of structures for apoA-II(18-30)+ in aqueous solution containing SDS was calculated using distance geometry/simulated annealing methods from 308 NOE-based distance restraints. The backbone (N-C-C = O) RMSD from the average structure of an ensemble of 15 out of 20 calculated structures was 0.54 +/- 0.16 A. Apart from some dynamic fraying at both termini, the distance geometry and simulated annealing calculations showed that apoA-II(18-30)+ adopted a well defined amphipathic helix with distinct hydrophobic and hydrophilic faces.

Amino Acid Sequence↗

Cluster synchronization modes in an ensemble of coupled chaotic oscillators.

Considering systems of diffusively coupled identical chaotic oscillators, an effective method to determine the possible states of cluster synchronization and ensure their stability is presented. The method, which may find applications in communication engineering and other fields of science and technology, is illustrated through concrete examples of coupled biological cell models.

Journal Article↗

Melting line of charged colloids from primitive model simulations.

We develop an efficient simulation method to study suspensions of charged spherical colloids using the primitive model. In this model, the colloids and the co- and counterions are represented by charged hard spheres, whereas the solvent is treated as a dielectric continuum. In order to speed up the simulations, we restrict the positions of the particles to a cubic lattice, which allows precalculation of the Coulombic interactions at the beginning of the simulation. Moreover, we use multiparticle cluster moves that make the Monte Carlo sampling more efficient. The simulations are performed in the semigrand canonical ensemble, where the chemical potential of the salt is fixed. Employing our method, we study a system consisting of colloids carrying a charge of 80 elementary charges and monovalent co- and counterions. At the colloid densities of our interest, we show that lattice effects are negligible for sufficiently fine lattices. We determine the fluid-solid melting line in a packing fraction eta-inverse screening length kappa plane and compare it with the melting line of charged colloids predicted by the Yukawa potential of the Derjaguin-Landau-Verwey-Overbeek theory. We find qualitative agreement with the Yukawa results, and we do not find any effects of many-body interactions. We discuss the difficulties involved in the mapping between the primitive model and the Yukawa model at high colloid packing fractions (eta>0.2).

Journal Article↗

Tandem machine learning for the identification of genes regulated by transcription factors.

BACKGROUND: The identification of promoter regions that are regulated by a given transcription factor has traditionally relied upon the identification and distributions of binding sites recognized by the factor. In this study, we have developed a tandem machine learning approach for the identification of regulatory target genes based on these parameters and on the corresponding binding site information contents that measure the affinities of the factor for these cognate elements. RESULTS: This method has been validated using models of DNA binding sites recognized by the xenobiotic-sensitive nuclear receptor, PXR/RXRalpha, for target genes within the human genome. An information theory-based weight matrix was first derived and refined from known PXR/RXRalpha binding sites. The promoter region of candidate genes was scanned with the weight matrix. A novel information density-based clustering algorithm was then used to identify clusters of information rich sites. Finally, transformed data representing metrics of location, strength and clustering of binding sites were used for classification of promoter regions using an ensemble approach involving neural networks, decision trees and Naïve Bayesian classification. The method was evaluated on a set of 24 known target genes and 288 genes known not to be regulated by PXR/RXRalpha. We report an average accuracy (proportion of correctly classified promoter regions) of 71%, sensitivity of 73%, and specificity of 70%, based on multiple cross-validation and the leave-one-out strategy. The performance on a test set of 13 genes showed that 10 were correctly classified. CONCLUSION: We have developed a machine learning approach for the successful detection of gene targets for transcription factors with high accuracy. The method has been validated for the transcription factor PXR/RXRalpha and has the potential to be extended to other transcription factors.

Algorithms↗

Mesoscale model of polymer melt structure: self-consistent mapping of molecular correlations to coarse-grained potentials.

Development and application of coarse-graining methods to condensed phases of macromolecules is an active area of research. Multiscale modeling of polymeric systems using coarse-graining methods presents unique challenges. Here we apply a coarse-graining method that self-consistently maps structural correlations from detailed molecular dynamics (MD) simulations of alkane oligomers onto coarse-grained potentials using a combination of MD and inverse Monte Carlo methods. Once derived, the coarse-grained potentials allow computationally efficient sampling of ensemble of conformations of significantly longer polyethylene chains. Conformational properties derived from coarse-grained simulations are in excellent agreement with experiments. The level of coarse graining provides a control over the balance of computational efficiency and retention of chemical identity of the underlying polymeric system. Challenges to extension and application of this and similar structure-based coarse-graining methods to model dynamics and phase behavior in polymeric systems are briefly discussed.

Journal Article↗