Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

An efficient method for detecting connectivity in neural ensembles.

Modern technology is allowing researchers to collect data from neural ensembles with a large number of units, and the analysis of interaction between these units can be very time consuming. Estimation of pairwise connectivity is the most common method of determining the neural 'network' but usually necessitates the production of numerous histograms for each pair considered. We present a method which will indicate which pairs in a network represent potential connections and thereby simplify the postexperimental analysis. The technique uses cross-interval information to create an n x n matrix which represents all possible connections in an n neuron ensemble and can be calculated recursively on-line. The performance of this technique is analyzed with respect to data size and strength of the connections. It is compared to 2 similar techniques that are also presented here, one in which perfect knowledge of the timing of the excitation is known, and one in which the timing can be bounded.

Computer Simulation↗

Ensemble competitive learning neural networks with reduced input dimension.

Conventional neural networks utilize all the dimensions of the original input patterns for training and classification. However, a particular attribute of the input patterns does not necessarily contribute to classification and may even cause misclassification in certain cases. A new ensemble competitive learning method using the reduced input dimension is proposed. In contrast to the previous ensemble neural networks which adjust learning parameters, the proposed method takes advantage of the information in each dimension of the input patterns. Since the degree of contribution of each attribute to classification is not known beforehand, the different input data sets with one dimension reduced are presented to multiple neural networks. The classification information from each competitive learning neural network is then combined to make a final decision for classification. In order to improve classification accuracy, the ambiguous output neurons are eliminated which cannot be assigned to any class after training. We use three consensus schemes to judge the classification using ensemble neural networks. The experimental results with remote sensing and speech data indicate the improved performance of the proposed method.

Algorithms↗

Determination of the chemical potential using energy-biased sampling.

An energy-biased method to evaluate ensemble averages requiring test-particle insertion is presented. The method is based on biasing the sampling within the subdomains of the test-particle configurational space with energies smaller than a given value freely assigned. These energy wells are located via unbiased random insertion over the whole configurational space and are sampled using the so-called Hit-and-Run algorithm, which uniformly samples compact regions of any shape immersed in a space of arbitrary dimensions. Because the bias is defined in terms of the energy landscape it can be exactly corrected to obtain the unbiased distribution. The test-particle energy distribution is then combined with the Bennett relation for the evaluation of the chemical potential. We apply this protocol to a system with relatively small probability of low-energy test-particle insertion, liquid argon at high density and low temperature, and show that the energy-biased Bennett method is around five times more efficient than the standard Bennett method. A similar performance gain is observed in the reconstruction of the energy distribution.

Journal Article↗

Miniature carrier with six independently moveable electrodes for recording of multiple single-units in the cerebellar cortex of awake rats.

Ensemble recording in cerebellar cortex of awake rats presents unique methodological challenges not encountered when recording from the cerebral cortex or from deep brain structures with more homogeneous cell populations. Compared to the cerebral cortex, removal of dura over the cerebellum evokes pronounced swelling, and insertion of multiple closely spaced electrodes in the cerebellar cortex causes considerable dimpling (Welsh JP, Schwartz C. Multielectrode recording from the cerebellum. In: Nicolelis MAL, editor. Methods for Neural Ensemble Recordings, CRC Methods in Neuroscience Series. Boca Raton, FL: CRC Press LLC, 1999, pp. 79-100). Also, a repetitious and well-defined neural circuit characterizes the cerebellar cortex across its entire surface. With conventional multi-electrode methods, such as chronically implanted bundles or arrays of microwires, the risk of disrupting the cerebellar cytoarchitecture is high. In most conventional multi-electrode systems, electrodes have rather low impedance and cannot be moved independently after implantation. These limitations make proper unit isolation, necessary to identify each of the recorded cerebellar units, very difficult. We designed a lightweight (14 g), miniature (base plate: 19 x 23 mm; total height: 16 mm) multi-electrode system to allow for the chronic implantation of six independently moveable sharp electrodes with high impedance, in the cerebellar cortex. The six electrodes are arranged in a 2 x 3 matrix (inter-electrode distance: 0.6 mm). At any time after the implantation the vertical position of each individual electrode can be adjusted by screwing spring-loaded electrode heads up or down. The system preserves the integrity of the cerebellar cytoarchitecture, and enables easy isolation and identification of individual cerebellar units in awake, freely moving rats.

Animals↗

Simulation studies of the fidelity of biomolecular structure ensemble recreation.

We examine the ability of Bayesian methods to recreate structural ensembles for partially folded molecules from averaged data. Specifically we test the ability of various algorithms to recreate different transition state ensembles for folding proteins using a multiple replica simulation algorithm using input from "gold standard" reference ensembles that were first generated with a Go-like Hamiltonian having nonpairwise additive terms. A set of low resolution data, which function as the "experimental" phi values, were first constructed from this reference ensemble. The resulting phi values were then treated as one would treat laboratory experimental data and were used as input in the replica reconstruction algorithm. The resulting ensembles of structures obtained by the replica algorithm were compared to the gold standard reference ensemble, from which those "data" were, in fact, obtained. It is found that for a unimodal transition state ensemble with a low barrier, the multiple replica algorithm does recreate the reference ensemble fairly successfully when no experimental error is assumed. The Kolmogorov-Smirnov test as well as principal component analysis show that the overlap of the recovered and reference ensembles is significantly enhanced when multiple replicas are used. Reduction of the multiple replica ensembles by clustering successfully yields subensembles with close similarity to the reference ensembles. On the other hand, for a high barrier transition state with two distinct transition state ensembles, the single replica algorithm only samples a few structures of one of the reference ensemble basins. This is due to the fact that the phi values are intrinsically ensemble averaged quantities. The replica algorithm with multiple copies does sample both reference ensemble basins. In contrast to the single replica case, the multiple replicas are constrained to reproduce the average phi values, but allow fluctuations in phi for each individual copy. These fluctuations facilitate a more faithful sampling of the reference ensemble basins. Finally, we test how robustly the reconstruction algorithm can function by introducing errors in phi comparable in magnitude to those suggested by some authors. In this circumstance we observe that the chances of ensemble recovery with the replica algorithm are poor using a single replica, but are improved when multiple copies are used. A multimodal transition state ensemble, however, turns out to be more sensitive to large errors in phi (if appropriately gauged) and attempts at successful recreation of the reference ensemble with simple replica algorithms can fall short.

Algorithms↗

Crystallization of confined non-Brownian spheres by vibrational annealing.

We introduce an experimental method to crystallize ensembles of non-Brownian spheres confined in narrow containers. The method is based on programmed vibrations and a cooling procedure (annealing). Starting with a granular gas, the system slowly relaxes into a solid ordered structure: Body-centered-tetragonal and face-centered-cubic single crystals are obtained depending on the dimensions of the capillaries. Dry and lubricated beads behave differently, indicating that a sticking coefficient between the particles is important in the dynamics of the crystallization.

Journal Article↗

The significance of neural ensemble codes during behavior and cognition.

The development of techniques to record from populations of neurons has made it possible to ask questions concerning the encoding of task-relevant information in awake, behaving animals. The issue of how groups of neurons within different brain structures register and retrieve representations of behaviorally significant events can now be addressed using multineuron-recording techniques. This review examines recent studies employing simultaneous recording of ten or more individual neurons in the mammalian brain. A major issue discussed is whether ensemble information content reconstructed from single-neuron recordings may be underestimated if compared to ensembles where those same neurons were recorded simultaneously. The mechanics of ensemble information encoding in the hippocampus is illustrated from population statistical analyses of ensemble activity during performance of a delay task. Detailed descriptions of methods of extracting ensemble information, as well as cross-correlational analyses, are discussed in the context of emergent issues regarding interpretation of ensemble data.

Animals↗

The classification of cancer based on DNA microarray data that uses diverse ensemble genetic programming.

OBJECT: The classification of cancer based on gene expression data is one of the most important procedures in bioinformatics. In order to obtain highly accurate results, ensemble approaches have been applied when classifying DNA microarray data. Diversity is very important in these ensemble approaches, but it is difficult to apply conventional diversity measures when there are only a few training samples available. Key issues that need to be addressed under such circumstances are the development of a new ensemble approach that can enhance the successful classification of these datasets. MATERIALS AND METHODS: An effective ensemble approach that does use diversity in genetic programming is proposed. This diversity is measured by comparing the structure of the classification rules instead of output-based diversity estimating. RESULTS: Experiments performed on common gene expression datasets (such as lymphoma cancer dataset, lung cancer dataset and ovarian cancer dataset) demonstrate the performance of the proposed method in relation to the conventional approaches. CONCLUSION: Diversity measured by comparing the structure of the classification rules obtained by genetic programming is useful to improve the performance of the ensemble classifier.

Artificial Intelligence↗

Comparing systemic properties of ensembles of biological networks by graphical and statistical methods.

MOTIVATION: When dealing with questions that concern a general class of models for biological networks, large numbers of distinct models within the class can be grouped into an ensemble that gives a statistical view of the properties for the general class. Comparing properties of different ensembles through the use of point measures (e.g. medians, standard deviations, correlation coefficients) can mask inhomogeneities in the correlations between properties. We are therefore motivated to develop strategies that allow these inhomogeneities to be more easily detected. RESULTS: Methods are described for constructing ensembles of models within the context of a Mathematically Controlled Comparison. A Density of Ratios Plot for a given systemic property is then defined as follows: the y axis represents the value of the systemic property in a reference model divided by the value in the alternative model, and the x axis represents the value of the systemic property in the reference model. Techniques involving moving quantiles are introduced to generate secondary plots in which correlations and inhomogeneities in correlations are more easily detected. Several examples that illustrate the advantages of these techniques are presented and discussed.

Biometry↗

Structural interpretation of hydrogen exchange protection factors in proteins: characterization of the native state fluctuations of CI2.

Protection factors obtained from equilibrium hydrogen exchange experiments are an important source of structural information on both native and nonnative states of proteins. We present a method for determining ensembles of protein structures by using hydrogen exchange data as restraints in molecular dynamics simulations in conjunction with an empirical force-field. The method is applied to determine the ensemble of structures representing the native state of chymotrypsin inhibitor 2 (CI2), including the rare, large fluctuations responsible for hydrogen exchange.

Chymotrypsin↗

Ensemble reactions of neurons in the reticular formation of the cat--categorization of inputs and temporal influences.

A method of recording ensemble reactions of reticular neurons was applied in anesthetized cats to obtain information about categorization of inputs and the influence of the time factor upon a pattern of responses, reactions to stimulus intensity changes and to complex stimuli, and temporal influences with their relation to plastic changes induced by iterative stimulation. In spite of some variability of the pattern of responses evoked in the reticular formation (RF) by electrical stimulation of the same source, reactions of the parenchyma of the RF differentiate heterotopic stimuli; reactions evoked from receptor areas situated closely to each other exhibited a similar pattern whereas less resemblance was found in the case of responses evoked from sources with a heterosegmental projection and from contralateral sources. Stimulus intensity changes did not influence substantially the pattern of recordings in the RF. The noticeable effect resulting from the increase of stimulus intensity was enlargement of responses. A complex stimulus (two heterotopic stimuli applied simultaneously) was manifested by a reaction which, in comparison with reactions evoked from each of both sources separately, had a different pattern. A trend of changes in the size of responses evoked in the RF by continuous stimulation (60 stimuli, 0.3 Hz) was observed only in some cases. Abrupt alterations of the stimulation rate (0.3 Hz in equilibrium 1 Hz, 0.3 Hz in equilibrium 2 Hz) resulted in a prompt change in responses. With constant stimulation conditions, this activity was fairly constant with an absence of trends. Persistent changes were observed only in a minority of cases. The neuronal apparatus of the RF computes ongoing effects on the basis of ability to discriminate different inputs and to readily change its function. The extent to which the RF holds information about the antecedent "history" of its functioning and the extent to which intrinsic mechanisms of the RF by itself are responsible for behavioral habituation remains problematical.

Animals↗

Simulation estimates of cloud points of polydisperse fluids.

We describe two distinct approaches to obtaining the cloud-point densities and coexistence properties of polydisperse fluid mixtures by Monte Carlo simulation within the grand-canonical ensemble. The first method determines the chemical potential distribution mu(sigma) (with the polydisperse attribute) under the constraint that the ensemble average of the particle density distribution rho(sigma) match a prescribed parent form. Within the region of phase coexistence (delineated by the cloud curve) this leads to a distribution of the fluctuating overall particle density n, p(n), that necessarily has unequal peak weights in order to satisfy a generalized lever rule. A theoretical analysis shows that as a consequence, finite-size corrections to estimates of coexistence properties are power laws in the system size. The second method assigns mu(sigma) such that an equal-peak-weight criterion is satisfied for p(n) for all points within the coexistence region. However, since equal volumes of the coexisting phases cannot satisfy the lever rule for the prescribed parent, their relative contributions must be weighted appropriately when determining mu(sigma). We show how to ascertain the requisite weight factor operationally. A theoretical analysis of the second method suggests that it leads to finite-size corrections to estimates of coexistence properties which are exponentially small in the system size. The scaling predictions for both methods are tested via Monte Carlo simulations of a polydisperse lattice-gas model near its cloud curve, the results showing excellent quantitative agreement with the theory.

Journal Article↗

Ab initio computational modeling of loops in G-protein-coupled receptors: lessons from the crystal structure of rhodopsin.

With the help of the crystal structure of rhodopsin an ab initio method has been developed to calculate the three-dimensional structure of the loops that connect the transmembrane helices (TMHs). The goal of this procedure is to calculate the loop structures in other G-protein coupled receptors (GPCRs) for which only model coordinates of the TMHs are available. To mimic this situation a construct of rhodopsin was used that only includes the experimental coordinates of the TMHs while the rest of the structure, including the terminal domains, has been removed. To calculate the structure of the loops a method was designed based on Monte Carlo (MC) simulations which use a temperature annealing protocol, and a scaled collective variables (SCV) technique with proper structural constraints. Because only part of the protein is used in the calculations the usual approach of modeling loops, which consists of finding a single, lowest energy conformation of the system, is abandoned because such a single structure may not be a representative member of the native ensemble. Instead, the method was designed to generate structural ensembles from which the single lowest free energy ensemble is identified as representative of the native folding of the loop. To find the native ensemble a successive series of SCV-MC simulations are carried out to allow the loops to undergo structural changes in a controlled manner. To increase the chances of finding the native funnel for the loop, some of the SCV-MC simulations are carried out at elevated temperatures. The native ensemble can be identified by an MC search starting from any conformation already in the native funnel. The hypothesis is that native structures are trapped in the conformational space because of the high-energy barriers that surround the native funnel. The existence of such ensembles is demonstrated by generating multiple copies of the loops from their crystal structures in rhodopsin and carrying out an extended SCV-MC search. For the extracellular loops e1 and e3, and the intracellular loop i1 that were used in this work, the procedure resulted in dense clusters of structures with Calpha-RMSD approximately 0.5 angstroms. To test the predictive power of the method the crystal structure of each loop was replaced by its extended conformations. For e1 and i1 the procedure identifies native clusters with Calpha-RMSD approximately 0.5 angstroms and good structural overlap of the side chains; for e3, two clusters were found with Calpha-RMSD approximately 1.1 angstroms each, but with poor overlap of the side chains. Further searching led to a single cluster with lower Calpha-RMSD but higher energy than the two previous clusters. This discrepancy was found to be due to the missing elements in the constructs available from experiment for use in the calculations. Because this problem will likely appear whenever parts of the structural information are missing, possible solutions are discussed.

Computer Simulation↗

Solution structure of the Grb2 N-terminal SH3 domain complexed with a ten-residue peptide derived from SOS: direct refinement against NOEs, J-couplings and 1H and 13C chemical shifts.

Refined ensembles of solution structures have been calculated for the N-terminal SH3 domain of Grb2 (N-SH3) complexed with the ac-VPPPVPPRRR-nh2 peptide derived from residues 1135 to 1144 of the mouse SOS-1 sequence. NMR spectra obtained from different combinations of both 13C-15N-labeled and unlabeled N-SH3 and SOS peptide fragment were used to obtain stereo-assignments for pro-chiral groups of the peptide, angle restraints via heteronuclear coupling constants, and complete 1H, 13C, and 15N resonance assignments for both molecules. One ensemble of structures was calculated using conventional methods while a second ensemble was generated by including additional direct refinements against both 1H and 13C(alpha)/13C(beta) chemical shifts. In both ensembles, the protein:peptide interface is highly resolved, reflecting the inclusion of 110 inter-molecular nuclear Overhauser enhancement (NOE) distance restraints. The first and second peptide-binding sub-sites of N-SH3 interact with structurally well-defined portions of the peptide. These interactions include hydrogen bonds and extensive hydrophobic contacts. In the third highly acidic sub-site, the conformation of the peptide Arg8 side-chain is partially ordered by a set of NOE restraints to the Trp36 ring protons. Overall, several lines of evidence point to dynamical averaging of peptide and N-SH3 side-chain conformations in the third subsite. These conformations are characterized by transient charge stabilized hydrogen bond interactions between the peptide arginine side-chain hydrogen bond donors and either single, or possibly multiple, acceptor(s) in the third peptide-binding sub-site.

Adaptor Proteins, Signal Transducing↗

Ring closure probabilities for DNA fragments by Monte Carlo simulation.

The rate of ligation of DNA molecules into circular forms depends on the ring closure probability, commonly called the j-factor, which is a sensitive measure of the extent to which thermal fluctuations contribute to bending and twisting of DNA molecules in solution. We present a theoretical treatment of the cyclization equilibria of DNA that employs a special Monte Carlo method for generating large ensembles of model DNA chains. Using this method, the chain length dependence of the j-factor was calculated for molecules. in the size range 250 to 2000 base-pairs. The Monte Carlo results are compared with recent analytical theory and experimental data. We show that a value of 475 A for the persistence length of DNA, close to values measured by a number of other methods, is in excellent agreement with the cyclization results. Preliminary applications of the Monte Carlo method to the problem of systematically bent DNA molecules are presented. The calculated j-factor is shown to be very sensitive to the amount of bending in these fragments. This fact suggests that ligase closure measurements of systematically bent DNA molecules should be a useful method for studying sequence-directed bending in DNA.

Cyclization↗

Temperature and density extrapolations in canonical ensemble monte carlo simulations

We show how to use the multiple histogram method to combine canonical ensemble Monte Carlo simulations made at different temperatures and densities. The method can be applied to study systems of particles with arbitrary interaction potential and to compute the thermodynamic properties over a range of temperatures and densities. The calculation of the Helmholtz free energy relative to some thermodynamic reference state enables us to study phase coexistence properties. We test the method on the Lennard-Jones fluids for which many results are available.

Journal Article↗

Automated clustering of ensembles of alternative models in protein structure databases.

Experimentally determined protein structures have been classified in different public databases according to their structural and evolutionary relationships. Frequently, alternative structural models, determined using X-ray crystallography or NMR spectroscopy, are available for a protein. These models can present significant structural dissimilarity. Currently there is no classification available for these alternative structures. In order to classify them, we developed STRuster, an automated method for clustering ensembles of structural models according to their backbone structure. The method is based on the calculation of carbon alpha (Calpha) distance matrices. Two filters are applied in the calculation of the dissimilarity measure in order to identify both large and small (but significant) backbone conformational changes. The resulting dissimilarity value is used for hierarchical clustering and partitioning around medoids (PAM). Hierarchical clustering reflects the hierarchy of similarities between all pairs of models, while PAM groups the models into the 'optimal' number of clusters. The method has been applied to cluster the structures in each SCOP species level and can be easily applied to any other sets of conformers. The results are available at: http://bioinf.mpi-sb.mpg.de/projects/struster/.

Aldehyde-Lyases↗

Conditioned spikes: a simple and fast method to represent rates and temporal patterns in multielectrode recordings.

Increasing evidence suggests that the brain utilizes distributed codes that can only be analyzed by simultaneously recording the activity of multiple neurons. This paper introduces a new methodology for studying neural ensemble recordings. The method uses a novel representation to provide complementary information about the stimuli which are contained in the temporal pattern of the spike sequence. By using this procedure, a high correlation of synchronized events with stimuli times is apparent. To quantify the results and to compare the performance of this method against the most traditional raster plot, we have used Fano factor and cross-correlation analysis. Our results suggest that several consecutive spikes from different neurons within an extended time window may encode behaviorally relevant information. We propose that this new representation, in addition to the other approaches currently used (standard raster plots, multivariate statistical methods, neuronal networks, information theory, etc.), can be a useful procedure to describe population spike dynamics.

Action Potentials↗