Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Enhanced sampling in numerical path integration: an approximation for the quantum statistical density matrix based on the nonextensive thermostatistics.

Here we examine a proposed approximation for the quantum statistical density matrix motivated by the nonextensive thermostatistics of Tsallis and co-workers. The approximation involves replacing the physical potential energy with an effective one, corresponding to a generalized nonextensive statistical ensemble. We examine the convergence properties of averages calculated using the effective potential, and introduce a related method for enhanced sampling in numerical path integration. As a necessary measure, path integral energy estimators are introduced for potentials that involve explicit temperature dependence. This sampling method is found to be effective for path integral simulations involving broken ergodicity.

Journal Article↗

Interplay between structure and size in a critical crystal nucleus.

We study the kinetics of crystal nucleation of an undercooled Lennard-Jones liquid using various path-sampling methods. We obtain the rate constant and elucidate the pathways for crystal nucleation. Analysis of the path ensemble reveals that crystal nucleation occurs along many different pathways, in which critical solid nuclei can be small, compact, and face centered cubic, but also large, less ordered, and more body centered cubic. The reaction coordinate thus includes, besides the cluster size, also the quality of the crystal structure.

Journal Article↗

Ensemble tracking.

We consider tracking as a binary classification problem, where an ensemble of weak classifiers is trained online to distinguish between the object and the background. The ensemble of weak classifiers is combined into a strong classifier using AdaBoost. The strong classifier is then used to label pixels in the next frame as either belonging to the object or the background, giving a confidence map. The peak of the map and, hence, the new position of the object, is found using mean shift. Temporal coherence is maintained by updating the ensemble with new weak classifiers that are trained online during tracking. We show a realization of this method and demonstrate it on several video sequences.

Algorithms↗

Relating HIV-1 sequence variation to replication capacity via trees and forests.

The problem of relating genotype (as represented by amino acid sequence) to phenotypes is distinguished from standard regression problems by the nature of sequence data. Here we investigate an instance of such a problem where the phenotype of interest is HIV-1 replication capacity and contiguous segments of protease and reverse transcriptase sequence constitutes genotype. A variety of data analytic methods have been proposed in this context. Shortcomings of select techniques are contrasted with the advantages afforded by tree-structured methods. However, tree-structured methods, in turn, have been criticized on grounds of only enjoying modest predictive performance. A number of ensemble approaches (bagging, boosting, random forests) have recently emerged, devised to overcome this deficiency. We evaluate random forests as applied in this setting, and detail why prediction gains obtained in other situations are not realized. Other approaches including logic regression, support vector machines and neural networks are also applied. We interpret results in terms of HIV-1 reverse transcriptase structure and function.

Journal Article↗

[The structure of the background EEG of the rabbit in the high-frequency portion of the spectrum].

With the purpose of studying the character and structure of high frequency bioelectric activity of rabbits cerebral cortex in the state of calm alertness, the EEG ensembles of different areas of the cortex (sensorimotor, visual, acoustic) and dorsal hippocampus were studied with FFT method. A supposition was made about the presence of systemic organization of the background EEG in rabbits cerebral cortex, reflected, in particular, in the presence of determined components both of chaotic and rhythmic character having different degrees of manifestation. Heterogeneity was revealed in distribution of energies of spectral EEG components in the studied frequency ranges from 14.7 to 100 Hz with predominance of total specific energy value in the band of 14.7-60 Hz. In coherence functions of all the studied pairs of EEG leads rhythmic component, stable in time, was absent. Functions of the mean EEG coherence in the band of 61-100 Hz had significantly greater values in comparison with the values in the band of 14.7-40 Hz.

Animals↗

Role of spike timing in the forelimb somatosensory cortex of the rat.

The aim of this study was to test the hypothesis that the significance of spike timing in somatosensory processing is not a specific feature of the whisker cortex but a more general characteristic of the primary somatosensory cortex. We recorded ensembles of neurons using microwire arrays implanted in the deep layers of the forelimb region of the rat primary somatosensory cortex in response to step stimuli delivered to the cutaneous surface of the contralateral body. We used a recently developed peristimulus time histogram (PSTH)-based classification method to investigate the temporal precision of the code by evaluating how changing the bin size (from 40 to 1 msec) would affect the ability of the ensemble responses to discriminate stimulus location on a single-trial basis. The information related to the discrimination was redundantly distributed within the ensembles, and the ability to discriminate stimulus location increased when decreasing the bin size, reaching a maximum at 4 msec. In our experiment, at 4 msec bin size the first spike per neuron after the stimulus conveyed almost as much information as the entire responses, so the temporal precision of the code was preserved in the first spikes. Subsequent spikes were less frequent but conveyed more information per spike. Finally, not only the trials correctly classified but also the trials incorrectly classified conveyed information about stimulus location with a similar temporal precision. We conclude that the role of spike timing in cortical somatosensory processing is not an exclusive feature of the highly specialized rat trigeminal system, but a more general property of the rat primary somatosensory cortex.

Action Potentials↗

Adaptive clutter filtering via blind source separation for two-dimensional ultrasonic blood velocity measurement.

A method for adaptive clutter rejection via blind source separation (BSS) using principal and independent component analyses is presented in application to blood velocity measurement in the carotid artery. In particular, the filtering method's efficacy for eliminating clutter and preserving lateral blood flow signal components is presented. The performance of IIR filters is compromised by shorth data ensembles (10 to 20 temporal samples) as implemented for color-flow and high frame-rate imaging due to initialization requirements. Further, the ultrasonic imaging system's transfer function maps axial wall and lateral blood motion to overlapping spectra. As such, frequency domain-based approaches to wall filtering are ineffective for distinguishing wall from blood motion signals. Rather than operating in the frequency domain. BSS performs clutter rejection by decomposing the input data ensemble into N constitutive source signals in time, where N is the ensemble length. Source signal energy coupled with respective signal depth and time course profiles reveal which source signals correspond to blood, noise and clutter components. Clutter components may then be removed without disruption of lateral blood flow information needed for two-dimensional blood velocity measurement. A simplistic data simulation is employed to offer an intuitive understanding of BSS methods for signal separation. The adaptive BSS filter is further demonstrated using a Field II simulation of blood flow through the carotid artery including tissue motion. BSS clutter filter performance is compared to the performance of FIR, IIR and polynomial regression clutter filters. Finally, the filter is employed for clinical application using a Siemens Elegra scanner, carotid artery data with lateral blood flow collected from healthy volunteers, and Speckle Tracking; velocity magnitude and angle profiles are shown. Once again, the BSS clutter filter is contrasted to FIR, IIR and polynomial regression clutter filters using clinical examples. Velocities computed with Speckle Tracking after BSS wall filtering are highest in the center of the artery and diminish to low velocities near the vessel walls, with velocity magnitudes consistent with physiological expectations. These results demonstrate that the BSS adaptive filter sufficiently suppresses wall motion signal for clinical lateral blood velocity measurement using data ensembles suitable for color-flow and high frame-rate imaging.

Blood Flow Velocity↗

Bias annealing: a method for obtaining transition paths de novo.

Computational studies of dynamics in complex systems require means for generating reactive trajectories with minimum knowledge about the processes of interest. Here, we introduce a method for generating transition paths when an existing one is not already available. Starting from biased paths obtained from steered molecular dynamics, we use a Monte Carlo procedure in the space of whole trajectories to shift gradually to sampling an ensemble of unbiased paths. Application to basin-to-basin hopping in a two-dimensional model system and nucleotide-flipping by a DNA repair protein demonstrates that the method can efficiently yield unbiased reactive trajectories even when the initial steered dynamics differ significantly. The relation of the method to others and the physical basis for its success are discussed.

Journal Article↗

Generalized simulated tempering realized on expanded ensembles of non-Boltzmann weights.

A generalized version of the simulated tempering operated in the expanded ensembles of non-Boltzmann weights has been proposed to mitigate a quasiergodicity problem occurring in simulations of rough energy landscapes. In contrast to conventional simulated tempering employing the Boltzmann weight, our method utilizes a parametrized, generalized distribution as a workhorse for stochastic exchanges of configurations and subensembles transitions, which allows a considerable enhancement for the rate of convergence of Monte Carlo and molecular dynamics simulations using delocalized weights. A feature of our method is that the exploration of the parameter space encouraging subensembles transitions is greatly accelerated using the dynamic update scheme for the weight via the average guide specific to the energy distribution. The performance and characteristic feature of our method have been validated in the liquid-solid transition of Lennard-Jones clusters and the conformational sampling of alanine dipeptide by taking two types of Tsallis [C. Tsallis, J. Stat. Phys. 52, 479 (1988)] expanded ensembles associated with different parametrization schemes.

Computer Simulation↗

Methods for peptide identification by spectral comparison.

BACKGROUND: Tandem mass spectrometry followed by database search is currently the predominant technology for peptide sequencing in shotgun proteomics experiments. Most methods compare experimentally observed spectra to the theoretical spectra predicted from the sequences in protein databases. There is a growing interest, however, in comparing unknown experimental spectra to a library of previously identified spectra. This approach has the advantage of taking into account instrument-dependent factors and peptide-specific differences in fragmentation probabilities. It is also computationally more efficient for high-throughput proteomics studies. RESULTS: This paper investigates computational issues related to this spectral comparison approach. Different methods have been empirically evaluated over several large sets of spectra. First, we illustrate that the peak intensities follow a Poisson distribution. This implies that applying a square root transform will optimally stabilize the peak intensity variance. Our results show that the square root did indeed outperform other transforms, resulting in improved accuracy of spectral matching. Second, different measures of spectral similarity were compared, and the results illustrated that the correlation coefficient was most robust. Finally, we examine how to assemble multiple spectra associated with the same peptide to generate a synthetic reference spectrum. Ensemble averaging is shown to provide the best combination of accuracy and efficiency. CONCLUSION: Our results demonstrate that when combined, these methods can boost the sensitivity and specificity of spectral comparison. Therefore they are capable of enhancing and complementing existing tools for consistent and accurate peptide identification.

Journal Article↗

An automated method for modeling proteins on known templates using distance geometry.

We present an automated method incorporated into a software package, FOLDER, to fold a protein sequence on a given three-dimensional (3D) template. Starting with the sequence alignment of a family of homologous proteins, tertiary structures are modeled using the known 3D structure of one member of the family as a template. Homologous interatomic distances from the template are used as constraints. For nonhomologous regions in the model protein, the lower and the upper bounds for the interatomic distances are imposed by steric constraints and the globular dimensions of the template, respectively. Distance geometry is used to embed an ensemble of structures consistent with these distance bounds. Structures are selected from this ensemble based on minimal distance error criteria, after a penalty function optimization step. These structures are then refined using energy optimization methods. The method is tested by simulating the alpha-chain of horse hemoglobin using the alpha-chain of human hemoglobin as the template and by comparing the generated models with the crystal structure of the alpha-chain of horse hemoglobin. We also test the packing efficiency of this method by reconstructing the atomic positions of the interior side chains beyond C beta atoms of a protein domain from a known 3D structure. In both test cases, models retain the template constraints and any additionally imposed constraints while the packing of the interior residues is optimized with no short contacts or bond deformations. To demonstrate the use of this method in simulating structures of proteins with nonhomologous disulfides, we construct a model of murine interleukin (IL)-4 using the NMR structure of human IL-4 as the template. The resulting geometry of the nonhomologous disulfide in the model structure for murine IL-4 is consistent with standard disulfide geometry.

Algorithms↗

Improving the accuracy of protein pKa calculations: conformational averaging versus the average structure.

Several methods for including the conformational flexibility of proteins in the calculation of titration curves are compared. The methods use the linearized Poisson-Boltzmann equation to calculate the electrostatic free energies of solvation and are applied to bovine pancreatic trypsin inhibitor (BPTI) and hen egg-white lysozyme (HEWL). An ensemble of conformations is generated by a molecular dynamics simulation of the proteins with explicit solvent. The average titration curve of the ensemble is calculated in three different ways: an average structure is used for the pKa calculation; the electrostatic interaction free energies are averaged and used for the pKa calculation; and the titration curve for each structure is calculated and the curves are averaged. The three averaging methods give very similar results and improve the pKa values to approximately the same degree. This suggests, in contrast to implications from other work, that the observed improvement of pKa values in the present studies is due not to averaging over an ensemble of structures, but rather to the generation of a single properly averaged structure for the pKa calculation.

Aprotinin↗

Using ensemble classifier to identify membrane protein types.

Predicting membrane protein type is both an important and challenging topic in current molecular and cellular biology. This is because knowledge of membrane protein type often provides useful clues for determining, or sheds light upon, the function of an uncharacterized membrane protein. With the explosion of newly-found protein sequences in the post-genomic era, it is in a great demand to develop a computational method for fast and reliably identifying the types of membrane proteins according to their primary sequences. In this paper, a novel classifier, the so-called "ensemble classifier", was introduced. It is formed by fusing a set of nearest neighbor (NN) classifiers, each of which is defined in a different pseudo amino acid composition space. The type for a query protein is determined by the outcome of voting among these constituent individual classifiers. It was demonstrated through the self-consistency test, jackknife test, and independent dataset test that the ensemble classifier outperformed other existing classifiers widely used in biological literatures. It is anticipated that the idea of ensemble classifier can also be used to improve the prediction quality in classifying other attributes of proteins according to their sequences.

Algorithms↗

Ensemble attribute profile clustering: discovering and characterizing groups of genes with similar patterns of biological features.

BACKGROUND: Ensemble attribute profile clustering is a novel, text-based strategy for analyzing a user-defined list of genes and/or proteins. The strategy exploits annotation data present in gene-centered corpora and utilizes ideas from statistical information retrieval to discover and characterize properties shared by subsets of the list. The practical utility of this method is demonstrated by employing it in a retrospective study of two non-overlapping sets of genes defined by a published investigation as markers for normal human breast luminal epithelial cells and myoepithelial cells. RESULTS: Each genetic locus was characterized using a finite set of biological properties and represented as a vector of features indicating attributes associated with the locus (a gene attribute profile). In this study, the vector space models for a pre-defined list of genes were constructed from the Gene Ontology (GO) terms and the Conserved Domain Database (CDD) protein domain terms assigned to the loci by the gene-centered corpus LocusLink. This data set of GO- and CDD-based gene attribute profiles, vectors of binary random variables, was used to estimate multiple finite mixture models and each ensuing model utilized to partition the profiles into clusters. The resultant partitionings were combined using a unanimous voting scheme to produce consensus clusters, sets of profiles that co-occurred consistently in the same cluster. Attributes that were important in defining the genes assigned to a consensus cluster were identified. The clusters and their attributes were inspected to ascertain the GO and CDD terms most associated with subsets of genes and in conjunction with external knowledge such as chromosomal location, used to gain functional insights into human breast biology. The 52 luminal epithelial cell markers and 89 myoepithelial cell markers are disjoint sets of genes. Ensemble attribute profile clustering-based analysis indicated that both lists contained groups of genes with the functional properties of membrane receptor biology/signal transduction and nucleic acid binding/transcription. A subset of the luminal markers was associated with metabolic and oxidoreductase activities, whereas a subset of myoepithelial markers was associated with protein hydrolase activity. CONCLUSION: Given a set of genes and/or proteins associated with a phenomenon, process or system of interest, ensemble attribute profile clustering provides a simple method for collating and sythesizing the annotation data pertaining to them that are present in text-based, gene-centered corpora. The results provide information about properties common and unique to subsets of the list and hence insights into the biology of the problem under investigation.

Algorithms↗

Comparison of rotation models for describing DNA conformations: application to static and polymorphic forms.

A new method, based on a space-fixed rotation axis, or local helix axis, is proposed for the calculation of the relative orientation variables for a sequence of base pairs. With this method, orientation variables are determined through the rotation of a base pair about this axis. These variables uniquely determine a set of helical variables, similar to the roll, tilt, and twist, commonly used for a description of spatial orientations of internally rigid base pairs. The proposed identification of roll and tilt with the direction cosines of the space-fixed rotation axis agrees well with their customary definitions as the openings of the angles between adjoining base pairs toward the minor groove and toward the ascending (5' to 3') backbone strand, respectively. These new variables permit a more direct physical comprehension of DNA conformations and also the behavior of self-complementary sequences. These direction cosines, together with the rotation angle about the space-fixed axis, form a set of three independent orientation variables of the bases that afford some advantages over the variously defined twist, roll, and tilt angles, either for static or average forms. An example for the static form of these variables is shown through their use to interpret crystal coordinates. An example for the average of orientation variables is based on statistical calculations. In this example, the orientation variables, together with the translational variables that describe the relative displacements of a pair of adjacent base pairs, form a canonically distributed ensemble in phase space spanned by these variables. Two sets of conformational variables are generated by using two different methods for performing rotation operations on the sequences of base pairs. The first method is based on the new single rotation about a space-fixed axis of rotation. This space-fixed axis of rotation is, in fact, the local helical axis as constructed previously by others. The second method is based on three consecutive rotations by Euler angles. Because of large flexibilities and anisotropies along various conformational variables of DNA base pairs, the two sets of generated conformational variables, based on these two different methods of performing rotation operations, lead to slightly different sets of structurally different, but energetically equivalent, spatial arrangements of the base pairs.

Base Composition↗

Molecular modeling of antibody combining sites.

Each of the six CDRs of Gloop2 is shown with the modeled structure in. Overall, the results obtained using the combined algorithm are similar in accuracy to those achieved using the canonical method of Chothia et al. However, the canonical method is limited to those loops where the key residues identified by Chothia are present. With the number of antibody structures currently available, it is not possible to classify CDR-H3 into canonical ensembles. Additionally, a small percentage of examples in the remaining CDRs do not match the current canonical classifications and the protein engineer may well wish to mutate the key residues, precluding the use of Chothia's method for modeling the resulting conformation. Thus the best approach appears to be to use Chothia's method (at least to model the backbone conformation) when the loop to be modeled is represented in the database of canonical structures. Any other loops, either unrepresented among the known canonicals (including CDR-H3), or where mutations have been made to the key residues, may then be modeled by the combined algorithm presented here.

Algorithms↗

Spectral analysis of full field digital mammography data.

The spectral content of mammograms acquired from using a full field digital mammography (FFDM) system are analyzed. Fourier methods are used to show that the FFDM image power spectra obey an inverse power law; in an average sense, the images may be considered as 1/f fields. Two data representations are analyzed and compared (1) the raw data, and (2) the logarithm of the raw data. Two methods are employed to analyze the power spectra (1) a technique based on integrating the Fourier plane with octave ring sectioning developed previously, and (2) an approach based on integrating the Fourier plane using rings of constant width developed for this work. Both methods allow theoretical modeling. Numerical analysis indicates that the effects due to the transformation influence the power spectra measurements in a statistically significant manner in the high frequency range. However, this effect has little influence on the inverse power law estimation for a given image regardless of the data representation or the theoretical analysis approach. The analysis is presented from two points of view (1) each image is treated independently with the results presented as distributions, and (2) for a given representation, the entire image collection is treated as an ensemble with the results presented as expected values. In general, the constant ring width analysis forms the foundation for a spectral comparison method for finding spectral differences, from an image distribution sense, after applying a nonlinear transformation to the data. The work also shows that power law estimation may be influenced due to the presence of noise in the higher frequency range, which is consistent with the known attributes of the detector efficiency. The spectral modeling and inverse power law determinations obtained here are in agreement with that obtained from the analysis of digitized film-screen images presented previously. The form of the power spectrum for a given image is approximately l/f2beta with beta approximately 1.4-1.5.

Biophysical Phenomena↗

MegaPlantTF: a machine learning framework for comprehensive identification and classification of plant transcription factors.

MOTIVATION: Understanding the role of transcription factors (TFs) in plants is essential for the study of gene regulation and various biological processes. However, both TF detection and classification remain challenging due to the great diversity and complexity of these proteins. Conventional approaches, such as BLAST, often suffer from high computational complexity and limited performance on less common TF families. RESULTS: We introduce MegaPlantTF, the first comprehensive machine learning and deep learning framework for the prediction (TF versus non-TF) and classification (family-level) of plant TFs. Our method employs k-mer-based protein representations and a two-stage architecture combining a deep feed-forward neural network with a stacking ensemble classifier. To ensure robust performance assessment, we report micro-, macro-, and weighted-average performance metrics, providing a holistic evaluation of both frequent and underrepresented TF families. Additionally, we employ threshold-based evaluation to calibrate confidence in TF detection. The results show that MegaPlantTF achieves strong accuracy and precision, particularly with a k-mer size of 3 and a classification threshold of 0.5, and maintains stable performance even under stringent thresholds. In addition to the standard cross-validation tests, a use case study on Sorghum bicolor confirms that our method performs strongly in the genome-wide analysis, making it highly suitable for large-scale TF identification and classification tasks. MegaPlantTF represents a novel contribution by integrating k-mer encoding, binary family-specific classifiers, and a two-stage stacking ensemble into a unified, reproducible framework for large-scale plant TF identification and classification. AVAILABILITY AND IMPLEMENTATION: MegaPlantTF is freely accessible through a public web server available at https://bioinformatics.um6p.ma/MegaPlantTF. The complete source code, including pretrained models and example datasets, is available at https://github.com/Bioinformatics-UM6P/MegaPlantTF.

Transcription Factors↗