Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Evaluation of a novel shape-based computational filter for lead evolution: application to thrombin inhibitors.

A novel shape-feature-based computational method is described and used to rapidly filter compound libraries. The computational model, built using three-dimensional conformations of active and inactive molecules, consists of a collection of whole molecule shapes and chemical feature positions that are ranked according to their correlation with activity. A small ensemble of these shapes and features is used to filter virtual compound libraries. The method is applied to two thrombin data sets and is shown to be efficient in identifying novel scaffolds with enhanced hit rates.

Combinatorial Chemistry Techniques↗

Bias-free separation of internal and overall motion of biomolecules.

Collective internal motions are known to be important for the function of biological macromolecules. It has been discussed in the past whether the application of superimposing algorithms to remove the overall motion from a structural ensemble introduces artificial correlations between distant atoms. Here we present a new method to eliminate residual rotation and translation from cartesian modes derived from a normal mode analysis or from a principal component analysis. Bias-free separation is based on the idea that the addition of modes of pure rotation/translation can compensate the residual overall motion. Removal of overall motion must reduce the "total amount of motion" (TAM) in the mode. Our algorithm allows to back-calculate revised covariance matrices. The approach was applied to two model systems that show residual overall motion, when analyzed using all atoms as reference for the superimposing algorithm. In both cases, our algorithm was capable of eliminating residual covariances caused by the overall motion, while retaining internal covariances even for very distant atoms. A structural ensemble obtained for a 13-ns molecular dynamics simulation of the protein Ribonuclease T1 showed a covariance matrix of the corrected modes with significantly sharper contours after applying the bias-free separation.

Algorithms↗

Recapitulation of protein family divergence using flexible backbone protein design.

We use flexible backbone protein design to explore the sequence and structure neighborhoods of naturally occurring proteins. The method samples sequence and structure space in the vicinity of a known sequence and structure by alternately optimizing the sequence for a fixed protein backbone using rotamer based sequence search, and optimizing the backbone for a fixed amino acid sequence using atomic-resolution structure prediction. We find that such a flexible backbone design method better recapitulates protein family sequence variation than sequence optimization on fixed backbones or randomly perturbed backbone ensembles for ten diverse protein structures. For the SH3 domain, the backbone structure variation in the family is also better recapitulated than in randomly perturbed backbones. The potential application of this method as a model of protein family evolution is highlighted by a concerted transition to the amino acid sequence in the structural core of one SH3 domain starting from the backbone coordinates of an homologous structure.

Evolution, Molecular↗

Calculated pH-dependent population and protonation of carbon-monoxy-myoglobin conformers.

X-ray structures of carbonmonoxymyoglobin (MbCO) are available for different pH values. We used conventional electrostatic continuum methods to calculate the titration behavior of MbCO in the pH range from 3 to 7. For our calculations, we considered five different x-ray structures determined at pH values of 4, 5, and 6. We developed a Monte Carlo method to sample protonation states and conformations at the same time so that we could calculate the population of the considered MbCO structures at different pH values and the titration behavior of MbCO for an ensemble of conformers. To increase the sampling efficiency, we introduced parallel tempering in our Monte Carlo method. The calculated population probabilities show, as expected, that the x-ray structures determined at pH 4 are most populated at low pH, whereas the x-ray structure determined at pH 6 is most populated at high pH, and the population of the x-ray structures determined at pH 5 possesses a maximum at intermediate pH. The calculated titration behavior is in better agreement with experimental results compared to calculations using only a single conformation. The most striking feature of pH-dependent conformational changes in MbCO-the rotation of His-64 out of the CO binding pocket-is reproduced by our calculations and is correlated with a protonation of His-64, as proposed earlier.

Animals↗

HOPPSIGEN: a database of human and mouse processed pseudogenes.

Processed pseudogenes result from reverse transcribed mRNAs. In general, because processed pseudogenes lack promoters, they are no longer functional from the moment they are inserted into the genome. Subsequently, they freely accumulate substitutions, insertions and deletions. Moreover, the ancestral structure of processed pseudogenes could be easily inferred using the sequence of their functional homologous genes. Owing to these characteristics, processed pseudogenes represent good neutral markers for studying genome evolution. Recently, there is an increasing interest for these markers, particularly to help gene prediction in the field of genome annotation, functional genomics and genome evolution analysis (patterns of substitution). For these reasons, we have developed a method to annotate processed pseudogenes in complete genomes. To make them useful to different fields of research, we stored them in a nucleic acid database after having annotated them. In this work, we screened both mouse and human complete genomes from ENSEMBL to find processed pseudogenes generated from functional genes with introns. We used a conservative method to detect processed pseudogenes in order to minimize the rate of false positive sequences. Within processed pseudogenes, some are still having a conserved open reading frame and some have overlapping gene locations. We designated as retroelements all reverse transcribed sequences and more strictly, we designated as processed pseudogenes, all retroelements not falling in the two former categories (having a conserved open reading or overlapping gene locations). We annotated 5823 retroelements (5206 processed pseudogenes) in the human genome and 3934 (3428 processed pseudogenes) in the mouse genome. Compared to previous estimations, the total number of processed pseudogenes was underestimated but the aim of this procedure was to generate a high-quality dataset. To facilitate the use of processed pseudogenes in studying genome structure and evolution, DNA sequences from processed pseudogenes, and their functional reverse transcribed homologs, are now stored in a nucleic acid database, HOPPSIGEN. HOPPSIGEN can be browsed on the PBIL (Pole Bioinformatique Lyonnais) World Wide Web server (http://pbil.univ-lyon1.fr/) or fully downloaded for local installation.

Animals↗

Molecular dynamics simulations of isolated transmembrane helices of potassium channels.

In the middle of the S6 helix in voltage-gated potassium channels there is a highly conserved Pro-Val-Pro motif, while the equivalent M2 helix of inward rectifier potassium channels contains a conserved glycine residue in a comparable position. The structural implications of these conserved motifs are of interest given the evidence that S6 and M2 are components of the lining of their respective pores. Multiple sequence alignment and TM helix prediction methods were used to define consensus regions for S6 and M2. Ensembles of 50 structures for each helix were generated by simulated annealing and restrained molecular dynamics. Time-dependent fluctuations of S6 and M2 were investigated by long time scale molecular dynamics simulations on representative members of each ensemble carried out in vacuo in the presence and absence of a hydrophobic potential that mimics a lipid bilayer. The results are discussed in terms of the structural basis of the kink in S6 and M2 and of a putative functional role for flexible helices as "molecular swivels."

Amino Acid Sequence↗

Local structure and thermodynamics of a core-softened potential fluid: theory and simulation.

Phase behavior and structural properties of homogeneous and inhomogeneous core-softened (CS) fluid consisting of particles interacting via the potential, which combines the hard-core repulsion and double attractive well interaction, are investigated. The vapour-liquid coexistence curves and critical points for various interaction ranges of the potential are determined by discrete molecular dynamics simulations to provide guidance for the choice of the bulk density and potential parameters for the study of homogeneous and inhomogeneous structures. Spatial correlations in the homogeneous CS system are studied by the Ornstein-Zernike integral equation in combination with the modified hypernetted chain (MHNC) approximation. The local structure of CS fluid subjected to diverse external fields maintaining the equilibrium with the bulk CS fluid are studied on the basis of a recently proposed third order+second order perturbation density functional approximation (DFA). The accuracy of DFA predictions is tested against the results of a grand canonical ensemble Monte Carlo simulation. Reasonable agreement between the results of both methods proves that the DFA theory applied in this work is a convenient theoretical tool for the investigation of the CS fluid, which is practically applicable for modeling numerous real systems.

Journal Article↗

Protein nanoarray on Prolinker surface constructed by atomic force microscopy dip-pen nanolithography for analysis of protein interaction.

Protein nanoarrays are addressable ensembles of nano-scale protein domain on solid surfaces. This method can serve as a useful platform for ultraminiaturized bioanalysis. In this study, we investigated single molecular nanopatterning and molecular interaction of proteins that were immobilized on Prolinker surface of gold-coated silicon wafer by using dip-pen nanolithography (DPN) method. Contact force and humidity were optimized at 0.01 nN and 80%, respectively. The domain features of protein nanoarrays were developed at the contact time of 5 s. The optimized conditions for the nanoarray process were applied to create protein nanoarray using integrin alpha(v)beta3 and angiogenin. Constructed protein nanoarrays using integrin alpha(v)beta3 have single molecular monolayer with regular domain shape (height 15 +/- 5 nm). The changed height value due to the single molecular interaction between integrin alpha(v)beta3 and vitronectin was approximately 30 +/- 5 nm on Prolinker surface as measured with atomic force microscopy tip. Taken together, these results suggest that protein nanoarray on Prolinker surface fabricated by well-controlled DPN process can be used to analyze single molecular interaction of protein.

Gold↗

Improving the prediction of protein secondary structure in three and eight classes using recurrent neural networks and profiles.

Secondary structure predictions are increasingly becoming the workhorse for several methods aiming at predicting protein structure and function. Here we use ensembles of bidirectional recurrent neural network architectures, PSI-BLAST-derived profiles, and a large nonredundant training set to derive two new predictors: (a) the second version of the SSpro program for secondary structure classification into three categories and (b) the first version of the SSpro8 program for secondary structure classification into the eight classes produced by the DSSP program. We describe the results of three different test sets on which SSpro achieved a sustained performance of about 78% correct prediction. We report confusion matrices, compare PSI-BLAST to BLAST-derived profiles, and assess the corresponding performance improvements. SSpro and SSpro8 are implemented as web servers, available together with other structural feature predictors at: http://promoter.ics.uci.edu/BRNN-PRED/.

Algorithms↗

Combined infrared multiphoton dissociation and electron capture dissociation with a hollow electron beam in Fourier transform ion cyclotron resonance mass spectrometry.

An electron injection system based on an indirectly heated ring-shaped dispenser cathode has been developed and installed in a 7 Tesla Fourier transform ion cyclotron resonance (FTICR) mass spectrometer. This new hardware design allows high-rate electron capture dissociation (ECD) to be carried out by a hollow electron beam coaxial with the ion cyclotron resonance (ICR) trap. Infrared multiphoton dissociation (IRMPD) can also be performed with an on-axis IR-laser beam passing through a hole at the centre of the dispenser cathode. Electron and photon irradiation times of the order of 100 ms are required for efficient ECD and IRMPD, respectively. As ECD and IRMPD generate fragments of different types (mostly c, z and b, y, respectively), complementary structural information that improves the characterization of peptides and proteins by FTICR mass spectrometry can be obtained. The developed technique enables the consecutive or simultaneous use of the ECD and IRMPD methods within a single FTICR experimental sequence and on the same ensemble of trapped ions in multistage tandem (MS/MS/MS or MS(n)) mass spectrometry. Flexible changing between ECD and IRMPD should present advantages for the analysis of protein digests separated by liquid chromatography prior to FTICRMS. Furthermore, ion activation by either electron or laser irradiation prior to, as well as after, dissociation by IRMPD or ECD increases the efficiency of ion fragmentation, including the w-type fragment ion formation, and improves sequencing of peptides with multiple disulfide bridges. The developed instrumental configuration is essential for combined ECD and IRMPD on FTICR mass spectrometers with limited access into the ICR trap.

Animals↗

Prestroke dementia in patients with atrial fibrillation. Frequency and associated factors.

BACKGROUND AND PURPOSE: Prestroke dementia is frequent but usually not identified. Non-valvular atrial fibrillation (NVAF) is independently associated with an increased risk for dementia. However, the frequency and determinants of prestroke dementia in patients with NVAF have never been evaluated. OBJECTIVE: The aim of this study was to determine the frequency of prestroke dementia and associated factors in patients with a previously known NVAF. METHODS: This is an ancillary study of Stroke in Atrial Fibrillation Ensemble II (SAFE II), an observational study conducted in patients with a previously known NVAF, consecutively admitted for an acute stroke in French and Italian centers. Prestroke dementia was evaluated by the IQCODE in patients with a reliable informant. Patients were considered as demented before stroke when their IQCODE score was > or = 104. RESULTS: of 204 patients, 39 (19.1%; 95% confidence interval [CI]: 13.7%-24.5%) patients met criteria for prestroke dementia. The only variable independently associated with prestroke dementia was increasing age (adjusted odds ratio for 1 year increase in age: 1.10; 95 % CI: 1.04-1.17), and there was a nonsignificant tendency for previous ischemic stroke or TIA and arterial hypertension. CONCLUSION: One fifth of stroke patients with a previously known NVAF were already demented before stroke. The main determinant of prestroke dementia is increasing age. A large cohort is necessary to identify other determinants.

Age Factors↗

An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer.

This study presents a novel integrative approach for the analysis of high-dimensional gene expression data, leveraging the complementary strengths of unsupervised clustering and supervised classification. Using K-means clustering, the data set is stratified into three distinct clusters, revealing intrinsic biological patterns and relationships. The resulting cluster assignments are subsequently used as pseudolabels to train machine learning models, including support vector machines, random forest, and a stacking ensemble classifier. To validate and enhance the robustness of clustering, complementary methods, such as hierarchical clustering and density-based spatial clustering of applications with noise (DBSCAN), are used, with results visualized through principal component analysis-driven dimensionality reduction. The high predictive accuracy achieved by the classifiers underlines the separability and reliability of the identified clusters. Furthermore, feature importance analysis highlighted key genetic determinants within each cluster, offering actionable insights into potential biomarkers and critical genomic features. This framework bridges the gap between exploratory unsupervised learning and predictive supervised modeling, providing a scalable and interpretable method for analyzing complex genomic data sets. Its applicability extends to biomarker discovery, patient stratification, and other precision medicine applications, emphasizing its utility in advancing genomic research and clinical practice.

Humans↗

Differentiating temporal electromyographic waveforms between those with chronic low back pain and healthy controls.

OBJECTIVES: Temporal activation patterns from abdominal and lumbar muscles were compared between healthy control subjects and those with chronic low back pain. STUDY DESIGN: A cross-sectional comparative study. BACKGROUND: Synergist and antagonist coactivity has been considered an important neuromuscular control strategy to maintain spinal stability. Differences in onset times and amplitudes have been reported from trunk muscle EMG recordings between healthy subjects and those with low back pain;however, evaluating temporal EMG waveforms should demonstrate whether differences exist in the ability of those with and those without low back pain to respond to changing perturbations. METHODS: The Karhunen-Loève expansion was applied to the ensemble-average EMG profiles recorded from four abdominal and three trunk extensor muscle sites while subjects performed a leg-lifting task aimed at challenging lumbar spine stability. The principal patterns were derived and the weighting coefficients for each pattern were the main dependent variables in a series of two-factor (group and muscle) mixed ANOVA models. RESULTS: Three principal patterns explained 96% of the variance in the temporal EMG profiles. The ANOVAs revealed statistically significant group and muscle main effects (P<0.05) for the principal pattern and significant group by muscle interactions (P<0.05) for patterns two and three. Post hoc analysis showed that patterns were not different among all muscle sites for the healthy controls, but differences were significant for the low back pain group. CONCLUSIONS: The healthy group coactivated all seven sites with the same temporal pattern of activation. The low back pain group used different activation patterns indicative of a lack of synergistic coactivitation among the muscle sites examined. RELEVANCE: These results provide a foundation for developing a diagnostic classifier of neuromuscular impairment associated with low back pain, that could be used to evaluate the effectiveness of therapeutic interventions to improve muscle coactivation.

Abdominal Muscles↗

Reconstructing the engram: simultaneous, multisite, many single neuron recordings.

Little is known about the physiological principles that govern large-scale neuronal interactions in the mammalian brain. Here, we describe an electrophysiological paradigm capable of simultaneously recording the extracellular activity of large populations of single neurons, distributed across multiple cortical and subcortical structures in behaving and anesthetized animals. Up to 100 neurons were simultaneously recorded after 48 microwires were implanted in the brain stem, thalamus, and somatosensory cortex of rats. Overall, 86% of the implanted microwires yielded single neurons, and an average of 2.3 neurons were discriminated per microwire. Our population recordings remained stable for weeks, demonstrating that this method can be employed to investigate the dynamic and distributed neuronal ensemble interactions that underlie processes such as sensory perception, motor control, and sensorimotor learning in freely behaving animals.

Animals↗

Grand canonical Monte Carlo simulation of ligand-protein binding.

A new application of the grand canonical thermodynamics ensemble to compute ligand-protein binding is described. The described method is sufficiently rapid that it is practical to compute ligand-protein binding free energies for a large number of poses over the entire protein surface, thus identifying multiple putative ligand binding sites. In addition, the method computes binding free energies for a large number of poses. The method is demonstrated by the simulation of two protein-ligand systems, thermolysin and T4 lysozyme, for which there is extensive thermodynamic and crystallographic data for the binding of small, rigid ligands. These low-molecular-weight ligands correspond to the molecular fragments used in computational fragment-based drug design. The simulations correctly identified the experimental binding poses and rank ordered the affinities of ligands in each of these systems.

Binding Sites↗

Fluorescence studies of single biomolecules.

Single-molecule fluorescence has the capability to detect properties buried in ensemble measurements and, hence, provides new insights about biological processes. Ratiometric methods are normally used to reduce the effects of excitation beam inhomogeneity. Fluorescence resonance energy transfer is widely used but there are problems in inserting the fluorophores in the correct position on the biomolecule, particularly if the structure is not known. We have recently developed two-colour coincidence single-molecule fluorescence that addresses this problem. This method can be used to determine quantitatively the multimerization states of biomolecules, in solution without separation. The future prospects of single-molecule fluorescence as applied to biological molecules are discussed.

Fluorescence Resonance Energy Transfer↗

Stable method for the calculation of partition functions in the superconfiguration approach.

A general method for the calculation of the partition function of a canonical ensemble of noninteracting bound electrons is presented. It consists in a doubly recursive procedure with respect to the number of electrons and the number of orbitals. Contrary to existing approaches, this recursion relation contains no alternate summation of positive and negative numbers, which was the main source of numerical uncertainties. It is accompanied with a normalization of partition function through the determination of a free parameter consistent with the zeroth-order saddle-point approximation. The recursion relation allows one to calculate accurately partition functions for ions with a large number of orbitals, and is therefore important for calculations relying on the superconfiguration approximation.

Journal Article↗

Thermal Casimir effect in lipid bilayer tubules.

We calculate the thermal Casimir effect for a dielectric tube of radius R and thickness delta formed from a membrane in water. The method uses a field-theoretic approach in the grand canonical ensemble. The leading contribution to the Casimir free energy behaves as -k(B)TLkappa(c)/R giving rise to an attractive force which tends to contract the tube. We find that kappa(c) approximately 0.3 for the case of typical lipid membrane t tubules. We conclude that except in the case of a very soft membrane this force is insufficient to stabilize such tubes against the bending stress which tends to increase the radius.

Computer Simulation↗