Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Efficient RMSD measures for the comparison of two molecular ensembles. Root-mean-square deviation.

Quantitative measures are presented for comparing the conformations of two molecular ensembles. The measures are based on Kabsch's formula for the root-mean-square deviation (RMSD) and the covariance matrix of atomic positions of isotropically distributed ensembles (IDE). By using a Taylor series expansion, it is shown that the RMSD can be expressed solely in terms of the IDE matrices. A fast approximate method is introduced for the pairwise RMSD determination whose computational cost scales linearly with the number of structures. A similarity measure for two structural ensembles that is based on the trace metric of the differences of powers of the IDE matrices is presented. The measures are illustrated for conformational ensembles generated by a molecular dynamics computer simulation of a partially folded A-state analog of ubiquitin.

Computational Biology↗

Annotating nucleic acid-binding function based on protein structure.

Many of the targets of structural genomics will be proteins with little or no structural similarity to those currently in the database. Therefore, novel function prediction methods that do not rely on sequence or fold similarity to other known proteins are needed. We present an automated approach to predict nucleic-acid-binding (NA-binding) proteins, specifically DNA-binding proteins. The method is based on characterizing the structural and sequence properties of large, positively charged electrostatic patches on DNA-binding protein surfaces, which typically coincide with the DNA-binding-sites. Using an ensemble of features extracted from these electrostatic patches, we predict DNA-binding proteins with high accuracy. We show that our method does not rely on sequence or structure homology and is capable of predicting proteins of novel-binding motifs and protein structures solved in an unbound state. Our method can also distinguish NA-binding proteins from other proteins that have similar, large positive electrostatic patches on their surfaces, but that do not bind nucleic acids.

Amino Acid Motifs↗

Cyclic enkephalin and dermorphin analogues containing a carbonyl bridge.

Four cyclic enkephalin analogues and four cyclic dermorphin analogues have been synthesized. Cyclization of linear peptides containing basic amino acid residues of various side chain length in position 2 and 5 (enkephalin analogues) or 2 and 4 (dermorphin analogues) was achieved by treatment with bis-(4-nitrophenyl) carbonate to form a urea unit. The peptides were tested in the guinea-pig ileum (GPI) and mouse vas deferens (MVD) assays. Diverse activity was observed, depending on the size of the ring and the location of the urea unit. The conformation of two dermorphin analogues has been studied: one of high activity (IC(50) = 4.15 nM in the GPI assay) and a second of low activity (IC(50) = 6700 nM in the GPI assay). The conformational space of these peptides was examined using the EDMC method. Using data from the NMR spectra, each peptide was described as an ensemble of conformers. Biological activity was discussed in light of the structural data.

Amino Acid Sequence↗

Quantitative study of the structure-retention index relationship in the imine family.

The Kováts retention index is one of the most popular descriptors of the performance of organic compounds in gas chromatography (GC). The mathematical modeling of this index is an interesting and open problem in analytical chemistry. In this paper, two models for the prediction of the Kováts retention index are presented. Topologic, topographic and quantum-chemical descriptors were used as structural descriptors. Multiple linear regression (MLR) analysis provides the first model using the forward stepwise procedure for the variable selection. For the second one, an ensemble of artificial neural network (ANN) was constructed using the pruning algorithm. Both methods were validated by an external set of compounds, by the Golbraikh and Tropsha method and by the leave-one-out (LOO) and the leave many out (LMO) procedures. The R2, RMScv and Q2, values for the training sets were 0.884, 0.589 and 0.830 for NN and 0.974, 0.417 and 0.970 for MLR models, respectively. The robustness of both models was demonstrated. Both portrait the chromatographic performance of the sample but in this case, the results of MLR equation are better than the NN ones. The MLR model is recommended because of its simplicity.

Imines↗

Ultrasonic technique for imaging tissue vibrations: preliminary results.

We propose an ultrasound (US)-based technique for imaging vibrations in the blood vessel walls and surrounding tissue caused by eddies produced during flow through narrowed or punctured arteries. Our approach is to utilize the clutter signal, normally suppressed in conventional color flow imaging, to detect and characterize local tissue vibrations. We demonstrate the feasibility of visualizing the origin and extent of vibrations relative to the underlying anatomy and blood flow in real-time and their quantitative assessment, including measurements of the amplitude, frequency and spatial distribution. We present two signal-processing algorithms, one based on phase decomposition and the other based on spectral estimation using eigen decomposition for isolating vibrations from clutter, blood flow and noise using an ensemble of US echoes. In simulation studies, the computationally efficient phase-decomposition method achieved 96% sensitivity and 98% specificity for vibration detection and was robust to broadband vibrations. Somewhat higher sensitivity (98%) and specificity (99%) could be achieved using the more computationally intensive eigen decomposition-based algorithm. Vibration amplitudes as low as 1 mum were measured accurately in phantom experiments. Real-time tissue vibration imaging at typical color-flow frame rates was implemented on a software-programmable US system. Vibrations were studied in vivo in a stenosed femoral bypass vein graft in a human subject and in a punctured femoral artery and incised spleen in an animal model.

Algorithms↗

Lattice measurement and alloy compositions in metal and bimetallic nanoparticles.

A new reliable method for determining the lattice spacings of metallic and bimetallic nanoparticles in phase contrast high resolution electron microscopy (HREM) images was developed. In this study, we discuss problems in applying HREM techniques to single metal (Pt and Au) and bimetallic (AuPd) nanoparticles of unknown shapes and random orientations. Errors arising from particle tilt and edge effects are discussed and analysis criteria are presented to reduce these errors in measuring the lattice parameters of nanoparticles. The accuracy of an individual particle lattice measurement is limited by an effective standard deviation which depends on the size of the individual nanoparticle. For example, the standard deviation for 20-30 A Pt or Au nanoparticles is about 1.5%. To increase the accuracy in determining the lattice spacings of nanoparticles, statistical methods have to be used to obtain the average lattice spacing of an ensemble of nanoparticles. We measured approximately 100 nanoparticles with sizes in the range of 20-30 A and found that the mean lattice spacing can be determined to within 0.2%. By applying Vegard's law to the AuPd bimetallic systems we successfully detected the presence of alloying. For 30 A nanoparticles, the estimated ultimate error in determining the composition of the AuPd alloy is about 3% provided that at least 100 particles are measured. Finally, the challenges in determining the presence of more than one alloy phases in bimetallic nanoparticle systems were also discussed.

Journal Article↗

Chromosome elasticity and mitotic polar ejection force measured in living Drosophila embryos by four-dimensional microscopy-based motion analysis.

BACKGROUND: Mitosis involves the interaction of many different components, including chromatin, microtubules, and motor proteins. Dissecting the mechanics of mitosis requires methods of studying not just each component in isolation, but also the entire ensemble of components in its full complexity in genetically tractable model organisms. RESULTS: We have developed a mathematical framework for analyzing motion in four-dimensional microscopy data sets that allows us to measure elasticity, viscosity, and forces by tracking the conformational movements of mitotic chromosomes. We have used this approach to measure, for the first time, the basic biophysical parameters of mitosis in wild-type Drosophila melanogaster embryos. We found that Drosophila embryo chromosomes are significantly less rigid than the much larger chromosomes of vertebrates. Anaphase kinetochore force and nucleoplasmic viscosity were comparable with previous estimates in other species. Motion analysis also allowed us to measure the magnitude of the polar ejection force exerted on chromosome arms during metaphase by individual microtubules. We find the magnitude of this force to be approximately 1 pN, a number consistent with force generation either by collision of growing microtubules with chromosomes or by single kinesin motors. CONCLUSIONS: Motion analysis allows noninvasive mechanical measurements to be made in complex systems. This approach should allow the functional effects of Drosophila mitotic mutants on chromosome condensation, kinetochore forces, and the polar ejection force to be determined.

Algorithms↗

Correlations in single molecule photon statistics: renewal indicator.

Multiple time scales are the intrinsic nature of complex systems and can be revealed through single molecule photon statistical analysis. The standard Poisson indicator defined by the time-averaged initial condition measures photon bunching and antibunching but cannot be directly related to multiple time scales. A new indicator defined by the event-averaged initial condition is proposed to detect the deviation from the renewal behavior and to directly probe the effects of conformational fluctuations. Detailed calculations of modulated two-level systems are carried out using the transfer matrix method to demonstrate the difference between the two indicators. The relationship between ensemble-averaged survival probabilities and photon statistics is also explored in the context of single molecule measurements.

Journal Article↗

Geometrical capillary network analysis.

BACKGROUND: Skin microcirculation, especially the superficial network, can be assessed by a computer capillary video microscope system. The study of morphology and dynamics of microcirculation must include all dynamic and cooperative processes between the capillaries. For characterizing capillary ensembles, the statistical and geometrical properties of the network need to be explored. METHODS: The microvaculature of the skin and the microcirculation were investigated by combining videocapillaroscopy (VCP) and image processing techniques based on computational geometry and graph theory. Our goal was to characterize the capillary network in noisy pictures of the scalp. Different geometric methods were developed, based on proximity parameters (distance and surface) in order to circumscribe and construct this network. RESULTS: By studying the distribution of these parameters, extreme values or outliers, which usually correspond to artifact subregions in the pictures could be eliminated. Different algorithms were developed and has been implemented in an image processing software (Capilab Toolbox). CONCLUSION: This computerized system is capable of real-time processings, increasing the quality of videocapillaroscope images and minimizing the disturbance of artifacts. The algorithms presented here are easy to implement and can process any kind of images of the skin, even in the scalp. In association with an example-based detection system, this method can be generalized to other stimuli in the same conditions.

Capillaries↗

Quantification of modelling uncertainties in a large ensemble of climate change simulations.

Comprehensive global climate models are the only tools that account for the complex set of processes which will determine future climate change at both a global and regional level. Planners are typically faced with a wide range of predicted changes from different models of unknown relative quality, owing to large but unquantified uncertainties in the modelling process. Here we report a systematic attempt to determine the range of climate changes consistent with these uncertainties, based on a 53-member ensemble of model versions constructed by varying model parameters. We estimate a probability density function for the sensitivity of climate to a doubling of atmospheric carbon dioxide levels, and obtain a 5-95 per cent probability range of 2.4-5.4 degrees C. Our probability density function is constrained by objective estimates of the relative reliability of different model versions, the choice of model parameters that are varied and their uncertainty ranges, specified on the basis of expert advice. Our ensemble produces a range of regional changes much wider than indicated by traditional methods based on scaling the response patterns of an individual simulation.

Journal Article↗

Fully flexible unit cell simulation with recursive thermostat chains.

The recursive thermostat chained fully flexible cell molecular dynamic simulation (NsigmaT ensemble) is performed. The ensemble is based on the metric tensor, whose components are used as extended variables. These variables are combined with Nosé-Poincaré recursive thermostat chains. This extended Hamiltonian approach preserves Hamiltonian in structure, and the partition function satisfies the NsigmaT ensemble state in phase space. In the present study, the generalized leap frog method was employed for time integration. The resulting molecular dynamics simulation was performed for bulk and thin film solid materials in the face-centered-cubic crystal structure. Uniaxial tension test and simple shear test are performed to predict the behaviors of a solid material in the bulk state and nanoscale thin film state. The proposed flexible cell method should serve as a powerful tool for the prediction of mechanical and thermal properties of solid materials including nanoscale behavior.

Journal Article↗

Reconciliation and interpretation of Big Bend National Park particulate sulfur source apportionment: results from the Big Bend Regional Aerosol and Visibility Observational Study--part I.

The Big Bend Regional Aerosol and Visibility Observational (BRAVO) study was an intensive monitoring study from July through October 1999 followed by extensive assessments to determine the causes and sources of haze in Big Bend National Park, located in Southwestern Texas. Particulate sulfate compounds are the largest contributor of haze at Big Bend, and chemical transport models (CTMs) and receptor models were used to apportion the sulfate concentrations at Big Bend to North American source regions and the Carbón power plants, located 225 km southeast of Big Bend in Mexico. Initial source attribution methods had contributions that varied by a factor of > or =2. The evaluation and comparison of methods identified opposing biases between the CTMs and receptor models, indicating that the ensemble of results bounds the true source attribution results. The reconciliation of these differences led to the development of a hybrid receptor model merging the CTM results and air quality data, which allowed a nearly daily source apportionment of the sulfate at Big Bend during the BRAVO study. The best estimates from the reconciliation process resulted in sulfur dioxide (SO2) emissions from U.S. and Mexican sources contributing approximately 55% and 38%, respectively, of sulfate at Big Bend. The distribution among U.S. source regions was Texas, 16%; the Eastern United States, 30%; and the Western United States, 9%. The Carbón facilities contributed 19%, making them the largest single contributing facility. Sources in Mexico contributed to the sulfate at Big Bend on most days, whereas contributions from Texas and Eastern U.S. sources were episodic, with their largest contributions during Big Bend sulfate episodes. On the 20% of the days with the highest sulfate concentrations, U.S. and Mexican sources contributed approximately 71% and 26% of the sulfate, respectively. However, on the 20% of days with the lowest sulfate concentrations, Mexico contributed 48% compared with 40% for the United States.

Aerosols↗

Large-scale phylogenies and measuring the performance of phylogenetic estimators.

Performance measures of phylogenetic estimation methods such as accuracy, consistency, and power are an attempt at summarizing an ensemble of a given estimator's behavior. These summaries characterize an ensemble behavior with a single number, leading to a variety of definitions. In particular, the relationships between different performance measures such as accuracy and consistency or accuracy and error depend on the exact definition of these measures. In addition, it is relatively common to use large-sample behavior to infer similar behavior for small samples. In fact, large-sample results such as the claimed asymptotic efficiency of the maximum-likelihood estimator are often uninformative for small samples. Conversely, small-sample behavior using simulations is sometimes used to imply large-sample behavior such as consistency. However, such extrapolation is often difficult. How the performance of a phylogenetic estimator scales with the addition of taxa must be qualified with respect to whether the whole tree is being estimated or a fixed subset of taxa is being estimated. It must also be qualified with respect to how tree models are sampled. Over the ensemble of all possible trees of a given size, the performance of the estimators for the whole tree estimate suffers when the tree size becomes larger. However, under certain models of cladogenesis, the estimate can improve with the addition of taxa. In fact, at all numbers of taxa there are subsets of tree models that are easier to estimate than others. This suggests that with judicious addition or subtraction of taxa we can move from tree models that are more difficult to estimate at one number of taxa to those that are easier to estimate at another number of taxa.

Analysis of Variance↗

Improving prediction of protein secondary structure using structured neural networks and multiple sequence alignments.

The prediction of protein secondary structure by use of carefully structured neural networks and multiple sequence alignments has been investigated. Separate networks are used for predicting the three secondary structures alpha-helix, beta-strand, and coil. The networks are designed using a priori knowledge of amino acid properties with respect to the secondary structure and the characteristic periodicity in alpha-helices. Since these single-structure networks all have less than 600 adjustable weights, overfitting is avoided. To obtain a three-state prediction of alpha-helix, beta-strand, or coil, ensembles of single-structure networks are combined with another neural network. This method gives an overall prediction accuracy of 66.3% when using 7-fold cross-validation on a database of 126 nonhomologous globular proteins. Applying the method to multiple sequence alignments of homologous proteins increases the prediction accuracy significantly to 71.3% with corresponding Matthew's correlation coefficients C alpha = 0.59, C beta = 0.52, and Cc = 0.50. More than 72% of the residues in the database are predicted with an accuracy of 80%. It is shown that the network outputs can be interpreted as estimated probabilities of correct prediction, and, therefore, these numbers indicate which residues are predicted with high confidence.

Amino Acid Sequence↗

KemaDom: a web server for domain prediction using kernel machine with local context.

Predicting domains of proteins is an important and challenging problem in computational biology because of its significant role in understanding the complexity of proteomes. Although many template-based prediction servers have been developed, ab initio methods should be designed and further improved to be the complementarity of the template-based methods. In this paper, we present a novel domain prediction system KemaDom by ensembling three kernel machines with the local context information among neighboring amino acids. KemaDom, an alternative ab initio predictor, can achieve high performance in predicting the number of domains in proteins. It is freely accessible at http://www.iipl.fudan.edu.cn/lschen/kemadom.htm and http://www.iipl.fudan.edu.cn/~lschen/kemadom.htm.

Artificial Intelligence↗

Geometry of river networks. II. Distributions of component size and number.

The structure of a river network may be seen as a discrete set of nested subnetworks built out of individual stream segments. These network components are assigned an integral stream order via a hierarchical and discrete ordering method. Exponential relationships, known as Horton's laws, between stream order and ensemble-averaged quantities pertaining to network components are observed. We extend these observations to incorporate fluctuations and all higher moments by developing functional relationships between distributions. The relationships determined are drawn from a combination of theoretical analysis, analysis of real river networks including the Mississippi, Amazon, and Nile, and numerical simulations on a model of directed, random networks. Underlying distributions of stream segment lengths are identified as exponential. Combinations of these distributions form single-humped distributions with exponential tails, the sums of which are in turn shown to give power-law distributions of stream lengths. Distributions of basin area and stream segment frequency are also addressed. The calculations identify a single length scale as a measure of size fluctuations in network components. This article is the second in a series of three addressing the geometry of river networks.

Ecosystem↗

Effects of cross-phase modulation on phase jitter in soliton systems with constant dispersion.

In wavelength-division-multiplexed communication systems phase jitter is driven by amplifier noise and mediated by cross-phase modulation. The variational method is used to derive formulas for the absolute phase variance of an ensemble of solitons and the relative phase variance of ensembles of neighboring solitons. The predictions of these formulas are consistent with the results of numerical simulations.

Journal Article↗

Development of a PCR-based technique for genotyping UGT1A1 gene and distribution of rs3064744 alleles in the Russian population.

BACKGROUND: Accurate determination of tandem thymine-adenine (TA) repeat numbers in the UGT1A1 promoter region (rs3064744) is essential for diagnosing Gilbert's syndrome and personalizing therapy with toxic agents like irinotecan and atazanavir. However, traditional polymerase chain reaction (PCR) assays face severe limitations due to the AT-rich sequence and overlapping melting temperatures (Tm) of the highly homologous 7TA and 8TA alleles. In this context, melting curve analysis (MCA) employing fluorophore-quencher systems has emerged as a promising alternative. The purpose of this study was to develop a novel genotyping approach combining optimized aPCR-MCA analysis with an automated classifier to overcome the limitations posed by the differentiation of highly homologous alleles and to demonstrate its practical application, providing the distribution of rs3064744 genotypes across four regional cohorts of the Russian population. METHODS: A specialized Dual Head 1D-convolutional neural network (1D-CNN) ensemble with Test-Time Augmentation (TTA) was developed. The model was trained and internally validated on 1,620 engineered plasmid samples, and independently evaluated on an external clinical test set of 440 unique patient genomic DNA specimens. Real-time PCR was performed on CFX96 and DTprime platforms. Additionally, population-wide screening was conducted on 997 archival clinical samples from Moscow, Sakha (Yakutia), Dagestan, and Rostov regions. RESULTS: While 5TA and 6TA alleles were easily separated, absolute Tm distributions of 7TA and 8TA alleles overlapped significantly, and non-uniform Tm shifts of 0.8 °C-1.4 °C occurred across platforms. Conventional absolute Tm thresholding was therefore inadequate. By assessing relative morphological curve divergence against co-amplified 7TA/7TA and 7TA/8TA reference anchors, the 1D-CNN ensemble neutralized instrument noise. It achieved 100% accuracy on internal validation and 100% concordance (440/440) with clinical reference pyrosequencing. Population screening revealed that Dagestan, Yakutia, and Rostov cohorts closely align with the European population. Rare 5TA and 8TA alleles were detected at low frequencies in Yakutia and Moscow. CONCLUSION: Combining LNA-modified aPCR-MCA with a comparative 1D-CNN model successfully circumvents thermodynamic limitations and eliminates human operator bias. This integrated system offers an accessible, high-throughput, and clinically valid solution for routine UGT1A1 pharmacogenetic testing.

1D-CNN↗