Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Supercoiled DNA energetics and dynamics by computer simulation.

A new formulation is presented for investigating supercoiled DNA configurations by deterministic techniques. Thus far, the computational difficulties involved in applying deterministic methods to supercoiled DNA studies have generally limited computer simulations to stochastic approaches. While stochastic methods, such as simulated annealing and Metropolis-Monte Carlo sampling, are successful at generating a large number of configurations and estimating thermodynamic properties of topoisomer ensembles, deterministic methods offer an accurate characterization of the minima and a systematic following of their dynamics. To make this feasible, we model circular duplex DNA compactly by a B-spline ribbon-like model in terms of a small number of control vertices. We associate an elastic deformation energy composed of bending and twisting integrals and represent intrachain contact by a 6-12 Lennard Jones potential. The latter is parameterized to yield an energy minimum at the observed DNA-helix diameter inclusive of a hydration shell. A penalty term to ensure fixed contour length is also included. First and second partial derivatives of the energy function have been derived by using various mathematical simplifications. First derivatives are essential for Newton-type minimization as well as molecular dynamics, and partial second-derivative information can significantly accelerate minimization convergence through preconditioning. Here we apply a new large-scale truncated-Newton algorithm for minimization and a Langevin/implicit-Euler scheme for molecular dynamics. Our truncated-Newton method exploits the separability of potential energy functions into terms of differing complexity. It relies on a preconditioned conjugate gradient method that is efficient for large-scale problems to solve approximately for the search direction at every step. Our dynamics algorithm is numerically stable over large time steps. It also introduces a frequency-discriminating mechanism so that vibrational modes with frequencies greater than a chosen cutoff frequency are essentially frozen by the method. With these tools, we rapidly identify corresponding circular and interwound energy minima for small DNA rings for a series of imposed linking-number differences. These structures are consistent with available electron microscopy data. The energetic exchange of stability between the circle and the figure-8, in very good agreement with analytical results, is also detailed. Molecular dynamics trajectories at 100 femtosecond time steps then reveal the rapid folding of the unstable circular state into supercoiled forms. Significant bending and twisting motions of the interwound structures are also observed. Such information may be useful for understanding transition states along the folding pathway and the role of enzymes that regulate supercoiling.(ABSTRACT TRUNCATED AT 400 WORDS)

Algorithms↗

A correlation-based method for the enhancement of scoring functions on funnel-shaped energy landscapes.

A correlation-based approach is introduced for enhancing the ability of structure-scoring methods to identify and distinguish native-like conformations. The proposed method relies on a funnel-shaped scoring function that decreases steadily toward the native state. It takes advantage of the idea that the structure from a given ensemble that is closest to the native basin leads to the highest correlation coefficient between a given score and distance to that structure as an approximation of the native state for the entire ensemble. The method is applied successfully to a number of different test cases that demonstrate substantial improvements in the correlation of the score with the distance from the true native state but also result in the selection of more native-like structures compared to the original score.

Algorithms↗

Dynamic spectral analysis of event-related EEG data.

A method for analysing the time course of power spectra of event-related EEG data is presented. A sequence of autoregressive models is fitted to segments of the EEG within which the data exhibit local stationarity. For parameter estimation a method involving ensemble averages is introduced. Besides investigating the evolution of power spectra, time courses of peak frequency, bandwidth and power of alpha (mu) and beta rhythms are traced. The method is applied to EEG recorded over the primary motor area during self-paced finger movements.

Electroencephalography↗

A baseline detection method for analyzing transient electrophysiological events.

A baseline detection method has been developed that identifies and defines event transitions for whole-cell voltage or current ('transient') events that are produced by the activity of ion channel ensembles. The method utilizes a variety of iterative techniques that independently determine, for each event, several output parameters that are ultimately referrable to the mean and variance of each pre-event baseline. Examination of miniature endplate potentials using the baseline detection method provided the following output parameters for each transient event: pre-event mean and variance; rise time; peak amplitude and duration; a determination of whether the decay phase was best fit by a one- or two-component negative exponential function; time constants for the slow and/or fast decay components; percent contribution of the slow component to the decay phase; and the predicted peak amplitude determined by extrapolation of the least squares fit to the decay phase. Joint probability density representations involving the rise time and peak amplitudes of miniature endplate potentials indicated the power of this multivariate approach in identifying and isolating specific event classes. The baseline detection method is particularly advantageous for analyzing records containing multiple classes of event amplitudes, and provides a reproducible statistical standard for the analyses of transient events that are characteristic of whole-cell electrophysiological recordings.

Animals↗

Phase equilibria in carbon dioxide expanded solvents: Experiments and molecular simulations.

We present complementary molecular simulations and experimental results of phase equilibria for carbon dioxide expanded acetonitrile, methanol, ethanol, acetone, acetic acid, toluene, and 1-octene. The volume expansion measurements were done using a high-pressure Jerguson view cell. Molecular simulations were performed using the Gibbs ensemble Monte Carlo method. Calculations in the canonical ensemble (NVT) were performed to determine the coexistence curve of the pure solvent systems. Binary mixtures were simulated in the isobaric-isothermal distribution (NPT). Predictions of vapor-liquid equilibria of the pure components agree well with experimental data. The simulations accurately reproduced experimental data on saturated liquid and vapor densities for carbon dioxide, methanol, ethanol, acetone, acetic acid, toluene, and 1-octene. In all carbon dioxide expanded liquids (CXL's) studied, the molecular simulation results for the volume expansion of these binary mixtures were found to be as good as, and in many cases superior to, predictions based on the Peng-Robinson equation of state, demonstrating the utility of molecular simulation in the prediction of CXL phase equilibria.

Carbon Dioxide↗

Protein-resistant polymer coatings on silicon oxide by surface-initiated atom transfer radical polymerization.

The modification of silicon oxide with poly(ethylene glycol) to effectively eliminate protein adsorption has proven to be technically challenging. In this paper, we demonstrate that surface-initiated atom transfer radical polymerization (SI-ATRP) of oligo(ethylene glycol) methyl methacrylate (OEGMA) successfully produces polymer coatings on silicon oxide that have excellent protein resistance in a biological milieu. The level of serum adsorption on these coatings is below the detection limit of ellipsometry. We also demonstrate a new soft lithography method via which SI-ATRP is integrated with microcontact printing to create micropatterns of poly(OEGMA) on glass that can spatially direct the adsorption of proteins on the bare regions of the substrate. This ensemble of methods will be useful in screening biological interactions where nonspecific binding must be suppressed to discern low probability binding events from a complex mixture and to pattern anchorage-dependent cells on glass and silicon oxide.

Adsorption↗

Apparent stresses in disturbed pulsatile flows.

Traditional attempts at decomposing measured velocities into repeatable and random components are examined for a set of velocity data measured under pulsatile flow conditions distal to a 90% axisymmetric constriction. The Reynolds numbers, which are typical of those found in the human carotid artery, are such that transitional phenomena occur during portions of the pulsatile cycle at several axial stations. The implications of the method selected for velocity decomposition upon the computation of fluctuating or 'apparent' stresses is a point of major focus. It is shown that the usual estimation of Reynolds stresses in a pulsatile flow by subtracting the ensemble-averaged velocity from the instantaneous velocity leads to an underestimation of the apparent stress when coherent or repeatable disturbances exist in the flow. An alternative decomposition using a frequency domain approach is presented which combines both random and coherent stresses into a single apparent stress, and it is proposed that this approach is preferable to the traditional ensemble averaging method when estimating fluctuating stresses in arterial flows.

Blood Flow Velocity↗

Information coding in artificial olfaction multisensor arrays.

High-density sensor arrays were prepared with microbead vapor sensors to explore and compare the information coded in sensor response profiles following odor stimulus. The coded information in the sensor-odor response profiles, which is used for odor discrimination purposes, was extracted from the microsensor arrays via two different approaches. In the first approach, the responses from individual microsensors were separated (decoded array) and independently processed. In the second approach, response profiles from all microsensors within the entire array, i.e., the sensor ensemble, were combined to create one response per odor stimulus (nondecoded array). Although the amount of response data is markedly reduced in the second approach, the system shows comparable odor discrimination rates for the two signal extraction methods. The ensemble approach streamlines system resources without decreasing system performance. These signal compression approaches may simulate or parallel information coding in the mammalian olfactory system.

Animals↗

Noninvasive quantitation of blood flow turbulence in patients with aortic valve disease using online digital computer analysis of Doppler velocity data.

Previous experimental studies have demonstrated that aortic valve disease is associated with significant downstream turbulence (T). In this study, we developed a noninvasive method on the basis of Doppler velocity recording for quantitating aortic blood flow T in patients with aortic valve disease. The instantaneous blood velocity at a point in the aorta is equal to the sum of a mean periodic velocity component with a random or turbulent velocity component. According to the ensemble average method, time mean absolute T intensity is the root-mean-square value of turbulent velocity averaged over time and T is better quantitated by the relative T intensity (TIr), which is the ratio of absolute T intensity to the ensemble average velocity averaged over time. We computed TIr in 18 patients with mild to severe aortic stenosis and in 13 healthy volunteers from instantaneous modal velocities of 70 cycle length-matched heart beats recorded in the proximal part of the descending aorta by pulsed Doppler using an ultrasound system with an output port for online digital data transfer into a microcomputer. TIr was greater in patients with aortic valve disease (18.4 +/- 5.1%, range 11.2%-28.9%) than in control patients (7.9 +/- 1.9%, range 4.8%-9.8%; P =.0001). In patients with aortic valve disease, TIr was better linearly related to the ratio of postvalvular aorta to valvular orifice cross-sectional areas (r = 0.89, P =.0001) than to other parameters of valve restriction: transvalvular pressure gradient (r = 0.78, P =.0001); valve area (r = -0.56, P =.01); and valve resistance (r = 0.72, P =.0002). Thus, T that can be computed noninvasively from direct digital transfer of Doppler velocity data appears to be linearly related to indices of aortic valve restriction. Our data support the concept of the postvalvular aorta to valvular orifice cross-sectional areas ratio as a new important hemodynamic parameter in patients with aortic valve disease.

Adult↗

Prediction of genomewide conserved epitope profiles of HIV-1: classifier choice and peptide representation.

Identification of peptides binding to Major Histocompatibility Complex (MHC) molecules is important for accelerating vaccine development and improving immunotherapy. Accordingly, a wide variety of prediction methods have been applied in this context. In this paper, we introduce (tree-based) ensemble classifiers for such problems and contrast their predictive performance with forefront existing methods for both MHC class I and class II molecules. In addition, we investigate the impact of differing peptide representation schemes on performance. Finally, classifier predictions are used to conduct genomewide scans of a diverse collection of HIV-1 strains, enabling assessment of epitope conservation. We investigated all combinations of six classification methods (classification trees, artificial neural networks, support vector machines, as well as the more recently devised ensemble methods (bagging, random forests, boosting) with four peptide representation schemes (amino acid sequence, select biophysical properties, select quantitative structure-activity relationship (QSAR) descriptors, and the combination of the latter two) in predicting peptide binding to an MHC class I molecule (HLA-A2) and MHC class II molecule (HLA-DR4). Our results show that the ensemble methods are consistently more accurate than the other three alternatives. Furthermore, they are robust with respect to parameter tuning. Among the four representation schemes, the amino acid sequence representation gave consistently (across classifiers) best results. This finding obviates the need for feature selection strategies incurred by use of biophysical and/or QSAR properties. We obtained, and aligned, a diverse set of 32 HIV-1 genomes and pursued genomewide HLA-DR4 epitope profiling by querying with respect to classifier predictions, as obtained under each of the four peptide representation schemes. We validated those epitopes conserved across strains against known T-cell epitopes. Once again, amino acid sequence representation was at least as effective as using properties. Assessment of novel epitope predictions awaits experimental verification.

Journal Article↗

Determination of solvation free energies by adaptive expanded ensemble molecular dynamics.

A new method of calculating absolute free energies is presented. It was developed as an extension to the expanded ensemble molecular dynamics scheme and uses probability density estimation to continuously optimize the expanded ensemble parameters. The new method is much faster as it removes the time-consuming and expertise-requiring step of determining balancing factors. Its efficiency and accuracy are demonstrated for the dissolution of three qualitatively very different chemical species in water: methane, ionic salts, and benzylamine. A recently suggested optimization scheme by Wang and Landau [Phys. Rev. Lett. 86, 2050 (2001)] was also implemented and found to be computationally less efficient than the proposed adaptive expanded ensemble method.

Journal Article↗

Interactions between spherical colloids mediated by a liquid crystal: a molecular simulation and mesoscale study.

Monte Carlo simulations and dynamic field theory (DyFT) are used to study the interactions between dilute spherical particles, dispersed in nematic and isotropic phases of a liquid crystal. A recently developed simulation method (expanded ensemble density of states) was used to determine the potential of mean force (PMF) between the two spheres as a function of their separation and size. The PMF was also calculated by a dynamic field theory that describes the evolution of the local tensor order parameter. Both methods reveal an overall attraction between the colloids in the nematic phase; in the isotropic phase, the overall attraction between the colloids is much weaker, whereas the repulsion at short range is stronger. In addition, both methods predict a new topology of the disclination lines, which arises when the particles approach each other. The theory is found to describe the results of simulations remarkably well, down to length scales comparable to the size of the molecules. At separations corresponding to the width of individual molecular layers on the particles' surface, the two methods yield different defect structures. We attribute this difference to the neglect of density inhomogeneities in the DyFT. We also investigate the effects of the size of spherical colloids on their interactions.

Biosensing Techniques↗

Anion recognition: synthetic receptors for anions and their application in sensors.

Important contributions to the field of anion sensing include electrochemical lipophilic uranyl salophene receptors incorporated into membranes that act as fluoride-selective potentiometric microsensors. A promising optical-based sensor, selective for cyclic AMP, involves a preorganized, molecularly imprinted polymer employing an intrinsic fluorophore. Competition methods using ensembles of recognition units and external indicators have been used to sense citrate in highly competitive media and micromolar concentrations of inositol(tris)phosphate in water. In addition, DNA dendrimers immobilized on a quartz-crystal microbalance acted as an elegant biosensor for Cryptosporidium DNA. These designs display the varied methods of anion detection currently being pursued.

Anions↗

PredIL13: Stacking a variety of machine and deep learning methods with ESM-2 language model for identifying IL13-inducing peptides.

Interleukin (IL)-13 has emerged as one of the recently identified cytokine. Since IL-13 causes the severity of COVID-19 and alters crucial biological processes, it is urgent to explore novel molecules or peptides capable of including IL-13. Computational prediction has received attention as a complementary method to in-vivo and in-vitro experimental identification of IL-13 inducing peptides, because experimental identification is time-consuming, laborious, and expensive. A few computational tools have been presented, including the IL13Pred and iIL13Pred. To increase prediction capability, we have developed PredIL13, a cutting-edge ensemble learning method with the latest ESM-2 protein language model. This method stacked the probability scores outputted by 168 single-feature machine/deep learning models, and then trained a logistic regression-based meta-classifier with the stacked probability score vectors. The key technology was to implement ESM-2 and to select the optimal single-feature models according to their absolute weight coefficient for logistic regression (AWCLR), an indicator of the importance of each single-feature model. Especially, the sequential deletion of single-feature models based on the iterative AWCLR ranking (SDIWC) method constructed the meta-classifier consisting of the top 16 single-feature models, named PredIL13, while considering the model's accuracy. The PredIL13 greatly outperformed the-state-of-the-art predictors, thus is an invaluable tool for accelerating the detection of IL13-inducing peptide within the human genome.

Humans↗

Ensemble recordings of human subcortical neurons as a source of motor control signals for a brain-machine interface.

OBJECTIVE: Patients with severe neurological injury, such as quadriplegics, might benefit greatly from a brain-machine interface that uses neuronal activity from motor centers to control a neuroprosthetic device. Here, we report an implementation of this strategy in the human intraoperative setting to assess the feasibility of using neurons in subcortical motor areas to drive a human brain-machine interface. METHODS: Acute ensemble recordings from subthalamic nucleus and thalamic motor areas (ventralis oralis posterior [VOP]/ventralis intermediate nucleus [VIM]) were obtained in 11 awake patients during deep brain stimulator surgery by use of a 32-microwire array. During extracellular neuronal recordings, patients simultaneously performed a visual feedback hand-gripping force task. Offline analysis was then used to explore the relationship between neuronal modulation and gripping force. RESULTS: Individual neurons (n = 28 VOP/VIM, n = 119 subthalamic nucleus) demonstrated a variety of modulation responses both before and after onset of changes in gripping force of the contralateral hand. Overall, 61% of subthalamic nucleus neurons and 81% of VOP/VIM neurons modulated with gripping force. Remarkably, ensembles of 3 to 55 simultaneously recorded neurons were sufficiently information-rich to predict gripping force during 30-second test periods with considerable accuracy (up to R = 0.82, R(2) = 0.68) after short training periods. Longer training periods and larger neuronal ensembles were associated with improved predictive accuracy. CONCLUSION: This initial feasibility study bridges the gap between the nonhuman primate laboratory and the human intraoperative setting to suggest that neuronal ensembles from human subcortical motor regions may be able to provide informative control signals to a future brain-machine interface.

Electrodes, Implanted↗

Development and external validation of an explainable machine learning model for predicting chronic kidney disease progression in the Korean population.

BACKGROUND: Current risk stratification models, such as the Kidney Failure Risk Equation (KFRE), exhibit variable performance across ethnic groups and fail to capture dynamic clinical trajectories. This study aimed to develop and validate a Korean-specific machine learning (ML) model for predicting chronic kidney disease (CKD) progression using an ensemble approach. METHODS: We used electronic health records from Seoul National University Hospital for model development (n = 28,209) and the Korean Genome and Epidemiology Study (KoGES) CKD cohort for external validation (n = 3,960). The primary outcome was a composite of ≥40% decline in estimated glomerular filtration rate (eGFR) or progression to end-stage renal disease within 2 years. A soft-voting ensemble of four ML algorithms (XGBoost, LightGBM, CatBoost, and Random Forest) was developed. RESULTS: The ensemble model demonstrated robust discrimination in internal validation (area under the receiver operating characteristic curve [AUROC], 0.939; 95% confidence interval [CI], 0.934-0.944), significantly exceeding the KFRE (AUROC, 0.879-0.884). External validation in the KoGES cohort showed comparable discrimination (AUROC, 0.859; 95% CI, 0.798-0.914) versus KFRE (four-variable AUROC, 0.882; 95% CI, 0.818-0.935). Shapley Additive exPlanations (SHAP) analysis identified baseline eGFR, serum creatinine, eGFR slope, albumin, and hemoglobin as key prognostic features, supporting a complementary framework using KFRE for community screening and the ML model for hospital-based risk stratification. CONCLUSION: The ensemble ML model accurately predicts short-term CKD progression in Korean patients. By incorporating longitudinal features and ensemble learning, it provides a precise alternative to Western-derived equations, particularly in tertiary care settings.

Chronic kidney failure↗

Automated multiple structure alignment and detection of a common substructural motif.

While a number of approaches have been geared toward multiple sequence alignments, to date there have been very few approaches to multiple structure alignment and detection of a recurring substructural motif. Among these, none performs both multiple structure comparison and motif detection simultaneously. Further, none considers all structures at the same time, rather than initiating from pairwise molecular comparisons. We present such a multiple structural alignment algorithm. Given an ensemble of protein structures, the algorithm automatically finds the largest common substructure (core) of C(alpha) atoms that appears in all the molecules in the ensemble. The detection of the core and the structural alignment are done simultaneously. Additional structural alignments also are obtained and are ranked by the sizes of the substructural motifs, which are present in the entire ensemble. The method is based on the geometric hashing paradigm. As in our previous structural comparison algorithms, it compares the structures in an amino acid sequence order-independent way, and hence the resulting alignment is unaffected by insertions, deletions and protein chain directionality. As such, it can be applied to protein surfaces, protein-protein interfaces and protein cores to find the optimally, and suboptimally spatially recurring substructural motifs. There is no predefinition of the motif. We describe the algorithm, demonstrating its efficiency. In particular, we present a range of results for several protein ensembles, with different folds and belonging to the same, or to different, families. Since the algorithm treats molecules as collections of points in three-dimensional space, it can also be applied to other molecules, such as RNA, or drugs.

Algorithms↗

Probability assessment of conformational ensembles: sugar repuckering in a DNA duplex in solution.

Conformational flexibility of molecules in solution implies that different conformers contribute to the NMR signal. This may lead to internal inconsistencies in the 2D NOE-derived interproton distance restraints and to conflict with scalar coupling-based torsion angle restraints. Such inconsistencies have been revealed and analyzed for the DNA octamer GTATAATG.CATATTAC, containing the Pribnow box consensus sequence. A number of subsets of distance restraints were constructed and used in the restrained Monte Carlo refinement of different double-helical conformers. The probabilities of conformers were then calculated by a quadratic programming algorithm, minimizing a relaxation rate-base residual index. The calculated distribution of conformers agrees with the experimental NOE data as an ensemble better than any single structure. A comparison with the results of this procedure, which we term PARSE (Probability Assessment via Relaxation rates of a Structural Ensemble), to an alternative method to generate solution ensembles showed, however, that the detailed multi-conformational description of solution DNA structure remains ambiguous at this stage. Nevertheless, some ensemble properties can be deduced with confidence, the most prominent being a distribution of sugar puckers with minor populations in the N-region and major populations in the S-region. Importantly, such a distribution is in accord with the analysis of independent experimental data--deoxyribose proton-proton scalar coupling constants.

Algorithms↗