Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Lung cancer cell identification based on artificial neural network ensembles.

An artificial neural network ensemble is a learning paradigm where several artificial neural networks are jointly used to solve a problem. In this paper, an automatic pathological diagnosis procedure named Neural Ensemble-based Detection (NED) is proposed, which utilizes an artificial neural network ensemble to identify lung cancer cells in the images of the specimens of needle biopsies obtained from the bodies of the subjects to be diagnosed. The ensemble is built on a two-level ensemble architecture. The first-level ensemble is used to judge whether a cell is normal with high confidence where each individual network has only two outputs respectively normal cell or cancer cell. The predictions of those individual networks are combined by a novel method presented in this paper, i.e. full voting which judges a cell to be normal only when all the individual networks judge it is normal. The second-level ensemble is used to deal with the cells that are judged as cancer cells by the first-level ensemble, where each individual network has five outputs respectively adenocarcinoma, squamous cell carcinoma, small cell carcinoma, large cell carcinoma, and normal, among which the former four are different types of lung cancer cells. The predictions of those individual networks are combined by a prevailing method, i.e. plurality voting. Through adopting those techniques, NED achieves not only a high rate of overall identification, but also a low rate of false negative identification, i.e. a low rate of judging cancer cells to be normal ones, which is important in saving lives due to reducing missing diagnoses of cancer patients.

Adenocarcinoma↗

Inverse Monte Carlo procedure for conformation determination of macromolecules.

A novel numerical method for determining the conformational structure of macromolecules is applied to idealized biomacromolecules in solution. The method computes effective inter-residue interaction potentials solely from the corresponding radial distribution functions, such as would be obtained from experimental data. The interaction potentials generate conformational ensembles that reproduce thermodynamic properties of the macromolecule (mean energy and heat capacity) in addition to the target radial distribution functions. As an evaluation of its utility in structure determination, we apply the method to a homopolymer and a heteropolymer model of a three-helix bundle protein [Zhou, Y.; Karplus, M. Proc Natl Acad Sci USA 1997, 94, 14429; Zhou, Y. et al. J Chem Phys 1997, 107, 10691] at various thermodynamic state points, including the ordered globule, disordered globule, and random coil states.

Journal Article↗

Rapid determination of complex mixtures by dual-column gas chromatography with a novel stationary phase combination and spectrometric detection.

Fast GC separations of a broad range of analytes are demonstrated using a capillary column coated with a novel immobilized ionic liquid (IIL) stationary phase. Both completely cross-linked and partially cross-linked columns were evaluated, yielding approximately 1600 and approximately 2000 theoretical plates per meter, respectively. Enhanced separation is demonstrated using a dual-column ensemble comprised of an IIL column, a commercially coated Rtx-1 column, and a pneumatic valve connecting the inlet to the junction point between the two columns. Enhanced separation of 20 components, with two sets of co-eluting peaks is shown in approximately 150 s, while sacrificing only a length of time equivalent to the sum of the stop flow pulses, or about 15.5 s. A novel application of a band trajectory model that shows band position as a function of analysis time as analytes move through the column ensemble is employed to determine pulse application times. The model predicts component retention times within a few seconds. Another method of selectivity enhancement of the IIL stationary phase-coated columns is demonstrated using a differential mobility spectrometer (DMS) that provides a second dimension separation based on ion mobility in a high-frequency electrical field. The DMS is able to separate all but one set of co-eluting components from the IIL column. The separation of 13 components found in the headspace above U.S. currency is demonstrated using the IIL column in a dual-column ensemble as well as with the DMS.

Chromatography, Gas↗

Evaluating the conformational entropy of macromolecules using an energy decomposition approach.

We have developed a novel method to compute the conformational entropy of any molecular system via conventional simulation techniques. This method only requires that the total energy of the system is available and that the Hamiltonian is separable, with individual energy terms for the various degrees of freedom. Consequently the method, which we call the energy decomposition (Edcp) approach, is general and applicable to any large polymer in implicit solvent. Edcp is applied to estimate the entropy differences due to the peptide and ester groups in polyalanine and polyalanil ester. Ensembles over a wide range of temperatures were generated by replica exchange molecular dynamics, and densities of states were estimated using the weighted histogram analysis method. The results are compared with those obtained via evaluating the P ln P integral or employing the quasiharmonic approximation, other approaches widely employed to evaluate the entropy of molecular systems. Unlike the former method, Edcp can accommodate the correlations present between separate degrees of freedom. In addition, the Edcp model assumes no specific form for the underlying fluctuations present in the system, in contrast to the quasiharmonic approximation. For the molecules studied, the quasiharmonic approximation is observed to produce a good estimate of the vibrational entropy, but not of the conformational entropy. In contrast, our energy decomposition approach generates reasonable estimates for both of these entropy terms. We suggest that this approach embodies a simple yet effective solution to the problem of evaluating the conformational entropy of large macromolecules in implicit solvent.

Chemistry, Physical↗

Ensemble variance in free energy calculations by thermodynamic integration: theory, optimal "Alchemical" path, and practical solutions.

Thermodynamic integration is a widely used method to calculate and analyze the effect of a chemical modification on the free energy of a chemical or biochemical process, for example, the impact of an amino acid substitution on protein association. Numerical fluctuations can introduce large uncertainties, limiting the domain of application of the method. The parametric energy function describing the chemical modification in the thermodynamic integration, the "Alchemical path," determines the amplitudes of the fluctuations. In the present work, I propose a measure of the fluctuations in the thermodynamic integration and an approach to search for a parametric energy path minimizing that measure. The optimal path derived with this approach is very close to the theoretical minimum of the measure, but produces nonergodic sampling. Nevertheless, this path is used to guide the design of a practical and efficient path producing correct sampling. The convergence with this practical path is evaluated on test cases, and compares favorably with that of other methods such as power or polynomial path, soft-core van der Waals, and some other approaches presented in the literature.

Journal Article↗

Exploring brain circuitry with neurotropic viruses: new horizons in neuroanatomy.

There have been substantial advances in methods for defining connections among neurons over the past quarter century. However, most tracers have been limited in their ability to define populations of functionally related neurons that contribute to a multisynaptic circuit because they are not transported across synapses. As a result, the large body of literature that has employed these tracers has established regional associations between regions that must be further explored with electron microscopy and electrophysiological methods to define the synaptic relations among constituent neurons. Recently, neurotropic alpha herpesviruses have been used to visualize ensembles of neurons that contribute to polysynaptic networks. These pathogens invade permissive cells, replicate, and pass transynaptically to infect other neurons. In effect, the viruses become self-amplifying tracers whose natural tropism and invasiveness define populations of functionally related neurons. The recent increase in the use of this experimental approach has emerged from advances in our understanding of the life cycle of these viruses and the resulting evidence in support of specific transynaptic passage of progeny virus rather than infection by lytic release into the extracellular space. This article reviews the advances that have made this a viable experimental approach and considers ways in which this method has been creatively used to illuminate aspects of nervous system circuit organization that could not be defined with conventional tracers.

Animals↗

Long- and short-range interactions in native protein structures are consistent/minimally frustrated in sequence space.

We show that long- and short-range interactions in almost all protein native structures are actually consistent with each other for coarse-grained energy scales; specifically we mean the long-range inter-residue contact energies and the short-range secondary structure energies based on peptide dihedral angles, which are potentials of mean force evaluated from residue distributions observed in protein native structures. This consistency is observed at equilibrium in sequence space rather than in conformational space. Statistical ensembles of sequences are generated by exchanging residues for each of 797 protein native structures with the Metropolis method. It is shown that adding the other category of interaction to either the short- or long-range interactions decreases the means and variances of those energies for essentially all protein native structures, indicating that both interactions consistently work by more-or-less restricting sequence spaces available to one of the interactions. In addition to this consistency, independence by these interaction classes is also indicated by the fact that there are almost no correlations between them when equilibrated using both interactions and significant but small, positive correlations at equilibrium using only one of the interactions. Evidence is provided that protein native sequences can be regarded approximately as samples from the statistical ensembles of sequences with these energy scales and that all proteins have the same effective conformational temperature. Designing protein structures and sequences to be consistent and minimally frustrated among the various interactions is a most effective way to increase protein stability and foldability.

Amino Acid Sequence↗

Predicting enhancer-promoter interactions using a stacking-based ensemble strategy.

MOTIVATION: Enhancer-promoter interactions (EPIs) are essential for gene regulation and disease progression. Recent studies have shown that distal enhancers can regulate target genes through interactions with nearby promoters, providing important insights into transcriptional regulation mechanisms. Although high-throughput experimental techniques have enabled large-scale identification of EPIs, these methods are often costly and time-consuming. In addition, existing computational approaches still face challenges in effectively integrating heterogeneous feature representations from different cell lines. RESULTS: We propose a stacked ensemble framework for EPI prediction that integrates feature representations from diverse cell line datasets using multiple machine learning algorithms. The extracted complementary patterns are further combined by an XGBoost classifier to improve robustness against overfitting. Experiments on six independent datasets show that the proposed method achieves superior accuracy and generalization compared with existing EPI prediction models, with an average AUROC of 0.909 while maintaining computational efficiency. AVAILABILITY: The source code and its archived release are available at GitHub and Zenodo. The Zenodo archive provides a versioned snapshot of the repository: https://zenodo.org/records/19952998.

Promoter Regions, Genetic↗

Noninvasive detection and differentiation of gastric malignancy using cell-free DNA biomarkers.

INTRODUCTION: Gastric cancer remains a major global health burden, with high mortality driven by late-stage diagnoses that limit treatment options and reduce survival. Current diagnostic methods such as endoscopy and biopsy are invasive, resource-intensive, and impractical for large-scale early detection. OBJECTIVES: This study aimed to develop and validate an ensemble machine learning model integrating four cell-free DNA (cfDNA) fragmentomic feature classes derived from 5 × whole genome sequencing (WGS) data to non-invasively differentiate malignant gastric cancer from benign gastric lesions in high-risk or symptomatic patients. METHODS: A total of 681 plasma samples were prospectively collected, comprising 329 from patients with gastric cancer or high-grade intraepithelial neoplasia (HGIN) and 352 from individuals with benign gastric conditions. The dataset was divided into a training cohort (n = 333) and a temporally independent validation cohort (n = 348). An external validation cohort of 305 participants was also included. RESULTS: The ensemble model achieved an AUROC of 0.920 in cross-validation testing on the training cohort, 0.912 in the independent validation cohort, and 0.896 (95% CI 0.860-0.932) in the external cohort. At a pre-specified prediction threshold of 0.402, the model demonstrated 93.3% sensitivity and 71.9% specificity in the validation cohort, yielding a PPV of 71.3% and an NPV of 93.5%. In the external cohort, sensitivity and specificity were 91.7% and 69.1%, respectively (PPV 75.7%, NPV 88.8%). Model scores correlated with clinical stage, tumor grade, and histopathological subtype. Approximately 71% of non-cancer patients could have been spared unnecessary endoscopy. CONCLUSIONS: The cfDNA fragmentomics-based ensemble model enables accurate, non-invasive differentiation between gastric cancer and benign gastric lesions in high-risk or symptomatic patients. This approach demonstrates strong potential as a pre-endoscopy triage tool, supporting earlier detection and more efficient use of diagnostic resources.

Humans↗

Theoretical calculations of infrared absorption, vibrational circular dichroism, and two-dimensional vibrational spectra of acetylproline in liquids water and chloroform.

Infrared absorption, vibrational circular dichroism, and two-dimensional infrared pump-probe and photon echo spectra of acetylproline solutions are theoretically calculated and directly compared with experiments. In order to quantitatively determine interpeptide interaction-induced amide I mode frequency shifts, high-level quantum chemistry calculations were performed. The solvatochromic amide I mode frequency shift and fluctuation were taken into account by carrying out molecular dynamics simulations of acetylproline dissolved in liquids water and chloroform and by using the extrapolation method developed recently. We first studied correlation time scales of the two amide I vibrational frequency fluctuations, cross correlation between the two fluctuating local mode frequencies, ensemble averaged conformations of the acetylproline molecule in liquids water and chloroform. The corresponding conformations of the acetylproline in liquids water and chloroform are close to the ideal 3(10) helix and the C(7) structure, respectively. A few methods proposed to determine the angle between the two transition dipoles associated with the amide I vibrations were tested and their limitations are discussed.

Journal Article↗

Decomposing stimulus and response component waveforms in ERP.

Event-related potentials (ERPs) are evoked brain potentials that are averaged across many trial repetitions with individual trials aligned (i.e. time-locked) to a specific behavioral event, typically the onset of the stimulus (s-lock) or the onset of the behavioral response (r-lock). These evoked potential averages may reflect brain activities during the stimulus encoding/analyzing stage (stimulus component waveform, or 'S-component'), during the response preparation/production stage (response component waveform, or 'R-component'), or a combination thereof. In the stimulus-locked average of the ensemble of the recorded waveforms (i.e. in the s-locked ERP), the contribution of an R-component will be convoluted, due to the trial-by-trial variance in reaction time (RT): so will an S-component in the r-locked ERP. It is shown here that the knowledge of (1) the s-locked and r-locked ERP waveforms constructed from the same ensemble of trials and (2) the RT distribution of this ensemble allows us to determine whether the recorded potential results from a single S-component, a single R-component, or a single intermediate ('decisional' or D-) component related to the transition of the two stochastically independent stages. If it can be assumed that the evoked potential is the result of a linear summation of an S-component and an R-component, then there is a unique recovery into these two components, such that the reconstructed waveform on an individual trial is a superposition of the two components with their relative offset determined by the RT of that trial and the ensemble average is the experimentally obtained s-locked and r-locked ERP waveforms. Two independent methods can be used to recover those components, one based on Fourier transform techniques which was first proposed by Hansen (1983) in the context of ERP component isolation and the other based on a recursive iteration approach through which the contamination of the R or S-component is successively removed from the s-locked or r-locked ERP waveforms, respectively. The iterative procedure is analytically proven to converge to the Fourier-based solution, demonstrating the equivalence of the two approaches. Finally, if the condition of a single intermediate D-component is satisfied, then one can recover this component waveform along with the probability distributions of the relative durations of the two underlying linear stages; however, there is always an equivalent pair of S- and R-component which also satisfy the same data set (s-locked and r-locked ERP waveforms and the overall RT distribution). In this case, the S/R-component assumption and the D-component assumption cannot be distinguished solely on the ground of the available data set. The technique developed here outlines the assumptions and the boundary conditions upon which ensemble ERP waveforms are to be analyzed and interpreted in terms of processing mechanisms related to stimulus, to response, or to the transition between the two.

Algorithms↗

Activated sampling in complex materials at finite temperature: the properly obeying probability activation-relaxation technique.

While the dynamics of many complex systems is dominated by activated events, there are very few simulation methods that take advantage of this fact. Most of these procedures are restricted to relatively simple systems or, as with the activation-relaxation technique (ART), sample the conformation space efficiently at the cost of a correct thermodynamical description. We present here an extension of ART, the properly obeying probability ART (POP-ART), that obeys detailed balance and samples correctly the thermodynamic ensemble. Testing POP-ART on two model systems, a vacancy and an interstitial in crystalline silicon, we show that this method recovers the proper thermodynamical weights associated with the various accessible states and is significantly faster than molecular dynamics in the simulations of a vacancy below 700 K.

Journal Article↗

Ascertaining the importance of neurons to develop better brain-machine interfaces.

In the design of brain-machine interface (BMI) algorithms, the activity of hundreds of chronically recorded neurons is used to reconstruct a variety of kinematic variables. A significant problem introduced with the use of neural ensemble inputs for model building is the explosion in the number of free parameters. Large models not only affect model generalization but also put a computational burden on computing an optimal solution especially when the goal is to implement the BMI in low-power, portable hardware. In this paper, three methods are presented to quantitatively rate the importance of neurons in neural to motor mapping, using single neuron correlation analysis, sensitivity analysis through a vector linear model, and a model-independent cellular directional tuning analysis for comparisons purpose. Although, the rankings are not identical, up to sixty percent of the top 10 ranking cells were in common. This set can then be used to determine a reduced-order model whose performance is similar to that of the ensemble. It is further shown that by pruning the initial ensemble neural input with the ranked importance of cells, a reduced sets of cells (between 40 and 80, depending upon the methods) can be found that exceed the BMI performance levels of the full ensemble.

Action Potentials↗

Uncertainties associated with parameter estimation in atmospheric infrasound arrays.

This study describes a method for determining the statistical confidence in estimates of direction-of-arrival and trace velocity stemming from signals present in atmospheric infrasound data. It is assumed that the signal source is far enough removed from the infrasound sensor array that a plane-wave approximation holds, and that multipath and multiple source effects are not present. Propagation path and medium inhomogeneities are assumed not to be known at the time of signal detection, but the ensemble of time delays of signal arrivals between array sensor pairs is estimable and corrupted by uncorrelated Gaussian noise. The method results in a set of practical uncertainties that lend themselves to a geometric interpretation. Although quite general, this method is intended for use by analysts interpreting data from atmospheric acoustic arrays, or those interested in designing and deploying them. The method is applied to infrasound arrays typical of those deployed as a part of the International Monitoring System of the Comprehensive Nuclear-Test-Ban Treaty Organization.

Journal Article↗

Influence of multiple well defined conformations on small-angle scattering of proteins in solution.

A common structural motif for many proteins comprises rigid domains connected by a flexible hinge or linker. The flexibility afforded by these domains is important for proper function and such proteins may be able to adopt more than one conformation in solution under equilibrium conditions. Small-angle scattering of proteins in solution samples all conformations that exist in the sampled volume during the time of the measurement, providing an ensemble-averaged intensity. In this paper, the influence of sampling an ensemble of well defined protein structures on the small-angle solution scattering intensity profile is examined through common analysis methods. Two tests were performed using simulated data: one with the extended and collapsed states of the bilobal calcium-binding protein calmodulin and the second with the catalytic subunit of protein kinase A, which has two globular domains connected by a glycine hinge. In addition to analyzing the simulated data for the radii of gyration Rg, distance distribution function P(r) and particle volume, shape restoration was applied to the simulated data. Rg and P(r) of the ensemble profiles could be easily mistaken for a single intermediate state. The particle volumes and models of the ensemble intensity profiles show that some indication of multiple conformations exists in the case of calmodulin, which manifests an enlarged volume and shapes that are clear superpositions of the conformations used. The effect on the structural parameters and models is much more subtle in the case of the catalytic subunit of protein kinase A. Examples of how noise influences the data and analyses are also presented. These examples demonstrate the loss of the indications of multiple conformations in cases where even broad distributions of structures exist. While the tests using calmodulin show that the ensemble states remain discernible from the other ensembles tested or a single partially collapsed state, the tests performed using the simulated catalytic subunit of protein kinase A with noise added demonstrate that it can mask out the ensemble-dependent effects observed for the noiseless profiles.

Amino Acid Motifs↗

Monte Carlo-minimization approach to the multiple-minima problem in protein folding.

A Monte Carlo-minimization method has been developed to overcome the multiple-minima problem. The Metropolis Monte Carlo sampling, assisted by energy minimization, surmounts intervening barriers in moving through successive discrete local minima in the multidimensional energy surface. The method has located the lowest-energy minimum thus far reported for the brain pentapeptide [Met5]enkephalin in the absence of water. Presumably it is the global minimum-energy structure. This supports the concept that protein folding may be a Markov process. In the presence of water, the molecules appear to exist as an ensemble of different conformations.

Enkephalin, Methionine↗

Optical methods for exploring dynamics of single copies of green fluorescent protein.

Single copies of four different phenolate ion mutants of the green fluorescent protein (GFP) exhibit a complex blinking and fluctuating behavior, a phenomenon that is hidden in measurements on large ensembles. Both total internal reflection microscopy and scanning confocal microscopy can be used to study the blinking dynamics, and autocorrelation analysis yields histograms of the correlation times for many individual molecules. While the total internal reflection method can follow several single molecules simultaneously, the confocal method offers higher time resolution at the expense of parallelism. We compare and contrast the two methods in terms of the ability to follow the complex dynamics of this system.

Green Fluorescent Proteins↗

Reconstruction of the postsubiculum head direction signal from neural ensembles.

Head direction cells change their firing rates as a function of the orientation of an animal within an environment. Typically, these cells display a unimodal tuning curve with maximal firing at the cell's preferred direction. As different cells have different preferred directions, the population of cells has been hypothesized to represent the orientation of the animal within the environment. Previous research has shown that pairs of simultaneously recorded head direction cells respond similarly to cue manipulations, suggesting that a population of head direction cells acts in concert to represent the animal's orientation within its environment. Ensembles of head direction cells were recorded from the postsubiculum from rats foraging in an open field. Directional responses of each cell were quantified by the nonparametric Watson's U2 statistic, a measure which makes no explicit assumptions of tuning curve shape. Directionally responsive cells were then used to reconstruct each animal's orientation within the open field using population vector, optimal-linear estimator, and Bayesian methods. The results indicated that postsubiculum contained a complete representation of the animal's orientation. The internal consistency of a neural ensemble can be assessed by comparing the ensemble activity to the expected activity given the reconstructed orientation. This has been termed the "coherency" of the neural ensemble. Reconstruction error decreased as the coherency of the orientation representation increased, indicating that coherency could be used to measure a level of confidence in the representation quality. Because coherency is a linear measure dependent only on internal variables, coherency may be a behaviorally relevant measure used to ascertain the animal's confidence in its representation of orientation.

Action Potentials↗