Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Human-mouse gene identification by comparative evidence integration and evolutionary analysis.

The identification of genes in the human genome remains a challenge, as the actual predictions appear to disagree tremendously and vary dramatically on the basis of the specific gene-finding methodology used. Because the pattern of conservation in coding regions is expected to be different from intronic or intergenic regions, a comparative computational analysis can lead, in principle, to an improved computational identification of genes in the human genome by using a reference, such as mouse genome. However, this comparative methodology critically depends on three important factors: (1) the selection of the most appropriate reference genome. In particular, it is not clear whether the mouse is at the correct evolutionary distance from the human to provide sufficiently distinctive conservation levels in different genomic regions, (2) the selection of comparative features that provide the most benefit to gene recognition, and (3) the selection of evidence integration architecture that effectively interprets the comparative features. We address the first question by a novel evolutionary analysis that allows us to explicitly correlate the performance of the gene recognition system with the evolutionary distance (time) between the two genomes. Our simulation results indicate that there is a wide range of reference genomes at different evolutionary time points that appear to deliver reasonable comparative prediction of human genes. In particular, the evolutionary time between human and mouse generally falls in the region of good performance; however, better accuracy might be achieved with a reference genome further than mouse. To address the second question, we propose several natural comparative measures of conservation for identifying exons and exon boundaries. Finally, we experiment with Bayesian networks for the integration of comparative and compositional evidence.

Animals↗

Predicting the operon structure of Bacillus subtilis using operon length, intergene distance, and gene expression information.

We predict the operon structure of the Bacillus subtilis genome using the average operon length, the distance between genes in base pairs, and the similarity in gene expression measured in time course and gene disruptant experiments. By expressing the operon prediction for each method as a Bayesian probability, we are able to combine the four prediction methods into a Bayesian classifier in a statistically rigorous manner. The discriminant value for the Bayesian classifier can be chosen by considering the associated cost of misclassifying an operon or a non-operon gene pair. For equal costs, an overall accuracy of 88.7% was found in a leave-one-out analysis for the joint Bayesian classifier, whereas the individual information sources yielded accuracies of 58.1%, 83.1%, 77.3%, and 71.8% respectively. The predicted operon structure based on the joint Bayesian classifier is available from the DBTBS database (http://dbtbs.hgc.jp).

Bacillus subtilis↗

Wavelets and functional magnetic resonance imaging of the human brain.

The discrete wavelet transform (DWT) is widely used for multiresolution analysis and decorrelation or "whitening" of nonstationary time series and spatial processes. Wavelets are naturally appropriate for analysis of biological data, such as functional magnetic resonance images of the human brain, which often demonstrate scale invariant or fractal properties. We provide a brief formal introduction to key properties of the DWT and review the growing literature on its application to fMRI. We focus on three applications in particular: (i) wavelet coefficient resampling or "wavestrapping" of 1-D time series, 2- to 3-D spatial maps and 4-D spatiotemporal processes; (ii) wavelet-based estimators for signal and noise parameters of time series regression models assuming the errors are fractional Gaussian noise (fGn); and (iii) wavelet shrinkage in frequentist and Bayesian frameworks to support multiresolution hypothesis testing on spatially extended statistic maps. We conclude that the wavelet domain is a rich source of new concepts and techniques to enhance the power of statistical analysis of human fMRI data.

Algorithms↗

A novel neural network-based survival analysis model.

A feedforward neural network architecture aimed at survival probability estimation is presented which generalizes the standard, usually linear, models described in literature. The network builds an approximation to the survival probability of a system at a given time, conditional on the system features. The resulting model is described in a hierarchical Bayesian framework. Experiments with synthetic and real world data compare the performance of this model with the commonly used standard ones.

Bayes Theorem↗

Bayesian internal dosimetry calculations using Markov Chain Monte Carlo.

A new numerical method for solving the inverse problem of internal dosimetry is described. The new method uses Markov Chain Monte Carlo and the Metropolis algorithm. Multiple intake amounts, biokinetic types, and times of intake are determined from bioassay data by integrating over the Bayesian posterior distribution. The method appears definitive, but its application requires a large amount of computing time.

Algorithms↗

Extension of a local backbone description using a structural alphabet: a new approach to the sequence-structure relationship.

Protein Blocks (PBs) comprise a structural alphabet of 16 protein fragments, each 5 Calpha long. They make it possible to approximate and correctly predict local protein three-dimensional (3D) structures. We have selected the 72 most frequent sequences of five PBs, which we call Structural Words (SWs). Analysis of four different protein data banks shows that SWs cover 92% of the amino acids in them and provide a good structural approximation for residues (i.e., sequences) 9 Calpha long. We present most of them in a simple network that describes 90% of the overall residues and, interestingly, includes more than 80% of the amino acids present in coils. Analysis of the network shows the specificity and quality of the 3D descriptions as well as a new type of relation between local folds and amino acid distribution. The results show that the 3D structure of these protein data banks can be easily described by a combination of subgraphs included in the network. Finally, a Bayesian probabilistic approach improved the prediction rate by 4%.

Amino Acid Sequence↗

Automatic identification of Mycobacterium tuberculosis by Gaussian mixture models.

Tuberculosis and other kinds of mycobacteriosis are serious illnesses for which early diagnosis is critical for disease control. Sputum sample analysis is a common manual technique employed for bacillus detection but current sample-analysis techniques are time-consuming, very tedious, subject to poor specificity and require highly trained personnel. Image-processing and pattern-recognition techniques are appropriate tools for improving the manual screening of samples. Here we present a new technique for sputum image analysis that combines invariant shape features and chromatic channel thresholding. Some feature descriptors were extracted from an edited bacillus data set to characterize their shape. They were statistically represented by using a Gaussian mixture model representation and a minimal error Bayesian classification procedure was employed for the last identification stage. This technique constitutes a step towards automating the process and providing a high specificity.

Bacteriological Techniques↗

Bayesian inference in multipoint gene mapping.

The problem of ordering and mapping genes on the basis of recombinant data and radiation hybrid data is formulated as a problem of Bayesian inference for an unknown permutation. The challenging computational problems posed by this approach are shown to be resolvable using Markov chain Monte Carlo methods.

Algorithms↗

On classification with incomplete data.

We address the incomplete-data problem in which feature vectors to be classified are missing data (features). A (supervised) logistic regression algorithm for the classification of incomplete data is developed. Single or multiple imputation for the missing data is avoided by performing analytic integration with an estimated conditional density function (conditioned on the observed data). Conditional density functions are estimated using a Gaussian mixture model (GMM), with parameter estimation performed using both Expectation-Maximization (EM) and Variational Bayesian EM (VB-EM). The proposed supervised algorithm is then extended to the semisupervised case by incorporating graph-based regularization. The semisupervised algorithm utilizes all available data-both incomplete and complete, as well as labeled and unlabeled. Experimental results of the proposed classification algorithms are shown.

Algorithms↗

Bayesian gene/species tree reconciliation and orthology analysis using MCMC.

MOTIVATION: Comparative genomics in general and orthology analysis in particular are becoming increasingly important parts of gene function prediction. Previously, orthology analysis and reconciliation has been performed only with respect to the parsimony model. This discards many plausible solutions and sometimes precludes finding the correct one. In many other areas in bioinformatics probabilistic models have proven to be both more realistic and powerful than parsimony models. For instance, they allow for assessing solution reliability and consideration of alternative solutions in a uniform way. There is also an added benefit in making model assumptions explicit and therefore making model comparisons possible. For orthology analysis, uncertainty has recently been addressed using parsimonious reconciliation combined with bootstrap techniques. However, until now no probabilistic methods have been available. RESULTS: We introduce a probabilistic gene evolution model based on a birth-death process in which a gene tree evolves 'inside' a species tree. Based on this model, we develop a tool with the capacity to perform practical orthology analysis, based on Fitch's original definition, and more generally for reconciling pairs of gene and species trees. Our gene evolution model is biologically sound (Nei et al., 1997) and intuitively attractive. We develop a Bayesian analysis based on MCMC which facilitates approximation of an a posteriori distribution for reconciliations. That is, we can find the most probable reconciliations and estimate the probability of any reconciliation, given the observed gene tree. This also gives a way to estimate the probability that a pair of genes are orthologs. The main algorithmic contribution presented here consists of an algorithm for computing the likelihood of a given reconciliation. To the best of our knowledge, this is the first successful introduction of this type of probabilistic methods, which flourish in phylogeny analysis, into reconciliation and orthology analysis. The MCMC algorithm has been implemented and, although not yet being in its final form, tests show that it performs very well on synthetic as well as biological data. Using standard correspondences, our results carry over to allele trees as well as biogeography.

Algorithms↗

Comparison of site-specific rate-inference methods for protein sequences: empirical Bayesian methods are superior.

The degree to which an amino acid site is free to vary is strongly dependent on its structural and functional importance. An amino acid that plays an essential role is unlikely to change over evolutionary time. Hence, the evolutionary rate at an amino acid site is indicative of how conserved this site is and, in turn, allows evaluation of its importance in maintaining the structure/function of the protein. When using probabilistic methods for site-specific rate inference, few alternatives are possible. In this study we use simulations to compare the maximum-likelihood and Bayesian paradigms. We study the dependence of inference accuracy on such parameters as number of sequences, branch lengths, the shape of the rate distribution, and sequence length. We also study the possibility of simultaneously estimating branch lengths and site-specific rates. Our results show that a Bayesian approach is superior to maximum-likelihood under a wide range of conditions, indicating that the prior that is incorporated into the Bayesian computation significantly improves performance. We show that when branch lengths are unknown, it is better first to estimate branch lengths and then to estimate site-specific rates. This procedure was found to be superior to estimating both the branch lengths and site-specific rates simultaneously. Finally, we illustrate the difference between maximum-likelihood and Bayesian methods when analyzing site-conservation for the apoptosis regulator protein Bcl-x(L).

Animals↗

Bayesian dynamic modeling of latent trait distributions.

Studies of latent traits often collect data for multiple items measuring different aspects of the trait. For such data, it is common to consider models in which the different items are manifestations of a normal latent variable, which depends on covariates through a linear regression model. This article proposes a flexible Bayesian alternative in which the unknown latent variable density can change dynamically in location and shape across levels of a predictor. Scale mixtures of underlying normals are used in order to model flexibly the measurement errors and allow mixed categorical and continuous scales. A dynamic mixture of Dirichlet processes is used to characterize the latent response distributions. Posterior computation proceeds via a Markov chain Monte Carlo algorithm, with predictive densities used as a basis for inferences and evaluation of model fit. The methods are illustrated using data from a study of DNA damage in response to oxidative stress.

Algorithms↗

Bayesian statistical theory in the preoperative diagnosis of pulmonary lesions.

We used a computerized Bayesian algorithm to assist in the preoperative diagnosis of pulmonary lesions. One hundred consecutive patients who were undergoing exploratory thoracotomy for newly discovered pulmonary lesions were prospectively evaluated. The Bayesian model used a total of 44 preoperative clinical and roentgenographic factors to categorize the lesions as benign or malignant. The Bayesian algorithm correctly categorized 96 of the 100 lesions, thereby providing an accuracy of 96 percent. The sensitivity of the model was 98 percent and the specificity was 87 percent. All but two of the 85 malignant lesions were correctly categorized and 13 of the 15 benign lesions were correctly analyzed by the model. These results indicate that computer-assisted diagnosis using the Theorem of Bayes may provide valuable preoperative information for the management of selected patients.

Adolescent↗

Population pharmacokinetics. Theory and clinical application.

Good therapeutic practice should always be based on an understanding of pharmacokinetic variability. This ensures that dosage adjustments can be made to accommodate differences in pharmacokinetics due to genetic, environmental, physiological or pathological factors. The identification of the circumstances in which these factors play a significant role depends on the conduct of pharmacokinetic studies throughout all stages of drug development. Advances in pharmacokinetic data analysis in the last 10 years have opened up a more comprehensive approach to this subject: early traditional small group studies may now be complemented by later population-based studies. This change in emphasis has been largely brought about by the development of appropriate computer software (NONMEM: Nonlinear Mixed Effects Model) and its successful application to the retrospective analysis of clinical data of a number of commonly used drugs, e.g. digoxin, phenytoin, gentamicin, procainamide, mexiletine and lignocaine (lidocaine). Success has been measured in terms of the provision of information which leads to increased efficiency in dosage adjustment, usually based on a subsequent Bayesian feedback procedure. The application of NONMEM to new drugs, however, raises a number of interesting questions, e.g. 'what experimental design strategies should be employed?' and 'can kinetic parameter distributions other than those which are unimodal and normal be identified?' An answer to the later question may be provided by an alternative non-parametric maximum likelihood (NPML) approach. Population kinetic studies generate a considerable amount of demographic and concentration-time data; the effort involved may be wasted unless sufficient attention is paid to the organisation and storage of such information. This is greatly facilitated by the creation of specially designed clinical pharmacokinetic data bases, conveniently stored on microcomputers. A move towards the adoption of population pharmacokinetics as a routine procedure during drug development should now be encouraged. A number of studies have shown that it is possible to organise existing, routine data in such a way that valuable information on pharmacokinetic variability can be obtained. It should be relatively easy to organise similar studies prospectively during drug development and, where appropriate, proceed to the establishment of control systems based on Bayesian feedback.

Demography↗

Estimation of distribution algorithms with Kikuchi approximations.

The question of finding feasible ways for estimating probability distributions is one of the main challenges for Estimation of Distribution Algorithms (EDAs). To estimate the distribution of the selected solutions, EDAs use factorizations constructed according to graphical models. The class of factorizations that can be obtained from these probability models is highly constrained. Expanding the class of factorizations that could be employed for probability approximation is a necessary step for the conception of more robust EDAs. In this paper we introduce a method for learning a more general class of probability factorizations. The method combines a reformulation of a probability approximation procedure known in statistical physics as the Kikuchi approximation of energy, with a novel approach for finding graph decompositions. We present the Markov Network Estimation of Distribution Algorithm (MN-EDA), an EDA that uses Kikuchi approximations to estimate the distribution, and Gibbs Sampling (GS) to generate new points. A systematic empirical evaluation of MN-EDA is done in comparison with different Bayesian network based EDAs. From our experiments we conclude that the algorithm can outperform other EDAs that use traditional methods of probability approximation in the optimization of functions with strong interactions among their variables.

Algorithms↗

Bayesian modeling for linking causally related observations in chest X-ray reports.

Our natural language understanding system outputs a list of diseases, findings, and appliances found in a chest x-ray report. The system described in this paper links those diseases and findings that are causally related. Using Bayesian networks to model the conceptual and diagnostic information found in a chest x-ray we are able to infer more specific information about the findings that are linked to diseases.

Algorithms↗

A hierarchical aggregate data model with spatially correlated disease rates.

The aggregate data study design (Prentice and Sheppard, 1995, Biometrika 82, 113-125) estimates individual-level exposure effects by regressing population-based disease rates on covariate data from survey samples in each population group. In this work, we further develop the aggregate data model to allow for residual spatial correlation among disease rates across populations. Geographical variation that is not explained by model predictors and has a spatial component often arises in studies of rare chronic diseases, such as breast cancer. We combine the aggregate and Bayesian disease-mapping models to provide an intuitive approach to the modeling of spatial effects while drawing correct inference regarding the exposure effect. Based on the results of simulation studies, we suggest guidelines for use of the proposed model.

Bayes Theorem↗

Space complexity of estimation of distribution algorithms.

In this paper, we investigate the space complexity of the Estimation of Distribution Algorithms (EDAs), a class of sampling-based variants of the genetic algorithm. By analyzing the nature of EDAs, we identify criteria that characterize the space complexity of two typical implementation schemes of EDAs, the factorized distribution algorithm and Bayesian network-based algorithms. Using random additive functions as the prototype, we prove that the space complexity of the factorized distribution algorithm and Bayesian network-based algorithms is exponential in the problem size even if the optimization problem has a very sparse interaction structure.

Algorithms↗