Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

Detection of quantitative trait loci in outbred populations with incomplete marker data.

Augmentation of marker genotypes for ungenotyped individuals is implemented in a Bayesian approach via the use of Markov chain Monte Carlo techniques. Marker data on relatives and phenotypes are combined to compute conditional posterior probabilities for marker genotypes of ungenotyped individuals. The presented procedure allows the analysis of complex pedigrees with ungenotyped individuals to detect segregating quantitative trait loci (QTL). Allelic effects at the QTL were assumed to follow a normal distribution with a covariance matrix based on known QTL position and identity by descent probabilities derived from flanking markers. The Bayesian approach estimates variance due to the single QTL, together with polygenic and residual variance. The method was empirically tested through analyzing simulated data from a complex granddaughter design. Ungenotyped dams were related to one or more sons or grandsires in the design. Heterozygosity of the marker loci and size of QTL were varied. Simulation results indicated a significant increase in power when ungenotyped dams were included in the analysis.

Animals↗

Local computation of angular velocity in rotational visual motion.

Retinal images evolve continuously over time owing to self-motions and to movements in the world. Such an evolving image, also known as optic flow, if arising from natural scenes can be locally decomposed in a Bayesian manner into several elementary components, including translation, expansion, and rotation. To take advantage of this decomposition, the brain has neurons tuned to these types of motions. However, these neurons typically have large receptive fields, often spanning tens of degrees of visual angle. Can neurons such as these compute elementary optic-flow components sufficiently locally to achieve a reasonable decomposition? We show that human discrimination of angular velocity is local. Local discrimination of angular velocity requires an accurate estimation of the center of rotation within the optic-flow field. Inaccuracies in estimating the center of rotation result in a predictable systematic error when one is estimating local angular velocity. Our results show that humans make the predicted errors. We discuss how the brain might estimate the elementary components of the optic flow locally by using large receptive fields.

Brain↗

Practical Bayesian inference using mixtures of mixtures.

Discrete mixtures of normal distributions are widely used in modeling amplitude fluctuations of electrical potentials at synapses of human and other animal nervous systems. The usual framework has independent data values yj arising as yj = mu j + xn0 + j, where the means mu j come from some discrete prior G(mu) and the unknown xno + j's and observed xj, j = 1,...,n0, are Gaussian noise terms. A practically important development of the associated statistical methods is the issue of nonnormality of the noise terms, often the norm rather than the exception in the neurological context. We have recently developed models, based on convolutions of Dirichlet process mixtures, for such problems. Explicitly, we model the noise data values xj as arising from a Dirichlet process mixture of normals, in addition to modeling the location prior G(mu) as a Dirichlet process itself. This induces a Dirichlet mixture of mixtures of normals, whose analysis may be developed using Gibbs sampling techniques. We discuss these models and their analysis, and illustrate them in the context of neurological response analysis.

Animals↗

Development of quantitative structure-activity relationships for the toxicity of aromatic compounds to Tetrahymena pyriformis: comparative assessment of the methodologies.

The purpose of this study was to develop quantitative structure-activity relationships (QSARs) for the toxicity of 268 aromatic compounds in the Tetrahymena pyriformis growth inhibition assay. The QSARs were developed using the response-surface (or two-parameter) approach, which was also modified using linear free-energy parameters to account for outliers. Subsequently, the data set was analyzed using partial least-squares (PLS). The results of the modeling using different methodologies were compared to the use of a Bayesian regularized neural network (BRANN) trained on the same data. Both response surface approaches, and PLS explained between 75 and 80% of the variance in the data; BRANN gave a higher statistical fit. In terms of the transparency of the approaches, the response surface clearly provides the simplest and easiest to use QSAR, it is readily interpreted in terms of mechanism of toxic action. PLS and BRANN are respectively less transparent. The use of atomistic and fragment-based indexes as descriptors in QSARs is assessed also, these are found not to be as useful as whole molecule parameters for the prediction of toxicity for molecules outside of the training set. The relative merits of the different approaches to the development of QSARs are described.

Animals↗

Shrinkage-based similarity metric for cluster analysis of microarray data.

The current standard correlation coefficient used in the analysis of microarray data was introduced by M. B. Eisen, P. T. Spellman, P. O. Brown, and D. Botstein [(1998) Proc. Natl. Acad. Sci. USA 95, 14863-14868]. Its formulation is rather arbitrary. We give a mathematically rigorous correlation coefficient of two data vectors based on James-Stein shrinkage estimators. We use the assumptions described by Eisen et al., also using the fact that the data can be treated as transformed into normal distributions. While Eisen et al. use zero as an estimator for the expression vector mean mu, we start with the assumption that for each gene, mu is itself a zero-mean normal random variable [with a priori distribution N(0,tau 2)], and use Bayesian analysis to obtain a posteriori distribution of mu in terms of the data. The shrunk estimator for mu differs from the mean of the data vectors and ultimately leads to a statistically robust estimator for correlation coefficients. To evaluate the effectiveness of shrinkage, we conducted in silico experiments and also compared similarity metrics on a biological example by using the data set from Eisen et al. For the latter, we classified genes involved in the regulation of yeast cell-cycle functions by computing clusters based on various definitions of correlation coefficients and contrasting them against clusters based on the activators known in the literature. The estimated false positives and false negatives from this study indicate that using the shrinkage metric improves the accuracy of the analysis.

Algorithms↗

An evaluation of Bayesian microcomputer predictions of theophylline concentrations in newborn infants.

Determination of appropriate theophylline maintenance doses in preterm infants is confounded by interpatient variability. This study evaluated the performance of an IBM PC computer program applying Bayesian regression before and during steady state in 37 preterm infants. Prior population estimates of clearance and distribution volume in preterm infants and Bayesian estimates of clearance and distribution volume based on one to three theophylline plasma concentrations were used to predict subsequent concentrations (drawn 1-17 days later). We assessed the accuracy and precision of the predictive performance of the Bayesian program with the mean prediction error and the mean absolute prediction error. The absolute prediction error (mean absolute error +/- SEM) significantly decreased with increasing feedback concentrations from 3.54 +/- 0.45 micrograms/ml (population estimates) to 2.74 +/- 0.42 (one feedback) and 2.02 +/- 0.35 micrograms/ml (two feedback concentrations). Mean prediction errors (+/- SEM) based on one to three feedbacks (-1.5 +/- 0.40 micrograms/ml) were significant improvements over population predictions (-2.63 +/- 0.72 micrograms/ml, p less than 0.05), although a small but significant average overprediction remained. Absolute prediction error was correlated with postconceptional and postnatal age when zero or one but not two feedback concentrations were available. Computer program predictions based on one measured feedback concentration were more accurate and precise than population-based predictions. Refinement of population parameters or two feedback concentrations further improved performance.

Bayes Theorem↗

Hunting drug targets by systems-level modeling of gene expression profiles.

Structural learning of Bayesian networks applied to sets of genome-wide expression patterns has been recently discovered as a potentially useful tool for the systems-level statistical description of gene interactions. We train and analyze Bayesian networks with the goal of inferring biological aspects of gene function. Our two-component approach focuses on supporting the drug discovery process by identifying genes with central roles for the network operation, which could act as drug targets. The first component, referred to as scale-free analysis, uses topological measures of the network-related to a high-traffic load of genes-as estimators for their functional importance. The second component, referred to as generative inverse modeling, is a method of estimating the effect of a simulated drug treatment or mutation on the global state of the network, as measured in the expression profile. We show for a dataset from acute lymphoblastic leukemia patients that both approaches are suitable for finding genes with central cellular functions. In addition, generative inverse modeling correctly identifies a known oncogene in a purely data-driven way.

Algorithms↗

QSAR models for predicting the activity of non-peptide luteinizing hormone-releasing hormone (LHRH) antagonists derived from erythromycin A using quantum chemical properties.

Multiple linear regression (MLR) combined with genetic algorithm (GA) and Bayesian-regularized Genetic Neural Networks (BRGNNs) were used to model the binding affinity (pK(I)) of 38 11,12-cyclic carbamate derivatives of 6-O-methylerythromycin A for the Human Luteinizing Hormone-Releasing Hormone (LHRH) receptor using quantum chemical descriptors. A multiparametric MLR equation with good statistical quality was obtained that describes the features relevant for antagonistic activity when the substituent at the position 3 of the erythronolide core was varied. In addition, four-descriptor linear and nonlinear models were established for the whole dataset. Such models showed high statistical quality. However, the BRGNN model was better than the linear model according to the external validation process. In general, our linear and nonlinear models reveal that the binding affinity of the compounds studied for the LHRH receptor is modulated by electron-related terms.

Animals↗

Phylogeny and evolution of the major intrinsic protein family.

BACKGROUND INFORMATION: MIPs (major intrinsic proteins) form channels across biological membranes that control recruitment of water and small solutes such as glycerol and urea in all living organisms. Because of their widespread occurrence and large number, MIPs are a sound model system to understand evolutionary mechanisms underlying the generation of protein structural and functional diversity. With the recent increase in genomic projects, there is a considerable increase in the quantity and taxonomic range of MIPs in molecular databases. RESULTS: In the present study, I compiled more than 450 non-redundant amino acid sequences of MIPs from NCBI databases. Phylogenetic analyses using Bayesian inference reconstructed a statistically robust tree that allowed the classification of members of the family into two main evolutionary groups, the GLPs (glycerol-uptake facilitators or aquaglyceroporins) and the water transport channels or AQPs (aquaporins). Separate phylogenetic analyses of each of the MIP subfamilies were performed to determine the main groups of orthology. In addition, comparative sequence analyses were conducted to identify conserved signatures in the MIP molecule. CONCLUSIONS: The earliest and major gene duplication event in the history of the MIP family led to its main functional split into GLPs and AQPs. GLPs show typically one single copy in microbes (eubacteria, archaea and fungi), up to four paralogues in vertebrates and they are absent from plants. AQPs are usually single in microbes and show their greatest numbers and diversity in angiosperms and vertebrates. Functional recruitment of NOD26-like intrinsic proteins to glycerol transport due to the absence of GLPs in plants was highly supported. Acquisition of other MIP functions such as permeability to ammonia, arsenite or CO2 is restricted to particular MIP paralogues. Up to eight fairly conserved boxes were inferred in the primary sequence of the MIP molecule. All of them mapped on to one side of the channel except the conserved glycine residues from helices 2 and 5 that were found in the opposite side.

Amino Acid Sequence↗

A Bayesian design and analysis for dose-response using informative prior information.

We wish to use prior information on an existing drug in the design and analysis of a dose-response study for a new drug candidate within the same pharmacological class. Using the Bayesian methodology, this prior information can be used quantitatively and the randomization can be weighted in favor of the new compound, where there is less information. An Emax model is used to describe the dose-response of the existing drug. The estimates from this model are used to provide informative prior information used for the design and analysis of the new study to establish the relative potency between the new compound and the existing drug therapy. The assumption is made that the data from previous trials and the new study are exchangeable. The impact of departures from this assumption can be quantified through simulations and by assessing the operating characteristics of various scenarios. Simulations show that relatively modest sample sizes can yield informative results about the magnitude of the relative potency using this approach. The operating characteristics are good when assessing model estimates against clinically important changes in relative potency.

Algorithms↗

Bayesian procedures for the estimation of mutation rates from fluctuation experiments.

Bayesian procedures are developed for estimating mutation rates from fluctuation experiments. Three Bayesian point estimators are compared with four traditional ones using the results of 10,000 simulated experiments. The Bayesian estimators were found to be at least as efficient as the best of the previously known estimators. The best Bayesian estimator is one that uses (1/m2) as the prior probability density function and a quadratic loss function. The advantage of using these estimators is most pronounced when the number of fluctuation test tubes is small. Bayesian estimation allows the incorporation of prior knowledge about the estimated parameter, in which case the resulting estimators are the most efficient. It enables the straightforward construction of confidence intervals for the estimated parameter. The increase of efficiency with prior information and the narrowing of the confidence intervals with additional experimental results are investigated. The results of the simulations show that any potential inaccuracy of estimation arising from lumping together all cultures with more than n mutants (the jackpots) almost disappears at n = 70 (provided that the number of mutations in a culture is low). These methods are applied to a set of experimental data to illustrate their use.

Bacteria↗

Statistical evaluation of local alignment features predicting allergenicity using supervised classification algorithms.

BACKGROUND: Recently, two promising alignment-based features predicting food allergenicity using the k nearest neighbor (kNN) classifier were reported. These features are the alignment score and alignment length of the best local alignment obtained in a database of known allergen sequences. METHODS: In the work reported here a much more comprehensive statistical evaluation of the potential of these features was performed, this time for the prediction of allergenicity in general. The evaluation consisted of the following four key components. (1) A new high quality database consisting of 318 carefully selected, non-redundant allergens and 1,007 sequences carefully selected to be non-allergens. (2) Three different supervised algorithms: the kNN classifier, the Bayesian linear Gaussian classifier, and the Bayesian quadratic Gaussian classifier. (3) A large set of local alignment procedures defined using the FASTA3 alignment program by means of a wide range of different parameter settings. (4) Novel performance curves, alternative to conventional receiver-operating characteristic curves, to display not only average behaviors but also statistical variations due to small data sets. RESULTS: The linear Gaussian classifier proved most useful among the tested supervised machine learning algorithms, closely followed by the quadratic Gaussian equivalent and kNN. The overall best classification results were obtained with a novel feature vector consisting of the combined alignment scores derived from local alignment procedures using different substitution matrices. CONCLUSIONS: The models reported here should be useful as a part of an integrated assessment scheme for potential protein allergenicity and for future comparisons with alternative bioinformatic approaches.

Algorithms↗

A Bayesian neural network approach for modelling censored data with an application to prognosis after surgery for breast cancer.

A Bayesian framework is introduced to carry out Automatic Relevance Determination (ARD) in feedforward neural networks to model censored data. A procedure to identify and interpret the prognostic group allocation is also described. These methodologies are applied to 1616 records routinely collected at Christie Hospital, in a monthly cohort study with 5-year follow-up. Two cohort studies are presented, for low- and high-risk patients allocated by standard clinical staging. The results of contrasting the Partial Logistic Artificial Neural Network (PLANN)-ARD model with the proportional hazards model are that the two are consistent, but the neural network may be more specific in the allocation of patients into prognostic groups. With automatic model selection, the regularised neural network is more conservative than the default stepwise forward selection procedure implemented by SPSS with the Akaike Information Criterion.

Bayes Theorem↗

Population pharmacokinetic analysis of didanosine (2',3'-dideoxyinosine) plasma concentrations obtained in phase I clinical trials in patients with AIDS or AIDS-related complex.

Plasma didanosine concentration data from 36 patients receiving once-a-day therapy and from 33 patients receiving twice-a-day therapy were subject to population pharmacokinetic analysis with the computer program NONMEM. Once- or twice-a-day regimens of didanosine were administered intravenously (i.v.) (dose: 0.8-33 mg/kg) during the first 2 weeks of therapy, and orally (dose: 1.6-66 mg/kg) for the remaining 4 weeks of therapy. Plasma pharmacokinetics were determined after the first and last (steady-state) i.v. and oral doses. Population pharmacokinetic parameters for the combined i.v. and oral steady-state data were (mean [%CV]): systemic clearance, CL, 0.70 (5.2) L/h/kg; central compartment volume, Vc, 0.18 (32) L/kg; steady-state distribution volume, Vdss, 0.84 (6.8) L/kg; first-order absorption rate constant, Ka, 1.3 (9.5) hr-1; and bioavailable fraction, F, 0.34 (8.5). Interindividual variability (omega) was (%CV) 22.3 and 71.0 for CL and Vc, respectively. Intraindividual (residual) variability (sigma) in plasma concentrations (%CV) was 50.2. Body weight, sex, and age did not account for the variability in either CL or Vc, and the use of alternate pharmacokinetic models did not reduce the value of intraindividual variability. Population parameters for the combined i.v. and oral first-dose data were generally similar to those for the steady-state data. The parameters can be used to design dosing regimens in patients using the Bayesian feedback approach.

AIDS-Related Complex↗

Bayesian hierarchical error model for analysis of gene expression data.

MOTIVATION: Analysis of genome-wide microarray data requires the estimation of a large number of genetic parameters for individual genes and their interaction expression patterns under multiple biological conditions. The sources of microarray error variability comprises various biological and experimental factors, such as biological and individual replication, sample preparation, hybridization and image processing. Moreover, the same gene often shows quite heterogeneous error variability under different biological and experimental conditions, which must be estimated separately for evaluating the statistical significance of differential expression patterns. Widely used linear modeling approaches are limited because they do not allow simultaneous modeling and inference on the large number of these genetic parameters and heterogeneous error components on different genes, different biological and experimental conditions, and varying intensity ranges in microarray data. RESULTS: We propose a Bayesian hierarchical error model (HEM) to overcome the above restrictions. HEM accounts for heterogeneous error variability in an oligonucleotide microarray experiment. The error variability is decomposed into two components (experimental and biological errors) when both biological and experimental replicates are available. Our HEM inference is based on Markov chain Monte Carlo to estimate a large number of parameters from a single-likelihood function for all genes. An F-like summary statistic is proposed to identify differentially expressed genes under multiple conditions based on the HEM estimation. The performance of HEM and its F-like statistic was examined with simulated data and two published microarray datasets-primate brain data and mouse B-cell development data. HEM was also compared with ANOVA using simulated data. AVAILABILITY: The software for the HEM is available from the authors upon request.

Algorithms↗

DIAMED: a probabilistic diagnostic aid system on the web.

DIAMED is a system to assist the physicians in the diagnostic process using probabilistic networks as knowledge representation. These networks make it possible to reason on medical data by applying Bayesian methods and to take into account uncertainties of the facts in the resolution of the clinical cases. The proposed model re-uses knowledge contained in an existing knowledge base (ADM). An interface of DIAMED developed on a Web server remotely assists the experts of each medical specialty in updating and validating the knowledge base. Most of the data processing is automated while being based on information preexistent in the ADM base : Constitution of lexicons starting from the existing dictionaries of the ADM system, are then used to work out the requests for selection and update of the knowledge base. One of its assets resides in its pseudo-segmented structure in several layers. The propagation of information is thus limited to only one part of the probabilistic network and calculations are therefore limited.

Artificial Intelligence↗

A comparison of imputation methods in a longitudinal randomized clinical trial.

It is common for longitudinal clinical trials to face problems of item non-response, unit non-response, and drop-out. In this paper, we compare two alternative methods of handling multivariate incomplete data across a baseline assessment and three follow-up time points in a multi-centre randomized controlled trial of a disease management programme for late-life depression. One approach combines hot-deck (HD) multiple imputation using a predictive mean matching method for item non-response and the approximate Bayesian bootstrap for unit non-response. A second method is based on a multivariate normal (MVN) model using PROC MI in SAS software V8.2. These two methods are contrasted with a last observation carried forward (LOCF) technique and available-case (AC) analysis in a simulation study where replicate analyses are performed on subsets of the originally complete cases. Missing-data patterns were simulated to be consistent with missing-data patterns found in the originally incomplete cases, and observed complete data means were taken to be the targets of estimation. Not surprisingly, the LOCF and AC methods had poor coverage properties for many of the variables evaluated. Multiple imputation under the MVN model performed well for most variables but produced less than nominal coverage for variables with highly skewed distributions. The HD method consistently produced close to nominal coverage, with interval widths that were roughly 7 per cent larger on average than those produced from the MVN model.

Aged↗

Reconstruction of gene networks using Bayesian learning and manipulation experiments.

MOTIVATION: The analysis of high-throughput experimental data, for example from microarray experiments, is currently seen as a promising way of finding regulatory relationships between genes. Bayesian networks have been suggested for learning gene regulatory networks from observational data. Not all causal relationships can be inferred from correlation data alone. Often several equivalent but different directed graphs explain the data equally well. Intervention experiments where genes are manipulated can help to narrow down the range of possible networks. RESULTS: We describe an active learning algorithm that suggests an optimized sequence of intervention experiments. Simulation experiments show that our selection scheme is better than an unguided choice of interventions in learning the correct network and compares favorably in running time and results with methods based on value of information calculations.

Algorithms↗