Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

Relationship estimation in affected sib pair analysis of late-onset diseases.

In linkage studies, errors in pedigree structure will often be uncovered through Mendelian inconsistencies. In affected sib pair analysis of diseases with late onset, however, such mistakes will usually go undetected since parental genotypes are commonly not known. Cases of nonpaternity, unrecorded adoption or accidental sample swap in the laboratory will then not be noticed. Typically, such relationship errors lead to a decrease in power for linkage. In this paper, a method is presented which allows verification of the relationship between stated sibs using their marker genotypes. The method is likelihood-based and incorporates a Bayesian approach to compute posterior relationship probabilities. It is shown that sibs, half-sibs and unrelated individuals can be distinguished from each other quite reliably using numbers of markers that should be available in most sib pair studies. It is demonstrated that elimination of false sib pairs increases the power to detect linkage in affected sib pair studies. The gain in power may be large if relationship errors occur quite frequently; the gain will be only moderate if relationship errors are very infrequent. Software for relationship estimation is provided.

Bayes Theorem↗

DIAS-NIDDM--a model-based decision support system for insulin dose adjustment in insulin-treated subjects with NIDDM.

A decision support system has been developed, Diabetes Insulin Advisory System for patients with non-insulin dependent diabetes mellitus (DIAS-NIDDM), assisting in the adjustment of insulin doses in insulin-treated subjects. DIAS-NIDDM uses a causal probabilistic network (CPN) model of carbohydrate metabolism to make stochastic predictions of blood glucose (BG) excursions. The CPN model is an extension of an existing model with an added component representing endogenous insulin secretion. A linear relationship between BG and insulin concentration due to BG stimulated insulin secretion is assumed. Model parameters (pancreatic sensitivity, insulin sensitivity, and time-to-peak of NPH insulin) are estimated by Bayesian probability updating from patient's specific data (food intake, insulin doses, BG measurements) recorded over a period of 4 days. The estimated parameters allow the system to be potentially used as a diagnostic tool to identify abnormalities of carbohydrate metabolism: impaired insulin secretion, insulin resistance and the severity of the impairments. DIAS-NIDDM was used to predict patient-specific BG profiles and advise on insulin doses during a pilot study in eight patients with NIDDM of whom five were treated with insulin. Compared to the administered insulin amount, daily insulin amount advised by DIAS-NIDDM was similar (within 4 U) in three patients, higher by 20% (19 U) in one patient and lower by 40% (18 U) and 50% (11 U) in two patients, respectively. The inter-day coefficient of variation of the daily insulin advice suggests that, at least according to DIAS-NIDDM criteria, day-to-day adjustment of insulin doses is necessary to maintain optimum control.

Computer Simulation↗

Bayesian approach to discovering pathogenic SNPs in conserved protein domains.

The success rate of association studies can be improved by selecting better genetic markers for genotyping or by providing better leads for identifying pathogenic single nucleotide polymorphisms (SNPs) in the regions of linkage disequilibrium with positive disease associations. We have developed a novel algorithm to predict pathogenic single amino acid changes, either nonsynonymous SNPs (nsSNPs) or missense mutations, in conserved protein domains. Using a Bayesian framework, we found that the probability of a microbial missense mutation causing a significant change in phenotype depended on how much difference it made in several phylogenetic, biochemical, and structural features related to the single amino acid substitution. We tested our model on pathogenic allelic variants (missense mutations or nsSNPs) included in OMIM, and on the other nsSNPs in the same genes (from dbSNP) as the nonpathogenic variants. As a result, our model predicted pathogenic variants with a 10% false-positive rate. The high specificity of our prediction algorithm should make it valuable in genetic association studies aimed at identifying pathogenic SNPs.

Algorithms↗

Putative ancestral origins of chromosomal segments in individual african americans: implications for admixture mapping.

Theoretically, markers that distinguish European from West African ancestry can be used to examine the origin of chromosomal segments in individual African Americans. In this study, putative ancestral origin was examined by using haplotypes estimated from genotyping 268 African Americans for 29 ancestry informative markers spaced over a 60-cM segment of chromosome 5. Analyses using a Bayesian algorithm (STRUCTURE) provided evidence that blocks of individual chromosomes derive from one or the other parental population. In addition, modeling studies were performed by using hidden real marker data to simulate patient and control populations under different genotypic risk ratios. Ancestry analysis showed significant results for a genotypic risk ratio of 2.5 in the African American population for modeled susceptibility genes derived from either putative parental population. These studies suggest that admixture mapping in the African American population can provide a powerful approach to defining genetic factors for some disease phenotypes.

Black or African American↗

Calculating the probability of multitaxon evolutionary trees: bootstrappers Gambit.

The reconstruction of multitaxon trees from molecular sequences is confounded by the variety of algorithms and criteria used to evaluate trees, making it difficult to compare the results of different analyses. A global method of multitaxon phylogenetic reconstruction described here, Bootstrappers Gambit, can be used with any four-taxon algorithm, including distance, maximum likelihood, and parsimony methods. It incorporates a Bayesian-Jeffreys'-bootstrap analysis to provide a uniform probability-based criterion for comparing the results from diverse algorithms. To examine the usefulness of the method, the origin of the eukaryotes has been investigated by the analysis of ribosomal small subunit RNA sequences. Three common algorithms (paralinear distances, Jukes-Cantor distances, and Kimura distances) support the eocyte topology, whereas one (maximum parsimony) supports the archaebacterial topology, suggesting that the eocyte prokaryotes are the closest prokaryotic relatives of the eukaryotes.

Algorithms↗

Bayesian estimation for species identification in single-molecule fluorescence microscopy.

In this article we describe a recursive Bayesian estimator for the identification of diffusing fluorophores using photon arrival-time data from a single spectral channel. We present derivations for all relevant diffusion and fluorescence models, and we use simulated diffusion trajectories and photon streams to evaluate the estimator's performance. We consider simplified estimation schemes that bin the photon counts within time intervals of fixed duration, and show that they can perform well in realistic parameter regimes. The latter results indicate the feasibility of performing identification experiments in real time. It will be straightforward to generalize our approach for use in more complicated scenarios, e.g., with multiple spectral channels or fast photophysical dynamics.

Algorithms↗

Advantages of terminating Zippy Estimation by Sequential Testing (ZEST) with dynamic criteria for white-on-white perimetry.

PURPOSE: A number of automated perimeters use the Zippy Estimation by Sequential Testing (ZEST) algorithm, which is an adaptive Bayesian method, for determining sensitivity measures. There are two popular rules for deciding when to terminate Bayesian procedures: (1) after a fixed number of presentations; or (2) when the probability density function (pdf) over all thresholds modified by the procedure becomes sufficiently narrow (a dynamic termination criterion). It has recently been argued that fixed termination criteria perform equally as well as dynamic criteria when applied in a fashion typical of laboratory-based visual psychophysics. Perimetry, however, has specific requirements; the tests must be very short, there is a wide range of possible sensitivities, and erroneous responses from the patient must be tolerated. This study used computer simulation to compare fixed and dynamic termination criteria for the ZEST algorithm using conditions typical of white-on-white perimetry. METHODS: Eight ZEST procedures were compared using the following termination criteria: fixed termination after 4, 5, 6, 7, and 8 presentations; dynamic termination when the standard deviation of the pdf was 1 dB, 1.5 dB, and 2 dB. Four patient error models were used: ideal, typical false-positive, typical false-negative, and unreliable patients. We also ran a version of ZEST that set the likelihood function exactly equal to the patient's frequency of seeing curve. RESULTS: The mean absolute error and standard deviation of error in threshold measurement was higher for the fixed termination criteria than for dynamic termination criteria of the same average number of presentations. CONCLUSIONS: The results of our simulations indicate that dynamic procedures have some distinct benefits over fixed termination procedures when a minimum of presentations are required and response errors are made as in a white-on-white perimetric setting. Dynamic termination criteria are at least partially successful in expending more presentations when required to enhance test precision.

Algorithms↗

A novel approach for clustering proteomics data using Bayesian fast Fourier transform.

MOTIVATION: Bioinformatics clustering tools are useful at all levels of proteomic data analysis. Proteomics studies can provide a wealth of information and rapidly generate large quantities of data from the analysis of biological specimens. The high dimensionality of data generated from these studies requires the development of improved bioinformatics tools for efficient and accurate data analyses. For proteome profiling of a particular system or organism, a number of specialized software tools are needed. Indeed, significant advances in the informatics and software tools necessary to support the analysis and management of these massive amounts of data are needed. Clustering algorithms based on probabilistic and Bayesian models provide an alternative to heuristic algorithms. The number of clusters (diseased and non-diseased groups) is reduced to the choice of the number of components of a mixture of underlying probability. The Bayesian approach is a tool for including information from the data to the analysis. It offers an estimation of the uncertainties of the data and the parameters involved. RESULTS: We present novel algorithms that can organize, cluster and derive meaningful patterns of expression from large-scaled proteomics experiments. We processed raw data using a graphical-based algorithm by transforming it from a real space data-expression to a complex space data-expression using discrete Fourier transformation; then we used a thresholding approach to denoise and reduce the length of each spectrum. Bayesian clustering was applied to the reconstructed data. In comparison with several other algorithms used in this study including K-means, (Kohonen self-organizing map (SOM), and linear discriminant analysis, the Bayesian-Fourier model-based approach displayed superior performances consistently, in selecting the correct model and the number of clusters, thus providing a novel approach for accurate diagnosis of the disease. Using this approach, we were able to successfully denoise proteomic spectra and reach up to a 99% total reduction of the number of peaks compared to the original data. In addition, the Bayesian-based approach generated a better classification rate in comparison with other classification algorithms. This new finding will allow us to apply the Fourier transformation for the selection of the protein profile for each sample, and to develop a novel bioinformatic strategy based on Bayesian clustering for biomarker discovery and optimal diagnosis.

Algorithms↗

Probabilistic motion estimation based on temporal coherence.

We develop a theory for the temporal integration of visual motion motivated by psychophysical experiments. The theory proposes that input data are temporally grouped and used to predict and estimate the motion flows in the image sequence. This temporal grouping can be considered a generalization of the data association techniques that engineers use to study motion sequences. Our temporal grouping theory is expressed in terms of the Bayesian generalization of standard Kalman filtering. To implement the theory, we derive a parallel network that shares some properties of cortical networks. Computer simulations of this network demonstrate that our theory qualitatively accounts for psychophysical experiments on motion occlusion and motion outliers. In deriving our theory, we assumed spatial factorizability of the probability distributions and made the approximation of updating the marginal distributions of velocity at each point. This allowed us to perform local computations and simplified our implementation. We argue that these approximations are suitable for the stimuli we are considering (for which spatial coherence effects are negligible).

Bayes Theorem↗

Visual extrapolation of contour geometry.

Computing the shapes of object boundaries from fragmentary image contours poses a formidable problem for the visual system. We investigated the extrapolation of contour shape by human vision. Measurements of extrapolation position and orientation were taken at six distances from the point of occlusion, thereby yielding a detailed representation of the extrapolated contours. Analyses of these measurements revealed that: (i) extrapolation curvature increases linearly with the curvature of the inducing contour, although there is individual bias in the slope; (ii) the precision with which an extrapolated contour is represented is roughly constant, in angular terms, with increasing distance from the point of occlusion; (iii) there is a substantial cost of curvature, in that the overall precision of an extrapolated contour decreases systematically with curvature; (iv) the shapes of visually extrapolated contours are characterized by a nonlinear decrease in curvature, asymptoting to zero; and (v) this decaying pattern of curvature is explained by a Bayesian model in which, with increasing distance from the point of occlusion, the prior tendency to minimize curvature gradually dominates the likelihood tendency to minimize variation in curvature.

Bayes Theorem↗

Bayesian neural network approaches to ovarian cancer identification from high-resolution mass spectrometry data.

MOTIVATION: The classification of high-dimensional data is always a challenge to statistical machine learning. We propose a novel method named shallow feature selection that assigns each feature a probability of being selected based on the structure of training data itself. Independent of particular classifiers, the high dimension of biodata can be fleetly reduced to an applicable case for consequential processing. Moreover, to improve both efficiency and performance of classification, these prior probabilities are further used to specify the distributions of top-level hyperparameters in hierarchical models of Bayesian neural network (BNN), as well as the parameters in Gaussian process models. RESULTS: Three BNN approaches were derived and then applied to identify ovarian cancer from NCI's high-resolution mass spectrometry data, which yielded an excellent performance in 1000 independent k-fold cross validations (k = 2,...,10). For instance, indices of average sensitivity and specificity of 98.56 and 98.42%, respectively, were achieved in the 2-fold cross validations. Furthermore, only one control and one cancer were misclassified in the leave-one-out cross validation. Some other popular classifiers were also tested for comparison. AVAILABILITY: The programs implemented in MatLab, R and Neal's fbm.2004-11-10.

Bayes Theorem↗

Bayesian sperm competition estimates.

We introduce a Bayesian method for estimating parameters for a model of multiple mating and sperm displacement from genotype counts of brood-structured data. The model is initially targeted for Drosophila melanogaster, but is easily adapted to other organisms. The method is appropriate for use with field studies where the number of mates and the genotypes of the mates cannot be controlled, but where unlinked markers have been collected for a set of females and a sample of their offspring. Advantages over previous approaches include full use of multilocus information and the ability to cope appropriately with missing data and ambiguities about which alleles are maternally vs. paternally inherited. The advantages of including X-linked markers are also demonstrated.

Animals↗

Proceedings of the SMBE Tri-National Young Investigators' Workshop 2005. Coalescent-based estimation of population parameters when the number of demes changes over time.

We expand a coalescent-based method that uses serially sampled genetic data from a subdivided population to incorporate changes to the number of demes and patterns of colonization. Often, when estimating population parameters or other parameters of interest from genetic data, the demographic structure and parameters are not constant over evolutionary time. In this paper, we develop a Bayesian Markov chain Monte Carlo method that allows for step changes in mutation, migration, and population sizes, as well as changing numbers of demes, where the times of these changes are also estimated. We show that in parameter ranges of interest, reliable estimates can often be obtained, including the historical times of parameter changes. However, posterior densities of migration rates can be quite diffuse and estimators somewhat biased, as reported by other authors.

Bayes Theorem↗

Fluorescence optical diffusion tomography.

A nonlinear, Bayesian optimization scheme is presented for reconstructing fluorescent yield and lifetime, the absorption coefficient, and the diffusion coefficient in turbid media, such as biological tissue. The method utilizes measurements at both the excitation and the emission wavelengths to reconstruct all unknown parameters. The effectiveness of the reconstruction algorithm is demonstrated by simulation and by application to experimental data from a tissue phantom containing the fluorescent agent Indocyanine Green.

Bayes Theorem↗

Flexible Bayesian methods for cancer phase I clinical trials. Dose escalation with overdose control.

We examine a large class of prior distributions to model the dose-response relationship in cancer phase I clinical trials. We parameterize the dose-toxicity model in terms of the maximum tolerated dose (MTD) gamma and the probability of dose limiting toxicity (DLT) at the initial dose rho(0). The MTD is estimated using the EWOC (escalation with overdose control) method of Babb et al. We show through simulations that a candidate joint prior for (rho0,gamma) with negative a priori correlation structure results in a safer trial than the one that assumes independent priors for these two parameters while keeping the efficiency of the estimate of the MTD essentially unchanged.

Antimetabolites, Antineoplastic↗

Comparison of recent methods for inference of variable influence in neural networks.

Neural networks (NNs) belong to 'black box' models and therefore 'suffer' from interpretation difficulties. Four recent methods inferring variable influence in NNs are compared in this paper. The methods assist the interpretation task during different phases of the modeling procedure. They belong to information theory (ITSS), the Bayesian framework (ARD), the analysis of the network's weights (GIM), and the sequential omission of the variables (SZW). The comparison is based upon artificial and real data sets of differing size, complexity and noise level. The influence of the neural network's size has also been considered. The results provide useful information about the agreement between the methods under different conditions. Generally, SZW and GIM differ from ARD regarding the variable influence, although applied to NNs with similar modeling accuracy, even when larger data sets sizes are used. ITSS produces similar results to SZW and GIM, although suffering more from the 'curse of dimensionality'.

Algorithms↗

Estimating gene regulatory networks and protein-protein interactions of Saccharomyces cerevisiae from multiple genome-wide data.

MOTIVATION: Biological processes in cells are properly performed by gene regulations, signal transductions and interactions between proteins. To understand such molecular networks, we propose a statistical method to estimate gene regulatory networks and protein-protein interaction networks simultaneously from DNA microarray data, protein-protein interaction data and other genome-wide data. RESULTS: We unify Bayesian networks and Markov networks for estimating gene regulatory networks and protein-protein interaction networks according to the reliability of each biological information source. Through the simultaneous construction of gene regulatory networks and protein-protein interaction networks of Saccharomyces cerevisiae cell cycle, we predict the role of several genes whose functions are currently unknown. By using our probabilistic model, we can detect false positives of high-throughput data, such as yeast two-hybrid data. In a genome-wide experiment, we find possible gene regulatory relationships and protein-protein interactions between large protein complexes that underlie complex regulatory mechanisms of biological processes.

Algorithms↗

A brief primer on automated signal detection.

BACKGROUND: Statistical techniques have traditionally been underused in spontaneous reporting systems used for postmarketing surveillance of adverse drug events. Regulatory agencies, pharmaceutical companies, and drug monitoring centers have recently devoted considerable efforts to develop and implement computer-assisted automated signal detection methodologies that employ statistical theory to enhance screening efforts of expert clinical reviewers. OBJECTIVE: To provide a concise state-of-the-art review of the most commonly used automated signal detection procedures, including the underlying statistical concepts, performance characteristics, and outstanding limitations, and issues to be resolved. DATA SOURCES: Primary articles were identified by MEDLINE search (1965-December 2002) and through secondary sources. STUDY SELECTION AND DATA EXTRACTION: All of the articles identified from the data sources were evaluated and all information deemed relevant was included in this review. DATA SYNTHESIS: Commonly used methods of automated signal detection are self-contained and involve screening large databases of spontaneous adverse event reports in search of interestingly large disproportionalities or dependencies between significant variables, usually single drug-event pairs, based on an underlying model of statistical independence. The models vary according to the underlying model of statistical independence and whether additional mathematical modeling using Bayesian analysis is applied to the crude measures of disproportionality. There are many potential advantages and disadvantages of these methods, as well as significant unresolved issues related to the application of these techniques, including lack of comprehensive head-to-head comparisons in a single large transnational database, lack of prospective evaluations, and the lack of gold standard of signal detection. CONCLUSIONS: Current methods of automated signal detection are nonclinical and only highlight deviations from independence without explaining whether these deviations are due to a causal linkage or numerous potential confounders. They therefore cannot replace expert clinical reviewers, but can help them to focus attention when confronted with the difficult task of screening huge numbers of drug-event combinations for potential signals. Important questions remain to be answered about the performance characteristics of these methods. Pharmacovigilance professionals should take the time to learn the underlying mathematical concepts in order to critically evaluate accumulating experience pertaining to the relative performance characteristics of these methods that are incompletely defined.

Automation↗