Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,615 records · Page 90Linked to original sources

A general method for the unbiased improvement of solution NMR structures by the use of related X-ray data, the AUREMOL-ISIC algorithm.

BACKGROUND: Rapid and accurate three-dimensional structure determination of biological macromolecules is mandatory to keep up with the vast progress made in the identification of primary sequence information. During the last few years the amount of data deposited in the protein data bank has substantially increased providing additional information for novel structure determination projects. The key question is how to combine the available database information with the experimental data of the current project ensuring that only relevant information is used and a correct structural bias is produced. For this purpose a novel fully automated algorithm based on Bayesian reasoning has been developed. It allows the combination of structural information from different sources in a consistent way to obtain high quality structures with a limited set of experimental data. The new ISIC (Intelligent Structural Information Combination) algorithm is part of the larger AUREMOL software package. RESULTS: Our new approach was successfully tested on the improvement of the solution NMR structures of the Ras-binding domain of Byr2 from Schizosaccharomyces pombe, the Ras-binding domain of RalGDS from human calculated from a limited set of NMR data, and the immunoglobulin binding domain from protein G from Streptococcus by their corresponding X-ray structures. In all test cases clearly improved structures were obtained. The largest danger in using data from other sources is a possible bias towards the added structure. In the worst case instead of a refined target structure the structure from the additional source is essentially reproduced. We could clearly show that the ISIC algorithm treats these difficulties properly. CONCLUSION: In summary, we present a novel fully automated method to combine strongly coupled knowledge from different sources. The combination with validation tools such as the calculation of NMR R-factors strengthens the impact of the method considerably since the improvement of the structures can be assessed quantitatively. The ISIC method can be applied to a large number of similar problems where the quality of the obtained three-dimensional structures is limited by the available experimental data like the improvement of large NMR structures calculated from sparse experimental data or the refinement of low resolution X-ray structures. Also structures may be refined using other available structural information such as homology models.

Algorithms↗

Effect of unsampled populations on the estimation of population sizes and migration rates between sampled populations.

Current estimators of gene flow come in two methods; those that estimate parameters assuming that the populations investigated are a small random sample of a large number of populations and those that assume that all populations were sampled. Maximum likelihood or Bayesian approaches that estimate the migration rates and population sizes directly using coalescent theory can easily accommodate datasets that contain a population that has no data, a so-called 'ghost' population. This manipulation allows us to explore the effects of missing populations on the estimation of population sizes and migration rates between two specific populations. The biases of the inferred population parameters depend on the magnitude of the migration rate from the unknown populations. The effects on the population sizes are larger than the effects on the migration rates. The more immigrants from the unknown populations that are arriving in the sample populations the larger the estimated population sizes. Taking into account a ghost population improves or at least does not harm the estimation of population sizes. Estimates of the scaled migration rate M (migration rate per generation divided by the mutation rate per generation) are fairly robust as long as migration rates from the unknown populations are not huge. The inclusion of a ghost population does not improve the estimation of the migration rate M; when the migration rates are estimated as the number of immigrants Nm then a ghost population improves the estimates because of its effect on population size estimation. It seems that for 'real world' analyses one should carefully choose which populations to sample, but there is no need to sample every population in the neighbourhood of a population of interest.

Bayes Theorem↗

PhyloGibbs: a Gibbs sampling motif finder that incorporates phylogeny.

A central problem in the bioinformatics of gene regulation is to find the binding sites for regulatory proteins. One of the most promising approaches toward identifying these short and fuzzy sequence patterns is the comparative analysis of orthologous intergenic regions of related species. This analysis is complicated by various factors. First, one needs to take the phylogenetic relationship between the species into account in order to distinguish conservation that is due to the occurrence of functional sites from spurious conservation that is due to evolutionary proximity. Second, one has to deal with the complexities of multiple alignments of orthologous intergenic regions, and one has to consider the possibility that functional sites may occur outside of conserved segments. Here we present a new motif sampling algorithm, PhyloGibbs, that runs on arbitrary collections of multiple local sequence alignments of orthologous sequences. The algorithm searches over all ways in which an arbitrary number of binding sites for an arbitrary number of transcription factors (TFs) can be assigned to the multiple sequence alignments. These binding site configurations are scored by a Bayesian probabilistic model that treats aligned sequences by a model for the evolution of binding sites and "background" intergenic DNA. This model takes the phylogenetic relationship between the species in the alignment explicitly into account. The algorithm uses simulated annealing and Monte Carlo Markov-chain sampling to rigorously assign posterior probabilities to all the binding sites that it reports. In tests on synthetic data and real data from five Saccharomyces species our algorithm performs significantly better than four other motif-finding algorithms, including algorithms that also take phylogeny into account. Our results also show that, in contrast to the other algorithms, PhyloGibbs can make realistic estimates of the reliability of its predictions. Our tests suggest that, running on the five-species multiple alignment of a single gene's upstream region, PhyloGibbs on average recovers over 50% of all binding sites in S. cerevisiae at a specificity of about 50%, and 33% of all binding sites at a specificity of about 85%. We also tested PhyloGibbs on collections of multiple alignments of intergenic regions that were recently annotated, based on ChIP-on-chip data, to contain binding sites for the same TF. We compared PhyloGibbs's results with the previous analysis of these data using six other motif-finding algorithms. For 16 of 21 TFs for which all other motif-finding methods failed to find a significant motif, PhyloGibbs did recover a motif that matches the literature consensus. In 11 cases where there was disagreement in the results we compiled lists of known target genes from the literature, and found that running PhyloGibbs on their regulatory regions yielded a binding motif matching the literature consensus in all but one of the cases. Interestingly, these literature gene lists had little overlap with the targets annotated based on the ChIP-on-chip data. The PhyloGibbs code can be downloaded from http://www.biozentrum.unibas.ch/~nimwegen/cgi-bin/phylogibbs.cgi or http://www.imsc.res.in/~rsidd/phylogibbs. The full set of predicted sites from our tests on yeast are available at http://www.swissregulon.unibas.ch.

Algorithms↗

Predicting carcinoid heart disease with the noisy-threshold classifier.

OBJECTIVE: To predict the development of carcinoid heart disease (CHD), which is a life-threatening complication of certain neuroendocrine tumors. To this end, a novel type of Bayesian classifier, known as the noisy-threshold classifier, is applied. MATERIALS AND METHODS: Fifty-four cases of patients that suffered from a low-grade midgut carcinoid tumor, of which 22 patients developed CHD, were obtained from the Netherlands Cancer Institute (NKI). Eleven attributes that are known at admission have been used to classify whether the patient develops CHD. Classification accuracy and area under the receiver operating characteristics (ROC) curve of the noisy-threshold classifier are compared with those of the naive-Bayes classifier, logistic regression, the decision-tree learning algorithm C4.5, and a decision rule, as formulated by an expert physician. RESULTS: The noisy-threshold classifier showed the best classification accuracy of 72% correctly classified cases, although differences were significant only for logistic regression and C4.5. An area under the ROC curve of 0.66 was attained for the noisy-threshold classifier, and equaled that of the physician's decision-rule. CONCLUSIONS: The noisy-threshold classifier performed favorably to other state-of-the-art classification algorithms, and equally well as a decision-rule that was formulated by the physician. Furthermore, the semantics of the noisy-threshold classifier make it a useful machine learning technique in domains where multiple causes influence a common effect.

Algorithms↗

Monitoring device safety in interventional cardiology.

OBJECTIVE: A variety of postmarketing surveillance strategies to monitor the safety of medical devices have been supported by the U.S. Food and Drug Administration, but there are few systems to automate surveillance. Our objective was to develop a system to perform real-time monitoring of safety data using a variety of process control techniques. DESIGN: The Web-based Data Extraction and Longitudinal Time Analysis (DELTA) system imports clinical data in real-time from an electronic database and generates alerts for potentially unsafe devices or procedures. The statistical techniques used are statistical process control (SPC), logistic regression (LR), and Bayesian updating statistics (BUS). MEASUREMENTS: We selected in-patient mortality following implantation of the Cypher drug-eluting coronary stent to evaluate our system. Data from the University of Michigan Consortium Bare-Metal Stent Study was used to calculate the event rate alerting boundaries. Data analysis was performed on local catheterization data from Brigham and Women's Hospital from July 1, 2003, shortly after the Cypher release, to December 31, 2004, including 2,270 cases with 27 observed deaths. RESULTS: The single-stratum SPC had alerts in months 4 and 10. The multistrata SPC had alerts in months 5, 10, and 18 in the moderate-risk stratum, and months 1, 4, 7, and 10 in the high-risk stratum. The only cumulative alerts were in the first month for the high-risk stratum of the multistrata SPC. The LR method showed no monthly or cumulative alerts. The BUS method showed an alert in the first month for the high-risk stratum. CONCLUSION: The system performed adequately within the Brigham and Women's Hospital Intranet environment based on the design goals. All three cumulative methods agreed that the overall observed event rates were not significantly higher for the new medical device than for a closely related medical device and were consistent with the observation that the initial concerns about this device dissipated as more data accumulated.

Bayes Theorem↗

The age of the angiosperms: a molecular timescale without a clock.

The age of the angiosperms has long been of interest to botanists and evolutionary biologists. Many early efforts to date the age of the angiosperms and evolutionary divergences within the angiosperm clade using a molecular clock have yielded age estimates that are grossly inconsistent with the fossil record. We investigated the age of angiosperms using Bayesian relaxed clock (BRC) and penalized likelihood (PL) approaches. Both of these methods allow the incorporation of multiple fossil constraints into the optimization procedure. The BRC method allows a range of values for among-lineage rate of substitution, from a nearly clocklike behavior to a condition in which each branch is allowed an optimal substitution rate, and also accounts for variation in molecular evolution across multiple genes. A topology derived from an analysis of genes from all three plant genomes for 71 taxa was used as a backbone. The effects on age estimates of different genes, single-gene versus concatenated datasets, and the inclusion and assumptions of fossils as age constraints were examined. In addition, the influence of prior distributions on estimates of divergence times was also explored. These results indicate that widely divergent age estimates can result from the different methods (198-139 million years ago), different sources of data (275-122 million years ago), and the inclusion of temporal constraints to topologies. Most dates, however, are between 180-140 million years ago, suggesting a Middle Jurassic-Early Cretaceous origin of flowering plants, predating the oldest unequivocal fossil angiosperms by about 45-5 million years. Nonetheless, these dates are consistent with other recent studies that have used methods that relax the assumption of a strict molecular clock and also agree with the hypothesis that the angiosperms may be somewhat older than the fossil record indicates.

Bayes Theorem↗

Intelligent machines in the twenty-first century: foundations of inference and inquiry.

The last century saw the application of Boolean algebra to the construction of computing machines, which work by applying logical transformations to information contained in their memory. The development of information theory and the generalization of Boolean algebra to Bayesian inference have enabled these computing machines, in the last quarter of the twentieth century, to be endowed with the ability to learn by making inferences from data. This revolution is just beginning as new computational techniques continue to make difficult problems more accessible. Recent advances in our understanding of the foundations of probability theory have revealed implications for areas other than logic. Of relevance to intelligent machines, we recently identified the algebra of questions as the free distributive algebra, which will now allow us to work with questions in a way analogous to that which Boolean algebra enables us to work with logical statements. In this paper, we examine the foundations of inference and inquiry. We begin with a history of inferential reasoning, highlighting key concepts that have led to the automation of inference in modern machine-learning systems. We then discuss the foundations of inference in more detail using a modern viewpoint that relies on the mathematics of partially ordered sets and the scaffolding of lattice theory. This new viewpoint allows us to develop the logic of inquiry and introduce a measure describing the relevance of a proposed question to an unresolved issue. Last, we will demonstrate the automation of inference, and discuss how this new logic of inquiry will enable intelligent machines to ask questions. Automation of both inference and inquiry promises to allow robots to perform science in the far reaches of our solar system and in other star systems by enabling them not only to make inferences from data, but also to decide which question to ask, which experiment to perform, or which measurement to take given what they have learned and what they are designed to understand.

Artificial Intelligence↗

Three quantitative approaches to the diagnosis of abdominal pain in children: practical applications of decision theory.

BACKGROUND/PURPOSE: The authors compared 3 quantitative methods for assisting clinicians in the differential diagnosis of abdominal pain in children, where the most common important endpoint is whether the patient has appendicitis. Pretest probability in different age and sex groups were determined to perform Bayesian analysis, binary logistic regression was used to determine which variables were statistically significantly likely to contribute to a diagnosis, and recursive partitioning was used to build decision trees with quantitative endpoints. METHODS: The records of all children (1,208) seen at a large urban emergency department (ED) with a chief complaint of abdominal pain were immediately reviewed retrospectively (24 to 72 hours after the encounter). Attempts were made to contact all the patients' families to determine an accurate final diagnosis. A total of 1,008 (83%) families were contacted. Data were analyzed by calculation of the posttest probability, recursive partitioning, and binary logistic regression. RESULTS: In all groups the most common diagnosis was abdominal pain (ICD-9 Code 789). After this, however, the order of the most common final diagnoses for abdominal pain varied significantly. The entire group had a pretest probability of appendicitis of 0.06. This varied with age and sex from 0.02 in boys 2 to 5 years old to 0.16 in boys older than 12 years. In boys age 5 to 12, recursive partitioning and binary logistic regression agreed on guarding and anorexia as important variables. Guarding and tenderness were important in girls age 5 to 12. In boys age greater than 12, both agreed on guarding and anorexia. Using sensitivities and specificities from the literature, computed tomography improved the posttest probability for the group from.06 to.33; ultrasound improved it from.06 to.48; and barium enema improved it from.06 to.58. CONCLUSIONS: Knowing the pretest probabilities in a specific population allows the physician to evaluate the likely diagnoses first. Other quantitative methods can help judge how much importance a certain criterion should have in the decision making and how much a particular test is likely to influence the probability of a correct diagnosis. It now should be possible to make these sophisticated quantitative methods readily available to clinicians via the computer.

Abdominal Pain↗

Optimization of Ga-67 imaging for detection and estimation tasks: dependence of imaging performance on spectral acquisition parameters.

UNLABELLED: We have compared the use of two (93 and 185 keV) and three (93, 185, and 300 keV) photopeaks for Ga-67 tumor imaging and optimized the placement of each energy window. METHODS: The bases for optimization and evaluation were ideal and Bayesian signal-to-noise ratios (SNR) for the detection of spheres embedded in a realistic anthropomorphic digital torso phantom and ideal SNR for the estimation of their size and activity concentration. Seven spheres of radii ranging from 1 to 3 cm, located at several sites in the torso, were simulated using a realistic Monte Carlo program. We also calculated the ideal SNR for the detection from simple phantom acquisitions. RESULTS: For detection and estimation tasks, the optimum windows were identical for all sphere sizes and locations. For the 93 keV photopeak, the optimal window was 84-102 keV for the detection and 87-102 keV for estimation; these windows are narrower than the 20% window often used in the clinic (83-101 keV). For the 185 keV photopeak, the optimal window was 170-220 keV for the detection and 170-215 keV for estimation; these are substantially different than the 15% window used in our clinic (171-199 keV). For the 300 keV photopeak, the optimal window for detection was 270-320 keV, and for estimation, 280-320 keV. Using the three optimized, rather than only the two lower-energy, windows yielded a 9% increase in the SNR for the detection of the 3 cm diam sphere (a 12% increase for a 2 cm diam sphere) and a 7% increase in the SNR for estimation of its size. For the acquired phantom data, detection also increased by 9%-12% when using three, rather than two, energy windows.

Abdomen↗

Bayesian and information theory analysis of MAS sideband patterns in spin 1/2 systems.

Bayesian statistics and information theory are used to analyze the reliability of extracting chemical shift parameters from spinning sideband patterns of spin 1/2 systems. Efficient code has been written to calculate the two-dimensional posterior probability as a function of the chemical shift anisotropy, delta, and the asymmetry parameter, eta, given the sideband intensities and the signal-to-noise ratio. This method has the advantage of assuming only that the noise in the sideband intensities is distributed as a Gaussian. It assumes nothing about the distribution of the values of parameters delta and eta, which are shown in some cases to be highly non-Gaussian. The utility of Bayesian analysis is demonstrated on 1D slow-spinning MAS spectra and on sideband patterns extracted from 2D PASS spectra. Previous investigations have shown that there is an optimal range of spinning frequencies for determining delta. In this study, information theory is used to determine the signal-to-noise ratio dependence of the entropy in delta, eta, and total entropy in spinning sideband spectra. The entropy is a measure of the information content of a probability distribution. When the entropy is zero, there is perfect information on a system, while if it is infinite, there is no information on the system. It is found that for all values of eta and for signal-to-noise ratios in the range 50-1000, an entropy minimum in nudelta/nur occurs for values 1.5<or=nudelta/nur<or=3. In the same range of signal-to-noise ratios, the entropy in eta is a monotonically decreasing function of nudelta/nur. The global information content of a spinning sideband pattern (i.e., the total entropy) is dependent on the signal-to-noise ratio and has an optimal value at nudelta/nur approximately 2 at a signal-to-noise ratio of 50 and increasing to approximately 2.5 for signal-to-noise ratios of 1000. Finally, the increase of information in a sideband pattern as a function of the number of sidebands used in the analysis is examined. Most of the information about delta and eta is contained in the five central sidebands; i.e., sidebands -2 to 2.

Algorithms↗

A parametric imaging approach for the segmentation of ultrasound data.

When an ultrasonic examination is performed, a segmentation tool would often be very useful, either for the detection of pathologies, the early diagnosis of cancer or the follow-up of the lesions. Such a tool must be both reliable and accurate. However, because of the relatively reduced quality of ultrasound images due to the speckled texture, the segmentation of ultrasound data is a difficult task. We have previously proposed to tackle the problem using a multiresolution Bayesian region-based algorithm. For computation time purposes, a multiresolution version has been implemented. In order to improve the quality of the segmentation, we propose to perform the segmentation not only from the envelope image but to combine more information about the properties of the tissues in the segmentation process. Several acoustical parameters have thus been computed, either directly from the images or from the radio-frequency (RF) signal. In a previous study, two parametric images were involved in the segmentation process. The parameter represented the integrated backscatter (IBS) and the mean central frequency (MCF), which is a measurement related to the attenuation of ultrasound waves in the media. In this study, parameters representative of the scattering conditions in the tissue are evaluated in the multiparametric segmentation process. They are extracted from the K-distribution (alpha,b) and the Nakagami distribution (m,Omega) and are related to the local density of scatterers (alpha,m), the size of the scatterers (b) and the backscattering properties of the medium (Omega). The acoustical features are calculated locally on a sliding window. This procedure allows to built parametric mapping representing the particular characteristics of the medium. To test the influence of the acoustical parameters in the segmentation process, a set of numerical phantoms has been computed using the Field software developed by J.A. Jensen. Each phantom consists in two regions with two different acoustical properties: the density of scatterers and the scattering amplitude. From both the simulated RF signals and envelope images, the parameters have been computed; their relevance to represent a particular characteristic of the medium is evaluated. The segmentation has been processed for each phantom. The ability of each parameter to improve the segmentation results is validated. A agar-gel phantom has also been created, in order to test the accuracy of the parameters in conditions closer to the in vivo ultrasound imaging. This phantom contains four inclusions with different concentrations of silica. A B&K ultrasound device provides the RF data. The quantification of the segmentation quality is based on the rate of correctly classified pixels and it has been computed for all the parameters either from the field images or the phantom images. The large improvement in the segmentation results obtained reveals that the multiparametric segmentation scheme proposed in this study can be a reliable tool for the processing of noisy ultrasound data.

Mathematics↗

The representational capacity of the distributed encoding of information provided by populations of neurons in primate temporal visual cortex.

It has been shown that it is possible to read, from the firing rates of just a small population of neurons, the code that is used in the macaque temporal lobe visual cortex to distinguish between different faces being looked at. To analyse the information provided by populations of single neurons in the primate temporal cortical visual areas, the responses of a population of 14 neurons to 20 visual stimuli were analysed in a macaque performing a visual fixation task. The population of neurons analysed responded primarily to faces, and the stimuli utilised were all human and monkey faces. Each neuron had its own response profile to the different members of the stimulus set. The mean response of each neuron to each stimulus in the set was calculated from a fraction of the ten trials of data available for every stimulus. From the remaining data, it was possible to calculate, for any population response vector, the relative likelihoods that it had been elicited by each of the stimuli in the set. By comparison with the stimuli actually shown, the mean percentage correct identification was computed and also the mean information about the stimuli, in bits, that the population of neurons carried on a single trial. When the decoding algorithm used for this calculation approximated an optimal, Bayesian estimate of the relative likelihoods, the percentage correct increased from 14% correct (chance was 5% correct) with one neuron to 67% with 14 neurons. The information conveyed by the population of neurons increased approximately linearly from 0.33 bits with one neuron to 2.77 bits with 14 neurons. This leads to the important conclusion that the number of stimuli that can be encoded by a population of neurons in this part of the visual system increases approximately exponentially as the number of cells in the sample increases (in that the log of the number of stimuli increases almost linearly). This is in contrast to a local encoding scheme (of "grandmother" cells), in which the number of stimuli encoded increases linearly with the number of cells in the sample. Thus one of the potentially important properties of distributed representations, an exponential increase in the number of stimuli that can be represented, has been demonstrated in the brain with this population of neurons. When the algorithm used for estimating stimulus likelihood was as simple as could be easily implemented by neurons receiving the population's output (based on just the dot product between the population response vector and each mean response vector), it was still found that the 14-neuron population produced 66% correct guesses and conveyed 2.30 bits of information, or 83% of the information that could be extracted with the nearly optimal procedure. It was also shown that, although there was some redundancy in the representation (with each neuron contributing to the information carried by the whole population 60% of the information it carried alone, rather than 100%), this is due to the fact that the number of stimuli in the set was limited (it was 20). The data are consistent with minimal redundancy for sufficiently large and diverse sets of stimuli. The implication for brain connectivity of the distributed encoding scheme, which was demonstrated here in the case of faces, is that a neuron can receive a great deal of information about what is encoded by a large population of neurons if it is able to receive its inputs from a random subset of these neurons, even of limited numbers (e.g. hundreds).

Algorithms↗

A Bayesian method for identifying missing enzymes in predicted metabolic pathway databases.

BACKGROUND: The PathoLogic program constructs Pathway/Genome databases by using a genome's annotation to predict the set of metabolic pathways present in an organism. PathoLogic determines the set of reactions composing those pathways from the enzymes annotated in the organism's genome. Most annotation efforts fail to assign function to 40-60% of sequences. In addition, large numbers of sequences may have non-specific annotations (e.g., thiolase family protein). Pathway holes occur when a genome appears to lack the enzymes needed to catalyze reactions in a pathway. If a protein has not been assigned a specific function during the annotation process, any reaction catalyzed by that protein will appear as a missing enzyme or pathway hole in a Pathway/Genome database. RESULTS: We have developed a method that efficiently combines homology and pathway-based evidence to identify candidates for filling pathway holes in Pathway/Genome databases. Our program not only identifies potential candidate sequences for pathway holes, but combines data from multiple, heterogeneous sources to assess the likelihood that a candidate has the required function. Our algorithm emulates the manual sequence annotation process, considering not only evidence from homology searches, but also considering evidence from genomic context (i.e., is the gene part of an operon?) and functional context (e.g., are there functionally-related genes nearby in the genome?) to determine the posterior belief that a candidate has the required function. The method can be applied across an entire metabolic pathway network and is generally applicable to any pathway database. The program uses a set of sequences encoding the required activity in other genomes to identify candidate proteins in the genome of interest, and then evaluates each candidate by using a simple Bayes classifier to determine the probability that the candidate has the desired function. We achieved 71% precision at a probability threshold of 0.9 during cross-validation using known reactions in computationally-predicted pathway databases. After applying our method to 513 pathway holes in 333 pathways from three Pathway/Genome databases, we increased the number of complete pathways by 42%. We made putative assignments to 46% of the holes, including annotation of 17 sequences of previously unknown function. CONCLUSIONS: Our pathway hole filler can be used not only to increase the utility of Pathway/Genome databases to both experimental and computational researchers, but also to improve predictions of protein function.

Amino Acid Oxidoreductases↗