Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian networks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Correction of temperature variations in kinetic-based determinations by use of pruning computational neural networks in conjunction with genetic algorithms.

The joint use of genetic algorithms and pruning computational neural networks is shown to be an effective means for selecting the number of inputs required to correct temperature variations in kinetic-based determinations. The genetic algorithm uses a pruning procedure based on Bayesian regularization and is highly efficient as a feature selector; it provides quite good results in the generalization process without the need to use a validation set. The fitness function is defined as the sum of two subfunctions: one controls the learning ability of the network and the other its complexity. The training, pruning, and generalization processes were initially tested with simulated data in order to acquire preliminary information for the ensuing work with real data. The performance of the proposed method was assessed by applying it to the determination of the amino acid L-glycine by its classical spectrophotometric reaction with ninhydrin. A straightforward network topology including temperature as input (40+T:2:1 with 19 connections after the pruning process) was used to estimate the L-glycine concentration from kinetic curves affected by temperature variations over the range 60-75 degrees C, using kinetic data acquired up to only 1.5 half-lives. The trained network estimates this concentration with a standard error of prediction for the testing set of ca. 8%, which is much smaller than those provided by a classical parametric method such as nonlinear regression (even if kinetic data acquired at longer half-lives are used). Finally, a kinetic interpretation of the pruning process is provided in order to better demonstrate its potential for kinetic analysis.

Algorithms↗

Consolidating the set of known human protein-protein interactions in preparation for large-scale mapping of the human interactome.

BACKGROUND: Extensive protein interaction maps are being constructed for yeast, worm, and fly to ask how the proteins organize into pathways and systems, but no such genome-wide interaction map yet exists for the set of human proteins. To prepare for studies in humans, we wished to establish tests for the accuracy of future interaction assays and to consolidate the known interactions among human proteins. RESULTS: We established two tests of the accuracy of human protein interaction datasets and measured the relative accuracy of the available data. We then developed and applied natural language processing and literature-mining algorithms to recover from Medline abstracts 6,580 interactions among 3,737 human proteins. A three-part algorithm was used: first, human protein names were identified in Medline abstracts using a discriminator based on conditional random fields, then interactions were identified by the co-occurrence of protein names across the set of Medline abstracts, filtering the interactions with a Bayesian classifier to enrich for legitimate physical interactions. These mined interactions were combined with existing interaction data to obtain a network of 31,609 interactions among 7,748 human proteins, accurate to the same degree as the existing datasets. CONCLUSION: These interactions and the accuracy benchmarks will aid interpretation of current functional genomics data and provide a basis for determining the quality of future large-scale human protein interaction assays. Projecting from the approximately 15 interactions per protein in the best-sampled interaction set to the estimated 25,000 human genes implies more than 375,000 interactions in the complete human protein interaction network. This set therefore represents no more than 10% of the complete network.

Algorithms↗

Genetic variants in migraine: a field synopsis and systematic re-analysis of meta-analyses.

OBJECTIVE: Numerous genetic variants from meta-analyses of observational studies and GWAS were reported to be associated with migraine susceptibility. However, due to the random errors in meta-analyses, the noteworthiness of the results showing statistically significant remains doubtful. Thus, we performed this field synopsis and re-analysis study to evaluate the noteworthiness using a Bayesian approach in hope of finding true associations. METHODS: Relevant meta-analyses from observational studies and GWAS examining correlation between all genetic variants and migraine risk were included in our study by a PubMed search. Identification of noteworthy associations were analyzed by false-positive rate probability (FPRP) and Bayesian false discovery probability (BFDP). Using noteworthy variants, GO enrichment analysis were conducted through DAVID online tool. Then, the PPI network and hub genes were performed using STRING database and CytoHubba software. RESULTS: As for 8 significant genetic variants from observational studies, none of which showed noteworthy at prior probability of 0.001. Out of 47 significant genetic variants in GWAS, 36 were noteworthy at prior probability of 0.000001 via FPRP or BFDP. We further found the pathways "positive regulation of cytosolic calcium ion concentration" and "inositol phosphate-mediated signaling" and hub genes including MEF2D, TSPAN2, PHACTR1, TRPM8 and PRDM16 related to migraine susceptibility. CONCLUSION: Herein, we have identified several noteworthy variants for migraine susceptibility in this field synopsis. We hope these data would help identify novel genetic biomarkers and potential therapeutic target for migraine.

Bayes Theorem↗

Dynamics of the evolution of learning algorithms by selection.

We study the evolution of artificial learning systems by means of selection. Genetic programming is used to generate populations of programs that implement algorithms used by neural network classifiers to learn a rule in a supervised learning scenario. In contrast to concentrating on final results, which would be the natural aim while designing good learning algorithms, we study the evolution process. Phenotypic and genotypic entropies, which describe the distribution of fitness and of symbols, respectively, are used to monitor the dynamics. We identify significant functional structures responsible for the improvements in the learning process. In particular, some combinations of variables and operators are useful in assessing performance in rule extraction and can thus implement annealing of the learning schedule. We also find combinations that can signal surprise, measured on a single example, by the difference between predicted and correct classification. When such favorable structures appear, they are disseminated on very short time scales throughout the population. Due to such abruptness they can be thought of as dynamical transitions. But foremost, we find a strict temporal order of such discoveries. Structures that measure performance are never useful before those for measuring surprise. Invasions of the population by such structures in the reverse order were never observed. Asymptotically, the generalization ability approaches Bayesian results.

Algorithms↗

Stochastic complexities of general mixture models in variational Bayesian learning.

In this paper, we focus on variational Bayesian learning of general mixture models. Variational Bayesian learning was proposed as an approximation of Bayesian learning. While it has provided computational tractability and good generalization in many applications, little has been done to investigate its theoretical properties. The asymptotic form was obtained for the stochastic complexity, or the free energy in the variational Bayesian learning of a mixture of exponential-family distributions, which is the main contribution this paper makes. We reveal that the stochastic complexities become smaller than those of regular statistical models, which implies that the advantages of Bayesian learning are still retained in variational Bayesian learning. Moreover, the derived bounds indicate what influence the hyperparameters have on the learning process, and the accuracy of the variational Bayesian approach as an approximation of true Bayesian learning.

Animals↗

Predicting protein secondary structure with probabilistic schemata of evolutionarily derived information.

We demonstrate the applicability of our previously developed Bayesian probabilistic approach for predicting residue solvent accessibility to the problem of predicting secondary structure. Using only single-sequence data, this method achieves a three-state accuracy of 67% over a database of 473 non-homologous proteins. This approach is more amenable to inspection and less likely to overlearn specifics of a dataset than "black box" methods such as neural networks. It is also conceptually simpler and less computationally costly. We also introduce a novel method for representing and incorporating multiple-sequence alignment information within the prediction algorithm, achieving 72% accuracy over a dataset of 304 non-homologous proteins. This is accomplished by creating a statistical model of the evolutionarily derived correlations between patterns of amino acid substitution and local protein structure. This model consists of parameter vectors, termed "substitution schemata," which probabilistically encode the structure-based heterogeneity in the distributions of amino acid substitutions found in alignments of homologous proteins. The model is optimized for structure prediction by maximizing the mutual information between the set of schemata and the database of secondary structures. Unlike "expert heuristic" methods, this approach has been demonstrated to work well over large datasets. Unlike the opaque neural network algorithms, this approach is physicochemically intelligible. Moreover, the model optimization procedure, the formalism for predicting one-dimensional structural features and our previously developed method for tertiary structure recognition all share a common Bayesian probabilistic basis. This consistency starkly contrasts with the hybrid and ad hoc nature of methods that have dominated this field in recent years.

Algorithms↗

Model evaluation and spatial interpolation by Bayesian combination of observations with outputs from numerical models.

Constructing maps of dry deposition pollution levels is vital for air quality management, and presents statistical problems typical of many environmental and spatial applications. Ideally, such maps would be based on a dense network of monitoring stations, but this does not exist. Instead, there are two main sources of information for dry deposition levels in the United States: one is pollution measurements at a sparse set of about 50 monitoring stations called CASTNet, and the other is the output of the regional scale air quality models, called Models-3. A related problem is the evaluation of these numerical models for air quality applications, which is crucial for control strategy selection. We develop formal methods for combining sources of information with different spatial resolutions and for the evaluation of numerical models. We specify a simple model for both the Models-3 output and the CASTNet observations in terms of the unobserved ground truth, and we estimate the model in a Bayesian way. This provides improved spatial prediction via the posterior distribution of the ground truth, allows us to validate Models-3 via the posterior predictive distribution of the CASTNet observations, and enables us to remove the bias in the Models-3 output. We apply our methods to data on SO2 concentrations, and we obtain high-resolution SO2 distributions by combining observed data with model output. We also conclude that the numerical models perform worse in areas closer to power plants, where the SO2 values are overestimated by the models.

Air Pollution↗

Historical demography of brown trout (Salmo trutta) in the Adriatic drainage including the putative S. letnica endemic to Lake Ohrid.

We explore the historical demography of the Adriatic lineage of brown trout and more explicitly the colonization and phylogenetic placement of Ohrid trout, based on variation at 12 microsatellite loci and the mtDNA control region. All Adriatic basin haplotypes reside in derived positions in a network that represents the entire lineage. The central presumably most ancestral haplotype in this network is restricted to the Iberian Peninsula, where it is very common, supporting a Western Mediterranean origin for the lineage. The expansion statistic R2, Bayesian based estimates of demographic parameters, and star-like genealogies support expansions on several geographic scales, whereas application of pairwise mismatch analysis was somewhat ambiguous. The estimated time since expansion (155,000 years ago) for the Adriatic lineage was supported by a narrow confidence interval compared to previous studies. Based on microsatellite and mtDNA sequence variation, the endemic Ohrid trout represents a monophyletic lineage isolated from other Adriatic basin populations, but nonetheless most likely evolving from within the Adriatic lineage of brown trout. Our results do not support the existence of population structuring within Lake Ohrid, even though samples included two putative intra-lacustrine forms. In the interests of protecting the unique biodiversity of this ancient ecosystem, we recommend retaining the taxonomic epithet Salmo letnica for the endemic Ohrid trout.

Animals↗

The Bayesian reader: explaining word recognition as an optimal Bayesian decision process.

This article presents a theory of visual word recognition that assumes that, in the tasks of word identification, lexical decision, and semantic categorization, human readers behave as optimal Bayesian decision makers. This leads to the development of a computational model of word recognition, the Bayesian reader. The Bayesian reader successfully simulates some of the most significant data on human reading. The model accounts for the nature of the function relating word frequency to reaction time and identification threshold, the effects of neighborhood density and its interaction with frequency, and the variation in the pattern of neighborhood density effects seen in different experimental tasks. Both the general behavior of the model and the way the model predicts different patterns of results in different tasks follow entirely from the assumption that human readers approximate optimal Bayesian decision makers.

Attention↗

Biochemical networks with uncertain parameters.

The modelling of biochemical networks becomes delicate if kinetic parameters are varying, uncertain or unknown. Facing this situation, we quantify uncertain knowledge or beliefs about parameters by probability distributions. We show how parameter distributions can be used to infer probabilistic statements about dynamic network properties, such as steady-state fluxes and concentrations, signal characteristics or control coefficients. The parameter distributions can also serve as priors in Bayesian statistical analysis. We propose a graphical scheme, the 'dependence graph', to bring out known dependencies between parameters, for instance, due to the equilibrium constants. If a parameter distribution is narrow, the resulting distribution of the variables can be computed by expanding them around a set of mean parameter values. We compute the distributions of concentrations, fluxes and probabilities for qualitative variables such as flux directions. The probabilistic framework allows the study of metabolic correlations, and it provides simple measures of variability and stochastic sensitivity. It also shows clearly how the variability of biological systems is related to the metabolic response coefficients.

Animals↗

A comparison of material classification techniques for ultrasound inverse imaging.

The conjugate gradient method with edge preserving regularization (CGEP) is applied to the ultrasound inverse scattering problem for the early detection of breast tumors. To accelerate image reconstruction, several different pattern classification schemes are introduced into the CGEP algorithm. These classification techniques are compared for a full-sized, two-dimensional breast model. One of these techniques uses two parameters, the sound speed and attenuation, simultaneously to perform classification based on a Bayesian classifier and is called bivariate material classification (BMC). The other two techniques, presented in earlier work, are univariate material classification (UMC) and neural network (NN) classification. BMC is an extension of UMC, the latter using attenuation alone to perform classification, and NN classification uses a neural network. Both noiseless and noisy cases are considered. For the noiseless case, numerical simulations show that the CGEP-BMC method requires 40% fewer iterations than the CGEP method, and the CGEP-NN method requires 55% fewer. The CGEP-BMC and CGEP-NN methods yield more accurate reconstructions than the CGEP method. A quantitative comparison of the CGEP-BMC, CGEP-NN, and GN-UMC methods shows that the CGEP-BMC and CGEP-NN methods are more robust to noise than the GN-UMC method, while all three are similar in computational complexity.

Breast Neoplasms↗

Bayesian ranking of sites for engineering safety improvements: decision parameter, treatability concept, statistical criterion, and spatial dependence.

In recent years, there has been a renewed interest in applying statistical ranking criteria to identify sites on a road network, which potentially present high traffic crash risks or are over-represented in certain type of crashes, for further engineering evaluation and safety improvement. This requires that good estimates of ranks of crash risks be obtained at individual intersections or road segments, or some analysis zones. The nature of this site ranking problem in roadway safety is related to two well-established statistical problems known as the small area (or domain) estimation problem and the disease mapping problem. The former arises in the context of providing estimates using sample survey data for a small geographical area or a small socio-demographic group in a large area, while the latter stems from estimating rare disease incidences for typically small geographical areas. The statistical problem is such that direct estimates of certain parameters associated with a site (or a group of sites) with adequate precision cannot be produced, due to a small available sample size, the rareness of the event of interest, and/or a small exposed population or sub-population in question. Model based approaches have offered several advantages to these estimation problems, including increased precision by "borrowing strengths" across the various sites based on available auxiliary variables, including their relative locations in space. Within the model based approach, generalized linear mixed models (GLMM) have played key roles in addressing these problems for many years. The objective of the study, on which this paper is based, was to explore some of the issues raised in recent roadway safety studies regarding ranking methodologies in light of the recent statistical development in space-time GLMM. First, general ranking approaches are reviewed, which include naïve or raw crash-risk ranking, scan based ranking, and model based ranking. Through simulations, the limitation of using the naïve approach in ranking is illustrated. Second, following the model based approach, the choice of decision parameters and consideration of treatability are discussed. Third, several statistical ranking criteria that have been used in biomedical, health, and other scientific studies are presented from a Bayesian perspective. Their applications in roadway safety are then demonstrated using two data sets: one for individual urban intersections and one for rural two-lane roads at the county level. As part of the demonstration, it is shown how multivariate spatial GLMM can be used to model traffic crashes of several injury severity types simultaneously and how the model can be used within a Bayesian framework to rank sites by crash cost per vehicle-mile traveled (instead of by crash frequency rate). Finally, the significant impact of spatial effects on the overall model goodness-of-fit and site ranking performances are discussed for the two data sets examined. The paper is concluded with a discussion on possible directions in which the study can be extended.

Accidents, Traffic↗

Proceedings of the SMBE Tri-National Young Investigators' Workshop 2005. Improved consensus network techniques for genome-scale phylogeny.

Although recent studies indicate that estimating phylogenies from alignments of concatenated genes greatly reduces the stochastic error, the potential for systematic error still remains, heightening the need for reliable methods to analyze multigene data sets. Consensus methods provide an alternative, more inclusive, approach for analyzing collections of trees arising from multiple genes. We extend a previously described consensus network method for genome-scale phylogeny (Holland, B. R., K. T. Huber, V. Moulton, and P. J. Lockhart. 2004. Using consensus networks to visualize contradictory evidence for species phylogeny. Mol. Biol. Evol. 21:1459-1461) to incorporate additional information. This additional information could come from bootstrap analysis, Bayesian analysis, or various methods to find confidence sets of trees. The new methods can be extended to include edge weights representing genetic distance. We use three data sets to illustrate the approach: 61 genes from 14 angiosperm taxa and one gymnosperm, 106 genes from eight yeast taxa, and 46 members of a gene family from 15 vertebrate taxa.

Animals↗

A decision-theoretic approach to identifying future high-cost patients.

OBJECTIVE: The objective of this study was to develop and evaluate a method of allocating funding for very-high-cost (VHC) patients among hospitals. RESEARCH DESIGN: Diagnostic cost groups (DCGs) were used for risk adjustment. The patient population consisted of 253,013 veterans who used Department of Veterans Affairs (VA) medical care services in fiscal year (FY) 2003 (October 1, 2002-September 30, 2003) in a network of 8 VA hospitals. We defined VHC as greater than 75,000 dollars (0.81%). The upper fifth percentile was also used for comparison. METHODS: A Bayesian decision rule for classifying patients as VHC/not VHC using DCGs was developed and evaluated. The method uses FY 2003 DCGs to allocate VHC funds for FY 2004. We also used FY 2002 DCGs to allocate VHC funds for FY 2003 for comparison. The resulting allocation was compared with using the allocation of VHC patients among the hospitals in the previous year. RESULTS: The decision rule identified DCG 17 as the optimal cutoff for identifying VHC patients for the next year. The previous year's allocation came closest to the actual distribution of VHC patients. CONCLUSIONS: The decision-theoretic approach may provide insight into the economic consequences of classifying a patient as VHC or not VHC. More research is needed into methods of identifying future VHC patients so that capitation plans can fairly reimburse healthcare systems for appropriately treating these patients.

Aged↗

Biogeographic patterns and phylogeography of dwarf chameleons (Bradypodion) in an African biodiversity hotspot.

The southern African landscape appears to have experienced frequent shifts in vegetation associated with climatic change through the mid-Miocene and Plio-Pleistocene. One group whose historical biogeography may have been affected by these fluctuations are the dwarf chameleons (Bradypodion), due to their associations with distinct vegetation types. Thus, this group provides an opportunity to investigate historical biogeography in light of climatic fluctuations. A total of 138 dwarf chameleons from the Cape Floristic Region of South Africa were sequenced for two mitochondrial genes (ND2 and 16S), and resulting phylogenetic analyses showed two well-supported clades that are distributed allopatrically. Within clades, diversity among some lineages was low, and haplotype networks showed patterns of reticulate evolution and incomplete lineage sorting, suggesting relatively recent origins for some of these lineages. A dispersal-vicariance analysis and a relaxed Bayesian clock suggest that vicariance between the two main clades occurred in the mid-Miocene, and that both dispersal and vicariance have played a role in shaping present-day distributions. These analyses also suggest that the most recent series of lineage diversification events probably occurred within the last 3-6 million years. This suggests that the origins of many present-day lineages were founded in the Plio-Pleistocene, a time period that corresponds to the reduction of forests in the region and the establishment of the fynbos biome.

Animals↗

Data mining in spontaneous reports.

The increasing size of spontaneous report data sets and the increasing capability for screening such data due to increases in computational power has led to a recent increase in interest and use of data mining on such data. While data mining plays an important role in the analysis of spontaneous reports, there is general debate on how and when data mining should be best performed. While the cornerstone principles for data mining of spontaneous reports have been in place since the 1960s, several significant changes have occurred to make their use widespread. Superficially the Bayesian methods seem unnecessarily complex, particularly given the nature of the data, but in practice implementation in Bayesian framework gives clear benefits. There are difficulties evaluating the performance of the methods, but they work and save resources in managing large data sets. The use of neural networks allows more sophisticated pattern recognition to be performed.

Adverse Drug Reaction Reporting Systems↗

Bayesian Genome-Wide Association Study of Feed Efficiency Traits in Pigs.

Feed efficiency traits are increasingly important in pig production for improving profitability and environmental sustainability. Understanding their genetic basis is crucial for uncovering underlying biological mechanisms and informing selection strategies. In this study, we analyzed residual feed intake (RFI), feed conversion ratio (FCR), and average daily feed intake (ADFI) in 201 animals. Three separate Bayesian GWASs were conducted using 29,844 SNPs in a case-control design, with the lowest and highest 15% of the phenotypic distribution selected as controls and cases (N = 30 per group), respectively, for each trait. The results confirmed the polygenic nature of the traits, identifying 4 SNPs for RFI on Sus scrofa chromosomes (SSC) 3, 13, and 15 with high posterior probability for the direction of their effects; 4 SNPs for FCR on SSC 8, 14, and 17; and 8 SNPs for ADFI on SSC 1, 2, 6, 8, and 11. A candidate gene search identified 41 potential genes involved in diverse biological processes, including feed efficiency, intestinal development, tissue remodeling and integrity, nutrient transport and absorption, metabolic homeostasis, cellular signaling, energy sensing, and neurological regulation. These genes formed a highly interconnected network, highlighting the complexity of feed efficiency and the interplay among multiple physiological, metabolic, and regulatory pathways.

Bayesian analysis↗