Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Shielding the First 24 Postnatal Months of Life: A Proposal for a Prospective Cohort Study of Early-Life Electromagnetic Exposure and Autism Risk.

BACKGROUND: Autism Spectrum Disorder (ASD) involves Mirror Neuron System (MNS) dysfunction, driving core social and imitative impairments. Systemic physiological alterations such as autonomic dysregulation, mitochondrial dysfunction and neuroinflammation are known to impair synchronization and plasticity of neuronal clusters. A less-evident environmental cofactor, coinciding with rising ASD prevalence, is the considerable world-wide increase in electromagnetic radiation (EMR) overall exposure among children. Experimental evidence shows how low-intensity EMR influences cellular processes, via voltage-gated calcium channels (VGCCs), oxidative stress, and mitochondrial metabolism. The Resonant Convergence framework, allow to predict how chronic EMR exposure during the first 24 postnatal months of life can act as a factor in ASD pathogenesis. The best candidate mechanism is chronic Ion Cyclotron Resonance (ICR) detuning the Ca2+-calmodulin pathway, thus disrupting MNS synchronization. METHODS AND ANALYSIS: A prospective observational pilot cohort study (24-month follow-up) proposes to enroll 1000 full-term newborns into two arms: an EMR-reduced cohort (n = 500, rest and sleep-phase Faraday shielding) and a standard exposure cohort (n = 500). Exposure is quantified via radiofrequency (RF)/extremely low frequency(ELF) measurements, proximity analysis, device inventories and wearable dosimetry. The primary endpoint is a continuous neurodevelopmental trajectory score (joint attention, language, electroencephalogram (EEG) mu-rhythm); binary ASD diagnosis (Autism Diagnostic Observation Schedule, Second Edition (ADOS-2), Autism Diagnostic Interview-Revised (ADI-R)) is a secondary, exploratory endpoint. Moreover, an optional genomic screening will evaluate gene-environment interactions within extremely low-frequency electromagnetic field (ELF-EMF) vulnerable pathways, including ASD-associated genes upregulated by RF via bromodomain and extraterminal protein (BET)-mediated epigenetic mechanisms. Analyses will employ risk ratios, Fisher's exact tests and logistic regression adjusted for confounders; mixed-effects and Bayesian modeling will evaluate longitudinal outcomes and exposure reduction effects. Given a 2-3% baseline prevalence, approximately 20-30 ASD cases are expected. The study is therefore powered for exploratory signal detection rather than definitive causal inference, providing the critical baseline data required to justify and design future confirmatory trials. Sex-stratified modeling will address the 4:1 male-to-female prevalence ratio. ETHICS AND DISSEMINATION: Ethics committee approval is not yet sought; full protocol review and approval will be obtained prior to the study initiation, in strict accordance with the Declaration of Helsinki. Written parental informed consent will be mandatory for all participants prior to enrollment. Study findings and methodological milestones will be disseminated through peer-reviewed international scientific publications. This protocol provides a structured methodological framework for the first prospective investigation of sleep-phase EMR reduction as a potential modulator of ASD incidence during early neurodevelopment. Results will inform adequately powered confirmatory trials in electromagnetic neurodevelopmental epidemiology.

autism spectrum disorder↗

Chromosome abnormalities in ovarian adenocarcinoma: III. Using breakpoint data to infer and test mathematical models for oncogenesis.

Cancer geneticists seek to identify genetic changes in tumor cells and to relate the genetic changes to tumor development. Because single changes can disrupt the cell cycle and promote other genetic changes, it is extremely hard to distinguish cause from effect. In this article we illustrate how 7 techniques from statistics, theoretical computer science, and phylogenetics can be used to infer and test possible models of tumor progression from single genome-wide descriptions of aberrations in a large sample of tumors. Specifically, we propose 4 tree models for tumor progression inferred from the large ovarian cancer data set described in the first 2 articles in this series. The models are derived from 2 different methods to select the non-random genetic aberrations and 2 different methods to infer the trees, given a set of events. Various aspects of the tree models are tested and extended by 5 methods: overall tests of independence, likelihood ratio tests, principal components analysis, directed acyclic graph modeling, and Bayesian survival analysis. All our methods lead to strikingly consistent conclusions about chromosomal breakpoints in ovarian adenocarcinoma, including (1) the non-random breakpoints in ovarian adenocarcinoma do not occur independently; (2) breakpoints in regions 1p3 and 11p1 are important early events and distinguish a class of tumors associated with poor prognosis; and (3) breakpoints in 1p1, 3p1, and 1q2 distinguish a class of ovarian tumors, and the breaks at 1p1 and 3p1 are associated with poor prognosis.

Adenocarcinoma↗

Using fold recognition to search for useful proteins: Bayesian approach to fold recognition.

The wealth of protein sequence and structure data is greater than ever, thanks to the ongoing Genomics and Structural Genomics projects. The information available through such efforts needs to be analysed by new methods that combine both databases. One important result of genomic sequence analysis is the inference of functional homology among proteins. Until recently sequence similarity comparison was the only method for homologue inference. The new fold recognition approach reviewed in this paper enhances sequence comparison methods by including structural information in the process of protein comparison. This additional information often allows for the detection of similarities that cannot be found by methods that only use sequence information.

Amino Acid Sequence↗

Selecting the right statistical model for analysis of insect count data by using information theoretic measures.

Researchers and regulatory agencies often make statistical inferences from insect count data using modelling approaches that assume homogeneous variance. Such models do not allow for formal appraisal of variability which in its different forms is the subject of interest in ecology. Therefore, the objectives of this paper were to (i) compare models suitable for handling variance heterogeneity and (ii) select optimal models to ensure valid statistical inferences from insect count data. The log-normal, standard Poisson, Poisson corrected for overdispersion, zero-inflated Poisson, the negative binomial distribution and zero-inflated negative binomial models were compared using six count datasets on foliage-dwelling insects and five families of soil-dwelling insects. Akaike's and Schwarz Bayesian information criteria were used for comparing the various models. Over 50% of the counts were zeros even in locally abundant species such as Ootheca bennigseni Weise, Mesoplatys ochroptera Stål and Diaecoderus spp. The Poisson model after correction for overdispersion and the standard negative binomial distribution model provided better description of the probability distribution of seven out of the 11 insects than the log-normal, standard Poisson, zero-inflated Poisson or zero-inflated negative binomial models. It is concluded that excess zeros and variance heterogeneity are common data phenomena in insect counts. If not properly modelled, these properties can invalidate the normal distribution assumptions resulting in biased estimation of ecological effects and jeopardizing the integrity of the scientific inferences. Therefore, it is recommended that statistical models appropriate for handling these data properties be selected using objective criteria to ensure efficient statistical inference.

Animals↗

Expert judgment and occupational hygiene: application to aerosol speciation in the nickel primary production industry.

In many situations characterized by sparse data, occupational hygienists have used subjective judgments that are claimed to be derived from their experience and knowledge. While this practice is widespread, there has been no systematic study of 'expert judgment' or the 'art' of occupational hygiene. Indeed, there is a need to address the question of whether there is such a thing as 'expert opinion' in occupational hygiene that is broadly shared by practicing professionals. This research, employing 11 experts who estimate an exposure parameter (the percentages of four nickel species) in 12 workplaces in a nickel primary production industry, provides a large dataset from which useful inferences can be drawn about the quality of expert judgments and the variability among the experts. A well-designed questionnaire that provided succinct information about the processes and baseline data served to calibrate the experts. The Bayesian framework has been used in this work to develop posterior means and standard deviations of the percentages of the four nickel species in the 12 workplaces of interest in the company. These estimates of the nickel speciation are at least as precise as--and most of the time more precise than--those provided by the sparse measurement data. There was a very high degree of agreement among the experts. A majority of the experts agreed among themselves 92% of the time, while almost two-thirds agreed 73% of the time. This, coupled with the fact that the experts came from varied backgrounds, seems to suggest that there is indeed some broad body of specialized knowledge that the experts are drawing on to reach similar judgments. It also seems that one type of expert is not necessarily any better than any other kind, and expertise does not necessarily require intimate familiarity with the workplace. In this example, the expert judgment exercise has indeed enhanced the quality of our knowledge of the exposure 'fingerprints' for the nickel industry workplaces studied and the combination of expert judgment and sparse data is better than the sparse data alone. For occupational hygiene exposure assessment, our experience suggests that such expert judgment methods can provide a cost-effective means to improve and refine information about workplace hazards. However, more study is warranted for situations where the domain of the quantity of interest has a much wider range of values, e.g. actual exposure values.

Air Pollutants, Occupational↗

Estimation of population growth or decline in genetically monitored populations.

This article introduces a new general method for genealogical inference that samples independent genealogical histories using importance sampling (IS) and then samples other parameters with Markov chain Monte Carlo (MCMC). It is then possible to more easily utilize the advantages of importance sampling in a fully Bayesian framework. The method is applied to the problem of estimating recent changes in effective population size from temporally spaced gene frequency data. The method gives the posterior distribution of effective population size at the time of the oldest sample and at the time of the most recent sample, assuming a model of exponential growth or decline during the interval. The effect of changes in number of alleles, number of loci, and sample size on the accuracy of the method is described using test simulations, and it is concluded that these have an approximately equivalent effect. The method is used on three example data sets and problems in interpreting the posterior densities are highlighted and discussed.

Alleles↗

Molecular systematics of Helicoma, Helicomyces and Helicosporium and their teleomorphs inferred from rDNA sequences.

Three genera of asexual, helical-spored fungi, Helicoma, Helicomyces and Helicosporium traditionally have been differentiated by the morphology of their conidia and conidiophores. In this paper we assessed their phylogenetic relationships from ribosomal sequences from ITS, 5.8S and partial LSU regions using maximum parsimony, maximum likelihood and Bayesian analysis. Forty-five isolates from the three genera were closely related and were within the teleomorphic genus Tubeufia sensu Barr (Tubeufiaceae, Ascomycota). Most of the species could be placed in one of the seven clades that each received 78% or greater bootstrap support. However none of the anamorphic genera were monophyletic and all but one of the clades contained species from more than one genus. The 15 isolates of Helicoma were scattered through the phylogeny and appeared in five of the clades. None of the four sections within the genus were monophyletic, although species from Helicoma sect. helicoma were concentrated in Clade A. The Helicosporium species also appeared in five clades. The four Helicomyces species were distributed among three clades. Most of the clades supported by sequence data lacked unifying morphological characters. Traditional characters such as the thickness of the conidial filament and whether conidiophores were conspicuous or reduced proved to be poor predictors of phylogenetic relationships. However some combinations of characters including conidium colour and the presence of lateral, tooth-like conidiogenous cells did appear to be predictive of genetic relationships.

Ascomycota↗

Simple Bayesian analysis in clinical trials: a tutorial.

In this tutorial paper we give a simple Bayesian analysis of data that arise in clinical trials. We consider the case when there are two treatment groups and the response in each group can be assumed to be binomially distributed. We also assume that prior beliefs about the rate parameter in each group can be adequately expressed by a Beta distribution. Using such a model approximate posterior inferences can then be made about the odds ratio between the two groups. We illustrate this methodology by analyzing a randomized trial to assess the benefits of treating patients with carcinoma of the pelvic region (rectum, bladder, colon, cervix) using high-energy fast neutrons as opposed to conventional megavoltage x-rays (photons). In this trial there was prior information about the relative efficacy of neutron therapy based on the beliefs of 10 clinicians. Some of the deficiencies of this simple approach are high-lighted and other approaches to analysis indicated. The paper facilitates practical consideration of a Bayesian approach without the complexities that a fuller analysis necessitates.

Bayes Theorem↗

Approximate maximum entropy joint feature inference consistent with arbitrary lower-order probability constraints: application to statistical classification

We propose a new learning method for discrete space statistical classifiers. Similar to Chow and Liu (1968) and Cheeseman (1983), we cast classification/inference within the more general framework of estimating the joint probability mass function (p.m.f.) for the (feature vector, class label) pair. Cheeseman's proposal to build the maximum entropy (ME) joint p.m.f. consistent with general lower-order probability constraints is in principle powerful, allowing general dependencies between features. However, enormous learning complexity has severely limited the use of this approach. Alternative models such as Bayesian networks (BNs) require explicit determination of conditional independencies. These may be difficult to assess given limited data. Here we propose an approximate ME method, which, like previous methods, incorporates general constraints while retaining quite tractable learning. The new method restricts joint p.m.f. support during learning to a small subset of the full feature space. Classification gains are realized over dependence trees, tree-augmented naive Bayes networks, BNs trained by the Kutato algorithm, and multilayer perceptrons. Extensions to more general inference problems are indicated. We also propose a novel exact inference method when there are several missing features.

Journal Article↗

Phylogenetic relationships in Nicotiana (Solanaceae) inferred from multiple plastid DNA regions.

For Nicotiana, with 75 naturally occurring species (40 diploids and 35 allopolyploids), we produced 4656bp of plastid DNA sequence for 87 accessions and various outgroups. The loci sequenced were trnL intron and trnL-F spacer, trnS-G spacer and two genes, ndhF and matK. Parsimony and Bayesian analyses yielded identical relationships for the diploids, and these are consistent with other data, producing the best-supported phylogenetic assessment currently available for the genus. For the allopolyploids, the line of maternal inheritance is traced via the plastid tree. Nicotiana and the Australian endemic tribe Anthocercideae form a sister pair. Symonanthus is sister to the rest of Anthocercideae. Nicotiana sect. Tomentosae is sister to the rest of the genus. The maternal parent of the allopolyploid species of N. sect. Polydicliae were ancestors of the same species, but the allopolyploids were produced at different times, thus making such sections paraphyletic to their extant diploid relatives. Nicotiana is likely to have evolved in southern South America east of the Andes and later dispersed to Africa, Australia, and southwestern North America.

Base Sequence↗

Toward a curse of dimensionality appropriate (CODA) asymptotic theory for semi-parametric models.

We argue, that due to the curse of dimensionality, there are major difficulties with any pure or smoothed likelihood-based method of inference in designed studies with randomly missing data when missingness depends on a high-dimensional vector of variables. We study in detail a semi-parametric superpopulation version of continuously stratified random sampling. We show that all estimators of the population mean that are uniformly consistent or that achieve an algebraic rate of convergence, no matter how slow, require the use of the selection (randomization) probabilities. We argue that, in contrast to likelihood methods which ignore these probabilities, inverse selection probability weighted estimators continue to perform well achieving uniform n 1/2-rates of convergence. We propose a curse of dimensionality appropriate (CODA) asymptotic theory for inference in non- and semi-parametric models in an attempt to formalize our arguments. We discuss whether our results constitute a fatal blow to the likelihood principle and study the attitude toward these that a committed subjective Bayesian would adopt. Finally, we apply our CODA theory to analyse the effect of the 'curse of dimensionality' in several interesting semi-parametric models, including a model for a two-armed randomized trial with randomization probabilities depending on a vector of continuous pretreatment covariates X. We provide substantive settings under which a subjective Bayesian would ignore the randomization probabilities in analysing the trial data. We then show that any statistician who ignores the randomization probabilities is unable to construct nominal 95 per cent confidence intervals for the true treatment effect that have both: (i) an expected length which goes to zero with increasing sample size; and (ii) a guaranteed expected actual coverage rate of at least 95 per cent over the ensemble of trials analysed by the statistician during his or her lifetime. However, we derive a new interval estimator, depending on the Randomization probabilities, that satisfies (i) and (ii).

Algorithms↗

Robust Bayesian clustering.

A new variational Bayesian learning algorithm for Student-t mixture models is introduced. This algorithm leads to (i) robust density estimation, (ii) robust clustering and (iii) robust automatic model selection. Gaussian mixture models are learning machines which are based on a divide-and-conquer approach. They are commonly used for density estimation and clustering tasks, but are sensitive to outliers. The Student-t distribution has heavier tails than the Gaussian distribution and is therefore less sensitive to any departure of the empirical distribution from Gaussianity. As a consequence, the Student-t distribution is suitable for constructing robust mixture models. In this work, we formalize the Bayesian Student-t mixture model as a latent variable model in a different way from Svensén and Bishop [Svensén, M., & Bishop, C. M. (2005). Robust Bayesian mixture modelling. Neurocomputing, 64, 235-252]. The main difference resides in the fact that it is not necessary to assume a factorized approximation of the posterior distribution on the latent indicator variables and the latent scale variables in order to obtain a tractable solution. Not neglecting the correlations between these unobserved random variables leads to a Bayesian model having an increased robustness. Furthermore, it is expected that the lower bound on the log-evidence is tighter. Based on this bound, the model complexity, i.e. the number of components in the mixture, can be inferred with a higher confidence.

Algorithms↗

Evidence from small-subunit ribosomal RNA sequences for a fungal origin of Microsporidia.

The phylum Microsporidia comprises a species-rich group of minute, single-celled, and intra-cellular parasites. Lacking normal mitochondria and with unique cytology, microsporidians have sometimes been thought to be a lineage of ancient eukaryotes. Although phylogenetic analyses using small-subunit ribosomal RNA (SSU-rRNA) genes almost invariably place the Microsporidia among the earliest branches on the eukaryotic tree, many other molecules suggest instead a relationship with fungi. Using maximum likelihood methods and a diverse SSU-rRNA data set, we have re-evaluated the phylogenetic affiliations of Microsporidia. We demonstrate that tree topologies used to estimate likelihood model parameters can materially affect phylogenetic searches. We present a procedure for reducing this bias: "tree-based site partitioning," in which a comprehensive set of alternative topologies is used to estimate sequence data partitions based on inferred evolutionary rates. This hypothesis-driven approach appears to be capable of utilizing phylogenetic information that is not available to standard likelihood implementations (e.g., approximation to a gamma distribution); we have employed it in maximum likelihood and Bayesian analysis. Applying our method to a phylogenetically diverse SSU-rRNA data set revealed that the early diverging ("deep") placement of Microsporidia typically found in SSU-rRNA trees is no better than a fungal placement, and that the likeliest placement of Microsporidia among non-long-branch eukaryotic taxa is actually within fungi. These results illustrate the importance of hypothesis testing in parameter estimation, provide a way to address certain problems in difficult data sets, and support a fungal origin for the Microsporidia.

Animals↗

Prediction of protein function using protein-protein interaction data.

Assigning functions to novel proteins is one of the most important problems in the postgenomic era. Several approaches have been applied to this problem, including the analysis of gene expression patterns, phylogenetic profiles, protein fusions, and protein-protein interactions. In this paper, we develop a novel approach that employs the theory of Markov random fields to infer a protein's functions using protein-protein interaction data and the functional annotations of protein's interaction partners. For each function of interest and protein, we predict the probability that the protein has such function using Bayesian approaches. Unlike other available approaches for protein annotation in which a protein has or does not have a function of interest, we give a probability for having the function. This probability indicates how confident we are about the prediction. We employ our method to predict protein functions based on "biochemical function," "subcellular location," and "cellular role" for yeast proteins defined in the Yeast Proteome Database (YPD, www.incyte.com), using the protein-protein interaction data from the Munich Information Center for Protein Sequences (MIPS, mips.gsf.de). We show that our approach outperforms other available methods for function prediction based on protein interaction data. The supplementary data is available at www-hto.usc.edu/~msms/ProteinFunction.

Bayes Theorem↗

Attributes and congruence of three molecular data sets: inferring phylogenies among Septoria-related species from woody perennial plants.

To improve our understanding of phylogenetic relationships within the anamorphic genus Septoria, three molecular data sets representing 2,417 bp of nuclear and mitochondrial genes were evaluated. Separate gene analyses and combined analyses were performed using first, the maximum parsimony criterion and second, a Bayesian framework. The homogeneity of data partitions was evaluated via a combination of homogeneity partition tests and tree topology incongruence tests before conducting combined analyses. A last incongruence re-evaluation using partitioned Bremer support was performed on the combined tree, which corroborated the previous estimates. After each separate data set attributes were examined, simple explanations were advocated as the causes of the significant incongruences detected. The analysis of multiple gene partitions showed unprecedented phylogenetic resolution within the genus Septoria that supported the results from previously published single gene phylogenies. Specifically, we have delimited distinct but closely related species representing monophyletic groups that frequently correlated with their respective host families. Conversely, the occurrence of well-supported groups including closely related but distinct molecular taxa sampled on unrelated host-plants allowed us to reject, in these particular cases, the co-evolutionary concept expected between a parasite and its host and to discuss alternative evolutionary models recently proposed for these pathogens.

Ascomycota↗

Finding optimal models for small gene networks.

Finding gene networks from microarray data has been one focus of research in recent years. Given search spaces of super-exponential size, researchers have been applying heuristic approaches like greedy algorithms or simulated annealing to infer such networks. However, the accuracy of heuristics is uncertain, which--in combination with the high measurement noise of microarrays--makes it very difficult to draw conclusions from networks estimated by heuristics. We present a method that finds optimal Bayesian networks of considerable size and show first results of the application to yeast data. Having removed the uncertainty due to the heuristic methods, it becomes possible to evaluate the power of different statistical models to find biologically accurate networks.

Algorithms↗

Prediction of protein function using protein-protein interaction data.

Assigning functions to novel proteins is one of the most important problems in the post-genomic era. Several approaches have been applied to this problem, including analyzing gene expression patterns, phylogenetic profiles, protein fusions and protein-protein interactions. We develop a novel approach that applies the theory of Markov random fields to infer a protein's functions using protein-protein interaction data and the functional annotations of its interaction protein partners. For each function of interest and a protein, we predict the probability that the protein has that function using Bayesian approaches. Unlike in other available approaches for protein annotation where a protein has or does not have a function of interest, we give a probability for having the function. This probability indicates how confident we are about the prediction. We apply our method to predict cellular functions (43 categories including a category "others") for yeast proteins defined in the Yeast Proteome Database (YPD), using the protein-protein interaction data from the Munich Information Center for Protein Sequences (MIPS, http://mips.gsf.de). We show that our approach outperforms other available methods for function prediction based on protein interaction data.

Amino Acid Sequence↗

Assumed and inferred spatial structure of populations: the Scandinavian brown bears revisited.

We reanalysed the spatial structure of the Scandinavian brown bear (Ursus arctos) population based on multilocus genotypes. We used data from a former study that had presumed a priori a specific population subdivision based on four subpopulations. Using two independent methods (neighbour-joining trees and Bayesian assignment tests), we analysed the data without any prior presumption about the spatial structure. A subdivision of the population into three subpopulations emerged from our study. The genetic pattern of these subpopulations matched the three geographical clusters of individuals present in the population. We recommend considering the Scandinavian brown bear population as consisting of three (instead of four) subpopulations. Our results underline the importance of determining genetic structure from the data, without presupposing a structure, even when there seems to be good reason to do so.

Animals↗