Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Divergent gene copies in the asexual class Bdelloidea (Rotifera) separated before the bdelloid radiation or within bdelloid families.

Rotifers of the asexual class Bdelloidea are unusual in possessing two or more divergent copies of every gene that has been examined. Phylogenetic analysis of the heat-shock gene hsp82 and the TATA-box-binding protein gene tbp in multiple bdelloid species suggested that for each gene, each copy belonged to one of two lineages that began to diverge before the bdelloid radiation. Such gene trees are consistent with the two lineages having descended from former alleles that began to diverge after meiotic segregation ceased or from subgenomes of an alloploid ancestor of the bdelloids. However, the original analyses of bdelloid gene-copy divergence used only a single outgroup species and were based on parsimony and neighbor joining. We have now used maximum likelihood and Bayesian inference methods and, for hsp82, multiple outgroups in an attempt to produce more robust gene trees. Here we report that the available data do not unambiguously discriminate between gene trees that root the origin of hsp82 and tbp copy divergence before the bdelloid radiation and those which indicate that the gene copies began to diverge within bdelloid families. The remarkable presence of multiple diverged gene copies in individual genomes is nevertheless consistent with the loss of sex in an ancient ancestor of bdelloids.

Animals↗

Isotropic probability measures in infinite-dimensional spaces.

Every isotropic probability measure on the space R(infinity) of real sequences x = (x(1), x(2),...) is a convex combination of the measure concentrated at 0 and a member of I(0)(R(infinity)), the set of all isotropic probability measures p(infinity) on R(infinity) with p(infinity)({0}) = 0. Each p(infinity) [unk] I(0)(R(infinity)) is completely determined by any one of its finite-dimensional marginal distributions p(n). Each p(n) has a density function f(n) with dp(n)(x(1),..., x(n)) = dx(1)... dx(n)f(n)(x(1) (2) +... + x(n) (2)). Each f(n) is completely monotone in 0 < xi < infinity (hence analytic in the right complex xi half-plane), and pi(n/2)Gamma(n/2)(-1) (0) (infinity)dxi xi(n/2-1)f(n)(xi) = 1. Every f that satisfies these two conditions is f(n) for a unique p(infinity) [unk] I(0)(R(infinity)). Hence the equation pi(xi) (infinity)dzeta f(2)(zeta) = (0) (infinity)dmu (t)e(-txi) defines a bijection between I(0)(R(infinity)) and the set of all probability measures mu on 0 </= t < infinity. If p(infinity) [unk] I(0)(R(infinity)) then p(infinity)({x: Sigma(i=1) (infinity)x(i) (2) < infinity}) = 0, so p(infinity) is not a "softened" or "fuzzy" version of the inequality Sigma(i=1) (infinity)x(i) (2) </= 1. If the prior information in a linear inverse problem consists of this inequality and nothing else, stochastic inversion and Bayesian inference are both unsuitable inversion techniques.

Journal Article↗

Automated high resolution optical mapping using arrayed, fluid-fixed DNA molecules.

New mapping approaches construct ordered restriction maps from fluorescence microscope images of individual, endonuclease-digested DNA molecules. In optical mapping, molecules are elongated and fixed onto derivatized glass surfaces, preserving biochemical accessibility and fragment order after enzymatic digestion. Measurements of relative fluorescence intensity and apparent length determine the sizes of restriction fragments, enabling ordered map construction without electrophoretic analysis. The optical mapping system reported here is based on our physical characterization of an effect using fluid flows developed within tiny, evaporating droplets to elongate and fix DNA molecules onto derivatized surfaces. Such evaporation-driven molecular fixation produces well elongated molecules accessible to restriction endonucleases, and notably, DNA polymerase I. We then developed the robotic means to grid DNA spots in well defined arrays that are digested and analyzed in parallel. To effectively harness this effect for high-throughput genome mapping, we developed: (i) machine vision and automatic image acquisition techniques to work with fixed, digested molecules within gridded samples, and (ii) Bayesian inference approaches that are used to analyze machine vision data, automatically producing high-resolution restriction maps from images of individual DNA molecules. The aggregate significance of this work is the development of an integrated system for mapping small insert clones allowing biochemical data obtained from engineered ensembles of individual molecules to be automatically accumulated and analyzed for map construction. These approaches are sufficiently general for varied biochemical analyses of individual molecules using statistically meaningful population sizes.

Animals↗

True and false gharials: a nuclear gene phylogeny of crocodylia.

The phylogeny of Crocodylia offers an unusual twist on the usual molecules versus morphology story. The true gharial (Gavialis gangeticus) and the false gharial (Tomistoma schlegelii), as their common names imply, have appeared in all cladistic morphological analyses as distantly related species, convergent upon a similar morphology. In contrast, all previous molecular studies have shown them to be sister taxa. We present the first phylogenetic study of Crocodylia using a nuclear gene. We cloned and sequenced the c-myc proto-oncogene from Alligator mississippiensis to facilitate primer design and then sequenced an 1,100-base pair fragment that includes both coding and noncoding regions and informative indels for one species in each extant crocodylian genus and six avian outgroups. Phylogenetic analyses using parsimony, maximum likelihood, and Bayesian inference all strongly agreed on the same tree, which is identical to the tree found in previous molecular analyses: Gavialis and Tomistoma are sister taxa and together are the sister group of Crocodylidae. Kishino-Hasegawa tests rejected the morphological tree in favor of the molecular tree. We excluded long-branch attraction and variation in base composition among taxa as explanations for this topology. To explore the causes of discrepancy between molecular and morphological estimates of crocodylian phylogeny, we examined puzzling features of the morphological data using a priori partitions of the data based on anatomical regions and investigated the effects of different coding schemes for two obvious morphological similarities of the two gharials.

Alligators and Crocodiles↗

Complex biogeographic patterns in Androsace (Primulaceae) and related genera: evidence from phylogenetic analyses of nuclear internal transcribed spacer and plastid trnL-F sequences.

We conducted phylogenetic analyses of Androsace and the closely related genera Douglasia, Pomatosace, and Vitaliana using DNA sequences of the nuclear internal transcribed spacer (ITS) and the plastid trnL-F region. Analyses using maximum parsimony and Bayesian inference yield congruent relationships among several major lineages found. These lineages largely disagree with previously recognized taxonomic groups. Most notably, (1) Androsace sect. Andraspis, comprising the short-lived taxa, is highly polyphyletic; (2) Pomatosace constitutes a separate phylogenetic lineage within Androsace; and (3) Douglasia and Vitaliana nest within Androsace sect. Aretia. Our results suggest multiple origins of the short-lived lifeform and a possible reversal from annual or biennial to perennial habit at the base of a group that now contains mostly perennial high mountain or arctic taxa. The group containing Androsace sect. Aretia, Douglasia, and Vitaliana includes predominantly high alpine and arctic taxa with an arctic-alpine distribution, but is not found in the European and northeastern American Arctic or in Central and East Asia. This group probably originated in Europe in the Pliocene, from where it reached the amphi-Beringian region in the Pleistocene or late Pliocene.

Base Sequence↗

Phylogeny of Eunicida (Annelida) and exploring data congruence using a partition addition bootstrap alteration (PABA) approach.

Even though relationships within Annelida are poorly understood, Eunicida is one of only a few major annelid lineages well supported by morphology. The seven recognized eunicid families possess sclerotized jaws that include mandibles and a maxillary apparatus. The maxillary apparatuses vary in shape and number of elements, and three main types are recognized in extant taxa: ctenognath, labidognath, and prionognath. Ctenognath jaws are usually considered to represent the plesiomorphic state of Eunicida, whereas taxa with labidognath and prionognath are thought to form a derived monophyletic assemblage. However, this hypothesis has never been tested in a statistical framework even though it holds considerable importance for understanding annelid phylogeny and possibly lophotrochozoan evolution because Eunicida has the best annelid fossil record. Therefore, we used maximum likelihood and Bayesian inference approaches to reconstruct Eunicida phylogeny using sequence data from nuclear 18S and 28S rDNA genes and mitochondrial 16S rDNA and cytochrome c oxidase subunit I genes. Additionally, we conducted three different tests to investigate suitability of combining data sets. Incongruence length difference (ILD) and Shimodaira-Hasegawa (SH) test comparisons of resultant trees under different data partitions have been widely used previously but do not give a good indication as to which nodes may be causing the conflict. Thus, we developed a partition addition bootstrap alteration (PABA) approach that evaluates congruence or conflict for any given node by determining how bootstrap scores are altered when different data partitions are added. PABA shows the contribution of each partition to the phylogeny obtained in the combined analysis. Generally, the ILD test performed worse than the other approaches in detecting incongruence. Both PABA and the SH approach indicated the 28S and COI data sets add conflicting signal, but PABA is more informative for elucidating which data partition may be misleading at a given node. All our analyses indicate that the monophyly of the labidognath/prionognath taxa and even a labidognath clade (i.e., a "Eunicidae"/Onuphidae/Lumbrineridae clade) is significantly rejected. We show that the definition of both the labidognath and ctenognath jaw type does not address adequately the variation within Eunicida and thus misleads our current evolutionary understanding. Based on the presented results a symmetric maxillary apparatus with a carrier and four to six maxillae is most likely the plesiomorphic condition for Eunicida. [COI; conflicting data; fossil record; ILD; Jaw Evolution; molecular phylogeny; rDNA; SH test.].

Animals↗

Mitochondrial genome characteristics and phylogenetic analysis of Ramaria longispora.

This study, for the first time, assembled and annotated the complete mitochondrial genome of R.&#xa0;longispora using high-throughput sequencing technology. The genome is a circular molecule with a total length of 157,712&#x2009;bp and a GC content of 31.55%. It encodes 71 genes, including 15 core protein-coding genes (PCGs), 25 transfer RNA (tRNA) genes, 2 ribosomal RNA (rRNA) genes, 5 free-stranding open reading frames (ORFs), and 24 intronic ORFs. Among these, most free-stranding ORFs have unknown functions but include a DNA polymerase gene, while the intronic ORFs primarily encode LAGLIDADG and GIY-YIG endonucleases. The mitochondrial genome contains 39 introns. Phylogenetic analyses based on 15 core PCGs using Bayesian inference (BI) and maximum likelihood (ML) methods revealed that this R. longispora is most closely related to Ramaria flavescens and Ramaria ichnusensis. This study provides foundational data for mitochondrial genome research in the Ramaria genus and offers important references for taxonomic and evolutionary studies of this group.

Mitochondrial genome↗

Monophyletic relationship between severe acute respiratory syndrome coronavirus and group 2 coronaviruses.

Although primary genomic analysis has revealed that severe acute respiratory syndrome coronavirus (SARS CoV) is a new type of coronavirus, the different protein trees published in previous reports have provided no conclusive evidence indicating the phylogenetic position of SARS CoV. To clarify the phylogenetic relationship between SARS CoV and other coronaviruses, we compiled a large data set composed of 7 concatenated protein sequences and performed comprehensive analyses, using the maximum-likelihood, Bayesian-inference, and maximum-parsimony methods. All resulting phylogenetic trees displayed an identical topology and supported the hypothesis that the relationship between SARS CoV and group 2 CoVs is monophyletic. Relationships among all major groups were well resolved and were supported by all statistical analyses.

Animals↗

Genetic analysis of rubella viruses found in the United States between 1966 and 2004: evidence that indigenous rubella viruses have been eliminated.

Wild-type rubella viruses are genetically classified into 2 clades and 10 intraclade genotypes, of which 3 are provisional. The genotypes of 118 viruses from the United States were determined by sequencing part of the E1 coding region of these viruses and comparing the resulting sequences with reference sequences for each genotype, using the Bayesian inference program MRBAYES. Three genotypes of rubella viruses were found in the United States too infrequently to be considered for indigenous transmission. A fourth genotype was found frequently until 1981, and a fifth genotype was found frequently until 1988, but neither was obtained from nonimported cases after 1988. A sixth genotype was found frequently during 1996-2000, likely because of multiple importations from neighboring countries. The results of the present genetic analysis of rubella viruses found in the United States are consistent with elimination of indigenous viruses by 2001, the year when rubella was considered to be eliminated on the basis of epidemiological evidence.

Bayes Theorem↗

Spatiotemporal noise covariance estimation from limited empirical magnetoencephalographic data.

The performance of parametric magnetoencephalography (MEG) and electroencephalography (EEG) source localization approaches can be degraded by the use of poor background noise covariance estimates. In general, estimation of the noise covariance for spatiotemporal analysis is difficult mainly due to the limited noise information available. Furthermore, its estimation requires a large amount of storage and a one-time but very large (and sometimes intractable) calculation or its inverse. To overcome these difficulties, noise covariance models consisting of one pair or a sum of multi-pairs of Kronecker products of spatial covariance and temporal covariance have been proposed. However, these approaches cannot be applied when the noise information is very limited, i.e., the amount of noise information is less than the degrees of freedom of the noise covariance models. A common example of this is when only averaged noise data are available for a limited prestimulus region (typically at most a few hundred milliseconds duration). For such cases, a diagonal spatiotemporal noise covariance model consisting of sensor variances with no spatial or temporal correlation has been the common choice for spatiotemporal analysis. In this work, we propose a different noise covariance model which consists of diagonal spatial noise covariance and Toeplitz temporal noise covariance. It can easily be estimated from limited noise information, and no time-consuming optimization and data-processing are required. Thus, it can be used as an alternative choice when one-pair or multi-pair noise covariance models cannot be estimated due to lack of noise information. To verify its capability we used Bayesian inference dipole analysis and a number of simulated and empirical datasets. We compared this covariance model with other existing covariance models such as conventional diagonal covariance, one-pair and multi-pair noise covariance models, when noise information is sufficient to estimate them. We found that our proposed noise covariance model yields better localization performance than a diagonal noise covariance, while it performs slightly worse than one-pair or multi-pair noise covariance models - although these require much more noise information. Finally, we present some localization results on median nerve stimulus empirical MEG data for our proposed noise covariance model.

Algorithms↗

Bayesian segmentation of protein secondary structure.

We present a novel method for predicting the secondary structure of a protein from its amino acid sequence. Most existing methods predict each position in turn based on a local window of residues, sliding this window along the length of the sequence. In contrast, we develop a probabilistic model of protein sequence/structure relationships in terms of structural segments, and formulate secondary structure prediction as a general Bayesian inference problem. A distinctive feature of our approach is the ability to develop explicit probabilistic models for alpha-helices, beta-strands, and other classes of secondary structure, incorporating experimentally and empirically observed aspects of protein structure such as helical capping signals, side chain correlations, and segment length distributions. Our model is Markovian in the segments, permitting efficient exact calculation of the posterior probability distribution over all possible segmentations of the sequence using dynamic programming. The optimal segmentation is computed and compared to a predictor based on marginal posterior modes, and the latter is shown to provide significant improvement in predictive accuracy. The marginalization procedure provides exact secondary structure probabilities at each sequence position, which are shown to be reliable estimates of prediction uncertainty. We apply this model to a database of 452 nonhomologous structures, achieving accuracies as high as the best currently available methods. We conclude by discussing an extension of this framework to model nonlocal interactions in protein structures, providing a possible direction for future improvements in secondary structure prediction accuracy.

Algorithms↗

DMLE+: Bayesian linkage disequilibrium gene mapping.

SUMMARY: The program DMLE+ allows Bayesian inference of the location of a gene carrying a mutation influencing a discrete trait (such as a disease) and/or other parameters of interest (such as mutation age) based on the observed linkage disequilibrium at multiple genetic markers. DMLE+ uses either individual marker genotypes, or haplotypes, integrates over uncertain population allele frequencies, and can incorporate prior information about gene location from an annotated human genome sequence. AVAILABILITY: DMLE+ is available in both Windows GUI and portable UNIX command line versions at http://dmle.org.

Bayes Theorem↗

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs↗

Biomarker discovery in microarray gene expression data with Gaussian processes.

MOTIVATION: In clinical practice, pathological phenotypes are often labelled with ordinal scales rather than binary, e.g. the Gleason grading system for tumour cell differentiation. However, in the literature of microarray analysis, these ordinal labels have been rarely treated in a principled way. This paper describes a gene selection algorithm based on Gaussian processes to discover consistent gene expression patterns associated with ordinal clinical phenotypes. The technique of automatic relevance determination is applied to represent the significance level of the genes in a Bayesian inference framework. RESULTS: The usefulness of the proposed algorithm for ordinal labels is demonstrated by the gene expression signature associated with the Gleason score for prostate cancer data. Our results demonstrate how multi-gene markers that may be initially developed with a diagnostic or prognostic application in mind are also useful as an investigative tool to reveal associations between specific molecular and cellular events and features of tumour physiology. Our algorithm can also be applied to microarray data with binary labels with results comparable to other methods in the literature.

Algorithms↗

Proper multivariate conditional autoregressive models for spatial data analysis.

In the past decade conditional autoregressive modelling specifications have found considerable application for the analysis of spatial data. Nearly all of this work is done in the univariate case and employs an improper specification. Our contribution here is to move to multivariate conditional autoregressive models and to provide rich, flexible classes which yield proper distributions. Our approach is to introduce spatial autoregression parameters. We first clarify what classes can be developed from the family of Mardia (1988) and contrast with recent work of Kim et al. (2000). We then present a novel parametric linear transformation which provides an extension with attractive interpretation. We propose to employ these models as specifications for second-stage spatial effects in hierarchical models. Two applications are discussed; one for the two-dimensional case modelling spatial patterns of child growth, the other for a four-dimensional situation modelling spatial variation in HLA-B allele frequencies. In each case, full Bayesian inference is carried out using Markov chain Monte Carlo simulation.

Alleles↗

Numerical equivalence of imputing scores and weighted estimators in regression analysis with missing covariates.

Imputation, weighting, direct likelihood, and direct Bayesian inference (Rubin, 1976) are important approaches for missing data regression. Many useful semiparametric estimators have been developed for regression analysis of data with missing covariates or outcomes. It has been established that some semiparametric estimators are asymptotically equivalent, but it has not been shown that many are numerically the same. We applied some existing methods to a bladder cancer case-control study and noted that they were the same numerically when the observed covariates and outcomes are categorical. To understand the analytical background of this finding, we further show that when observed covariates and outcomes are categorical, some estimators are not only asymptotically equivalent but also actually numerically identical. That is, although their estimating equations are different, they lead numerically to exactly the same root. This includes a simple weighted estimator, an augmented weighted estimator, and a mean-score estimator. The numerical equivalence may elucidate the relationship between imputing scores and weighted estimation procedures.

Case-Control Studies↗

Pilot study of an expert system adviser for controlling general anaesthesia.

RESAC (real time expert system for advice and control), has been developed to advise on the concentration of inhaled volatile anaesthetics during anaesthesia. It merges clinical information and on-line measurements using Bayesian inference and fuzzy logic. This paper describes a clinical trial in seven patients after initial development of the knowledge base and the human computer interface. Evaluation of the performance of RESAC included analysis of questionnaire responses. Anaesthetists were confident enough to follow the dosage advice given by RESAC in most of the patients.

Adult↗

A compound poisson process for relaxing the molecular clock.

The molecular clock hypothesis remains an important conceptual and analytical tool in evolutionary biology despite the repeated observation that the clock hypothesis does not perfectly explain observed DNA sequence variation. We introduce a parametric model that relaxes the molecular clock by allowing rates to vary across lineages according to a compound Poisson process. Events of substitution rate change are placed onto a phylogenetic tree according to a Poisson process. When an event of substitution rate change occurs, the current rate of substitution is modified by a gamma-distributed random variable. Parameters of the model can be estimated using Bayesian inference. We use Markov chain Monte Carlo integration to evaluate the posterior probability distribution because the posterior probability involves high dimensional integrals and summations. Specifically, we use the Metropolis-Hastings-Green algorithm with 11 different move types to evaluate the posterior distribution. We demonstrate the method by analyzing a complete mtDNA sequence data set from 23 mammals. The model presented here has several potential advantages over other models that have been proposed to relax the clock because it is parametric and does not assume that rates change only at speciation events. This model should prove useful for estimating divergence times when substitution rates vary across lineages.

Bayes Theorem↗