Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “variational inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Probabilistic inference of transcription factor concentrations and gene-specific regulatory activities.

MOTIVATION: Quantitative estimation of the regulatory relationship between transcription factors and genes is a fundamental stepping stone when trying to develop models of cellular processes. Recent experimental high-throughput techniques, such as Chromatin Immunoprecipitation (ChIP) provide important information about the architecture of the regulatory networks in the cell. However, it is very difficult to measure the concentration levels of transcription factor proteins and determine their regulatory effect on gene transcription. It is therefore an important computational challenge to infer these quantities using gene expression data and network architecture data. RESULTS: We develop a probabilistic state space model that allows genome-wide inference of both transcription factor protein concentrations and their effect on the transcription rates of each target gene from microarray data. We use variational inference techniques to learn the model parameters and perform posterior inference of protein concentrations and regulatory strengths. The probabilistic nature of the model also means that we can associate credibility intervals to our estimates, as well as providing a tool to detect which binding events lead to significant regulation. We demonstrate our model on artificial data and on two yeast datasets in which the network structure has previously been obtained using ChIP data. Predictions from our model are consistent with the underlying biology and offer novel quantitative insights into the regulatory structure of the yeast cell. AVAILABILITY: MATLAB code is available from http://umber.sbs.man.ac.uk/resources/puma

Algorithms↗

Interactions between diverse proteinoids and microspheres in simulation of primordial evolution.

Experiments demonstrating an incorporation of different enzymelike activities into a single preparation of proteinoid microspheres provide a conceptual basis for the primitive lengthening of protometabolic pathways. An enhancement of one enzymelike activity by another proteinoid in the same microsphere has been found. This effect, plus the pathway-lengthening propensity of combinations of microspheres, indicates selective advantages contributing to adaptive protoselection. Data reported in this paper also bring into purview the concept of internally controlled variation. Inferences are derived for the origin of protosexuality in protocells. When allowance is made for a closer relationship to the environment than that needed in contemporary selection, the fundamental mechanistic requirements of protoevolution are regarded as met by the proteinoid microsphere.

Catalysis↗

Comparative genetic mutation frequencies based on amino acid composition differences.

Genetic variation inferred from large-scale amino acid composition comparisons among genomes and chromosomes of several species, Saccharomyces cerevisiae, Drosophila melanogaster, Ceanorhabditis elegans, H. sapiens, is shown to be correlated (highest, r(2)=0.9855, p<0.01) with reported mutation rates for various genes in these species. This study, based largely on pseudogene data, helps to establish reference mutation frequencies that are likely to be representative of overall genome mutation rates in each of the species examined, and provides further insight into heterogeneity of mutation rates among genomes.

Amino Acid Sequence↗

Exploiting Omic Data to Advance Predictive Ecotoxicology.

Predicting species-specific chemical sensitivity using in silico approaches has the potential to transform environmental risk assessment, conservation, and biomonitoring, while reducing, and ultimately replacing, animal testing. Genomic and transcriptomic data capture extensive sensitivity-relevant variation, including differences in molecular targets, xenobiotic metabolism, and damage mitigation pathways. Large-scale sequencing initiatives therefore offer an unprecedented opportunity to address ecotoxicology's "too many species" problem. Although existing omic-based predictive tools provide proof of concept, they have so far been applied to a narrow set of relatively straightforward prediction scenarios. To achieve broader applicability, current and future tools must be firmly grounded in the diverse molecular mechanisms underlying differential chemical responses. Here, we critically evaluate the emerging field of predicting species sensitivity using molecular variation inferred from omic data. We analyze the strengths and limitations of current omic-based approaches and identify major sequence and ecotoxicological data gaps, as well as critical bioinformatic challenges. We then review the current knowledge of how molecular biology underlies differential chemical sensitivity, outlining research paths to allow the next generation of sensitivity prediction tools to exploit ever expanding omic data.

Ecotoxicology↗

NExON-Bayes: a Bayesian approach to network estimation informed by ordinal covariates.

MOTIVATION: In heterogeneous disease settings, accounting for intrinsic sample variability is crucial for obtaining reliable and interpretable omic network estimates. However, most graphical model analyses of biomedical data assume homogeneous conditional dependence structures, potentially leading to misleading conclusions. To address this, we propose a joint Gaussian graphical model that leverages sample-level ordinal covariates (e.g. disease stage) to account for heterogeneity and improve the estimation of partial correlation structures. RESULTS: Our modelling framework, called NExON-Bayes, extends the graphical spike-and-slab framework to account for ordinal covariates, jointly estimating their relevance to the graph structure and leveraging them to improve the accuracy of network estimation. To scale to high-dimensional omic settings, we develop an efficient variational inference algorithm tailored to our model. Through simulations, we demonstrate that our method outperforms the vanilla graphical spike-and-slab (with no covariate information), as well as other state-of-the-art network approaches which exploit covariate information. Applying our method to reverse phase protein array data from patients diagnosed with stage I, II or III breast carcinoma, we estimate the behaviour of proteomic networks as cancer progresses. Our model provides insights not only through inspection of the estimated proteomic networks, but also of the estimated ordinal covariate dependencies of key groups of proteins within those networks, offering a comprehensive understanding of how biological pathways shift across disease stages. AVAILABILITY AND IMPLEMENTATION: A user-friendly R package for NExON-Bayes with tutorials is available on Github at github.com/jf687/NExON, and archived at https://doi.org/10.5281/zenodo.20312938. The source of the dataset used is cited in the relevant section.

Bayes Theorem↗

Probe-level measurement error improves accuracy in detecting differential gene expression.

MOTIVATION: Finding differentially expressed genes is a fundamental objective of a microarray experiment. Numerous methods have been proposed to perform this task. Existing methods are based on point estimates of gene expression level obtained from each microarray experiment. This approach discards potentially useful information about measurement error that can be obtained from an appropriate probe-level analysis. Probabilistic probe-level models can be used to measure gene expression and also provide a level of uncertainty in this measurement. This probe-level measurement error provides useful information which can help in the identification of differentially expressed genes. RESULTS: We propose a Bayesian method to include probe-level measurement error into the detection of differentially expressed genes from replicated experiments. A variational approximation is used for efficient parameter estimation. We compare this approximation with MAP and MCMC parameter estimation in terms of computational efficiency and accuracy. The method is used to calculate the probability of positive log-ratio (PPLR) of expression levels between conditions. Using the measurements from a recently developed Affymetrix probe-level model, multi-mgMOS, we test PPLR on a spike-in dataset and a mouse time-course dataset. Results show that the inclusion of probe-level measurement error improves accuracy in detecting differential gene expression. AVAILABILITY: The MAP approximation and variational inference described in this paper have been implemented in an R package pplr. The MCMC method is implemented in Matlab. Both software are available from http://umber.sbs.man.ac.uk/resources/puma.

Algorithms↗

Southward migration of the intertropical convergence zone through the Holocene.

Titanium and iron concentration data from the anoxic Cariaco Basin, off the Venezuelan coast, can be used to infer variations in the hydrological cycle over northern South America during the past 14,000 years with subdecadal resolution. Following a dry Younger Dryas, a period of increased precipitation and riverine discharge occurred during the Holocene "thermal maximum." Since approximately 5400 years ago, a trend toward drier conditions is evident from the data, with high-amplitude fluctuations and precipitation minima during the time interval 3800 to 2800 years ago and during the "Little Ice Age." These regional changes in precipitation are best explained by shifts in the mean latitude of the Atlantic Intertropical Convergence Zone (ITCZ), potentially driven by Pacific-based climate variability. The Cariaco Basin record exhibits strong correlations with climate records from distant regions, including the high-latitude Northern Hemisphere, providing evidence for global teleconnections among regional climates.

Journal Article↗

The effects of rate variation on ancestral inference in the coalescent.

We describe a Markov chain Monte Carlo approach for assessing the role of site-to-site rate variation in the analysis of within-population samples of DNA sequences using the coalescent. Our framework is a Bayesian one. We discuss methods for assessing the goodness-of-fit of these models, as well as problems concerning the separate estimation of effective population size and mutation rate. Using a mitochondrial data set for illustration, we show that ancestral inference concerning coalescence times can be dramatically affected if rate variation is ignored.

Algorithms↗

Tasks in statistical inference for studying variation in medicine.

When studying variation in medicine, traditional hypothesis-testing procedures are too limited to obtain useful inferences except in special situations. More generally, full probability modelling is necessary. Even a relatively simple example can illustrate this point rather dramatically. The development and application of full-probability methods for medical problems comprise exciting areas for statistical and medical researchers, especially if working together.

Health Services Research↗

Computational anatomy and neuropsychiatric disease: probabilistic assessment of variation and statistical inference of group difference, hemispheric asymmetry, and time-dependent change.

Three components of computational anatomy (CA) are reviewed in this paper: (i) the computation of large-deformation maps, that is, for any given coordinate system representations of two anatomies, computing the diffeomorphic transformation from one to the other; (ii) the computation of empirical probability laws of anatomical variation between anatomies; and (iii) the construction of inferences regarding neuropsychiatric disease states. CA utilizes spatial-temporal vector field information obtained from large-deformation maps to assess anatomical variabilities and facilitate the detection and quantification of abnormalities of brain structure in subjects with neuropsychiatric disorders. Neuroanatomical structures are divided into two types: subcortical structures-gray matter (GM) volumes enclosed by a single surface-and cortical mantle structures-anatomically distinct portions of the cerebral cortical mantle layered between the white matter (WM) and cerebrospinal fluid (CSF). Because of fundamental differences in the geometry of these two types of structures, image-based large-deformation high-dimensional brain mapping (HDBM-LD) and large-deformation diffeomorphic metric matching (LDDMM) were developed for the study of subcortical structures and labeled cortical mantle distance mapping (LCMDM) was developed for the study of cortical mantle structures. Studies of neuropsychiatric disorders using CA usually require the testing of hypothesized group differences with relatively small numbers of subjects per group. Approaches that increase the power for testing such hypotheses include methods to quantify the shapes of individual structures, relationships between the shapes of related structures (e.g., asymmetry), and changes of shapes over time. Promising preliminary studies employing these approaches to studies of subjects with schizophrenia and very mild to mild Alzheimer's disease (AD) are presented.

Algorithms↗

Phylogeny of "Oxycanus" lineages of hepialid moths from New Zealand inferred from sequence variation in the mtDNA COI and II gene regions.

The phylogeny of the New Zealand hepialid moths was estimated from a 527-bp nucleotide sequence from the mitochondrial DNA cytochrome oxidase subunit I and II gene regions. New haplotypes were identified for Wiseana cervinata, W. copularis, and W. signata. Phylogenetic reconstructions using maximum parsimony and maximum likelihood methods indicated that the four hepialid lineages Aenetus, Aoraia, "Oxycanus" Cladoxycanus, and "Oxycanus" s. str. hypothesized by Dugdale (1994) based on a morphological taxonomic revision were monophyletic within New Zealand. Addition of exemplars from the Australian genera Fraus, Jeana, Oxycanus, and Trictena to the data set tentatively support the monophyly of the New Zealand "Oxycanus" lineages. Estimated times of divergence for the genus Wiseana taxa fitted well with known geological events and suggest that the genus may have diverged 1-1.5 mya.

Animals↗

Phenotypic evolution and hidden speciation in Candidula unifasciata ssp. (Helicellinae, Gastropoda) inferred by 16S variation and quantitative shell traits.

In an effort to link quantitative morphometric information with molecular data on the population level, we have analysed 19 populations of the conchologically variable land snail Candidula unifasciata from across the species range for variation in quantitative shell traits and at the mitochondrial 16S ribosomal (r)DNA locus. In genetic analysis, including 21 additional populations, we observed two fundamental haplotype clades with an average pairwise sequence divergence of 0.209 +/- 0.009 between clades compared to 0.017 +/- 0.012 within clades, suggesting the presence of two different evolutionary lineages. Integrating additional shell material from the Senckenberg Malacological Collection, a highly significant discriminant analysis on the morphological shell traits with fundamental haplotype clades as grouping variable suggested that the less frequent haplotype corresponds to the described subspecies C. u. rugosiuscula, which we propose to regard as a distinct species. Both taxa were highly subdivided genetically (FST = 0.648 and 0.777 P < 0.001). This was contrasted by the partition of morphological variance, where only 29.6% and 21.9% of the variance were distributed among populations, respectively. In C. unifasciata, no significant association between population pairwise FST estimates and corresponding morphological fixation indices could be detected, indicating independent evolution of the two character sets. Partial least square analysis of environmental factors against shell trait variables in C. u. unifasciata revealed significant correlations between environmental factors and certain quantitative shell traits, whose potential adaptational values are discussed.

Animals↗

Evolutionary inferences from DNA variation at the 6-phosphogluconate dehydrogenase locus in natural populations of drosophila: selection and geographic differentiation.

Several allozyme-coding genes in Drosophila melanogaster show patterns suggesting that polymorphisms at these loci are targets of balancing selection. An important question is whether these genes have similar distributions of underlying DNA sequence variation which would indicate similar evolutionary processes occurring in this class of loci. One such locus, 6-phosphogluconate dehydrogenase (Pgd), has previously been shown to exhibit clinal variation for Fast/Slow electromorph variation in the United States and Australia, unusually large electromorph frequency differences between the United States and Africa, and other patterns indicative of selection. We measured four-cutter DNA restriction site and allozyme variation at Pgd among 142 D. melanogaster X chromosomes collected from several geographic regions including North Carolina, California, and Zimbabwe (Africa). We also sequenced a representative sample of 13 D. melanogaster Pgd genes collected in North Carolina and a single copy of Pgd from the sibling species, Drosophila simulans. While some population genetic models predict excess DNA polymorphism in genes which are targets of balancing selection, the D. melanogaster samples from the United States had significantly reduced levels of DNA polymorphism and extraordinarily high levels of linkage disequilibrium, providing evidence of hitchhiking effects of advantageous mutants at Pgd or at linked sites. Therefore, while selection has probably influenced the distribution of DNA variation at Pgd, the precise nature of these selective events remains obscure. Since the Pgd region appears to have low rates of crossing over, the reduced level of variation at this locus supports the idea that recombination rates are important determinants of levels of DNA polymorphism in natural populations. Furthermore, while patterns of allozyme variation are very similar at Pgd and Adh, the DNA data show that the evolutionary histories of these genes are dramatically different. We observed extensive differences in the amount and distribution of variation in D. melanogaster Pgd samples from the United States and Zimbabwe which cannot be explained by differential selection on the Fast/Slow polymorphism in these two geographic regions. Thus, genetic drift among partially isolated populations has also been an important factor in determining the distribution of variation at Pgd in D. melanogaster. Finally, we assayed four-cutter variation at Pgd in a sample of 19 D. simulans X chromosomes and observed reduced levels of DNA variability and high levels of linkage disequilibrium. These patterns are consistent with predictions of some hitchhiking models.

Animals↗

Population dynamics inferred from temporal variation at microsatellite loci in the selfing snail Bulinus truncatus.

We analyzed short-term forces acting on the genetics of subdivided populations based on a temporal survey of the microsatellite variability in the hermaphrodite freshwater snail Bulinus truncatus. This species inhabits temporary habitats, has a short generation time and exhibits variable rates of selfing. We studied the variability over three sampling dates in 12 Sahelian populations (1161 individuals). Classical genetic parameters (estimators of Ho, He, f, selfing rate and Fst) showed limited change over time whereas important temporal changes of allelic frequencies were detected for 10 of the ponds studied. These variations are not easily explained by selection, sampling drift and genetic drift alone and may be due to periodic migration. Indeed the habitats occupied by the populations studied are subject to large temporal fluctuations owing to annual cycles of drought and flood. In such ponds our results support a demographic model of population expansions and contractions under which available habitats, after the rainy season, are colonized by individuals originating from a smaller number of refuges (areas that never dry out in the deepest parts of the ponds). In contrast, selfing appeared to be an important force affecting the genetic structure in permanent ponds.

Animals↗

Toward reconciling inferences concerning genetic variation in senescence in Drosophila melanogaster.

Standard models for senescence predict an increase in the additive genetic variance for log mortality rate late in the life cycle. Variance component analysis of age-specific mortality rates of related cohorts is problematic. The actual mortality rates are not observable and can be estimated only crudely at early ages when few individuals are dying and at late ages when most are dead. Therefore, standard quantitative genetic analysis techniques cannot be applied with confidence. We present a novel and rigorous analysis that treats the mortality rates as missing data following two different parametric senescence models. Two recent studies of Drosophila melanogaster, the original analyses of which reached different conclusions, are reanalyzed here. The two-parameter Gompertz model assumes that mortality rates increase exponentially with age. A related but more complex three-parameter logistic model allows for subsequent leveling off in mortality rates at late ages. We find that while additive variance for mortality rates increases for late ages under the Gompertz model, it declines under the logistic model. The results from the two studies are similar, with differences attributable to differences between the experiments.

Aging↗

The relationship between allozyme and chromosomal polymorphism inferred from nucleotide variation at the Acph-1 gene region of Drosophila subobscura.

The Acph-1 gene region was sequenced in 51 lines of Drosophila subobscura. Lines differ in their chromosomal arrangement for segment I of the O chromosome (O(st) and O(3+4)) and in the Acph-1 electrophoretic allele (Acph-1(100), Acph-1(054), and Acph-1(>100)). The ACPH-1 protein exhibits much more variation than previously detected by electrophoresis. The amino acid replacements responsible for the Acph-1(054) and Acph-1(>100) electrophoretic variants are different within O(st) and within O(3+4), which invalidates all previous studies on linkage disequilibrium between chromosomal and allozyme polymorphisms at this locus. The Acph-1(>100) allele within O(3+4) has a recent origin, while both Acph-1(054) alleles are rather old. Levels of nucleotide variation are higher within the O(3+4) than within the O(st) arrangement except for nonsynonymous sites. The McDonald and Kreitman test shows a significant excess of nonsynonymous polymorphisms within O(st) when D. guanche is used as the outgroup. According to the nearly neutral model of molecular evolution, this excess is consistent with a smaller effective size of O(st) relative to O(3+4) arrangements. A smaller population size, a lower recombination, and a more recent bottleneck might be contributing to the smaller effective size of O(st).

Acid Phosphatase↗

The historical biogeography of two Caribbean butterflies (Lepidoptera: Heliconiidae) as inferred from genetic variation at multiple loci.

Mitochondrial DNA and allozyme variation was examined in populations of two Neotropical butterflies, Heliconius charithonia and Dryas iulia. On the mainland, both species showed evidence of considerable gene flow over huge distances. The island populations, however, revealed significant genetic divergence across some, but not all, ocean passages. Despite the phylogenetic relatedness and broadly similar ecologies of these two butterflies, their intraspecific biogeography clearly differed. Phylogenetic analyses of mitochondrial DNA sequences revealed that populations of D. iulia north of St. Vincent are monophyletic and were probably derived from South America. By contrast, the Jamaican subspecies of H. charithonia rendered West Indian H. charithonia polyphyletic with respect to the mainland populations; thus, H. charithonia seems to have colonized the Greater Antilles on at least two separate occasions from Central America. Colonization velocity does not correlate with subsequent levels of gene flow in either species. Even where range expansion seems to have been instantaneous on a geological timescale, significant allele frequency differences at allozyme loci demonstrate that gene flow is severely curtailed across narrow ocean passages. Stochastic extinction, rapid (re)colonization, but low gene flow probably explain why, in the same species, some islands support genetically distinct and nonexpanding populations, while nearby a single lineage is distributed across several islands. Despite the differences, some common biogeographic patterns were evident between these butterflies and other West Indian taxa; such congruence suggests that intraspecific evolution in the West Indies has been somewhat constrained by earth history events, such as changes in sea level.

Animals↗

Phylogeography of the endangered Cathaya argyrophylla (Pinaceae) inferred from sequence variation of mitochondrial and nuclear DNA.

Cathaya argyrophylla is an endangered conifer restricted to subtropical mountains of China. To study phylogeographical pattern and demographic history of C. argyrophylla, species-wide genetic variation was investigated using sequences of maternally inherited mtDNA and biparentally inherited nuclear DNA. Of 15 populations sampled from all four distinct regions, only three mitotypes were detected at two loci, without single region having a mixed composition (G(ST) = 1). Average nucleotide diversity (theta(ws) = 0.0024; pi(s) = 0.0029) across eight nuclear loci is significantly lower than those found for other conifers (theta(ws) = 0.003 approximately 0.015; pi(s) = 0.002 approximately 0.012) based on estimates of multiple loci. Because of its highest diversity among the eight nuclear loci and evolving neutrally, one locus (2009) was further used for phylogeographical studies and eight haplotypes resulting from 12 polymorphic sites were obtained from 98 individuals. All the four distinct regions had at least four haplotypes, with the Dalou region (DL) having the highest diversity and the Bamian region (BM) the lowest, paralleling the result of the eight nuclear loci. An AMOVA revealed significant proportion of diversity attributable to differences among regions (13.4%) and among populations within regions (8.9%). F(ST) analysis also indicated significantly high differentiation among populations (F(ST) = 0.22) and between regions (F(ST) = 0.12-0.38). Non-overlapping distribution of mitotypes and high genetic differentiation among the distinct geographical groups suggest the existence of at least four separate glacial refugia. Based on network and mismatch distribution analyses, we do not find evidence of long distance dispersal and population expansion in C. argyrophylla. Ex situ conservation and artificial crossing are recommended for the management of this endangered species.

Cell Nucleus↗