Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “variational inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Structural and evolutionary inference from molecular variation in Neisseria porins.

The porin proteins of the pathogenic Neisseria species, Neisseria gonorrhoeae and Neisseria meningitidis, are important as serotyping antigens, putative vaccine components, and for their proposed role in the intracellular colonization of humans. A three-dimensional structural homology model for Neisseria porins was generated from Escherichia coli porin structures and N. meningitidis PorA and PorB sequences. The Neisseria sequences were readily assembled into the 16-strand beta-barrel fold characteristic of porins, despite relatively low sequence identity with the Escherichia proteins. The model provided information on the spatial relationships of variable regions of peptide sequences in the PorA and PorB trimers and insights relevant to the use of these proteins in vaccines. The nucleotide sequences of the porin genes from a number of other Neisseria species were obtained by PCR direct sequencing and from GenBank. Alignment and analysis of all available Neisseria porin sequences by use of the structurally conserved regions derived from the PorA and PorB structural models resulted in the recovery of an improved phylogenetic signal. Phylogenetic analyses were consistent with an important role for horizontal genetic exchange in the emergence of different porin classes and confirmed the close evolutionary relationships of the porins from N. meningitidis, N. gonorrhoeae, Neisseria lactamica, and Neisseria polysaccharea. Only members of this group contained three conserved lysine residues which form a potential GTP binding site implicated in pathogenesis. The model placed these residues on the inside of the pore, in close proximity, consistent with their role in regulating pore function when inserted into host cells.

Amino Acid Sequence↗

Phylogenetic relationships among two species of golden monkey and three species of leaf monkey inferred from rDNA variation.

Restriction maps of rDNA repeats of five species of Colobinae and three outgroup taxa, Hylobates leucogenys, Macaca mulatta, and Macaca irus, were constructed using 15 restriction endonucleases and cloned 18S and 28S rRNA gene probes. The site variation between Rhinopithecus roxellana and Rhinopithecus bieti is comparable to that between Presbytis françoisi and Preshytis phayrei, implying that R. bieti is a valid species rather than a subspecies of R. roxellana. Phylogenetic analysis on the 47 informative sites supports the case for Rhinopithecus being an independent genus and closely related to Presbytis. Furthermore, branch lengths of the tree seem to support the hypothesis that the leaf monkeys share some ancestral traits as well as some automorphic characters.

Animals↗

Genetic structure of korean wild populations of the Medaka Oryzias latipes inferred from allozymic variation.

Previous allozymic studies have revealed that Korean wild populations of Oryzias latipes have differentiated regionally, and are composed of two distinct groups, the East Korean Population and the China-West Korean Population. Recently, mitochondrial DNA (mtDNA) sequencing and restriction fragment length polymorphism (RFLP) analyses have confirmed these two groups, and shown that the distribution ranges of the two groups overlap in western Korea. In order to describe the detailed distributions of the two groups and the gene flow between them, genotypes of 13 allozymic loci were determined in 444 specimens from 96 localities in Korea. The two major groups were supported by remarkable allele frequency differences at six diagnostic loci: ACP*, AMY*, CK-A*, LDH-A*, PGM* and TF*. Individuals with the typical "eastern" genotype were mainly distributed in eastern and southern areas. In contrast, fish with the "western" genotype were predominant in the western area, and were further divided into two subgroups (the Han River and Geum River Subpopulations) by unique alleles at the ADH* locus. In the western coast, two distinct (eastern and western) genotypes were distributed in a mosaic fashion. This distribution pattern was identical to those from mtDNA analyses. Although the distribution patterns of the alleles at three loci (GPI-A*, LDH-C* and SOD*) showed introgressive conditions between the two groups, each population was nearly fixed as either the eastern or western genotype at all six diagnostic loci despite the proximity among samples. Therefore, it is suggested that some reproductive isolation mechanisms exist between the two groups in natural habitats.

Animals↗

Variational mixture of Bayesian independent component analyzers.

There has been growing interest in subspace data modeling over the past few years. Methods such as principal component analysis, factor analysis, and independent component analysis have gained in popularity and have found many applications in image modeling, signal processing, and data compression, to name just a few. As applications and computing power grow, more and more sophisticated analyses and meaningful representations are sought. Mixture modeling methods have been proposed for principal and factor analyzers that exploit local gaussian features in the subspace manifolds. Meaningful representations may be lost, however, if these local features are nongaussian or discontinuous. In this article, we propose extending the gaussian analyzers mixture model to an independent component analyzers mixture model. We employ recent developments in variational Bayesian inference and structure determination to construct a novel approach for modeling nongaussian, discontinuous manifolds. We automatically determine the local dimensionality of each manifold and use variational inference to calculate the optimum number of ICA components needed in our mixture model. We demonstrate our framework on complex synthetic data and illustrate its application to real data by decomposing functional magnetic resonance images into meaningful-and medically useful-features.

Bayes Theorem↗

5S rDNA variation and its phylogenetic inference in the genus Leporinus (Characiformes: Anostomidae).

5S rDNA sequences have proven to be valuable as genetic markers to distinguish closely related species and also in the understanding of the dynamic of repetitive sequences in the genomes. In the aim to contribute to the knowledge of the evolutionary history of Leporinus (Anostomidae) and also to contribute to the understanding of the 5S rDNA sequences organization in the fish genome, analyses of 5S rDNA sequences were conducted in seven species of this genus. The 5S rRNA gene sequence was highly conserved among Leporinus species, whereas NTS exhibit high levels of variations related to insertions, deletions, microrepeats, and base substitutions. The phylogenetic analysis of the 5S rDNA sequences clustered the species into two clades that are in agreement with cytogenetic and morphological data.

Animals↗

ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing.

MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1-5×, sWGS data display a pronounced zero‑inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.

DNA Copy Number Variations↗

Genetic variation among world populations: inferences from 100 Alu insertion polymorphisms.

We examine the distribution and structure of human genetic diversity for 710 individuals representing 31 populations from Africa, East Asia, Europe, and India using 100 Alu insertion polymorphisms from all 22 autosomes. Alu diversity is highest in Africans (0.349) and lowest in Europeans (0.297). Alu insertion frequency is lowest in Africans (0.463) and higher in Indians (0.544), E. Asians (0.557), and Europeans (0.559). Large genetic distances are observed among African populations and between African and non-African populations. The root of a neighbor-joining network is located closest to the African populations. These findings are consistent with an African origin of modern humans and with a bottleneck effect in the human populations that left Africa to colonize the rest of the world. Genetic distances among all pairs of populations show a significant product-moment correlation with geographic distances (r = 0.69, P < 0.00001). F(ST), the proportion of genetic diversity attributable to population subdivision is 0.141 for Africans/E. Asians/Europeans, 0.047 for E. Asians/Indians/Europeans, and 0.090 for all 31 populations. Resampling analyses show that approximately 50 Alu polymorphisms are sufficient to obtain accurate and reliable genetic distance estimates. These analyses also demonstrate that markers with higher F(ST) values have greater resolving power and produce more consistent genetic distance estimates.

Africa↗

Phylogeographic inferences from the mtDNA variation of the three-toed skink, Chalcides chalcides (Reptilia: Scincidae).

Genetic diversity was analyzed in Chalcides chalcides populations from peninsular Italy, Sardinia, Sicily and Tunisia by sequencing 400 bp at the 5' end of the mitochondrial gene encoding cytochrome b (cyt b) and by restriction fragment length polymorphism (RFLP) analysis of two mitochondrial DNA segments (ND-1/2 and ND-3/4). The results of the phylogenetic analysis highlighted the presence of three main clades corresponding with three of the four main geographical areas (Tunisia, Sicily and the Italian peninsula), while Sardinia proved to be closely related to Tunisian haplotypes suggesting a colonization of this island from North Africa by human agency in historical times. On the contrary, the splitting times estimated on the basis of cyt b sequence data seem to indicate a more ancient colonization of Sicily and the Italian Peninsula, as a consequence of tectonic and climatic events that affected the Mediterranean Basin during the Pleistocene. Finally, the analysis of the genetic variability of C. chalcides populations showed a remarkable genetic homogeneity in Italian populations when compared to the Tunisian ones. This condition could be explained by a rapid post-glacial expansion from refugial populations that implied serial bottlenecking with progressive loss of haplotypes, resulting in a low genetic diversity in the populations inhabiting the more recently colonized areas.

Analysis of Variance↗

Human 18 S ribosomal RNA sequence inferred from DNA sequence. Variations in 18 S sequences and secondary modification patterns between vertebrates.

We have determined the DNA sequences encoding 18 S ribosomal RNA in man and in the frog, Xenopus borealis. We have also corrected the Xenopus laevis 18 S sequence: an A residue follows G-684 in the sequence. These and other available data provide a number of representative examples of variation in primary structure and secondary modification of 18 S ribosomal RNA between different groups of vertebrates. First, Xenopus laevis and Xenopus borealis 18 S ribosomal genes differ from each other by only two base substitutions, and we have found no evidence of intraspecies heterogeneity within the 18 S ribosomal DNA of Xenopus (in contrast to the Xenopus transcribed spacers). Second, the human 18 S sequence differs from that of Xenopus by approx. 6.5%. About 4% of the differences are single base changes; the remainder comprise insertions in the human sequence and other changes affecting several nucleotides. Most of these more extensive changes are clustered in a relatively short region between nucleotides 190 and 280 in the human sequence. Third, the human 18 S sequence differs from non-primate mammalian sequences by only about 1%. Fourth, nearly all of the 47 methyl groups in mammalian 18 S ribosomal RNA can be located in the sequence. The methyl group distribution corresponds closely to that in Xenopus, but there are several extra methyl groups in mammalian 18 S ribosomal RNA. Finally, minor revisions are made to the estimated numbers of pseudouridines in human and Xenopus 18 S ribosomal RNA.

Animals↗

Model-based inference of haplotype block variation.

The haplotype block structure of SNP variation in human DNA has been demonstrated by several recent studies. The presence of haplotype blocks can be used to dramatically increase the statistical power of genetic mapping. Several criteria have already been proposed for identifying these blocks, all of which require haplotypes as input. We propose a comprehensive statistical model of haplotype block variation and show how the parameters of this model can be learned from haplotypes and/or unphased genotype data. Using real-world SNP data, we demonstrate that our approach can be used to resolve genotypes into their constituent haplotypes with greater accuracy than previously known methods.

Algorithms↗

The speciation history of Drosophila pseudoobscura and close relatives: inferences from DNA sequence variation at the period locus.

Thirty-five period locus sequences from Drosophila pseudoobscura and its siblings species, D. p. bogotana, D. persimilis, and D. miranda, were studied. A large amount of variation was found within D. pseudoobscura and D. persimilis, consistent with histories of large effective population sizes. D. p. bogotana, however, has a severe reduction in diversity. Combined analysis of per with two other loci, in both D. p. bogotana and D. pseudoobscura, strongly suggest this reduction is due to recent directional selection at or near per within D. p. bogotana. Since D. p. bogotana is highly variable and shares variation with D. pseudoobscura at other loci, the low level of variation at per within D. p. bogotana can not be explained by a small effective population size or by speciation via founder event. Both D. pseudoobscura and D. persimilis have considerable intraspecific gene flow. A large portion of one D. persimilis sequence appears to have arisen via introgression from D. pseudoobscura. The time of this event appears to be well after the initial separation of these two species. The estimated times since speciation are one mya for D. pseudoobscura and D. persimilis and 2 mya since the formation of D. miranda.

Amino Acid Sequence↗

Genetic structure and evolutionary history of a diploid hybrid pine Pinus densata inferred from the nucleotide variation at seven gene loci.

Although homoploid hybridization is increasingly recognized as an important phenomenon in plant evolution, its evolutionary genetic mechanisms are poorly documented and understood. Pinus densata, a pine native to the Tibetan Plateau, represents a good example of a homoploid hybrid speciation facilitated by adaptation to extreme environment and ecological isolation from the parents. Its ecologically and reproductively stabilized nature offers excellent opportunity for studying genetic processes associated with hybrid speciation. In this study, we investigated the levels and patterns of nucleotide variation in P. densata and its putative parents. Haplotype composition, gene genealogies, and the levels and patterns of nucleotide variation gave further support to the hybrid nature of P. densata. Allelic history, as revealed by our data, suggests the ancient nature of the hybrid preceding elevation of the Tibetan Plateau. We detected more deviations from neutrality in P. densata than in the parental species. Thus, at least some of the evolutionary forces that have shaped the genetic variation in P. densata are likely to be different from those acting upon parental species. We speculate that when populations of P. densata invaded new territories, they had elevated rates of response to selection in order to develop traits that help them to survive and adapt in the new environments.

Diploidy↗

TreeFlow: Probabilistic Modelling and Automatic Differentiation for Phylogenetics.

Probabilistic modelling frameworks are powerful tools for statistical modelling and inference. They are not immediately generalizable to phylogenetic problems due to the particular computational properties of the phylogenetic tree object. TreeFlow is a software library for probabilistic modelling and automatic differentiation with phylogenetic trees. It embeds phylogenetic trees in the TensorFlow Probability framework, and implements inference algorithms for phylogenetic models given a fixed tree topology. We demonstrate how TreeFlow can be used to quickly implement and assess new models. We also show that it provides reasonable performance for gradient-based inference algorithms compared to specialized computational libraries for phylogenetics.

Bayesian inference↗

Glacial survival of the Norwegian lemming (Lemmus lemmus) in Scandinavia: inference from mitochondrial DNA variation.

In order to evaluate the biogeographical hypothesis that the Norwegian lemming (Lemmus lemmus) survived the last glacial period in some Scandinavian refugia, we examined variation in the nucleotide sequence of the mitochondrial control region (402 base pairs (bp)) and the cytochrome b (cyt b) region (633 bp) in Norwegian and Siberian (Lemmus sibiricus) lemmings. The phylogenetic distinction and cyt b divergence estimate of 1.8% between the Norwegian and Siberian lemmings suggest that their separation pre-dated the last glaciation and imply that the Norwegian lemming is probably a relic of the Pleistocene populations from Western Europe. The star-like control region phylogeny and low mitochondrial DNA diversity in the Norwegian lemming indicate a reduction in its historical effective size followed by population expansion. The average estimate of post-bottleneck time (19-21 kyr) is close to the last glacial maximum (18-22 kyr BP). Taking these findings and the fossil records into consideration, it seems likely that, after colonization of Scandinavia in the Late Pleistocene, the Norwegian lemming suffered a reduction in its population effective size and survived the last glacial maximum in some local Scandinavian refugia, as suggested by early biogeographical work.

Animals↗

Rapid miocene-pliocene dispersal and evolution of Mediterranean rajid fauna as inferred by mitochondrial gene variation.

Rajidae (colloquially known as skates and rays) experienced multiple and parallel adaptive radiations allowing high species diversity and great differences of species composition between regional faunas. Nevertheless, they show considerable conservation of bio-ecological, morphological and reproductive traits. The evolutionary history and dispersal of North-east Atlantic and Mediterranean rajid fauna were investigated throughout the sequence analysis of the control region and 16S rDNA mitochondrial genes. Molecular estimates of divergence times indicated recent origin and rapid dispersal of the present species. Compared with the ancient origin of the family (Late Cretaceous), the present species diversity arose in a relatively narrow time-window (12 Myr) from Middle Miocene to Early Pleistocene, likely by speciation processes related to dramatic geological and climatic events in the Mediterranean. Nucleotide substitution rates and phylogenetic relationships indicated Mediterranean endemic skates derived from sister species with wider distribution during Late Pliocene-Pleistocene. Skate phylogeny and systematics obtained using mitochondrial gene variation were largely consistent with those based on morpho-anatomical data.

Animals↗