Search PubMedSearch

SEARCH · Search PubMed

Results for “Quartet trees”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

7 recordsLinked to original sources

Phlag: scalable detection of genomics regions with unexplained phylogenetic heterogeneity.

MOTIVATION: Phylogenetic analyses of entire genomes (phylogenomics) have revealed abundant heterogeneity of evolutionary histories. While much has been done to model this heterogeneity and to infer species trees despite it, the current toolkit has a limitation. Most methods assume that gene trees across the genome differ but are all sampled from the same distribution, defined by models such as the multi-species coalescent (MSC), and parametrized consistently across the genome. Empirical data strongly suggest this assumption is often violated because the species tree, its parameters, or the process generating the gene trees can all change across the genome. Errors in the data can further compound this heterogeneity. RESULTS: To address this challenge, we define the problem of detecting what segments of the genome are inconsistent with a putative species tree, even after allowing discordance according to MSC. We model gene trees not as a set, but rather as a series (a realization of a stochastic process) along genomic positions. We propose a Hidden Markov Model (HMM) approach applied to quartet statistics measured from gene trees and tie the model to MSC using simulations. The combined use of these three ideas leads to a scalable method called Phlag. On simulated and real data, we show that Phlag can detect many cases of change in underlying evolutionary processes, including reduced recombination rates, population size changes, and admixture, all using the same algorithm. AVAILABILITY AND IMPLEMENTATION: Phlag is available at github.com/bo1929/phlag. All results and scripts can be found at github.com/bo1929/shared.phlag.

Phylogeny

Phlag: Scalable detection of genomics regions with unexplained phylogenetic heterogeneity.

MOTIVATION: Phylogenetic analyses of entire genomes (phylogenomics) have revealed abundant heterogeneity of evolutionary histories. While much has been done to model this heterogeneity and to infer species trees despite it, the current toolkit has a limitation. Most methods assume that gene trees across the genome differ but are all sampled from the same distribution , defined by models such as the multi-species coalescent (MSC), and parametrized consistently across the genome. Empirical data strongly suggest this assumption is often violated because the species tree, its parameters, or the process generating the gene trees can all change across the genome. Errors in the data can further compound this heterogeneity. RESULTS: To address this challenge, we define the problem of detecting what segments of the genome are inconsistent with a putative species tree, even after allowing discordance according to MSC. We model gene trees not as a set, but rather as a series (a realization of a stochastic process) along genomic positions. We propose a Hidden Markov Model (HMM) approach applied to quartet statistics measured from gene trees and tie the model to MSC using simulations. The combined use of these three ideas leads to a scalable method called Phlag. On simulated and real data, we show that Phlag can detect many cases of change in underlying evolutionary processes, including reduced recombination rates, population size changes, and admixture, all using the same algorithm. AVAILABILITY AND IMPLEMENTATION: Phlag is available at github.com/bo1929/phlag . All results and scripts can be found at github.com/bo1929/shared.phlag .

Journal Article

IQ-NET: fast and accurate quartet phylogenetic inference using deep learning trained on empirical DNA alignments.

Phylogenetic inference is fundamental to modern biology, with many applications including evolutionary biology, epidemiology, and comparative genomics. While maximum likelihood and Bayesian methods remain the gold standard for phylogenetic analysis, they rely on simplifying assumptions and are computationally intensive. Recent machine learning approaches for phylogenetics offer speed advantages, but have several limitations: exclusive reliance on simulated data for training, inadequate handling of gaps, and sensitivity to input sequence order. Here, we introduce IQ-NET (Intelligent Quartet NETwork), a deep learning framework that solves these limitations to infer four-taxon trees. IQ-NET estimates both tree topology and branch lengths directly from gapped alignments. IQ-NET outperforms existing machine learning methods in terms of accuracy, and obtained a 24-fold speedup compared with the widely used maximum likelihood software, IQ-TREE. We finally introduce a pipeline using IQ-NET and the ASTRAL software to reconstruct a larger species tree, i.e., with more than four taxa.

Empirical data training

Beyond Level-1: Identifiability of a Class of Galled Tree-Child Networks.

Inference of phylogenetic networks is of increasing interest in the genomic era. However, the extent to which phylogenetic networks are identifiable from various types of data remains poorly understood, despite its crucial role in justifying methods. This work obtains strong identifiability results for large sub-classes of galled tree-child semidirected networks. Some of the conditions our proofs require, such as the identifiability of a network's tree of blobs or the circular order of 4 taxa around a cycle in a level-1 network, are already known to hold for many data types. We show that all these conditions hold for quartet concordance factor data under various gene tree models, yielding the strongest results from 2 or more samples per taxon. Although the network classes we consider have topological restrictions, they include non-planar networks of any level and are substantially more general than level-1 networks - the only class previously known to enjoy identifiability from many data types. Our work establishes a route for proving future identifiability results for tree-child galled networks from data types other than quartet concordance factors, by checking that explicit conditions are met.

Mathematical Concepts

Nuclear single-copy orthologous genes as phylogenomic markers for resolving the closely related firefly genera Pteroptyx, Medeopteryx, and Trisinuata (Coleoptera: Lampyridae: Luciolinae).

Fireflies (Lampyridae) are bioluminescent beetles with broad ecological roles across temperate and tropical ecosystems, occupying diverse habitats including forests, wetlands, grasslands, mangroves, and riverine systems. The subfamily Luciolinae is primarily distributed across Asia and the Indo-Pacific. Phylogenetic relationships among three closely related Luciolinae genera - Medeopteryx, Pteroptyx, and Trisinuata - remain unresolved using mitochondrial genome data alone. This study used nuclear genome data to resolve relationships among these genera and identify a lighter-weight nuclear marker panel for expanding taxon sampling. Draft genomes were reconstructed for fifteen firefly species, eight from the focal genera, and analyzed with five published firefly genomes. Using BUSCO and OrthoFinder, 1,011 nuclear single-copy orthologs (SCOs) were identified for phylogenomic inference. Discordance between concatenation- and coalescence-based phylogenies indicated incomplete lineage sorting (ILS). The coalescence-based phylogeny recoveredPteroptyxas monophyletic and sister to a (Medeopteryx,Trisinuata) clade, with Trisinuata nested within a non-monophyletic Medeopteryx; however, quartet support at the base of Pteroptyx, particularly at Pt. valida, was low.Filtering for compositional homogeneity, clock-likeness, and species-tree concordance yielded 103 SCOs with a significantly higher proportion of parsimony-informative sites than non-selected loci, retaining the backbone topology with higher gene concordance support at scored clades, while ILS-driven discordance at Pt. valida persists - confirming that the reduced panel retains phylogenetic resolving power for future taxon sampling. These findings demonstrate a practical framework for using nuclear SCOs to resolve close phylogenetic relationships within Luciolinae. Future work should expand taxon sampling - especially forTrisinuata - alongside long-read assemblies, for a more robust phylogenomic framework.

Fireflies

SNaQ.jl: Improved scalability for level-1 phylogenetic network inference.

MOTIVATION: Phylogenetic networks represent complex biological scenarios that are overlooked in trees, such as hybridization and horizontal gene transfer. Although numerous methods have been developed for phylogenetic network inference, their scalability is severely limited by the computational demands of likelihood optimization and the vastness of network space. Composite (or pseudo-) likelihood approaches like SNaQ have improved computational tractability for network inference, but they remain inadequate for datasets of sizes routinely handled by tree inference methods. RESULTS: Here, we introduce SNaQ.jl, a new standalone Julia package with the composite likelihood inference originally implemented within PhyloNetworks.jl as well as new scalability features that enhance computational efficiency through (i) parallelization of quartet likelihood calculations during composite likelihood computation, (ii) weighted random selection of quartets, and (iii) probabilistic decision-making during network search. Through a simulation study and empirical data analysis, we show that this new version of SNaQ.jl (version 1.1) improves average runtimes by up to 499% on average with no change in function parameters or method accuracy. AVAILABILITY AND IMPLEMENTATION: SNaQ.jl is a new open source Julia package available at https://github.com/JuliaPhylo/SNaQ.jl.

Phylogeny

Beyond species trees: pervasive gene flow limits phylogenomic resolution in the diversification of Juniperus from the Qinghai-Tibet Plateau.

Understanding how lineages diversify despite persistent ancestral polymorphism and recurrent gene flow remains a central challenge in evolutionary biology. Juniperus distributed across the Qinghai-Tibet Plateau provide an ideal system for addressing this question because repeated geological uplift and climatic oscillations have likely promoted cycles of lineage divergence, range shifts, and secondary contact. Here, we combined approximately 1.08 million genome-wide SNPs from 164 individuals representing thirteen Juniperus lineages with phylogenomic datasets comprising 3,381 nuclear single-copy genes and nearly complete plastomes. We detected extensive phylogenomic discordance and cytonuclear incongruence across genomic datasets. Topology weighting, coalescent simulations, quartet-based tests, and analyses of gene flow and reticulation collectively support the interpretation that these patterns were shaped by the combined effects of prolonged incomplete lineage sorting and gene flow during lineage diversification. Ecological niche analyses further provide a spatial and climatic context in which environmentally similar lineages may have had greater opportunities for secondary contact during historical range shifts. Collectively, our results reveal that the evolutionary history of Qinghai-Tibet Plateau Juniperus is characterized by reticulate diversification rather than strictly bifurcating evolution, and demonstrate how genome-wide discordance can provide biological insights into the evolutionary processes underlying lineage diversification.

Gene Flow