Search PubMedSearch

SEARCH · Search PubMed

Results for “phylogenetic substitution models”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

12 recordsLinked to original sources

Viral genome sequence datasets display pervasive evidence of strand-specific substitution biases that are best described using non-reversible nucleotide substitution models.

Most phylogenetic trees are inferred using time-reversible evolutionary models that assume that the relative rates of substitution for any given pair of nucleotides are the same regardless of the direction of the substitutions. However, there is no reason to assume that the underlying biochemical mutational processes that cause substitutions are similarly symmetrical. We consider two non-reversible nucleotide substitution models: (1) a 6-rate non-reversible model (NREV6) that is applicable to analyzing mutational processes in double-stranded genomes in that complementary substitutions occur at identical rates; and (2) a 12-rate non-reversible model (NREV12) that is applicable to analyzing mutational processes in single-stranded (ss) genomes in that all substitution types are free to occur at different rates. Using likelihood ratio and Akaike Information Criterion-based model tests, we show that, surprisingly, NREV12 provided a significantly better fit than the General Time Reversible (GTR) and NREV6 models to 21/31 dsRNA and 20/30 dsDNA datasets. As expected, however, NREV12 provided a significantly better fit to 24/33 ssDNA and 40/47 ssRNA datasets. We tested how non-reversibility impacts the accuracy with which phylogenetic trees are inferred. As simulated degrees of non-reversibility (DNR) increased, the tree topology inferences using both NREV12 and GTR became more accurate, whereas inferred tree branch lengths became less accurate. We conclude that while non-reversible models should be helpful in the analysis of mutational processes in most virus species, there is no pressing need to use these models for routine phylogenetic inference.

Models of evolution

Robustness of Ancestral Sequence Reconstruction to Among-site and Among-lineage Evolutionary Heterogeneity.

Ancestral sequence reconstruction is typically performed using homogeneous evolutionary models, which assume that the same substitution propensities affect all sites and lineages. These assumptions are routinely violated: heterogeneous structural and functional constraints favor different amino acids at different sites, and these constraints often change among lineages as epistatic substitutions accrue at other sites. To evaluate how violations of the homogeneity assumption affect ancestral sequence reconstruction under realistic conditions, we developed site-specific substitution models and parameterized them using data from deep mutational scanning experiments on three protein families; we then used these models to perform ancestral sequence reconstruction on the empirical alignments and on alignments simulated under heterogeneous conditions derived from the experiments. Extensive among-site and -lineage heterogeneity is present in these datasets, but the sequences reconstructed from empirical alignments are almost identical when heterogeneous or homogeneous models are used for ancestral sequence reconstruction. Using models fit to deep mutational scanning data from distantly related proteins in which mutational effects are very different also has a minimal impact on ancestral sequence reconstruction. The rare differences occur primarily where phylogenetic signal is weak-at fast-evolving sites and nodes connected by long branches. When ancestral sequence reconstruction is performed on simulated data, errors in the reconstructed sequences become more likely as branch lengths increase, but incorporating heterogeneity into the model does not improve accuracy. These data establish that ancestral sequence reconstruction is robust to unincorporated realistic forms of evolutionary heterogeneity, because the primary determinant of ancestral sequence reconstruction is phylogenetic signal, not the substitution model. The best way to improve accuracy is therefore not to develop more elaborate models but to apply ancestral sequence reconstruction to densely sampled alignments that maximize phylogenetic signal at the nodes of interest.

Phylogeny

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework to address this challenge, which reduces runtime and memory requirements by orders of magnitude. In PSU, small subsamples of sites from a concatenated alignment are analyzed, which are expanded by upsampling before inference, and the resulting inferences are aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-alignment analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis of concatenated alignments. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical. We suggest that PSU is a general approach for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

Phylogeny

MCALIGN: stochastic alignment of noncoding DNA sequences based on an evolutionary model of sequence evolution.

A method is described for performing global alignment of noncoding DNA sequences based on an evolutionary model parameterized by the frequency distribution of lengths of insertion/deletion events (indels) and their rate relative to nucleotide substitutions. A stochastic hill-climbing algorithm is used to search for the most probable alignment between a pair of sequences or three sequences of known phylogenetic relationship. The performance of the procedure, parameterized according to the empirical distribution of indel lengths in noncoding DNA of Drosophila species, is investigated by simulation. We show that there is excellent agreement between true and estimated alignments over a wide range of sequence divergences, and that the method outperforms other available alignment methods.

Algorithms

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework, in which small subsamples of sites from a concatenated alignment are expanded by upsampling before inference, and the resulting analyses are then aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-data analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures, while reducing runtime and memory requirements by orders of magnitude. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical across diverse datasets. We suggest that PSU is a general strategy for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

confidence limits

Use of 3D chaos game representation to quantify DNA sequence similarity with applications for hierarchical clustering.

A 3D chaos game is shown to be a useful way for encoding DNA sequences. Since matching subsequences in DNA converge in space in 3D chaos game encoding, a DNA sequence's 3D chaos game representation can be used to compare DNA sequences without prior alignment and without truncating or padding any of the sequences. Two proposed methods inspired by shape-similarity comparison techniques show that this form of encoding can perform as well as alignment-based techniques for building phylogenetic trees. The first method uses the volume overlap of intersecting spheres and the second uses shape signatures by summarizing the coordinates, oriented angles, and oriented distances of the 3D chaos game trajectory. The methods are tested using: (1) the first exon of the beta-globin gene for 11 species, (2) mitochondrial DNA from four groups of primates, and (3) a set of synthetic DNA sequences. Simulations show that the proposed methods produce distances that reflect the number of mutation events; additionally, on average, distances resulting from deletion mutations are comparable to those produced by substitution mutations.

Animals

Crawling under the radar: Two novel Paulinella species expand knowledge about the ecology and evolution of a primary plastid-containing amoeba lineage.

The genus Paulinella represents a rare, independent case of primary endosymbiosis, providing a unique system to study the early stages of organelle evolution. Here, we expand current understanding of primary plastid endosymbiosis through the discovery and characterization of two novel photosynthetic amoebae, Paulinella marae sp. nov. and Paulinella murrayi sp. nov., isolated from a brackish water habitat in North Carolina, United States. Complete chromatophore genomes and mitochondrial data revealed conserved gene content but notable structural variation, including genome rearrangements and inversion events. Phylogenetic analyses uncovered significant discordance between nuclear and organelle datasets, likely driven by substitution saturation, limited taxon sampling, and differing evolutionary signals across loci. Ecological observations over multiple years indicate that both species are in low abundance but consistently present, and when coupled with hobbyist data, support the hypothesis that photosynthetic Paulinella species are globally distributed yet under-sampled. These results increase known species diversity within the clade from four to six and highlight the importance of integrating field-based observations with genomic approaches. Overall, this work advances Paulinella as a model for studying ongoing primary endosymbiosis, lineage divergence, and the ecological strategies of low-abundance microbial eukaryotes.

Paulinella

Antibody diversification in cartilaginous fishes: Mechanistic insights from the nurse shark and comparative perspectives across jawed vertebrates.

Antibody diversity in vertebrates arises through the coordinated actions of V(D)J recombination and somatic hypermutation (SHM). Cartilaginous fishes occupy a key phylogenetic position as the sister lineage to bony vertebrates and therefore provide important comparative insights into the evolution of adaptive immunity. This review focuses on the nurse shark (Ginglymostoma cirratum) as a representative model for examining antibody-diversification mechanisms in cartilaginous fishes. Shark immunoglobulin genes exhibit a multicluster organization, while immunoglobulin new antigen receptor (IgNAR), a heavy-chain-only isotype, contains a single variable domain with an extended complementarity-determining region 3 (CDR3) that can be stabilized by non-canonical disulfide bonds. These structural features, together with intracluster multi-D V(D)J recombination and distinctive SHM characterized by single and tandem substitutions and insertions/deletions, contribute to antibody diversification in sharks. By comparing cartilaginous fishes, ray-finned fishes, and mammals, this review highlights lineage-specific combinations of immunoglobulin gene organization, recombination, mutational processing, and affinity maturation. Within the heuristic framework proposed here, shark and mammalian systems are described as emphasizing "breadth-first" repertoire generation and "precision-first" affinity optimization, respectively. These terms indicate relative mechanistic emphases rather than mutually exclusive categories or sequential evolutionary stages, while ray-finned fishes exhibit a distinct combination of genomic organization and mutational features. Investigating antibody diversification in cartilaginous fishes not only advances our understanding of vertebrate immune evolution but also provides structural and mechanistic insights that may inform the development of engineered antibodies based on the IgNAR scaffold.

Antibody diversity

Spatiotemporal patterns of Rift Valley fever virus in Africa: a retrospective genomic epidemiology and phylodynamic modelling study.

BACKGROUND: Rift Valley fever virus (RVFV) is a mosquito-borne zoonotic pathogen causing outbreaks in humans and ruminants across Africa and the Arabian Peninsula. Originally restricted to the Great Rift Valley, RVFV has expanded geographically, prompting its classification by WHO as a pathogen of pandemic potential. We investigated the evolutionary and spatial dynamics of RVFV across Africa. METHODS: We used genomic data generated at the International Livestock Research Institute Nairobi genomic laboratory (BioProject PRJNA1106221) and combined with publicly available datasets retrieved from the National Center for Biotechnology (NCBI) GenBank nucleotide database. In retrieving RVFV genome sequences from the NCBI GenBank, we applied the search terms "Rift Valley fever virus segment L AND 6404[SLEN]", "Rift Valley fever virus segment M AND 3885[SLEN]", and "Rift Valley fever virus segment S AND 1520:1690[SLEN]" for L (Large), M (Medium), and S (Small) segments, respectively. For sequences without additional spatiotemporal information, we searched PubMed to extract the associated sequence metadata. We performed molecular clock analysis, phylogenetic inference, phylodynamic modelling (continuous phylogeographic reconstruction), and landscape phylogeography on the three RVFV genome segments (L, M, and S). We aimed to assess evolutionary rates, dispersal patterns, and environmental drivers. Focus was placed on lineage C, the most widely distributed variant. FINDINGS: The global dataset used in this study consisted of large (n=236), medium (n=237), and small (n=247), which were further filtered to exclude potential reassortants and vaccine strains. Genome sequences retrieved from NCBI GenBank database comprised large (n=180), medium (n=184), and small (n=202). The genome sequences from retrospective human and livestock isolates comprised large (n=56), medium (n=53), and small (n=45) collected in Burundi (2018), Kenya (2007, 2018, 2019, 2021, and 2022), and Rwanda (2018 and 2022). Our dataset revealed that RVFV exhibited low overall genetic diversity. Lineage C, however, showed evidence of active evolution, with substitution rates ranging from 3·58 × 10-4 to 9·76 × 10-4 substitutions per site per year. This lineage probably originated in Zimbabwe in the mid-1970s and has since expanded across eastern and southern Africa. Phylogeographic reconstructions revealed rapid spread, with diffusion coefficients exceeding 50 000 km2 per year. INTERPRETATION: Lineage C appears capable of establishing endemic transmission in new regions, with ongoing diversification observed during interepidemic periods. These observations reinforce the value of continuous genomic surveillance, particularly during cryptic transmission phases when adaptive mutations might emerge. Although further evidence is needed, observed trends in climate variability and land-use change point to the potential benefit of targeted surveillance in settings that could be at increased risk, including urban centres and wetlands. FUNDING: This work was supported by the German Federal Ministry for Economic Cooperation and Development, the Rockefeller Foundation, and the Africa Centres for Disease Control and Prevention.

Rift Valley fever virus

Genomic Footprints of Historical Introgression Between Ancient Lineages of Wild Oryza AA-Genome Species With Widely Separated Contemporary Distributions.

Phylogenetic incongruence is increasingly recognized as pervasive, yet the extent to which reticulate evolution occurs between groups separated by substantial geographical distances and deep phylogenetic divergence remains poorly characterized. In the Oryza AA-genome group-a model for plant speciation and domestication-the traditional bifurcation model posits that Australian Oryza meridionalis and African Oryza longistaminata occupy basal branches, distinct from the more recently diversified monophyletic clade comprising Asian and other African lineages, including major cultivars. However, recent evidence from endogenous viral sequences has hinted at unexpected genetic relatedness between African O. longistaminata and Asian Oryza sativa, which are geographically and phylogenetically distant. Here, we conducted a genome-wide survey across 11 Oryza species to systematically identify genomic regions exhibiting phylogenetic incongruence. Widespread phylogenetic discordance was observed, notably involving genomic segments in which O. longistaminata showed phylogenetic proximity to Asian species, contradicting their established deep divergence. To distinguish between introgression and incomplete lineage sorting, we performed four-taxon ABBA-BABA tests, which provided statistical support for introgression. Furthermore, divergence time estimates for these incongruent regions were younger than the species divergence times, suggesting historical introgression between the ancestors of lineages that are currently separated by vast geographical distances. Systematic assessments indicated that potential analytical artifacts, such as compositional bias and substitution saturation, were unlikely to explain the observations. These convergent lines of evidence suggest that ancient introgression had occurred between currently geographically separated and evolutionarily divergent Oryza lineages, leaving detectable footprints across their modern genomes.

Oryza

Detecting Introgression in Shallow Phylogenies: How Minor Molecular Clock Deviations Lead to Major Inference Errors.

Recent theoretical and algorithmic advances in introgression detection, coupled with the growing availability of genome-scale data, have highlighted the widespread occurrence of interspecific gene flow across the tree of life. However, current methods largely depend on the molecular clock assumption-a questionable premise given empirical evidence of substitution rate variation across lineages. While such rate heterogeneity is known to compromise gene flow detection among divergent lineages, its impact on closely related taxa at shallow evolutionary timescales remains poorly understood, likely because these taxa are often assumed to adhere to a molecular clock. To address this gap, we combine theoretical analyses and simulations to evaluate the robustness of widely used site pattern methods (D-statistic and HyDe) to rate variation across phylogenetic timescales. Our results demonstrate that both methods exhibit high sensitivity to even minor deviations from the molecular clock at shallow timescales, complementing previous findings at deeper scales. Specifically, in young phylogenies (with an age of 3 × 105 generations) with small population sizes, weak (17% difference) and moderate (33% difference) rate variation can inflate false-positive rates up to 35% and 100%, respectively, using site pattern counts from a 500 Mb genome. Employing a more distant outgroup intensifies these spurious signals. Our study demonstrates that summary tests for introgression are pervasively vulnerable to minor rate variations and underscores the critical need for advanced methodologies to disentangle genuine introgression from false signals generated by rate heterogeneity.

Phylogeny

A general approach to proving the minimality of phylogenetic trees illustrated by an example with a set of 23 vertebrates.

We have recently described a method of building phylogenetic trees and have outlined an approach for proving whether a particular tree is optimal for the data used. In this paper we describe in detail the method of establishing lower bounds on the length of a minimal tree by partitioning the data set into subsets. All characters that could be involved in duplications in the data are paired with all other such characters. A matching algorithm is then used to obtain the pairing of characters that reveals the most duplications in the data. This matching may still not account for all nucleotide substitutions on the tree. The structure of the tree is then used to help select subsets of three or more characters until the lower bound found by partitioning is equal to the length of the tree. The tree must then be a minimal tree since no tree can exist with a length less than that of the lower bound. The method is demonstrated using a set of 23 vertebrate cytochrome c sequences with the criterion of minimizing the total number of nucleotide substitutions. There are 131130 7045768798 96033440625 topologically distinct trees that can be constructed from this data set. The method described in this paper does identify 144 minimal tree variants. The method is general in the sense that it can be used for other data and other criteria of length. It need not however always be possible to prove a treee minimal but the method will give an upper and lower bound on the length of minimal trees.

Amino Acid Sequence