Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “ancestral sequence reconstruction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Robustness of Ancestral Sequence Reconstruction to Among-site and Among-lineage Evolutionary Heterogeneity.

Ancestral sequence reconstruction is typically performed using homogeneous evolutionary models, which assume that the same substitution propensities affect all sites and lineages. These assumptions are routinely violated: heterogeneous structural and functional constraints favor different amino acids at different sites, and these constraints often change among lineages as epistatic substitutions accrue at other sites. To evaluate how violations of the homogeneity assumption affect ancestral sequence reconstruction under realistic conditions, we developed site-specific substitution models and parameterized them using data from deep mutational scanning experiments on three protein families; we then used these models to perform ancestral sequence reconstruction on the empirical alignments and on alignments simulated under heterogeneous conditions derived from the experiments. Extensive among-site and -lineage heterogeneity is present in these datasets, but the sequences reconstructed from empirical alignments are almost identical when heterogeneous or homogeneous models are used for ancestral sequence reconstruction. Using models fit to deep mutational scanning data from distantly related proteins in which mutational effects are very different also has a minimal impact on ancestral sequence reconstruction. The rare differences occur primarily where phylogenetic signal is weak-at fast-evolving sites and nodes connected by long branches. When ancestral sequence reconstruction is performed on simulated data, errors in the reconstructed sequences become more likely as branch lengths increase, but incorporating heterogeneity into the model does not improve accuracy. These data establish that ancestral sequence reconstruction is robust to unincorporated realistic forms of evolutionary heterogeneity, because the primary determinant of ancestral sequence reconstruction is phylogenetic signal, not the substitution model. The best way to improve accuracy is therefore not to develop more elaborate models but to apply ancestral sequence reconstruction to densely sampled alignments that maximize phylogenetic signal at the nodes of interest.

Phylogeny↗

Ancestral sequence reconstruction in primate mitochondrial DNA: compositional bias and effect on functional inference.

Reconstruction of ancestral DNA and amino acid sequences is an important means of inferring information about past evolutionary events. Such reconstructions suggest changes in molecular function and evolutionary processes over the course of evolution and are used to infer adaptation and convergence. Maximum likelihood (ML) is generally thought to provide relatively accurate reconstructed sequences compared to parsimony, but both methods lead to the inference of multiple directional changes in nucleotide frequencies in primate mitochondrial DNA (mtDNA). To better understand this surprising result, as well as to better understand how parsimony and ML differ, we constructed a series of computationally simple "conditional pathway" methods that differed in the number of substitutions allowed per site along each branch, and we also evaluated the entire Bayesian posterior frequency distribution of reconstructed ancestral states. We analyzed primate mitochondrial cytochrome b (Cyt-b) and cytochrome oxidase subunit I (COI) genes and found that ML reconstructs ancestral frequencies that are often more different from tip sequences than are parsimony reconstructions. In contrast, frequency reconstructions based on the posterior ensemble more closely resemble extant nucleotide frequencies. Simulations indicate that these differences in ancestral sequence inference are probably due to deterministic bias caused by high uncertainty in the optimization-based ancestral reconstruction methods (parsimony, ML, Bayesian maximum a posteriori). In contrast, ancestral nucleotide frequencies based on an average of the Bayesian set of credible ancestral sequences are much less biased. The methods involving simpler conditional pathway calculations have slightly reduced likelihood values compared to full likelihood calculations, but they can provide fairly unbiased nucleotide reconstructions and may be useful in more complex phylogenetic analyses than considered here due to their speed and flexibility. To determine whether biased reconstructions using optimization methods might affect inferences of functional properties, ancestral primate mitochondrial tRNA sequences were inferred and helix-forming propensities for conserved pairs were evaluated in silico. For ambiguously reconstructed nucleotides at sites with high base composition variability, ancestral tRNA sequences from Bayesian analyses were more compatible with canonical base pairing than were those inferred by other methods. Thus, nucleotide bias in reconstructed sequences apparently can lead to serious bias and inaccuracies in functional predictions.

Animals↗

A new method that simultaneously aligns and reconstructs ancestral sequences for any number of homologous sequences, when the phylogeny is given.

Among the fundamental problems in molecular evolution and in the analysis of homologous sequences are alignment, phylogeny reconstruction, and the reconstruction of ancestral sequences. This paper presents a fast, combined solution to these problems. The new algorithm gives an approximation to the minimal history in terms of a distance function on sequences. The distance function on sequences is a minimal weighted path length constructed from substitutions and insertions-deletions of segments of any length. Substitutions are weighted with an arbitrary metric on the set of nucleotides or amino acids, and indels are weighted with a gap penalty function of the form gk = a + (bxk), where k is the length of the indel and a and b are two positive numbers. A novel feature is the introduction of the concept of sequence graphs and a generalization of the traditional dynamic sequence comparison algorithm to the comparison of sequence graphs. Sequence graphs ease several computational problems. They are used to represent large sets of sequences that can then be compared simultaneously. Furthermore, they allow the handling of multiple, equally good, alignments, where previous methods were forced to make arbitrary choices. A program written in C implemented this method; it was tested first on 22 5S RNA sequences.

Algorithms↗

Detecting compensatory covariation signals in protein evolution using reconstructed ancestral sequences.

When protein sequences divergently evolve under functional constraints, some individual amino acid replacements that reverse the charge (e.g. Lys to Asp) may be compensated by a replacement at a second position that reverses the charge in the opposite direction (e.g. Glu to Arg). When these side-chains are near in space (proximal), such double replacements might be driven by natural selection, if either is selectively disadvantageous, but both together restore fully the ability of the protein to contribute to fitness (are together "neutral"). Accordingly, many have sought to identify pairs of positions in a protein sequence that suffer compensatory replacements, often as a way to identify positions near in space in the folded structure. A "charge compensatory signal" might manifest itself in two ways. First, proximal charge compensatory replacements may occur more frequently than predicted from the product of the probabilities of individual positions suffering charge reversing replacements independently. Conversely, charge compensatory pairs of changes may be observed to occur more frequently in proximal pairs of sites than in the average pair. Normally, charge compensatory covariation is detected by comparing the sequences of extant proteins at the "leaves" of phylogenetic trees. We show here that the charge compensatory signal is more evident when it is sought by examining individual branches in the tree between reconstructed ancestral sequences at nodes in the tree. Here, we find that the signal is especially strong when the positions pairs are in a single secondary structural unit (e.g. alpha helix or beta strand) that brings the side-chains suffering charge compensatory covariation near in space, and may be useful in secondary structure prediction. Also, "node-node" and "node-leaf" compensatory covariation may be useful to identify the better of two equally parsimonious trees, in a way that is independent of the mathematical formalism used to construct the tree itself. Further, compensatory covariation may provide a signal that indicates whether an episode of sequence evolution contains more or less divergence in functional behavior. Compensatory covariation analysis on reconstructed evolutionary trees may become a valuable tool to analyze genome sequences, and use these analyses to extract biomedically useful information from proteome databases.

Amino Acid Sequence↗

The adaptive evolution database (TAED).

BACKGROUND: The Master Catalog is a collection of evolutionary families, including multiple sequence alignments, phylogenetic trees and reconstructed ancestral sequences, for all protein-sequence modules encoded by genes in GenBank. It can therefore support large-scale genomic surveys, of which we present here The Adaptive Evolution Database (TAED). In TAED, potential examples of positive adaptation are identified by high values for the normalized ratio of nonsynonymous to synonymous nucleotide substitution rates (KA/KS values) on branches of an evolutionary tree between nodes representing reconstructed ancestral sequences. RESULTS: Evolutionary trees and reconstructed ancestral sequences were extracted from the Master Catalog for every subtree containing proteins from the Chordata only or the Embryophyta only. Branches with high KA/KS values were identified. These represent candidate episodes in the history of the protein family when the protein may have undergone positive selection, where the mutant form conferred more fitness than the ancestral form. Such episodes are frequently associated with change in function. An unexpectedly large number of families (between 10% and 20% of those families examined) were found to have at least one branch with high KA/KS values above arbitrarily chosen cut-offs (1 and 0.6). Most of these survived a robustness test and were collected into TAED. CONCLUSIONS: TAED is a raw resource for bioinformaticists interested in data mining and for experimental evolutionists seeking candidate examples of adaptive evolution for further experimental study. It can be expanded to include other evolutionary information (for example changes in gene regulation or splicing) placed in a phylogenetic perspective.

Adaptation, Physiological↗

Functional diversification of lepidopteran opsins following gene duplication.

A comparative approach was taken for identifying amino acid substitutions that may be under positive Darwinian selection and are correlated with spectral shifts among orthologous and paralogous lepidopteran long wavelength-sensitive (LW) opsins. Four novel LW opsin fragments were isolated, cloned, and sequenced from eye-specific cDNAs from two butterflies, Vanessa cardui (Nymphalidae) and Precis coenia (Nymphalidae), and two moths, Spodoptera exigua (Noctuidae) and Galleria mellonella (Pyralidae). These opsins were sampled because they encode visual pigments having a naturally occurring range of lambda(max) values (510-530 nm), which in combination with previously characterized lepidopteran opsins, provide a complete range of known spectral sensitivities (510-575 nm) among lepidopteran LW opsins. Two recent opsin gene duplication events were found within the papilionid but not within the nymphalid butterfly families through neighbor-joining, maximum parsimony, and maximum likelihood phylogenetic analyses of 13 lepidopteran opsin sequences. An elevated rate of evolution was detected in the red-shifted Papilio Rh3 branch following gene duplication, because of an increase in the amino acid substitution rate in the transmembrane domain of the protein, a region that forms the chromophore-binding pocket of the visual pigment. A maximum likelihood approach was used to estimate omega, the ratio of nonsynonymous to synonymous substitutions per site. Branch-specific tests of selection (free-ratio) identified one branch with omega = 2.1044, but the small number of substitutions involved was not significantly different from the expected number of changes under the neutral expectation of omega = 1. Ancestral sequences were reconstructed with a high degree of certainty from these data. Reconstructed ancestral sequences revealed several instances of convergence to the same amino acid between butterfly and vertebrate cone pigments, and between independent branches of the butterfly opsin tree that are correlated with spectral shifts.

Amino Acid Sequence↗

Reconstruction of ancestral sequences by the inferential method, a tool for protein engineering studies.

This paper describes the inferential method, an approach for reconstructing protein and nucleotide sequences of ancestral species, starting from known, homologous, contemporary sequences. The method requires knowledge of the topology of the phylogenetic tree, whose nodes are the species to whom the reconstructed sequences belong. The method has been tested by computer simulation of speciation and nucleotide substitutions, starting from a single ancestral sequence, and by subsequent reconstruction of nodal sequences. Results have shown that reconstructions obtained by the inferential method are affected by limited error frequencies, which (1) are proportional to the squares of nucleotide substitution rates and of internodal distances, and (2) are little influenced by non-uniformity of transformation rates of nucleotides. Furthermore, good agreement of the results has been obtained by comparing protein-sequence reconstructions carried out with the inferential method with those obtained using the maximum parsimony method in two different cases: e.g., a reconstruction of simulated sequences and a reconstruction of mammalian ribonuclease sequences.

Amino Acid Sequence↗

Reconstruction of ancestral protein sequences and its applications.

BACKGROUND: Modern-day proteins were selected during long evolutionary history as descendants of ancient life forms. In silico reconstruction of such ancestral protein sequences facilitates our understanding of evolutionary processes, protein classification and biological function. Additionally, reconstructed ancestral protein sequences could serve to fill in sequence space thus aiding remote homology inference. RESULTS: We developed ANCESCON, a package for distance-based phylogenetic inference and reconstruction of ancestral protein sequences that takes into account the observed variation of evolutionary rates between positions that more precisely describes the evolution of protein families. To improve the accuracy of evolutionary distance estimation and ancestral sequence reconstruction, two approaches are proposed to estimate position-specific evolutionary rates. Comparisons show that at large evolutionary distances our method gives more accurate ancestral sequence reconstruction than PAML, PHYLIP and PAUP*. We apply the reconstructed ancestral sequences to homology inference and functional site prediction. We show that the usage of hypothetical ancestors together with the present day sequences improves profile-based sequence similarity searches; and that ancestral sequence reconstruction methods can be used to predict positions with functional specificity. CONCLUSIONS: As a computational tool to reconstruct ancestral protein sequences from a given multiple sequence alignment, ANCESCON shows high accuracy in tests and helps detection of remote homologs and prediction of functional sites. ANCESCON is freely available for non-commercial use. Pre-compiled versions for several platforms can be downloaded from ftp://iole.swmed.edu/pub/ANCESCON/.

Amino Acid Sequence↗

Probabilistic reconstruction of ancestral protein sequences.

Using a maximum-likelihood formalism, we have developed a method with which to reconstruct the sequences of ancestral proteins. Our approach allows the calculation of not only the most probable ancestral sequence but also of the probability of any amino acid at any given node in the evolutionary tree. Because we consider evolution on the amino acid level, we are better able to include effects of evolutionary pressure and take advantage of structural information about the protein through the use of mutation matrices that depend on secondary structure and surface accessibility. The computational complexity of this method scales linearly with the number of homologous proteins used to reconstruct the ancestral sequence.

Amino Acid Sequence↗

Molecular evolution of cytochrome c oxidase: rate variation among subunit VIa isoforms.

Cytochrome c oxidase (COX) consists of 13 subunits, 3 encoded in the mitochondrial genome and 10 in the nucleus. Little is known of the role of the nuclear-encoded subunits, some of which exhibit tissue-specific isoforms. Subunit VIa is unique in having tissue-specific isoforms in all mammalian species examined. We examined relative evolutionary rates for the COX6A heart (H) and liver (L) isoform genes along the length of the molecule, specifically in relation to the tissue-specific function(s) of the two isoforms. Nonsynonymous (amino acid replacement) substitutions in the COX6AH gene occurred more frequently than in the ubiquitously expressed COX6AL gene. Maximum-parsimony analysis and sequence divergences from reconstructed ancestral sequences revealed that after the ancestral COX6A gene duplicated to yield the genes for the H and L isoforms, the sequences encoding the mitochondrial matrix region of the COX VIa protein experienced an elevated rate of nonsynonymous substitutions relative to synonymous substitutions. This is expected for relaxed selective constraints after gene duplication followed by purifying selection to preserve the replacements with tissue-specific functions.

Amino Acid Sequence↗

A hominoid-specific nuclear insertion of the mitochondrial D-loop: implications for reconstructing ancestral mitochondrial sequences.

A nuclear integration of a mitochondrial control region sequence on human chromosome 9 has been isolated. PCR analyses with primers specific for the respective insertion-flanking nuclear regions showed that the insertion took place on the lineage leading to Hominoidea (gibbon, orangutan, gorilla, chimpanzee, and human) after the Old World monkey-Hominoidea split. The sequences of the control region integrations were determined for humans, chimpanzees, gorillas, orangutans, and siamangs. These sequences were then used to construct phylogenetic trees with different methods, relating them with several hominoid, Old Work monkey, and New World monkey mitochondrial control region sequences. Applying maximum-likelihood, neighbor-joining, and parsimony algorithms, the insertion clade was attached to the branch leading to the hominoid mitochondrial sequences as expected from the PCR-determined presence/absence of this integration. An unexpected long branch leading to the internal node that connects all insertion sequences was observed for the different phylogeny reconstruction procedures. This finding is not totally compatible with the lower evolutionary rate in the nucleus than in the mitochondrial compartment. We determined the unambiguous substitutions on the branch leading to the most recent common ancestor (MRCA) of the mitochondrial inserts according to the parsimony criterium. We propose that they are unlikely to have been caused by damage of the transposing nucleic acid and that they are probably due to a change in the evolutionary mode after the transposition.

Animals↗

The Adaptive Evolution Database (TAED).

BACKGROUND: Developing an understanding of the molecular basis for the divergence of species lies at the heart of biology. The Adaptive Evolution Database (TAED) serves as a starting point to link events that occur at the same time in the evolutionary history (tree of life) of species, based upon coding sequence evolution analyzed with the Master Catalog. The Master Catalog is a collection of evolutionary models, including multiple sequence alignments, phylogenetic trees, and reconstructed ancestral sequences, for all independently evolving protein sequence modules encoded by genes in GenBank [1]. RESULTS: We have estimated from these models the ratio of nonsynonymous to synonymous nucleotide substitution (Ka/Ks), for each branch in their respective evolutionary trees of every subtree containing only chordata or only embryophyta proteins. Branches with high Ka/Ks values represent candidate episodes in the history of the family where the protein may have undergone positive selection, a phenomenon in molecular evolution where the mutant form of a gene must have conferred more fitness than the ancestral form. Such episodes are frequently associated with change in function. We have found that an unexpectedly large number of families (between 10 and 20% of those families examined) have at least one branch with a notably high Ka/Ks value (putative adaptive evolution). As a resource for biologists wishing to understand the interaction between protein sequences and the Darwinian processes that shape these sequences, we have collected these into The Adaptive Evolution Database (TAED). CONCLUSIONS: Placed in a phylogenetic perspective, candidate genes that are undergoing evolution at the same time in the same lineage can be viewed together. This framework based upon coding sequence evolution can be readily expanded to include other types of evolution. In its present form, TAED provides a resource for bioinformaticists interested in data mining and for experimental evolutionists seeking candidate examples of adaptive evolution for further experimental study.

Animals↗

Algorithms to reconstruct past indels: The deletion-only parsimony problem.

Ancestral sequence reconstruction is an important task in bioinformatics, with applications ranging from protein engineering to the study of genome evolution. When sequences can only undergo substitutions, optimal reconstructions can be efficiently computed using well-known algorithms. However, accounting for indels in ancestral reconstructions is much harder. First, for biologically-relevant problem formulations, no polynomial-time exact algorithms are available. Second, multiple reconstructions are often equally parsimonious or likely, making it crucial to correctly display uncertainty in the results. Here, we consider a parsimony approach where only deletions are allowed, while addressing the aforementioned limitations. First, we describe an exact algorithm to obtain all the optimal solutions. The algorithm runs in polynomial time if only one solution is sought. Second, we show that all possible optimal reconstructions for a fixed node can be represented using a graph computable in polynomial time. While previous studies have proposed graph-based representations of ancestral reconstructions, this result is the first to offer a solid mathematical justification for this approach. Finally we provide arguments for the relevance of the deletion-only case for the general case.

Algorithms↗

A branch-and-bound algorithm for the inference of ancestral amino-acid sequences when the replacement rate varies among sites: Application to the evolution of five gene families.

MOTIVATION: We developed an algorithm to reconstruct ancestral sequences, taking into account the rate variation among sites of the protein sequences. Our algorithm maximizes the joint probability of the ancestral sequences, assuming that the rate is gamma distributed among sites. Our algorithm probably finds the global maximum. The use of 'joint' reconstruction is motivated by studies that use the sequences at all the internal nodes in a phylogenetic tree, such as, for instance, the inference of patterns of amino-acid replacement, or tracing the biochemical changes that occurred during the evolution of a given protein family. RESULTS: We give an algorithm that guarantees finding the global maximum. The efficient search method makes our method applicable to datasets with large number sequences. We analyze ancestral sequences of five gene families, exploring the effect of the amount of among-site-rate-variation, and the degree of sequence divergence on the resulting ancestral states. AVAILABILITY AND SUPPLEMENTARY INFORMATION: http://evolu3.ism.ac.jp/~tal/ CONTACT: tal@ism.ac.jp

Algorithms↗

Using ancestral sequences to uncover potential gene homologues.

Gene homologues between distantly related species can be difficult to identify. We test the idea that inferred ancestral sequences could aid in finding gene homologues. Ancestral sequences are inferred by aligning gene homologues on a known tree and estimating the most likely amino acid for each position at each node in that tree. BLAST(R) and HMMER are used separately and together with ancestral sequences to search the genome sequence databases of Encephalitozoon cuniculi, Entamoeba histolytica and Giardia lamblia for RNase P protein homologues. RNase P proteins (Pop4, Pop1, Pop5 and Rpp21) have been reported in humans and at least two other eukaryotic species but have yet to be identified in the above genomes. Using ancestral sequences reconstruction (ASR) for these proteins, we successfully identified putative homologues from E. cuniculi, Ent. histolytica and G. lamblia. In some cases, the use of ASR outperformed BLAST and HMMER. Overall, including ancestral sequences in searches with BLAST and/or HMMER was the most successful approach in the recovery of potential RNase P protein gene homologues, making this a useful technique in early homologue identification.

Algorithms↗

In Silico Reconstruction of the Viral Evolutionary Lineage Yields a Potent Gene Therapy Vector.

Adeno-associated virus (AAV) vectors have emerged as a gene-delivery platform with demonstrated safety and efficacy in a handful of clinical trials for monogenic disorders. However, limitations of the current generation vectors often prevent broader application of AAV gene therapy. Efforts to engineer AAV vectors have been hampered by a limited understanding of the structure-function relationship of the complex multimeric icosahedral architecture of the particle. To develop additional reagents pertinent to further our insight into AAVs, we inferred evolutionary intermediates of the viral capsid using ancestral sequence reconstruction. In-silico-derived sequences were synthesized de novo and characterized for biological properties relevant to clinical applications. This effort led to the generation of nine functional putative ancestral AAVs and the identification of Anc80, the predicted ancestor of the widely studied AAV serotypes 1, 2, 8, and 9, as a highly potent in vivo gene therapy vector for targeting liver, muscle, and retina.

Dependovirus↗

Detecting excess radical replacements in phylogenetic trees.

There are a few instances in which positive Darwinian selection has been convincingly demonstrated at the molecular level. In this study, we present a novel test for detecting excess of radical amino-acid replacements. Such excess is usually indicative of positive Darwinian selection, but may also be due to relaxed functional constraints or model misspecification. In our test, each amino-acid replacement is characterized in terms of a physicochemical distance, i.e., the degree of dissimilarity between the exchanged amino-acid residues. By using phylogenetic trees based on protein sequences, our test identifies statistically significant deviations of the mean physicochemical distance from the random expectation, either along a taxonomic lineage or across a subtree. The mean inferred distance is calculated as the average physicochemical distance over all possible ancestral sequence reconstructions weighted by their likelihood. Our method substantially improves over previous approaches by taking into account the stochastic process, tree phylogeny, among-site rate variation, and alternative ancestral reconstructions. We provide a fast linear time algorithm for applying this test to all branches and all subtrees of a given phylogenetic tree. We validate this approach by applying it to two well-studied datasets: the MHC class I glycoproteins serving as a positive control, and the house-keeping gene carbonic anhydrase I serving as a negative control.

Algorithms↗

Timing and reconstruction of the most recent common ancestor of the subtype C clade of human immunodeficiency virus type 1.

Human immunodeficiency virus type 1 (HIV-1) subtype C is responsible for more than 55% of HIV-1 infections worldwide. When this subtype first emerged is unknown. We have analyzed all available gag (p17 and p24) and env (C2-V3) subtype C sequences with known sampling dates, which ranged from 1983 to 2000. The majority of these sequences come from the Karonga District in Malawi and include some of the earliest known subtype C sequences. Linear regression analyses of sequence divergence estimates (with four different approaches) were plotted against sample year to estimate the year in which there was zero divergence from the reconstructed ancestral sequence. Here we suggest that the most recent common ancestor of subtype C appeared in the mid- to late 1960s. Sensitivity analyses, by which possible biases due to oversampling from one district were explored, gave very similar estimates.

Evolution, Molecular↗