Search PubMed⌕ Search

Biomedical subjects

D Sankoff

Publications and source records attributed to D Sankoff.

51 records · Page 3Linked to original sources

On the evolutionary descent of organisms and organelles: a global phylogeny based on a highly conserved structural core in small subunit ribosomal RNA.

To probe the earliest evolutionary events attending the origin of the five known genome types (archaebacterial, eubacterial, nuclear, mitochondrial and plastid), we have analyzed sequences corresponding to a ubiquitous, highly conserved core of secondary structure in small subunit rRNA. Our results support (i) the existence of three primary lineages (archaebacterial, eubacterial, and nuclear), (ii) a specific eubacterial ancestry for plastids and mitochondria (plant, animal, fungal), and (iii) an endosymbiotic, evolutionary origin of the two types of organelle from within distinct groups of eubacteria (blue-green algae (cyanobacteria) in the case of plastids, nonphotosynthetic aerobic bacteria in the case of mitochondria). In addition, our analysis suggests (iv) a biphyletic origin of mitochondria, with animal and fungal mitochondria branching together but separately from plant mitochondria, and (v) a monophyletic origin of plastids. The method described here provides a powerful and generally applicable molecular taxonomic approach towards a global phylogeny encompassing all organisms and organelles.

Animals↗

An algorithm for the display of nucleic acid secondary structure.

A simple algorithm is presented for the graphic display of nucleic acid secondary structure. Examples of secondary structure displays are given for tRNA, 5S RNA and part of the 16S RNA. Due to its speed, this algorithm could easily be used in conjunction with secondary structure programs which calculate various alternate structures.

Base Sequence↗

A strategy for sequence phylogeny research.

Minimal mutation trees, and almost minimal trees, are constructed from two data sets, one of phenylalanine tRNA sequences, and the other of 5S RNA sequences, from a diverse range of organisms. The two sets of results are mutually consistent. Trees representing previous evolutionary hypotheses are compared using a total weighted mutational distance criterion. The importance of sequence data from relatively little-studed phylogenetic lines is stressed. A procedure is illustrated which circumvents the computational difficulty of evaluating the astronomically large number of possible trees, without resorting to suboptimal methods.

Bacteria↗

The evolving tRNA molecule.

The study of tRNA molecular evolution is crucial to understanding the origin and establishment of the genetic code as well as the differentiation and refinement of the machinery of protein synthesis in prokaryotes, eukaryotes, organelles, and phage systems. The small size of the molecule and its critical involvement in a multiplicity of roles distinguish its study from classical protein molecular evolution with respect to goals and methods. Here, the authors assess available and missing data, existing and needed methodology, and the impact of tRNA studies on current theories both of genetic code evolution and of the evolution of species. They analyze mutational "hot spots", the role of base modification, synthetase recognition, codon-anticodon interactions and the status of organelle tRNA.

Animals↗

Convergence and minimal mutation criteria for evaluating early events in tRNA evolution.

The convergence of ancestral sequences independently constructed from different branches of a phylogenetic tree can be used as a test of homology of data sequences. This criterion has shown that all phenylalanine tRNAs are related to a common ancestor, whereas eukaryotic and prokaryotic tyrosine tRNAs may have independent origins. All glycine tRNAs share a common ancestor. The glycine tRNA family splits according to the purine or pyrimidine nature of the first anticodon base prior to the divergence of eukaryotes and prokaryotes. The structural similarity between some prokaryotic glycine and and valine tRNAs is the result of their derivation from a common ancestor that existed previous to the divergence of the different glycine tRNAs. These results support models of genetic code evolution involving the incremental elaboration of earlier, simpler codes.

Animals↗

Evolution of methionine initiator and phenylalanine transfer RNAs.

Sequence data from methionine initiator and phenylalanine transfer RNAs were used to construct phylogenetic trees by the maximum parsimony method. Although eukaryotes, prokaryotes and chloroplasts appear related to a common ancestor, no firm conclusion can be drawn at this time about mitochondrial-coded transfer RNAs. tRNA evolution is not appropriately described by random hit models, since the various regions of the molecule differ sharply in their mutational fixation rates. "Hot" mutational spots are identified in the Tpsic, the amino acceptor and the upper anticodon stems; the D arm and the loop areas on the other hand are highly conserved. Crucial tertiary interactions are thus essentially preserved while most of the double helical domain undergoes base pair interchange. Transitions are about half as costly as transversions, suggesting that base pair interchanges proceed mostly through G-U and A-C intermediates. There is a preponderance of replacements starting from G and C but this bias appears to follow the high G + C content of the easily mutated base paired regions.

Animals↗

Bacteriophage MS2 RNA: a correlation between the stability of the codon: anticodon interaction and the choice of code words.

The non-random distribution of degenerate code words in Bacteriophage MS2 RNA can be explained partially by considerations of the stability of the codon-anticodon complex in prokaryotic systems. Supporting this hypothesis we note that wobble codons are positively selected in codons having G and/or C in the first two positions. In contrast, wobble codons are statistically less likely in codons composed of A and U in the first two positions. Analyses of nucleotides adjacent to 5' and 3' ends of codons indicate a nonrandom distribution as well. It is thus likely that some elements of RNA evolution are independent of the structural needs of the RNA itself and of the translated protein product.

Anticodon↗

Frequency of insertion-deletion, transversion, and transition in the evolution of 5S ribosomal RNA.

The problem of choosing an alignment of two or more nucleotide sequences is particularly difficult for nucleic acids, such as 5S ribosomal RNA, which do not code for protein and for which secondary structure is unknown. Given a set of 'costs' for the various types of replacement mutations and for base insertion or deletion, we present a dynamic programming algorithm which finds the optimal (least costly) alignment for a set of N sequences simultaneously, where each sequence is associated with one of the N tips of a given evolutionary tree. Concurrently, protosequences are constructed corresponding to the ancestral nodes of the tree. A version of this algorithm, modified to be computationally feasible, is implemented to align the sequences of 5S RNA from nine organisms. Complete sets of alignments and protosequence reconstructions are done for a large number of different configurations of mutation costs. Examination of the family of curbes of total replacements inferred versus the ratio of transitions/transversions inferred, each curve corresponding to a given number of insertions-deletions inferred, provides a method for estimating relative costs and relative frequencies for these different types of mutations.

Base Sequence↗

Matching sequences under deletion-insertion constraints.

Given two finite sequences, we wish to find the longest common subsequences satisfying certain deletion/insertion constraints. Consider two successive terms in the desired subsequence. The distance between their positions must be the same in the two original sequences for all but a limited number of such pairs of successive terms. Needleman and Wunsch gave an algorithm for finding longest common subsequences without constraints. This is improved from the viewpoint of computational economy. An economical algorithm is then elaborated for finding subsequences satisfying deletion/insertion constraints. This result is useful in the study of genetic homology based on nucleotide or amino-acid sequences.

Amino Acid Sequence↗

The evolutionary relationships among known life forms.

Sequences of small subunit (SSU) and large subunit (LSU) ribosomal RNA genes from archaebacteria, eubacteria, and the nucleus, chloroplasts, and mitochondria of eukaryotes have been compared in order to identify the most conservative positions. Aligned sets of these positions for both SSU and LSU rRNA have been used to generate tree diagrams relating the source organisms/organelles. Branching patterns were evaluated using the statistical bootstrapping technique. The resulting SSU and LSU trees are remarkably congruent and show a high degree of similarity with those based on alternative data sets and/or generated by different techniques. In addition to providing insights into the evolution of prokaryotic and eukaryotic (nuclear) lineages, the analysis reported here provides, for the first time, an extensive phylogeny of the mitochondrial lineage.

Base Sequence↗

Phylogenetic invariants for genome rearrangements.

We review the combinatorial optimization problems in calculating edit distances between genomes and phylogenetic inference based on minimizing gene order changes. With a view to avoiding the computational cost and the "long branches attract" artifact of some tree-building methods, we explore the probabilization of genome rearrangement models prior to developing a methodology based on branch-length invariants. We characterize probabilistically the evolution of the structure of the gene adjacency set for reversals on unsigned circular genomes and, using a nontrivial recurrence relation, reversals on signed genomes. Concepts from the theory of invariants developed for the phylogenetics of homologous gene sequences can be used to derive a complete set of linear invariants for unsigned reversals, as well as for a mixed rearrangement model for signed genomes, though not for pure transposition or pure signed reversal models. The invariants are based on an extended Jukes-Cantor semigroup. We illustrate the use of these invariants to relate mitochondrial genomes from a number of invertebrate animals.

Algorithms↗