Search PubMedSearch

Biomedical subjects

D Sankoff

Publications and source records attributed to D Sankoff.

At least 19 recordsLinked to original sources

Gene order breakpoint evidence in animal mitochondrial phylogeny.

Multiple genome rearrangement methodology facilitates the inference of animal phylogeny from gene orders on the mitochondrial genome. The breakpoint distance is preferable to other, highly correlated but computationally more difficult, genomic distances when applied to these data. A number of theories of metazoan evolution are compared to phylogenies reconstructed by ancestral genome optimization, using a minimal total breakpoints criterion. The notion of unambiguously reconstructed segments is introduced as a way of extracting the invariant aspects of multiple solutions for a given ancestral genome; this enables a detailed reconstruction of the evolution of non-tRNA mitochondrial gene order.

Animals

Genome structure and gene content in protist mitochondrial DNAs.

Although the collection of completely sequenced mitochondrial genomes is expanding rapidly, only recently has a phylogenetically broad representation of mtDNA sequences from protists (mostly unicellular eukaryotes) become available. This review surveys the 23 complete protist mtDNA sequences that have been determined to date, commenting on such aspects as mitochondrial genome structure, gene content, ribosomal RNA, introns, transfer RNAs and the genetic code and phylogenetic implications. We also illustrate the utility of a comparative genomics approach to gene identification by providing evidence that orfB in plant and protist mtDNAs is the homolog of atp8 , the gene in animal and fungal mtDNA that encodes subunit 8 of the F0portion of mitochondrial ATP synthase. Although several protist mtDNAs, like those of animals and most fungi, are seen to be highly derived, others appear to be have retained a number of features of the ancestral, proto-mitochondrial genome. Some of these ancestral features are also shared with plant mtDNA, although the latter have evidently expanded considerably in size, if not in gene content, in the course of evolution. Comparative analysis of protist mtDNAs is providing a new perspective on mtDNA evolution: how the original mitochondrial genome was organized, what genes it contained, and in what ways it must have changed in different eukaryotic phyla.

Amino Acid Sequence

Counting on comparative maps.

Comparative maps record the history of chromosome rearrangements that have occurred during the evolution of plants and animals. Effective use of these maps in genetic and evolutionary studies relies on quantitative analyses of the patterns of segment conservation. We review the analytical methods that have been developed for characterizing these maps and evaluate their application to existing comparative maps mainly for plants and animals.

Animals

Multiple genome rearrangement and breakpoint phylogeny.

Multiple alignment of macromolecular sequences generalizes from N = 2 to N > or = 3 the comparison of N sequences which have diverged through the local processes of insertion, deletion and substitution. Gene-order sequences diverge through non-local genome rearrangement processes such as inversion (or reversal) and transposition. In this paper we show which formulations of multiple alignment have counterparts in multiple rearrangement. Based on difficulties inherent in rearrangement edit-distance calculation and interpretation, we argue for the simpler "breakpoint analysis." Consensus-based multiple rearrangement of N > or = 3 orders can be solved exactly through reduction to instances of the Travelling Salesman Problem (TSP). We propose a branch-and-bound solution to TSP particularly suited to these instances. Simulations show how non-uniqueness of the solution is attenuated with increasing numbers of data genomes. Tree-based multiple alignment can be achieved to a great degree of accuracy by decomposing the tree into a number of overlapping 3-stars centered on the non-terminal nodes, and solving the consensus-based problem iteratively for these nodes until convergence. Accuracy improves with very careful initializations at the non-terminal nodes. The degree of non-uniqueness of solutions depends on the position of the node in the tree in terms of path length to the terminal vertices.

Algorithms

An ancestral mitochondrial DNA resembling a eubacterial genome in miniature.

Mitochondria, organelles specialized in energy conservation reactions in eukaryotic cells, have evolved from eubacteria-like endosymbionts whose closest known relatives are the rickettsial group of alpha-proteobacteria. Because characterized mitochondrial genomes vary markedly in structure, it has been impossible to infer from them the initial form of the proto-mitochondrial genome. This would require the identification of minimally derived mitochondrial DNAs that better reflect the ancestral state. Here we describe such a primitive mitochondrial genome, in the freshwater protozoon Reclinomonas americana. This protist displays ultrastructural characteristics that ally it with the retortamonads, a protozoan group that lacks mitochondria. R. americana mtDNA (69,034 base pairs) contains the largest collection of genes (97) so far identified in any mtDNA, including genes for 5S ribosomal RNA, the RNA component of RNase P, and at least 18 proteins not previously known to be encoded in mitochondria. Most surprising are four genes specifying a multisubunit, eubacterial-type RNA polymerase. Features of gene content together with eubacterial characteristics of genome organization and expression not found before in mitochondrial genomes indicate that R. americana mtDNA more closely resembles the ancestral proto-mitochondrial genome than any other mtDNA investigated to date.

Animals

Conserved segment identification.

The quantitative study of evolution based on comparative map data is dependent on the definition and identification of conserved segments remaining after interchromosomal exchanges such as reciprocal translocation. Because of experimental error and, more important, extensive local intrachromosomal rearrangement, it is difficult to reconstruct the configuration of conserved segments produced by interchromosomal exchanges. We present a formula to evaluate possible conserved segments and an algorithm which seeks the partition of the genome into segments optimal under this evaluation. Application is made to the human-mouse comparison.

Algorithms

Synteny conservation and chromosome rearrangements during mammalian evolution.

An important problem in comparative genome analysis has been defining reliable measures of synteny conservation. The published analytical measures of synteny conservation have limitations. Nonindependence of comparisons, conserved and disrupted syntenies that are as yet unidentified, and redundant rearrangements lead to systematic errors that tend to overestimate the degree of conservation. We recently derived methods to estimate the total number of conserved syntenies within the genome, counting both those that have already been described and those that remain to be discovered. With this method, we show that approximately 65% of the conserved syntenies have already been identified for humans and mice, that rates of synteny disruption vary approximately 25-fold among mammalian lineages, and that despite strong selection against reciprocal translocations, inter-chromosome rearrangements occurred approximately fourfold more often than inversions and other intra-chromosome rearrangements, at least for lineages leading to humans and mice.

Animals

Comparable rates of gene loss and functional divergence after genome duplications early in vertebrate evolution.

Duplicated genes are an important source of new protein functions and novel developmental and physiological pathways. Whereas most models for fate of duplicated genes show that they tend to be rapidly lost, models for pathway evolution suggest that many duplicated genes rapidly acquire novel functions. Little empirical evidence is available, however, for the relative rates of gene loss vs. divergence to help resolve these contradictory expectations. Gene families resulting from genome duplications provide an opportunity to address this apparent contradiction. With genome duplication, the number of duplicated genes in a gene family is at most 2n, where n is the number of duplications. The size of each gene family, e.g., 1, 2, 3, ..., 2n, reflects the patterns of gene loss vs. functional divergence after duplication. We focused on gene families in humans and mice that arose from genome duplications in early vertebrate evolution and we analyzed the frequency distribution of gene family size, i.e., the number of families with two, three or four members. All the models that we evaluated showed that duplicated genes are almost as likely to acquire a new and essential function as to be lost through acquisition of mutations that compromise protein function. An explanation for the unexpectedly high rate of functional divergence is that duplication allows genes to accumulate more neutral than disadvantageous mutations, thereby providing more opportunities to acquire diversified functions and pathways.

Animals

Parametric genome rearrangement.

Algorithms inspired by comparative genomics calculate an edit distance between two linear orders based on elementary edit operations such as inversion, transposition and reciprocal translocation. All operations are generally assigned the same weight, simply by default, because no systematic empirical studies exist verifying whether algorithmic outputs involve realistic proportion of each. Nor do we have data on how weights should vary with the length of the inverted or transposed segment of the chromosome. In this paper, we present a rapid algorithm that allows each operation to take on a range of weights, producing an relatively tight upper bound on the distance between single-chromosome genomes, by means of a greedy search with look-ahead. The efficiency of this algorithm allows us to test random genomes for each parameter setting, to detect gene order similarity and to infer the parameter values most appropriate to the phylogenetic domain under study. We apply this method to genome segments in which the same gene order is conserved in Escherichia coli and Bacillus subtilis, as well as to the gene order in human versus Drosophila mitochondrial genomes. In both cases, we conclude that it is most appropriate to assign somewhat more than twice the weight to transpositions and inverted transpositions than to inversions. We also explore segment-length weighting for fungal mitochondrial gene orders.

Algorithms

Evolution of fragmented mitochondrial ribosomal RNA genes in Chlamydomonas.

The fragmented mitochondrial ribosomal RNAs (rRNAs) of the green algae Chlamydomonas eugametos and Chlamydomonas reinhardtii are discontinuously encoded in subgenic modules that are scrambled in order and interspersed with protein coding and tRNA genes. The mitochondrial rRNA genes of these two algae differ, however, in both the distribution and organization of rRNA coding information within their respective genomes. The objectives of this study were (1) to examine the phylogenetic relationships between the mitochondrial rRNA gene sequences of C. eugametos and C. reinhardtii and those of the conventional mitochondrial rRNA genes of the green alga, Prototheca wickerhamii, and land plants and (2) to attempt to deduce the evolutionary pathways that gave rise to the unusual mitochondrial rRNA gene structures in the genus Chlamydomonas. Although phylogenetic analysis revealed an affiliation between the mitochondrial rRNA gene sequences of the two Chlamydomonas taxa to the exclusion of all other mitochondrial rRNA gene sequences tested, no specific affiliation was noted between the Chlamydomonas sequences and P. wickerhamii or land plants. Calculations of the minimal number of transpositions required to convert hypothetical ancestral rRNA gene organizations to the arrangements observed for C. eugametos and C. reinhardtii mitochondrial rRNA genes, as well as a limited survey of the size of mitochondrial rRNAs in other members of the genus, lead us to propose that the last common ancestor of Chlamydomonas algae contained fragmented mitochondrial rRNA genes that were nearly co-linear with conventional rRNA genes.

Animals

A remarkable nonlinear invariant for evolution with heterogeneous rates.

A model for DNA or protein sequence evolution is proposed where each position belongs to one of two distinct classes. The two classes evolve at different rates. For a phylogeny on four species, we find a cubic function of 4-tuple occurrence frequencies that is nontrivially invariant no matter what the proportion of positions in each rate class. This result refutes the major criticism of nonlinear polynomial invariants.

DNA

Karyotype distributions in a stochastic model of reciprocal translocation.

A random process of reciprocal translocation for a fixed number k of chromosomes (or arms) will have an equilibrium distribution of chromosome lengths. In this paper we calculate this distribution, by analytical means for k = 2 and partially for k = 3, and simulate the means of the marginal distributions for higher k. We compare this with a random (i.e., ahistorical) distribution of genomic DNA among k chromosomes and to a selection of karyotypes of real organisms. The results motivate a revised model where translocations giving rise to undersize chromosomes are disadvantaged.

Karyotyping

Phylogenetic invariants for more general evolutionary models.

An invariant Q of a tree T under a k-state Markov model, where a generalized time parameter is identified with the E edges of T, allows us to recognize whether data on N observed species (usually, N DNA sequences, one from each species) can be associated with the N leaves of T in the sense of having been generated on T rather than on any other N-leaf tree. The form of the generalized time parameter is a positive determinant matrix in some semigroup S of Markov matrices. The invariance is with respect to the choice of the set of E matrices in S, one associated with each of the E edges of T. The parametric form of S represents a model of the evolutionary process. In this paper, we apply a general method of finding invariants of a parametrized functional form to find low-degree polynomial invariants for different models. Quadratic invariants are obtained for the Kimura two-parameter model, for a model allowing evolutionary dependence between positions in the sequences and for an asymmetric model that allows for A + T versus G + C asymmetries in DNA base composition. Those invariants are found for trees (unrooted in case of the Kimura model and rooted for the others) with N = 3 or N = 4 terminal vertices. We also find cubic invariants for a ten-parameter model with k = 4 states, for rooted trees with N = 4. In each case, we use implicit function theory to predict the number of algebraically independent invariants and then use this prediction to guide a systematic search for algebraic dependence within the set of invariants produced by our method.

Animals

Skewed base compositions, asymmetric transition matrices, and phylogenetic invariants.

Evolutionary inference methods that assume equal DNA base compositions and symmetric nucleotide substitution matrices, where these assumptions do not hold, are likely to group species on the basis of similar base compositions rather than true phylogenetic relationships. We propose an invariants-based method for dealing with this problem. An invariant QT of a tree T under a k-state Markov model, where a generalized time parameter is identified with the E edges of T, allows us to recognize whether data on N observed species can be associated with the N terminal vertices of T in the sense of having been generated on T rather than on any other tree with N terminals. The form of the generalized time parameter is a positive determinant matrix in some semigroup S of stochastic matrices. The invariance is with respect to the choice of the set of E matrices in S, one associated with each of the E edges of T. We apply a general "empirical" method of finding invariants of a parametrized functional form. It involves calculating the probability f of all KN data possibilities for each of m sets of E matrices in S to associate with the edges of T, then solving for the parameters using the m equations of form Q(f) = 0. We discuss the problems of finding asymmetric models satisfying the property of semigroup closure, of finding asymmetric models that admit invariants at all, and of the computational complexity of the method. We propose a class of semigroups Sc containing matrices of form [formula: see text] to account for A+T versus G+C asymmetries in DNA base composition. Quadratic invariants are obtained for rooted trees with three and with four terminals. In the latter case the smallest set of algebraically independent invariants is sought. These invariants are applied to data pertaining the fungal evolution and to the origin of mitochondria as bacterial endosymbionts.

Algorithms

Analytical approaches to genomic evolution.

We model the non-local mechanisms of genomic evolution and propose methods for studying the evolutionary divergence of species based on these models. Mechanisms include the movement of segments of genomes within a single chromosome (transpositions), the reciprocal translocation of segments between two chromosomes, and the inversion of segments. Each of these is studied in the context of a different type of genomic data. We introduce the theory of phylogenetic invariants for evolutionary inference based on very long macromolecular sequences.

Animals

Gene order comparisons for phylogenetic inference: evolution of the mitochondrial genome.

Detailed knowledge of gene maps or even complete nucleotide sequences for small genomes leads to the feasibility of evolutionary inference based on the macrostructure of entire genomes, rather than on the traditional comparison of homologous versions of a single gene in different organisms. The mathematical modeling of evolution at the genomic level, however, and the associated inferential apparatus are qualitatively different from the usual sequence comparison theory developed to study evolution at the level of individual gene sequences. We describe the construction of a database of 16 mitochondrial gene orders from fungi and other eukaryotes by using complete or nearly complete genomic sequences; propose a measure of gene order rearrangement based on the minimal set of chromosomal inversions, transpositions, insertions, and deletions necessary to convert the order in one genome to that of the other; report on algorithm design and the development of the DERANGE software for the calculation of this measure; and present the results of analyzing the mitochondrial data with the aid of this tool.

Biological Evolution