Search PubMed⌕ Search

Biomedical subjects

W F Doolittle

Publications and source records attributed to W F Doolittle.

At least 73 records · Page 4Linked to original sources

The Sulfolobus solfataricus P2 genome project.

Over 800 kbp of the 3-Mbp genome of Sulfolobus solfataricus have been sequenced to date. Our approach is to sequence subclones of mapped cosmids, followed by sequencing directly on cosmid templates with custom primers. Using a prototype automated system for genome-scale analysis, known as MAGPIE, along with other tools, we have discovered one open reading frame of at least 100 amino acids per kbp of sequence, and have been able to associate 50% of these with known genes through database searches. An examination of completely sequenced cosmids suggests a clustering of genes by function in the S. solfataricus genome.

Databases, Factual↗

A non-canonical genetic code in an early diverging eukaryotic lineage.

The nearly invariant nature of the 'Universal Genetic Code' attests to its early establishment in evolution and to the difficulty of altering it now, since so many molecules are required for, and depend upon, faithful translation. Nevertheless, variations on the universal code are known in a handful of genomes. We have found one such variant in diplomonads, an early-diverging eukaryotic lineage. Genes for alpha-tubulin, beta-tubulin and elongation factor 1 alpha (EF-1alpha) from two unclassified strains of Hexamitidae were found to contain TAA and TAG (TAR) triplets at positions suggesting a variant code in which TAR codes for glutamine. We found confirmation of this hypothesis by identifying genes encoding glutamine-tRNAs with CUA and UUA anticodons. The alpha-tubulin and EF-1alpha genes from two other diplomonads, Spironucleus muris and Hexamita inflata, were also sequenced and shown to contain no such non-canonical codons. However, tRNA genes with the anticodons UUA and CUA were found in H.inflata, suggesting that this diplomonad also uses these codons, albeit infrequently. The high GC content of these genomes and the presence of two isoaccepting tRNAs compound the difficulty of understanding how this variant code arose by strictly neutral means.

Amino Acid Sequence↗

Complete nucleotide sequence of the Sulfolobus islandicus multicopy plasmid pRN1.

The complete sequence of the 5350-bp plasmid pRN1 from the crenarchaeote Sulfolobus islandicus has been determined. This plasmid is the first to be sequenced from this group of thermoacidophilic archaebacteria (Archaea) and its high copy number and wide host range make it a good candidate for a cloning vector. pRN1 contains several open reading frames, including one that spans over half the plasmid and has significant similarity to the helicase domain of viral primase proteins. Directly upstream of this putative primase is a homologue of Cop, a family of small proteins from promiscuous eubacterial plasmids which control copy number by repressing the expression of the replication initiation protein. In eubacterial plasmids cop is found upstream of the replication initiator protein. The location of a cop homologue upstream of a primase-like gene in pRN1 suggests that it controls DNA replication in a manner similar to these eubacterial plasmids, but does so using a mixture of components from plasmids and viruses.

Amino Acid Sequence↗

Alpha-tubulin from early-diverging eukaryotic lineages and the evolution of the tubulin family.

The tubulin gene family, which includes alpha-,beta-, and gamma-tubulin subfamilies, is composed of highly conserved proteins which are the principle structural and functional components of eukaryotic microtubules. We are interested in (1) establishing when in eukaryotic evolution the duplications leading to paralogous alpha, beta, and gamma subfamilies occurred and (2) the possible utility of tubulin sequences in reconstructing organismal phylogeny. To broaden the taxonomic representation of alpha-tubulins so that it roughly equals that of beta-tubulins, alpha-tubulin genes from three Microsporidia (Encephalitozoon hellem, Nosema locustae, and Spraguea lophii), two Parabasalia (Monocercomonas sp. and Trichomitus batrachorum), and one Heterolobosean (Acrasis rosea) were sequenced. With these new genes, phylogenetic trees of alpha- and beta-tubulins were constructed and compared. Trees were congruent with each other, but incongruent with other molecular phylogenies. The agreement between alpha- and beta-tubulin trees could arise by the co-adaptation of one molecule to variants of the other as a result of their intimate steric association in microtubules. Thus, these trees may not be providing independent support for the phylogenetic results. However, one of these unexpected results, that microsporidia cluster with fungi, is supported by other circumstantial evidence, and may therefore reflect a real relationship despite the basal position usually assigned to microsporidia. Relationships between the three tubulins were also examined by constructing trees of all three types. These trees were found to be of limited value for determining the position of the root within each subfamily because of the great interfamily distances, but they do confirm the classification of all known genes into three monophyletic subfamilies. Divergent genes from Caenorhabditis elegans and Saccharomyces cerevisiae that have been proposed to represent the novel classes delta- and epsilon-tubulin were found to be specifically related to gamma-tubulins from animals and fungi respectively, and therefore are best seen as rapidly evolving orthologues of gamma-tubulin.

Amino Acid Sequence↗

Organizational characteristics and information content of an archaeal genome: 156 kb of sequence from Sulfolobus solfataricus P2.

We have initiated a project to sequence the 3 Mbp genome of the thermoacidophilic archaebacterium Sulfolobus solfataricus P2. Cosmids were selected from a provisional set of minimally overlapping clones, subcloned in pUC18, and sequenced using a hybrid (random plus directed) strategy to give two blocks of contiguous unique sequence, respectively, 100,389 and 56,105 bp. These two contigs contain a total of 163 open reading frames (ORFs) in 26-29 putative operons; 56 ORFs could be identified with reasonable certainty. Clusters of ORFs potentially encode proteins of glycogen biosynthesis, oxidative decarboxylation of pyruvate, ATP-dependent transport across membranes, isoprenoid biosynthesis, protein synthesis, and ribosomes. Putative promoters occur upstream of most ORFs. Thirty per cent of the predicted strong and medium-strength promoters can initiate transcription at the start codon or within 10 nucleotides upstream, indicating a process of initial mRNA-ribosome contact unlike that of most eubacterial genes. A novel termination motif is proposed to account for 15 additional terminations. The two contigs differ in densities of ORFs, insertion elements and repeated sequences; together they contain two copies of the previously reported insertion sequence ISC 1217, five additional IS elements representing four novel types, four classes of long non-IS repeated sequences, and numerous short, perfect repeats.

Chromosomes, Bacterial↗

Root of the universal tree of life based on ancient aminoacyl-tRNA synthetase gene duplications.

Universal trees based on sequences of single gene homologs cannot be rooted. Iwabe et al. [Iwabe, N., Kuma, K.-I., Hasegawa, M., Osawa, S. & Miyata, T. (1989) Proc. Natl. Acad. Sci. USA 86, 9355-9359] circumvented this problem by using ancient gene duplications that predated the last common ancestor of all living things. Their separate, reciprocally rooted gene trees for elongation factors and ATPase subunits showed Bacteria (eubacteria) as branching first from the universal tree with Archaea (archaebacteria) and Eucarya (eukaryotes) as sister groups. Given its topical importance to evolutionary biology and concerns about the appropriateness of the ATPase data set, an evaluation of the universal tree root using other ancient gene duplications is essential. In this study, we derive a rooting for the universal tree using aminoacyl-tRNA synthetase genes, an extensive multigene family whose divergence likely preceded that of prokaryotes and eukaryotes. An approximately 1600-bp conserved region was sequenced from the isoleucyl-tRNA synthetases of several species representing deep evolutionary branches of eukaryotes (Nosema locustae), Bacteria (Aquifex pyrophilus and Thermotoga maritima) and Archaea (Pyrococcus furiosus and Sulfolobus acidocaldarius). In addition, a new valyl-tRNA synthetase was characterized from the protist Trichomonas vaginalis. Different phylogenetic methods were used to generate trees of isoleucyl-tRNA synthetases rooted by valyl- and leucyl-tRNA synthetases. All isoleucyl-tRNA synthetase trees showed Archaea and Eucarya as sister groups, providing strong confirmation for the universal tree rooting reported by Iwabe et al. As well, there was strong support for the monophyly (sensu Hennig) of Archaea. The valyl-tRNA synthetase gene from Tr. vaginalis clustered with other eukaryotic ValRS genes, which may have been transferred from the mitochondrial genome to the nuclear genome, suggesting that this amitochondrial trichomonad once harbored an endosymbiotic bacterium.

Amino Acid Sequence↗

Concerted evolution in protists: recent homogenization of a polyubiquitin gene in Trichomonas vaginalis.

Ubiquitin is a 76-amino-acid protein with a remarkably high degree of conservation between all known sequences. Ubiquitin genes are almost always multicopy in eukaryotes, and often are found as polyubiquitin genes--fused tandem repeats which are coexpressed. Seventeen ubiquitin sequences from the amitochondrial protist Trichomonas vaginalis have been examined here, including an 11-repeat fragment of a polyubiquitin gene. These sequences reveal a number of interesting features that are not seen in other eukaryotes. The predicted amino acid sequences lack several universally conserved residues, and individual units do not always encode identical peptides as is usually the case. On the nucleotide level, these repeats are in general highly variable, but one region in the polyubiquitin is extremely homogeneous, with seven repeats absolutely identical. Such extended stretches of homogeneity have never been observed in ubiquitin genes and since substitutions are common in other coding units, it is likely that these repeats are the product of a very recent homogenization or amplification.

Amino Acid Sequence↗

Methods for evaluating exon-protein correspondences.

According to the exon theory of genes, protein-coding genes evolved originally by combinatorial assembly of mini-gene precursors of modern exons. If so, then exons should tend to encode discrete bits of protein structure, as first suggested by C.C.F. Blake. In order to assess the evidence for Blake's conjecture, we have developed methods for evaluating the significance of correspondences between split gene structure and protein structure, using computer programs for measuring observed correspondences and comparing them to random expectations. Initial results of applying these methods to data on ancient proteins have been presented elsewhere. Here we describe the algorithms in detail, and demonstrate their effectiveness in finding correlations in idealized test cases. The likely effects of deletion and putative displacement ('sliding') of introns on the ability to detect correlations are also examined.

Algorithms↗

Evolution. Archaea and eukaryotes versus bacteria?

The recent discovery of homologs of the eukaryotic transcription factor TATA-binding protein in archaea has been taken as support for the view that archaea and eukaryotes have a close phylogenetic relationship.

Archaea↗

Tempo, mode, the progenote, and the universal root.

Early cellular evolution differed in both mode and tempo from the contemporary process. If modern lineages first began to diverge when the phenotype-genotype coupling was still poorly articulated, then we might be able to learn something about the evolution of that coupling through comparing the molecular biologies of living organisms. The issue is whether the last common ancestor of all life, the cenancestor, was a primitive entity, a progenote, with a more rudimentary genetic information-transfer system. Thinking on this issue is still unsettled. Much depends on the placement of the root of the universal tree and on whether or not lateral transfer renders such rooting meaningless.

Animals↗

Testing the exon theory of genes: the evidence from protein structure.

A tendency for exons to correspond to discrete units of protein structure in protein-coding genes of ancient origin would provide clear evidence in favor of the exon theory of genes, which proposes that split genes arose not by insertion of introns into unsplit genes, but from combinations of primordial mini-genes (exons) separated by spacers (introns). Although putative examples of such correspondence have strongly influenced previous debate on the origin of introns, a general correspondence has not been rigorously proved. Objective methods for detecting correspondences were developed and applied to four examples that have been cited previously as evidence of the exon theory of genes. No significant correspondence between exons and units of protein structure was detected, suggesting that the putative correspondence does not exist and that the exon theory of genes is untenable.

Alcohol Dehydrogenase↗

Evolutionary relationships of bacterial and archaeal glutamine synthetase genes.

Glutamine synthetase (GS), an essential enzyme in ammonia assimilation and glutamine biosynthesis, has three distinctive types: GSI, GSII and GSIII. Genes for GSI have been found only in bacteria (eubacteria) and archaea (archaebacteria), while GSII genes only occur in eukaryotes and a few soil-dwelling bacteria. GSIII genes have been found in only a few bacterial species. Recently, it has been suggested that several lateral gene transfers of archaeal GSI genes to bacteria may have occurred. In order to study the evolution of GS, we cloned and sequenced GSI genes from two divergent archaeal species: the extreme thermophile Pyrococcus furiosus and the extreme halophile Haloferax volcanii. Our phylogenetic analysis, which included most available GS sequences, revealed two significant prokaryotic GSI subdivisions: GSI-alpha and GSI-beta. GSI-alpha-genes are found in the thermophilic bacterium, Thermotoga maritima, the low G+C Gram-positive bacteria, and the Euryarchaeota (includes methanogens, halophiles, and some thermophiles). GSI-beta-type genes occur in all other bacteria. GSI-alpha- and GSI-beta-type genes also differ with respect to a specific 25-amino-acid insertion and adenylylation control of GS enzyme activity, both absent in the former but present in the latter. Cyanobacterial genes lack adenylylation regulation of GS and may have secondarily lost it. The GSI gene of Sulfolobus solfataricus, a member of the Crenarchaeota (extreme thermophiles), is exceptional and could not be definitely placed in either subdivision.

Amino Acid Sequence↗

Archaebacterial genomes: eubacterial form and eukaryotic content.

Since the recognition of the uniqueness and coherence of the archaebacteria (sometimes called Archaea), our perception of their role in early evolution has been modified repeatedly. The deluge of sequence data and rapidly improving molecular systematic methods have combined with a better understanding of archaebacterial molecular biology to describe a group that in some ways appears to be very similar to the eubacteria, though in others is more like the eukaryotes. The structure and contents of archaebacterial genomes are examined here, with an eye to their meaning in terms of the evolution of cell structure and function.

Archaea↗

Construction of composite transposons for halophilic Archaea.

Transposons with selectable marker genes (e.g., antibiotic resistance) have been extremely useful tools in bacterial genetics but have not been found naturally in Archaea. We constructed synthetic transposons consisting of halobacterial ISH elements (ISH2, ISH26, or ISH28) flanking a mevinolin resistance determinant. Introduction of these constructs into Haloferax volcanii cells can produce drug-resistant transformants through homologous recombination between the plasmid hmgA gene and the chromosomal hmgA locus. This problem was overcome by using another host, Haloarcula hispanica, the hmgA gene of which shares little homology with that from Haloferax volcanii. Introduction of an ISH28-based transposon (ThD28) into Haloarcula hispanica cells produced numerous transformants. Each of these was shown to contain an ISH-flanked mevinolin resistance determinant integrated into the cellular DNA. Integration was not obviously site specific. Transposon ThD26 (based on ISH26a), was less mobile, relative to ThD28, and the ISH2-based construct (ThD22) did not transpose at all in these cells. The further development of halobacterial transposons may provide useful genetic tools allowing rapid isolation and analysis of halobacterial genes, particularly those with no selectable phenotype.

Base Sequence↗