Search PubMed⌕ Search

Biomedical subjects

T Gaasterland

Publications and source records attributed to T Gaasterland.

32 records · Page 2Linked to original sources

The complete genome of the hyperthermophilic bacterium Aquifex aeolicus.

Aquifex aeolicus was one of the earliest diverging, and is one of the most thermophilic, bacteria known. It can grow on hydrogen, oxygen, carbon dioxide, and mineral salts. The complex metabolic machinery needed for A. aeolicus to function as a chemolithoautotroph (an organism which uses an inorganic carbon source for biosynthesis and an inorganic chemical energy source) is encoded within a genome that is only one-third the size of the E. coli genome. Metabolic flexibility seems to be reduced as a result of the limited genome size. The use of oxygen (albeit at very low concentrations) as an electron acceptor is allowed by the presence of a complex respiratory apparatus. Although this organism grows at 95 degrees C, the extreme thermal limit of the Bacteria, only a few specific indications of thermophily are apparent from the genome. Here we describe the complete genome sequence of 1,551,335 base pairs of this evolutionarily and physiologically interesting organism.

Chromosome Mapping↗

Completing the sequence of the Sulfolobus solfataricus P2 genome.

The Sulfolobus solfataricus P2 genome collaborators are poised to sequence the entire 3-Mbp genome of this crenarchaeote archaeon. About 80% of the genome has been sequenced to date, with the rest of the sequence being assembled fast. In this publication we introduce the genomic sequencing and automated analysis strategy and present intial data derived from the sequence analysis. After an overview of the general sequence features, metabolic pathway studies are explained, using sugar metabolism as an example. The paper closes with an overview of repetitive elements in S. solfataricus.

Base Sequence↗

Constructing multigenome views of whole microbial genomes.

We have designed and implemented a system to carry out cross-genome comparisons of open reading frames (ORFs) from multiple genomes. This implementation includes a genome profiling system that allows us to explore pairwise comparisons at different levels of match similarity and ask biologically motivated queries involving number and identity of ORFs, their function, functional category, distribution in genomes or in biological domains, and statistics on their matches and match families. This analysis required precise definition of new classification terms and concepts. We define the terms genomic signature, summary signature, biologic domain signature, domain class, match level, match family, and extended match family, then use these terms to define concepts, including genomically universal proteins and proteins characteristics of sets of genomes. We initiate an analysis based on automated FASTA (Pearson, 1996) comparison of 22,419 conceptually translated protein sequences from nine microbial genomes.

Amino Acid Sequence↗

Microbial genescapes: phyletic and functional patterns of ORF distribution among prokaryotes.

We have implemented a statistically based approach to comparative genomics that allows us to define and characterize distributional patterns of conceptually translated open reading frames (ORFs) at different confidence levels based on pairwise FASTA matches. In this report, we apply this methodology to nine microbial genomes, focusing particularly on phyletic and functional patterns of ORF distribution within and between the two prokaryotic domains of life, Bacteria and Archaea. We examine patterns of presence and absence of matches, determine the universal ORF set, analyze features of genome specialization between closely related organisms, and present genomic evidence for the monophyly of Archaea. These analyses illustrate how a quantitative approach to comparative genomics can illuminate questions of fundamental biological significance.

Archaea↗

Microbial genescapes: a prokaryotic view of the yeast genome.

We examine the translated open reading frames (ORFs) of the yeast Saccharomyces cerevisiae, focusing on those that have FASTA matches in phyletically defined sets of completely sequenced genomes. On this basis, we identify archaeal yeast, bacterial yeast, universal yeast, and yeast ORFs that do not have a match in any of nine prokaryote genomes. Similarly, we examine the yeast mitochondrial genome and the subset of the yeast nuclear ORFs identified as being involved in mitochondrial biogenesis. For the yeast ORFs that match one or more ORFs in these prokaryote genomes, we examine the phyletic and functional distributions of these matches as a function of match strength. These results provide genome level insights into the origin of the eukaryotic cell and the origin of mitochondria. More generally, they exemplify how the growing database of prokaryote genome sequences can help us understand eukaryote genomes.

Archaea↗

The Sulfolobus solfataricus P2 genome project.

Over 800 kbp of the 3-Mbp genome of Sulfolobus solfataricus have been sequenced to date. Our approach is to sequence subclones of mapped cosmids, followed by sequencing directly on cosmid templates with custom primers. Using a prototype automated system for genome-scale analysis, known as MAGPIE, along with other tools, we have discovered one open reading frame of at least 100 amino acids per kbp of sequence, and have been able to associate 50% of these with known genes through database searches. An examination of completely sequenced cosmids suggests a clustering of genes by function in the S. solfataricus genome.

Databases, Factual↗

The metabolic pathway collection from EMP: the enzymes and metabolic pathways database.

The Enzymes and Metabolic Pathways database (EMP) is an encoding of the contents of over 10 000 original publications on the topics of enzymology and metabolism. This large body of information has been transformed into a queryable database. An extraction of over 1800 pictorial representations of metabolic pathways from this collection is freely available on the World Wide Web. We believe that this collection will play an important role in the interpretation of genetic sequence data, as well as offering a meaningful framework for the integration of many other forms of biological data.

Animals↗

Fully automated genome analysis that reflects user needs and preferences. A detailed introduction to the MAGPIE system architecture.

A system called MAGPIE (Multipurpose Automated Genome Project Investigation Environment) has been designed and implemented to meet the challenges of automated whole genome analysis. The system initiates large numbers of remote and local transactions, each depending on evolving criteria and on changing remote and local conditions. Transactions are requested from different types of remote and local resources. The remote request load is fairly balanced with other community demands on server resources. Local decision modules monitor and obey user preferences and combine evidence from multiple sources to formulate credible hypotheses about sequence function. Consistency checks from multiple types of data are integrated into the ongoing local analysis. The system performs reliably on local UNIX workstations and communicates with remote resources through standard networking protocols.

Automation↗

Organizational characteristics and information content of an archaeal genome: 156 kb of sequence from Sulfolobus solfataricus P2.

We have initiated a project to sequence the 3 Mbp genome of the thermoacidophilic archaebacterium Sulfolobus solfataricus P2. Cosmids were selected from a provisional set of minimally overlapping clones, subcloned in pUC18, and sequenced using a hybrid (random plus directed) strategy to give two blocks of contiguous unique sequence, respectively, 100,389 and 56,105 bp. These two contigs contain a total of 163 open reading frames (ORFs) in 26-29 putative operons; 56 ORFs could be identified with reasonable certainty. Clusters of ORFs potentially encode proteins of glycogen biosynthesis, oxidative decarboxylation of pyruvate, ATP-dependent transport across membranes, isoprenoid biosynthesis, protein synthesis, and ribosomes. Putative promoters occur upstream of most ORFs. Thirty per cent of the predicted strong and medium-strength promoters can initiate transcription at the start codon or within 10 nucleotides upstream, indicating a process of initial mRNA-ribosome contact unlike that of most eubacterial genes. A novel termination motif is proposed to account for 15 additional terminations. The two contigs differ in densities of ORFs, insertion elements and repeated sequences; together they contain two copies of the previously reported insertion sequence ISC 1217, five additional IS elements representing four novel types, four classes of long non-IS repeated sequences, and numerous short, perfect repeats.

Chromosomes, Bacterial↗

Reconstruction of metabolic networks using incomplete information.

This paper describes an approach that uses methods for automated sequence analysis (Gaasterland et al. August 1994) and multiple databases accessed through an object+attribute view of the data (Baehr et al. 1992), together with metabolic pathways, reaction equations, and compounds parsed into a logical representation from the Enzyme and Metabolic Pathway Database (Selkov, Yunus, & et.al. 1994), as the sources of data for automatically reconstructing a weighted partial metabolic network for a prokaryotic organism. Additional information can be provided interactively by the expert user to guide reconstruction.

Algorithms↗

Assigning function to CDS through qualified query answering: beyond alignment and motifs.

In this paper, we show how to use qualitative query answering to annotate CDS-to-function relationships with confidence in the score, confidence in the tool, and confidence in the decision about the function. The system, implemented in Prolog, provides users with a powerful tool to analyze large quantities of data that have been produce by multiple sequence analysis programs. Using qualified query answering techniques, users can easily change the criteria for how tools reinforce each other and for how numbers of occurrences of particular functions reinforce each other. They can also alter how different scores for different tools are categorized.

Animals↗