Search PubMed⌕ Search

Biomedical subjects

T Gaasterland

Publications and source records attributed to T Gaasterland.

At least 19 recordsLinked to original sources

A functional update of the Escherichia coli K-12 genome.

BACKGROUND: Since the genome of Escherichia coli K-12 was initially annotated in 1997, additional functional information based on biological characterization and functions of sequence-similar proteins has become available. On the basis of this new information, an updated version of the annotated chromosome has been generated. RESULTS: The E. coli K-12 chromosome is currently represented by 4,401 genes encoding 116 RNAs and 4,285 proteins. The boundaries of the genes identified in the GenBank Accession U00096 were used. Some protein-coding sequences are compound and encode multimodular proteins. The coding sequences (CDSs) are represented by modules (protein elements of at least 100 amino acids with biological activity and independent evolutionary history). There are 4,616 identified modules in the 4,285 proteins. Of these, 48.9% have been characterized, 29.5% have an imputed function, 2.1% have a phenotype and 19.5% have no function assignment. Only 7% of the modules appear unique to E. coli, and this number is expected to be reduced as more genome data becomes available. The imputed functions were assigned on the basis of manual evaluation of functions predicted by BLAST and DARWIN analyses and by the MAGPIE genome annotation system. CONCLUSIONS: Much knowledge has been gained about functions encoded by the E. coli K-12 genome since the 1997 annotation was published. The data presented here should be useful for analysis of E. coli gene products as well as gene products encoded by other genomes.

Bacterial Proteins↗

Microarray-based analysis of early development in Xenopus laevis.

In order to examine transcriptional regulation globally, during early vertebrate embryonic development, we have prepared Xenopus laevis cDNA microarrays. These prototype embryonic arrays contain 864 sequenced gastrula cDNA. In order to analyze and store array data, a microarray analysis pipeline was developed and integrated with sequence analysis and annotation tools. In three independent experimental settings, we demonstrate the power of these global approaches and provide optimized protocols for their application to molecular embryology. In the first set, by comparing maternal versus zygotic transcription, we document groups of genes that are temporally regulated. This analytical approach resulted in the discovery of novel temporally regulated genes. In the second, we examine changes in gene expression spatially during development by comparing dorsal and ventral mesoderm dissected from early gastrula embryos. We have discovered novel genes with spatial enrichment from these experiments. Finally, we use the prototype microarray to examine transcriptional responses from embryonic explants treated with activin. We selected genes (two of which are novel) regulated by activin for further characterization. All results obtained by the arrays were independently tested by RT-PCR or by in situ hybridization to provide a direct assessment of the accuracy and reproducibility of these approaches in the context of molecular embryology.

Animals↗

The complete genome of the crenarchaeon Sulfolobus solfataricus P2.

The genome of the crenarchaeon Sulfolobus solfataricus P2 contains 2,992,245 bp on a single chromosome and encodes 2,977 proteins and many RNAs. One-third of the encoded proteins have no detectable homologs in other sequenced genomes. Moreover, 40% appear to be archaeal-specific, and only 12% and 2.3% are shared exclusively with bacteria and eukarya, respectively. The genome shows a high level of plasticity with 200 diverse insertion sequence elements, many putative nonautonomous mobile elements, and evidence of integrase-mediated insertion events. There are also long clusters of regularly spaced tandem repeats. Different transfer systems are used for the uptake of inorganic and organic solutes, and a wealth of intracellular and extracellular proteases, sugar, and sulfur metabolizing enzymes are encoded, as well as enzymes of the central metabolic pathways and motility proteins. The major metabolic electron carrier is not NADH as in bacteria and eukarya but probably ferredoxin. The essential components required for DNA replication, DNA repair and recombination, the cell cycle, transcriptional initiation and translation, but not DNA folding, show a strong eukaryal character with many archaeal-specific features. The results illustrate major differences between crenarchaea and euryarchaea, especially for their DNA replication mechanism and cell cycle processes and their translational apparatus.

Cell Cycle Proteins↗

Functional annotation of a full-length mouse cDNA collection.

The RIKEN Mouse Gene Encyclopaedia Project, a systematic approach to determining the full coding potential of the mouse genome, involves collection and sequencing of full-length complementary DNAs and physical mapping of the corresponding genes to the mouse genome. We organized an international functional annotation meeting (FANTOM) to annotate the first 21,076 cDNAs to be analysed in this project. Here we describe the first RIKEN clone collection, which is one of the largest described for any organism. Analysis of these cDNAs extends known gene families and identifies new ones.

Animals↗

Whole-genome analysis: annotations and updates.

The most important advances in the field of genome annotation over the past two years involve the use of cDNA sequences, protein structures and gene expression data to predict genes. These types of information not only improve gene identification, but they also give insights into variation in gene structure and function.

Alternative Splicing↗

Homology-based annotation yields 1,042 new candidate genes in the Drosophila melanogaster genome.

The approach to annotating a genome critically affects the number and accuracy of genes identified in the genome sequence. Genome annotation based on stringent gene identification is prone to underestimate the complement of genes encoded in a genome. In contrast, over-prediction of putative genes followed by exhaustive computational sequence, motif and structural homology search will find rarely expressed, possibly unique, new genes at the risk of including non-functional genes. We developed a two-stage approach that combines the merits of stringent genome annotation with the benefits of over-prediction. First we identify plausible genes regardless of matches with EST, cDNA or protein sequences from the organism (stage 1). In the second stage, proteins predicted from the plausible genes are compared at the protein level with EST, cDNA and protein sequences, and protein structures from other organisms (stage 2). Remote but biologically meaningful protein sequence or structure homologies provide supporting evidence for genuine genes. The method, applied to the Drosophila melanogaster genome, validated 1,042 novel candidate genes after filtering 19,410 plausible genes, of which 12,124 matched the original 13,601 annotated genes. This annotation strategy is applicable to genomes of all organisms, including human.

Animals↗

Enolase from Trypanosoma brucei, from the amitochondriate protist Mastigamoeba balamuthi, and from the chloroplast and cytosol of Euglena gracilis: pieces in the evolutionary puzzle of the eukaryotic glycolytic pathway.

Genomic or cDNA clones for the glycolytic enzyme enolase were isolated from the amitochondriate pelobiont Mastigamoeba balamuthi, from the kinetoplastid Trypanosoma brucei, and from the euglenid Euglena gracilis. Clones for the cytosolic enzyme were found in all three organisms, whereas Euglena was found to also express mRNA for a second isoenzyme that possesses a putative N-terminal plastid-targeting peptide and is probably targeted to the chloroplast. Database searching revealed that Arabidopsis also possesses a second enolase gene that encodes an N-terminal extension and is likely targeted to the chloroplast. A phylogeny of enolase amino acid sequences from 6 archaebacteria, 24 eubacteria, and 32 eukaryotes showed that the Mastigamoeba enolase tended to branch with its homologs from Trypanosoma and from the amitochondriate protist Entamoeba histolytica. The compartment-specific isoenzymes in Euglena arose through a gene duplication independent of that which gave rise to the compartment-specific isoenzymes in Arabidopsis, as evidenced by the finding that the Euglena enolases are more similar to the homolog from the eubacterium Treponema pallidum than they are to homologs from any other organism sampled. In marked contrast to all other glycolytic enzymes studied to date, enolases from all eukaryotes surveyed here (except Euglena) are not markedly more similar to eubacterial than to archaebacterial homologs. An intriguing indel shared by enolase from eukaryotes, from the archaebacterium Methanococcus jannaschii, and from the eubacterium Campylobacter jejuni maps to the surface of the three-dimensional structure of the enzyme and appears to have occurred at the same position in parallel in independent lineages.

Amino Acid Sequence↗

MAGPIE/EGRET annotation of the 2.9-Mb Drosophila melanogaster Adh region.

Our challenge in annotating the 2.91-Mb Adh region of the Drosophila melanogaster genome was to identify genetic and genomic features automatically, completely, and precisely within a 6-week period. To do so, we augmented the MAGPIE microbial genome annotation system to handle eukaryotic genomic sequence data. The new configuration required the integration of eukaryotic gene-finding tools and DNA repeat tools into the automatic data collection module. It also required us to define in MAGPIE new strategies to combine data about eukaryotic exon predictions with functional data to refine the exon predictions. At the heart of the resulting new eukaryotic genome annotation system is a reverse comparison of public protein and complementary DNA sequences against the input genome to identify missing exons and to refine exon boundaries. The software modules that add eukaryotic genome annotation capability to MAGPIE are available as EGRET (Eukaryotic Genome Rapid Evaluation Tool).

Alcohol Dehydrogenase↗

Gene content and organization of a 281-kbp contig from the genome of the extremely thermophilic archaeon, Sulfolobus solfataricus P2.

The sequence of a 281-kbp contig from the crenarchaeote Sulfolobus solfataricus P2 was determined and analysed. Notable features in this region include 29 ribosomal protein genes, 12 tRNA genes (four of which contain archaeal-type introns), operons encoding enzymes of histidine biosynthesis, pyrimidine biosynthesis, and arginine biosynthesis, an ATPase operon, numerous genes for enzymes of lipopolysaccharide biosynthesis, and six insertion sequences. The content and organization of this contig are compared with sequences from crenarchaeotes, euryarchaeotes, bacteria, and eukaryotes.

Amino Acid Sequence↗

Archaeal genomics.

Four euryarchaeal genomes have been completely sequenced and are publicly available: Methanococcus jannaschii, Methanobacterium thermoautotrophicum, Pyrococcus horikoshii and Archaeoglobus fulgidus. Four more genome sequences, two crenarchaeal and two pyrococci, will soon be released. In addition, seven more archaeal genome sequencing projects are under way, including two halophiles, two Thermoplasma, and a methanogen. These projects cover all branches of the archaeal domain and will lead to new insights into archaeal metabolism, DNA processing, and evolutionary relationships with the Bacteria and Eukarya.

Archaea↗

Structural genomics: beyond the human genome project.

With access to whole genome sequences for various organisms and imminent completion of the Human Genome Project, the entire process of discovery in molecular and cellular biology is poised to change. Massively parallel measurement strategies promise to revolutionize how we study and ultimately understand the complex biochemical circuitry responsible for controlling normal development, physiologic homeostasis and disease processes. This information explosion is also providing the foundation for an important new initiative in structural biology. We are about to embark on a program of high-throughput X-ray crystallography aimed at developing a comprehensive mechanistic understanding of normal and abnormal human and microbial physiology at the molecular level. We present the rationale for creation of a structural genomics initiative, recount the efforts of ongoing structural genomics pilot studies, and detail the lofty goals, technical challenges and pitfalls facing structural biologists.

Computational Biology↗

Complete sequence of a 184-kilobase catabolic plasmid from Sphingomonas aromaticivorans F199.

The complete 184,457-bp sequence of the aromatic catabolic plasmid, pNL1, from Sphingomonas aromaticivorans F199 has been determined. A total of 186 open reading frames (ORFs) are predicted to encode proteins, of which 79 are likely directly associated with catabolism or transport of aromatic compounds. Genes that encode enzymes associated with the degradation of biphenyl, naphthalene, m-xylene, and p-cresol are predicted to be distributed among 15 gene clusters. The unusual coclustering of genes associated with different pathways appears to have evolved in response to similarities in biochemical mechanisms required for the degradation of intermediates in different pathways. A putative efflux pump and several hypothetical membrane-associated proteins were identified and predicted to be involved in the transport of aromatic compounds and/or intermediates in catabolism across the cell wall. Several genes associated with integration and recombination, including two group II intron-associated maturases, were identified in the replication region, suggesting that pNL1 is able to undergo integration and excision events with the chromosome and/or other portions of the plasmid. Conjugative transfer of pNL1 to another Sphingomonas sp. was demonstrated, and genes associated with this function were found in two large clusters. Approximately one-third of the ORFs (59 of them) have no obvious homology to known genes.

Bacterial Proteins↗

The complete genome of the hyperthermophilic bacterium Aquifex aeolicus.

Aquifex aeolicus was one of the earliest diverging, and is one of the most thermophilic, bacteria known. It can grow on hydrogen, oxygen, carbon dioxide, and mineral salts. The complex metabolic machinery needed for A. aeolicus to function as a chemolithoautotroph (an organism which uses an inorganic carbon source for biosynthesis and an inorganic chemical energy source) is encoded within a genome that is only one-third the size of the E. coli genome. Metabolic flexibility seems to be reduced as a result of the limited genome size. The use of oxygen (albeit at very low concentrations) as an electron acceptor is allowed by the presence of a complex respiratory apparatus. Although this organism grows at 95 degrees C, the extreme thermal limit of the Bacteria, only a few specific indications of thermophily are apparent from the genome. Here we describe the complete genome sequence of 1,551,335 base pairs of this evolutionarily and physiologically interesting organism.

Chromosome Mapping↗

Completing the sequence of the Sulfolobus solfataricus P2 genome.

The Sulfolobus solfataricus P2 genome collaborators are poised to sequence the entire 3-Mbp genome of this crenarchaeote archaeon. About 80% of the genome has been sequenced to date, with the rest of the sequence being assembled fast. In this publication we introduce the genomic sequencing and automated analysis strategy and present intial data derived from the sequence analysis. After an overview of the general sequence features, metabolic pathway studies are explained, using sugar metabolism as an example. The paper closes with an overview of repetitive elements in S. solfataricus.

Base Sequence↗

Constructing multigenome views of whole microbial genomes.

We have designed and implemented a system to carry out cross-genome comparisons of open reading frames (ORFs) from multiple genomes. This implementation includes a genome profiling system that allows us to explore pairwise comparisons at different levels of match similarity and ask biologically motivated queries involving number and identity of ORFs, their function, functional category, distribution in genomes or in biological domains, and statistics on their matches and match families. This analysis required precise definition of new classification terms and concepts. We define the terms genomic signature, summary signature, biologic domain signature, domain class, match level, match family, and extended match family, then use these terms to define concepts, including genomically universal proteins and proteins characteristics of sets of genomes. We initiate an analysis based on automated FASTA (Pearson, 1996) comparison of 22,419 conceptually translated protein sequences from nine microbial genomes.

Amino Acid Sequence↗