Microbial Genomes Conference IV: sequencing, functional characterization and comparative genomics. Chantilly, Virginia, USA. February 12-15, 2000. Abstracts.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
BACKGROUND: Analysis of any newly sequenced bacterial genome starts with the identification of protein-coding genes. Despite the accumulation of multiple complete genome sequences, which provide useful comparisons with close relatives among other organisms during the annotation process, accurate gene prediction remains quite difficult. A major reason for this situation is that genes are tightly packed in prokaryotes, resulting in frequent overlap. Thus, detection of translation initiation sites and/or selection of the correct coding regions remain difficult unless appropriate biological knowledge (about the structure of a gene) is imbedded in the approach. RESULTS: We have developed a new program that automatically identifies biologically significant candidate genes in a bacterial genome. Twenty-six complete prokaryotic genomes were analyzed using this tool, and the accuracy of gene finding was assessed by comparison with existing annotations. This analysis revealed that, despite the enormous effort of genome program annotators, a small but not negligible number of genes annotated within the framework of sequencing projects are likely to be partially inaccurate or plainly wrong. Moreover, the analysis of several putative new genes shows that, as expected, many short genes have escaped annotation. In most cases, these new genes revealed frameshifts that could be either artifacts or genuine frameshifts. Some entirely unexpected new genes have also been identified. This allowed us to get a more complete picture of prokaryotic genomes. The results of this procedure are progressively integrated into the SWISS-PROT reference databank. CONCLUSIONS: The results described in the present study show that our procedure is very satisfactory in terms of gene finding accuracy. Except in few cases, discrepancies between our results and annotations provided by individual authors can be accounted for by the nature of each annotation process or by specific characteristics of some genomes. This stresses that close cooperation between scientists, regular update and curation of the findings in databases are clearly required to reduce the level of errors in genome annotation (and hence in reducing the unfortunate spreading of errors through centralized data libraries).
In 1995, Haemophilus influenzae became the first free-living organism to have its entire genome sequence published. Since then, many similar projects have been started and, by the millennium, the genomes of a significant number of important human pathogens will have been sequenced. During this period of increasing access to microbial sequence data, parallel advances have occurred in techniques that allow the large-scale study of the entire genetic complement of micro-organisms. In the near future, these approaches will enable researchers to unravel further the complexity of microbial pathogenesis and identify new virulence determinants. Many of these will be suitable targets for development as diagnostic reagents, antimicrobial agents and vaccine candidates. Although it is difficult to predict the full impact that this almost overwhelming volume of information will have on the practice of microbiology, it is clear that it will result ultimately in new ways of diagnosing and combating infectious diseases.
We are interested in quantifying the contribution of gene acquisition, loss, expansion and rearrangements to the evolution of microbial genomes. Here, we discuss factors influencing microbial genome divergence based on pair-wise genome comparisons of closely related strains and species with different lifestyles. A particular focus is on intracellular pathogens and symbionts of the genera Rickettsia, Bartonella and BUCHNERA: Extensive gene loss and restricted access to phage and plasmid pools may provide an explanation for why single host pathogens are normally less successful than multihost pathogens. We note that species-specific genes tend to be shorter than orthologous genes, suggesting that a fraction of these may represent fossil-orfs, as also supported by multiple sequence alignments among species. The results of our genome comparisons are placed in the context of phylogenomic analyses of alpha and gamma proteobacteria. We highlight artefacts caused by different rates and patterns of mutations, suggesting that atypical phylogenetic placements can not a priori be taken as evidence for horizontal gene transfer events. The flexibility in genome structure among free-living microbes contrasts with the extreme stability observed for the small genomes of aphid endosymbionts, in which no rearrangements or inflow of genetic material have occurred during the past 50 millions years (1). Taken together, the results suggest that genomic stability correlate with the content of repeated sequences and mobile genetic elements, and thereby indirectly with bacterial lifestyles.
There is an urgent need to develop novel classes of antibiotics to counter the threat of the spread of multiply resistant bacterial pathogens. The availability of the complete genome sequence of many pathogenic microbes provides information on every potential drug target and is an invaluable resource in the search for novel compounds. Here, we review the approaches being taken to exploit the genome databases through a combination of bioinformatics, transcriptional analysis, and a further understanding of the molecular basis of the disease process. The emphasis is changing from compound screening to target hunting, as the latter offers flexible ways to design and optimize the next generation of broad-spectrum antibiotics.
We present models describing the acquisition and deletion of novel sequences in populations of microorganisms. We infer that most novel sequences are neutral. Thus, sequence duplications and gene transfer between organisms sharing the same environment are rarely expected to generate adaptive functions. Two classes of models are considered: (1) a homogeneous population with constant size, and (2) an island model in which the population is subdivided into patches that are in contact through slow migration. Distributions of gene frequencies are derived in a Moran model with overlapping generations. We find that novel, neutral or near-neutral coding sequences in microorganisms will not be fixed globally because they offer large target sizes for mutations and because the populations are so large. At most, such genes may have a transient presence in only a small fraction of the population. Consequently, a microbial population is expected to have a very large diversity of transient neutral gene content. Only sequences that are under strong selection, globally or in individual patches, can be expected to persist. We suggest that genome size is maintained in microorganisms by a quasi-steady state mechanism in which random fluctuations in the effective acquisition and deletion rates result in genome sizes that vary from patch to patch. We assign the genomic identity of a global population to those genes that are required for the participation of patches in the genetic sweeps that maintain the genomic coherence of the population. In contrast, we stress the influence of sequence loss on the isolation and the divergence (speciation) of novel patches from a global population.
It is suggested that genomes found in any form of cellular life contain potentially size-variable repetitive DNA moieties. In eukaryotes, large proportions of the multi-chromosomal genome consist of various classes of repetitive DNA. Also in archaeal genomes, repetitive DNA is encountered and, as is the case for the eukaryotes as well, little or no function is at present attributable to most of it. For prokaryotes, elegant experiments have highlighted so-called slipped strand nucleotide mispairing (SSM) as a basic and causal mechanism, giving rise to repeat unit number variation at a distinct locus. Illegitimate base pairing in regions of repetitive DNA during replication, in association with defective DNA repair and enhanced nuclease susceptibility of replication intermediates, in the end gives rise to deletion or addition of repeat units. Prokaryotic short sequence repeats (SSRs) harbour arrays of short repeat units, between one and approximately 20 nucleotides in length. SSRs are involved in various mechanisms of microbial gene expression regulation. Promoter strength can be affected by altering the spacing between important structural domains as can the integrity of open reading frames. In the present communication the literature on microbial SSRs harbouring repeat units that are five nucleotides in length will be briefly reviewed. Examples of these SSRs with discrete functionality are encountered in bacterial species such as Haemophilus influenzae, Neisseria gonorrhoeae, and Pasteurella haemolytica. In addition, several of the currently known bacterial and archaeal whole genome sequences were scanned for the presence of novel examples of potential five-nucleotide SSRs (and others) in order to gather additional knowledge on the propensity and putative functions of this type of potential genetic switch.
A high-throughput robotic workstation system was used for double-stranded plasmid DNA template preparation and sequencing reaction setup to streamline the sequencing process in genome projects. All 96-well miniprep kits that were tested provided high quality plasmid DNA suitable for fluorescent DNA sequencing. After quantitation in a 96-well UV spectrophotometer, the plasmid DNA was used as template to automatically set up sequencing reactions. The setup was controlled by spread sheets that were imported into the robotic system. We utilized this integrated system to prepare all necessary shotgun templates for our contributions to a number of large-scale genome projects as well as a full-length cDNA sequencing project.
We have conducted genome sequence analyses of seven prokaryotic microorganisms for which completely sequenced genomes are available (Escherichia coli, Haemophilus influenzae, Helicobacter pylori, Bacillus subtilis, Mycoplasma genitalium, Synechocystis PCC6803 and Methanococcus jannaschii). We report the distribution of encoded known and putative polytopic cytoplasmic membrane transport proteins within these genomes. Transport systems for each organism were classified according to (1) putative membrane topology, (2) protein family, (3) bioenergetics, and (4) substrate specificities. The overall transport capabilities of each organism were thereby estimated. Probable function was assigned to greater than 90% of the putative transport proteins identified. The results show the following: (1) Numbers of transport systems in eubacteria are approximately proportional to genome size and correspond to 9.7 to 10.8% of the total encoded genes except for H. pylori (5.4%), Synechocystis (4.7%) and M. jannaschii (3.5%) which exhibit substantially lower proportions. (2) The distribution of topological types is similar in all seven organisms. (3) Transport systems belonging to 67 families were identified within the genomes of these organisms, and about half of these families are also found in eukaryotes. (4) 12% of these families are found exclusively in Gram-negative bacteria, but none is found exclusively in Gram-positive bacteria, cyanobacteria or archaea. (5) Two superfamilies, the ATP-binding cassette (ABC) and major facilitator (MF) superfamilies account for nearly 50% of all transporters in each organism, but the relative representation of these two transporter types varies over a tenfold range, depending on the organism. (6) Secondary, pmf-dependent carriers are 1.5 to threefold more prevalent than primary ATP-dependent carriers in E. coli, H. influenzae, H. pylori and B. subtilis while primary carriers are about twofold more prevalent in M. genitalium and Synechocystis. M. jannaschii exhibits a slight preference for secondary carriers. (7) Bioenergetics of transport generally correlate with the primary forms of energy generated via available metabolic pathways but ecological niche and substrate availability may also be determining factors. (8) All organisms display a similar range of transport specificities with quantitative differences presumably reflective of disparate ecological niches. (9) M. jannaschii and Synechocystis have a two to threefold increased proportion of transporters for inorganic ions with a concomitant decrease in transporters for organic compounds. (10) 6 to 18% of all transporters in these bacteria probably function as drug export systems showing that these systems are prevalent in non-pathogenic as well as pathogenic organisms. (11) All seven prokaryotes examined encode proteins homologous to known channel proteins, but none of the channel types identified occurs in all of these organisms. (12) The phosphoenolpyruvate:sugar phosphotransferase system is prevalent in the large genome organisms, E. coli and B. subtilis, and is present in the small genome organisms, H. influenzae and M. genitalium, but is totally lacking in H. pylori, Synechocystis and M. jannaschii. Details of the information summarized in this article are available on our web sites, and this information will be periodically updated and corrected as new sequence and biochemical data become available.
Here, we present a comprehensive analysis of solute transport systems encoded within the completely sequenced genomes of 18 prokaryotic organisms. These organisms include four Gram-positive bacteria, seven Gram-negative bacteria, two spirochetes, one cyanobacterium and four archaea. Membrane proteins are analyzed in terms of putative membrane topology, and the recognized transport systems are classified into 76 families, including four families of channel proteins, four families of primary carriers, 54 families of secondary carriers, six families of group translocators, and eight unclassified families. These families are analyzed in terms of the paralogous and orthologous relationships of their protein members, the substrate specificities of their constituent transporters and their distributions in each of the 18 organisms studied. The families vary from large superfamilies with hundreds of represented members, to small families with only one or a few members. The mode of transport generally correlates with the primary mechanism of energy generation, and the numbers of secondary transporters relative to primary transporters are roughly proportional to the total numbers of primary H(+) and Na(+) pumps in the cell. The phosphotransferase system is less prevalent in the analyzed bacteria than previously thought (only six of 14 bacteria transport sugars via this system) and is completely lacking in archaea and eukaryotes. Escherichia coli is shown to be exceptionally broad in its transport capabilities and therefore, at a membrane transport level, does not appear representative of the bacteria thus far sequenced. Archaea and spirochetes exhibit fewer proteins with multiple transmembrane segments and fewer net transporters than most bacteria. These results provide insight into the relevance of transport to the overall physiology of prokaryotes.
Explore the source record for details and available documents.
Along the gene, nucleotides in various codon positions tend to exert a slight but observable influence on the nucleotide choice at neighboring positions. Such context biases are different in different organisms and can be used as genomic signatures. In this paper, we will focus specifically on the dinucleotide composed of a third codon position nucleotide and its succeeding first position nucleotide. Using the 16 possible dinucleotide combinations, we calculate how well individual genes conform to the observed mean dinucleotide frequencies of an entire genome, forming a distance measure for each gene. It is found that genes from different genomes can be separated with a high degree of accuracy, according to these distance values. In particular, we address the problem of recent horizontal gene transfer, and how imported genes may be evaluated by their poor assimilation to the host's context biases. By concentrating on the third- and succeeding first position nucleotides, we eliminate most spurious contributions from codon usage and amino-acid requirements, focusing mainly on mutational effects. Since imported genes are expected to converge only gradually to genomic signatures, it is possible to question whether a gene present in only one of two closely related organisms has been imported into one organism or deleted in the other. Striking correlations between the proposed distance measure and poor homology are observed when Escherichia coli genes are compared to Salmonella typhi, indicating that sets of outlier genes in E. coli may contain a high number of genes that have been imported into E. coli, and not deleted in S. typhi.
Unequal use of synonymous codons has been found in several prokaryotic and eukaryotic genomes. This bias has been associated with translational efficiency. The prevalence of this bias across lineages is currently unknown. Here, a new method (GCB) to measure codon usage bias is presented. It uses an iterative approach for the determination of codon scores and allows the computation of an index of codon bias suitable for interspecies comparison. A server to calculate GCB-values of individual genes as well as a list of compiled results are available at www.g21.bio.uni-goettingen.de. The method was applied to complete bacterial genomes. The relation of codon usage bias with amino acid composition and the choice of stop codons were determined and discussed.
The ubiquitous glyoxalase system, which is composed of two enzymes, removes cellular cytotoxic methylglyoxal (MG). In an effort to identify critical residues conserved in the evolution of the first enzyme in this system, glyoxalase I (GlxI), as well as the structural implications of sequence alterations in this enzyme, a search of the National Center for Biotechnology Information (NCBI) database of unfinished genomes was undertaken. Eleven putative GlxI sequences from pathogenic organisms were identified and analyses of these sequences in relation to the known and previously identified GlxI enzymes were performed. Several of these sequences show a very high similarity to the Escherichia coli GlxI sequence, most notably the 79% identity of the sequence identified from Yersinia pestis, the causative agent of bubonic plague. In addition to the conservation of residues critical to binding the catalytic metal in all of the proposed GlxI enzymes, four regions in the Homo sapiens GlxI enzyme are absent in all of the bacterial GlxI sequences, with the exception of Pseudomonas putida. Removal of these regions may alter the active-site conformation of the bacterial enzymes in relation to that of the H. sapiens. These differences may be targeted for the development of inhibitors selective to the bacterial enzymes.
A Brazilian consortium has unveiled the genomic DNA sequence of the purple-pigmented bacterium Chromobacterium violaceum, a dominant component of the tropical soil microbiota. The sequence provides insight into the abundant potential of this organism for biotechnological and pharmaceutical applications.
A set of rubber oxygenases was discovered through phylogenetic analysis and AI-based structural modeling of complexes of the putative enzymes with a substrate mimicking cis-1,4-polyisoprene. Sixteen candidate proteins were selected from thermophilic microorganisms, all sequence-related to the Latex clearing protein from Streptomyces sp. K30 (LcpK30). Sequence truncation and solubility tags were then evaluated to enhance protein expression, with the SUMO tag proving to be the most effective. Including LcpK30, nine heme-containing oxygenases were successfully expressed in E. coli NEB 10-beta cells, purified (35-157 mg L-1 yield) and characterized. Steady-state kinetics revealed significant rubber latex-degrading properties for six of them, with the truncated SUMO-fused LcpK30 (SUMO-LcpK30T) showing activity in agreement with literature. Notably, the catalytic efficiencies of all the expressed homologs lay within one order of magnitude and the oxygenase from Thermomonospora echinospora was found to be particularly promising in terms of activity, especially at high latex concentrations (more than 1% w/v). The analysis of reaction mixtures by both HPLC and HPLC-MS confirmed the oxidation of cis-1,4-polyisoprene to form the expected isoprenoid oligomers (n = 2-12), whose distribution was consistent with the usual endo-type cleavage pattern in all but one case. This bioprospecting effort afforded a platform of new rubber-degrading enzymes with diverse efficiencies and product profiles, capable of adapting to targeted applications.
Shine-Dalgarno (SD) sequence has been considered as one of the common features of 5' end untranslated region (5'UTR) of prokaryotic transcripts. However, more leaderless bacteria and archaea mRNAs are being increasingly reported in recent years. To understand the distribution of SD-led genes and non-SD-led genes, we have analyzed 162 completed prokaryotic genomes leading to various new conclusions and validations of previous smaller scale studies. The fact that the number of the SD-led genes among those genomes varies from 11.6% to 90.8% implies that the populations of non-SD-led genes as well as leaderless genes are significant. We found that there is a strong SD conserved region in genomes with high proportion of SD-led genes. Following a t-test we showed that SD sequence content (SDSC) has no correlation with GC content. We observed that the closely related phylogenetic microbes mostly possess a similar SDSC value, and archaeal nonleading genes possess higher SDSC. This study shows that the 5'UTR of prokaryotic genes are highly diverse, particularly when genomes of distantly related organisms are compared, suggesting that more flexible mechanisms are used for translation initiation process in various prokaryotes.
The quest for the discovery of novel natural products has entered a new chapter with the enormous wealth of genetic data that is now available. This information has been exploited by using whole-genome sequence mining to uncover cryptic pathways, or biosynthetic pathways for previously undetected metabolites. Alternatively, using known paradigms for secondary metabolite biosynthesis, genetic information has been 'fished out' of DNA libraries resulting in the discovery of new natural products and isolation of gene clusters for known metabolites. Novel natural products have been discovered by expressing genetic data from uncultured organisms or difficult-to-manipulate strains in heterologous hosts. Furthermore, improvements in heterologous expression have not only helped to identify gene clusters but have also made it easier to manipulate these genes in order to generate new compounds. Finally, and perhaps the most crucial aspect of the efficient and prosperous use of the abundance of genetic information, novel enzyme chemistry continues to be discovered, which has aided our understanding of how natural products are biosynthesized de novo, and enabled us to rework the current paradigms for natural product biosynthesis.