Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome, Microbial”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Insights on biology and evolution from microbial genome sequencing.

No field of research has embraced and applied genomic technology more than the field of microbiology. Comparative analysis of nearly 300 microbial species has demonstrated that the microbial genome is a dynamic entity shaped by multiple forces. Microbial genomics has provided a foundation for a broad range of applications, from understanding basic biological processes, host-pathogen interactions, and protein-protein interactions, to discovering DNA variations that can be used in genotyping or forensic analyses, the design of novel antimicrobial compounds and vaccines, and the engineering of microbes for industrial applications. Most recently, metagenomics approaches are allowing us to begin to probe complex microbial communities for the first time, and they hold great promise in helping to unravel the relationships between microbial species.

DNA, Bacterial↗

Sequence-based approach to finding functional lipases from microbial genome databases.

A sequence-based approach was used to retrieve functional lipases from microbial genome databases. Many novel genes assigned as putative lipases were tested using the criteria of the typical lipase sequence rule, based on a consensus sequence of a catalytic triad (Ser, Asp, His) and oxyanion hole sequence (HG). To obtain the lipase genes satisfying the sequence rule, PCR cloning was performed, while the lipase activities were tested using a tributyrin/tricaprylin plate and p-nitrophenyl caproate. Among nine putative lipases from four strains, five functional lipolytic proteins were obtained from Archaeoglobus fulgidus, Deinococcus radiodurans, and Agrobacterium tumefaciens. All five lipases exhibited a relatively low sequence similarity (less than 26.7%) with known lipases and turned out to belong to different lipase families. Accordingly, the current results indicate that the proposed strategic approach based on the microbial genome is an efficient and rapid method for finding novel and functional lipases.

Agrobacterium tumefaciens↗

Seven GC-rich microbial genomes adopt similar codon usage patterns regardless of their phylogenetic lineages.

Seven GC-rich (group I) and three AT-rich (group II) microbial genomes are analyzed in this paper. The seven microbes in group I belong to different phylogenetic lineages, even different domains of life. The common feature is that they are highly GC-rich organisms, with more than 60% genomic GC content. Group II includes three bacteria, which belong to the same subdivision as Pseudomonas aeruginosa in group I. The genomic GC content of the three bacteria is in the range of 26-50%. It is shown that although the phylogenetic lineages of the organisms in group I are remote, the common feature of highly genomic GC content forces them to adopt similar codon usage patterns, which constitutes the basis of an algorithm using a set of universal parameters to recognize known genes in the seven genomes. The common codon usage pattern of function known genes in the seven genomes is GGS type, where G, G, and S are the bases of G, non-G, and G/C, respectively. On the contrary, although the phylogenetic lineages of the three bacteria in group II are quite close, the codon usage patterns of function known genes in these genomes are obviously distinct. There are no universal parameters to identify known genes in the three genomes in group II. It can be deduced that the genomic GC content is more important than phylogenetic lineage in gene recognition programs. We hope that the work might be useful for understanding the common characteristics in the organization of microbial genomes.

Algorithms↗

Assignment of folds for proteins of unknown function in three microbial genomes.

Analysis of DNA sequences of several microbial genomes has revealed that a large fraction of predicted coding regions has no known protein function. Information about the three-dimensional folds of these proteins may provide insight into their possible functions. To predict the folds for protein sequences with little or no homology to proteins of known function, we used computational neural networks trained on the database of proteins with known three-dimensional structures. Global descriptions of protein sequences based on physical and structural properties of the constituent amino acids were used as inputs for neural networks. Of the 131, 498, and 868 protein sequences of unknown function from Mycoplasma genitalium, Haemophilus influenzae, and Methanococcus jannaschii (Fleischmann et al. 1995), we have made high-confidence fold assignments for 4, 10, and 19 sequences, respectively.

Amino Acid Sequence↗

Computational identification of operons in microbial genomes.

By applying graph representations to biochemical pathways, a new computational pipeline is proposed to find potential operons in microbial genomes. The algorithm relies on the fact that enzyme genes in operons tend to catalyze successive reactions in metabolic pathways. We applied this algorithm to 42 microbial genomes to identify putative operon structures. The predicted operons from Escherichia coli were compared with a selected metabolism-related operon dataset from the RegulonDB database, yielding a prediction sensitivity (89%) and specificity (87%) relative to this dataset. Several examples of detected operons are given and analyzed. Modular gene cluster transfer and operon fusion are observed. A further use of predicted operon data to assign function to putative genes was suggested and, as an example, a previous putative gene (MJ1604) from Methanococcus jannaschii is now annotated as a phosphofructokinase, which was regarded previously as a missing enzyme in this organism. GC content changes in the operon region and nonoperon region were examined. The results reveal a clear GC content transition at the boundaries of putative operons. We looked further into the conservation of operons across genomes. A trp operon alignment is analyzed in depth to show gene loss and rearrangement in different organisms during operon evolution.

Algorithms↗

Microbial genome evolution: sources of variability.

Comparative genome analyses of close relatives have yielded exciting insight into the sources of microbial genome variability with respect to gene content, gene order and evolution of genes with unknown functions. The genomes of free-living bacteria often carry phages and repetitive sequences that mediate genomic rearrangements in contrast to the small genomes of obligate host-associated bacteria. This suggests that genomic stability correlates with the genomic content of repeated sequences and movable genetic elements, and thereby with bacterial lifestyle. Genes with unknown functions present in a single species tend to be shorter than conserved, functional genes, indicating that the fraction of unique genes in microbial genomes has been overestimated.

Bacteria↗

The highest priority: what microbial genomes are telling us about immunity.

Study of microbial genomes has provided new insight into the functions that pathogens require for survival in the animal host. Small genome bacterial pathogens, defined as those < or = 1/3 the size of Escherichia coli, include chlamydiae, rickettsiae and ehrlichiae, mycoplasmas, and spirochetes. The small genome size is believed to result from reductive evolution, a process of initial mutation with loss of function followed by progressive accumulation of mutations and eventual gene deletion. This is most notable in the 1.1 Mb genome of Rickettsia prowazekki in which 24% of the genome is non-coding, as compared to approximately 10% in the 4.4 Mb E. coli. Consequently, these pathogens are thus presumed to retain only the most important functions for survival and propagation. There is consistent evidence from small genomes that the genetic deletion is primarily related to the loss of metabolic function and especially reduction of multiple overlapping pathways and duplicated genes. Thus, these pathogens undergo progressive reduction in their genomes yet maintain the ability to infect, survive within, and cause disease in animals. In the face of this reductive process, what genes and associated functions are maintained? Strikingly, these pathogens devote a high percentage of their genomes to paralogous families of polymorphic surface molecules. This retention suggests that evasion of the immune response is the highest priority of obligate microbial pathogens and provides a strategy for identifying protective antigens for vaccine development to control disease.

Adaptation, Physiological↗

Microbial genomes have over 72% structure assignment by the threading algorithm PROSPECTOR_Q.

The genome scale threading of five complete microbial genomes is revisited using our state-of-the-art threading algorithm, PROSPECTOR_Q. Considering that structure assignment to an ORF could be useful for predicting biochemical function as well as for analyzing pathways, it is important to assess the current status of genome scale threading. The fraction of ORFs to which we could assign protein structures with a reasonably good confidence level to each genome sequences is over 72%, which is significantly higher than earlier studies. Using the assigned structures, we have predicted the function of several ORFs through "single-function" template structures, obtained from an analysis of the relationship between protein fold and function. The fold distribution of the genomes and the effect of the number of homologous sequences on structure assignment are also discussed.

Algorithms↗

Genes linked by fusion events are generally of the same functional category: a systematic analysis of 30 microbial genomes.

Recent work in computational genomics has shown that a functional association between two genes can be derived from the existence of a fusion of the two as one continuous sequence in another genome. For each of 30 completely sequenced microbial genomes, we established all such fusion links among its genes and determined the distribution of links within and among 15 broad functional categories. We found that 72% of all fusion links related genes of the same functional category. A comparison of the distribution of links to simulations on the basis of a random model further confirmed the significance of intracategory fusion links. Where a gene of annotated function is linked to an unclassified gene, the fusion link suggests that the two genes belong to the same functional category. The predictions based on fusion links are shown here for Methanobacterium thermoautotrophicum, and another 661 predictions are available at http://fusion.bu.edu.

Gene Expression Regulation, Bacterial↗

Accuracy improvement for identifying translation initiation sites in microbial genomes.

MOTIVATION: At present the computational gene identification methods in microbial genomes have a high prediction accuracy of verified translation termination site (3' end), but a much lower accuracy of the translation initiation site (TIS, 5' end). The latter is important to the analysis and the understanding of the putative protein of a gene and the regulatory machinery of the translation. Improving the accuracy of prediction of TIS is one of the remaining open problems. RESULTS: In this paper, we develop a four-component statistical model to describe the TIS of prokaryotic genes. The model incorporates several features with biological meanings, including the correlation between translation termination site and TIS of genes, the sequence content around the start codon; the sequence content of the consensus signal related to ribosomal binding sites (RBSs), and the correlation between TIS and the upstream consensus signal. An entirely non-supervised training system is constructed, which takes as input a set of annotated coding open reading frames (ORFs) by any gene finder, and gives as output a set of organism-specific parameters (without any prior knowledge or empirical constants and formulas). The novel algorithm is tested on a set of reliable datasets of genes from Escherichia coli and Bacillus subtillis. MED-Start may correctly predict 95.4% of the start sites of 195 experimentally confirmed E.coli genes, 96.6% of 58 reliable B.subtillis genes. Moreover, the test results indicate that the algorithm gives higher accuracy for more reliable datasets, and is robust to the variation of gene length. MED-Start may be used as a postprocessor for a gene finder. After processing by our program, the improvement of gene start prediction of gene finder system is remarkable, e.g. the accuracy of TIS predicted by MED 1.0 increases from 61.7 to 91.5% for 854 E.coli verified genes, while that by GLIMMER 2.02 increases from 63.2 to 92.0% for the same dataset. These results show that our algorithm is one of the most accurate methods to identify TIS of prokaryotic genomes. AVAILABILITY: The program MED-Start can be accessed through the website of CTB at Peking University: http://ctb.pku.edu.cn/main/SheGroup/MED_Start.htm.

Algorithms↗

GWFASTA: server for FASTA search in eukaryotic and microbial genomes.

Similarity searches are a powerful method for solving important biological problems such as database scanning, evolutionary studies, gene prediction, and protein structure prediction. FASTA is a widely used sequence comparison tool for rapid database scanning. Here we describe the GWFASTA server that was developed to assist the FASTA user in similarity searches against partially and/or completely sequenced genomes. GWFASTA consists of more than 60 microbial genomes, eight eukaryote genomes, and proteomes of annotatedgenomes. Infact, it provides the maximum number of databases for similarity searching from a single platform. GWFASTA allows the submission of more than one sequence as a single query for a FASTA search. It also provides integrated post-processing of FASTA output, including compositional analysis of proteins, multiple sequences alignment, and phylogenetic analysis. Furthermore, it summarizes the search results organism-wise for prokaryotes and chromosome-wise for eukaryotes. Thus, the integration of different tools for sequence analyses makes GWFASTA a powerful toolfor biologists.

Computer Systems↗

DNA/DNA hybridization to microarrays reveals gene-specific differences between closely related microbial genomes.

DNA microarrays constructed with full length ORFs from Shewanella oneidensis, MR-1, were hybridized with genomic DNA from nine other Shewanella species and Escherichia coli K-12. This approach enabled visualization of relationships between organisms by comparing individual ORF hybridizations to 164 genes and is further amenable to high-density high-throughput analyses of complete microbial genomes. Conserved genes (arcA and ATP synthase) were identified among all species investigated. The mtr operon, which is involved in iron reduction, was poorly conserved among other known metal-reducing Shewanella species. Results were most informative for closely related organisms with small subunit rRNA sequence similarities greater than 93% and gyrB sequence similarities greater than 80%. At this level of relatedness, the similarity between hybridization profiles was strongly correlated with sequence divergence in the gyrB gene. Results revealed that two strains of S. oneidensis (MR-1 and DLM7) were nearly identical, with only 3% of the ORFs hybridizing poorly, in contrast to hybridizations with Shewanella putrefaciens, formerly considered to be the same species as MR-1, in which 63% of the ORFs hybridized poorly (log ratios below -0.75). Genomic hybridizations showed that genes in operons had consistent levels of hybridization across an operon in comparison to a randomly sampled data set, suggesting that similar applications will be informative for identification of horizontally acquired genes. The full value of microbial genomic hybridizations lies in providing the ability to understand and display specific differences between closely related organisms providing a window into understanding microheterogeneity, bacterial speciation, and taxonomic relationships.

Bacterial Proteins↗

Detection of lateral gene transfer among microbial genomes.

An increasingly comprehensive assessment is being developed of the extent and potential significance of lateral gene transfer among microbial genomes. Genomic sequences can be identified as being of putatively lateral origin by their unexpected phyletic distribution, atypical sequence composition, differential presence or absence in closely related genomes, or incongruent phylogenetic trees. These complementary approaches sometimes yield inconsistent results. Not only more data but also quantitative models and simulations are needed urgently.

Evolution, Molecular↗

Self-identification of protein-coding regions in microbial genomes.

A new method for predicting protein-coding regions in microbial genomic DNA sequences is presented. It uses an ab initio iterative Markov modeling procedure to automatically perform the partition of genomic sequences into three subsets shown to correspond to coding, coding on the opposite strand, and noncoding segments. In contrast to current methods, such as GENEMARK [Borodovsky, M. & McIninch, J. D. (1993) Comput. Chem. 17, 123-133], no training set or prior knowledge of the statistical properties of the studied genome are required. This new method tolerates error rates of 1-2% and can process unassembled sequences. It is thus ideal for the analysis of genome survey and/or fragmented sequence data from uncharacterized microorganisms. The method was validated on 10 complete bacterial genomes (from four major phylogenetic lineages). The results show that protein-coding regions can be identified with an accuracy of up to 90% with a totally automated and objective procedure.

Algorithms↗

Community structure and metabolism through reconstruction of microbial genomes from the environment.

Microbial communities are vital in the functioning of all ecosystems; however, most microorganisms are uncultivated, and their roles in natural systems are unclear. Here, using random shotgun sequencing of DNA from a natural acidophilic biofilm, we report reconstruction of near-complete genomes of Leptospirillum group II and Ferroplasma type II, and partial recovery of three other genomes. This was possible because the biofilm was dominated by a small number of species populations and the frequency of genomic rearrangements and gene insertions or deletions was relatively low. Because each sequence read came from a different individual, we could determine that single-nucleotide polymorphisms are the predominant form of heterogeneity at the strain level. The Leptospirillum group II genome had remarkably few nucleotide polymorphisms, despite the existence of low-abundance variants. The Ferroplasma type II genome seems to be a composite from three ancestral strains that have undergone homologous recombination to form a large population of mosaic genomes. Analysis of the gene complement for each organism revealed the pathways for carbon and nitrogen fixation and energy generation, and provided insights into survival strategies in an extreme environment.

Archaea↗

Tetranucleotide frequencies in microbial genomes.

A computational strategy for determining the variability of long DNA sequences in microbial genomes is described. Composite portraits of bacterial genomes were obtained by computing tetranucleotide frequencies of sections of genomic DNA, converting the frequencies to color images and arranging the images according to their genetic position. The resulting images revealed that the tetranucleotide frequencies of genomic DNA sequences are highly conserved. Sections that were visibly different from those of the rest of the genome contained ribosomal RNA, bacteriophage, or undefined coding regions and had corresponding differences in the variances of tetranucleotide frequencies and GC content. Comparison of nine completely sequenced bacterial genomes showed that there was a nonlinear relationship between variances of the tetranucleotide frequencies and GC content, with the highest variances occurring in DNA sequences with low GC contents (less than 0.30 mol). High variances were also observed in DNA sequences having high GC contents (greater than 0.60 mol), but to a much lesser extent than DNA sequences having low GC contents. Differences in the tetranucleotide frequencies may be due to the mechanisms of intercellular genetic exchange and/or processes involved in maintaining intracellular genetic stability. Identification of sections that were different from those of the rest of the genome may provide information on the evolution and plasticity of bacterial genomes.

Cytosine↗

Microbial genomics: all that you can't leave behind.

Projects designed to scan entire microbial genomes for essential genes have revealed a remarkably compact and conserved, but not universal, set of genes whose functions are necessary for survival or reproduction.

Bacillus subtilis↗