Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome, Microbial”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Microbial genomes--the untapped resource.

Although the 1990s have ushered in the genome, they have also exposed our limitations for deriving structural and functional information. In parallel, molecular phylogeny has demonstrated that the majority of microbial genomes are currently inaccessible. Key objectives for the next century are the development of techniques for accessing 'unculturable' genomes, exploiting their biotechnologically valuable genes and products, and linking genome-sequence data to molecular structure and function.

Eukaryota↗

Properties of overlapping genes are conserved across microbial genomes.

There are numerous examples from the genomes of viruses, mitochondria, and chromosomes that adjacent genes can overlap, sharing at least one nucleotide. Overlaps have been hypothesized to be involved in genome size minimization and as a regulatory mechanism of gene expression. Here we show that overlapping genes are a consistent feature (approximately one-third of all genes) across all microbial genomes sequenced to date, have homologs in more microbes than do non-overlapping genes, and are therefore likely more conserved. In addition, the size, phase (reading frame offset), and distribution, among other characteristics, of overlapping genes are most consistent with the hypothesis that overlaps function in the regulation of gene expression. The upstream sequences and conservation of overlapping orthologs of two model organisms from the genus Prochlorococcus that have significantly different GC-content, and therefore different nucleotide sequences for orthologs, are also consistent with small overlapping sequence regions and programmed shifts in reading frame as a common mechanism in the regulation of microbial gene expression.

Computational Biology↗

Microbial genomics - new targets, new drugs.

Genomics has changed our view of the biological world in the past decade, providing both new information and new tools to characterise biological systems. Over 100 microbial genomes - including many of substantial clinical importance - have been fully or partially sequenced, pushing the search for novel antimicrobial compounds into the post-genomic era. Genomic information and associated new technologies have the potential to revolutionise the drug discovery process. Genomic methods have created a wealth of potential new antimicrobial targets; strategies are evolving to provide validation for these targets before chemical inhibitors are identified. The ability to obtain large amounts of purified target proteins and advances in X-ray crystallography have caused significant increases in available protein structures, which may foreshadow an increased effort in structure-based drug design. The post-genomics strategies used in antimicrobial drug discovery may have application for small molecule drug discovery in numerous therapeutic areas.

Journal Article↗

Conservation of DNA regulatory motifs and discovery of new motifs in microbial genomes.

Regulatory motifs can be found by local multiple alignment of upstream regions from coregulated sets of genes, or regulons. We searched for regulatory motifs using the program AlignACE together with a set of filters that helped us choose the motifs most likely to be biologically relevant in 17 complete microbial genomes. We searched the upstream regions of potentially coregulated genes grouped by three methods: (1) genes that make up functional pathways; (2) genes homologous to regulons from a well-studied species (Escherichia coli); and (3) groups of genes derived from conserved operons. This last group is based on the observation that genes making up homologous regulons in different species are often assorted into coregulated operons in different combinations. This allows partial reconstruction of regulons by looking at operon structure across several species. Unlike other methods for predicting regulons, this method does not depend on the availability of experimental data other than the genome sequence and the locations of genes. New, statistically significant motifs were found in the genome sequence of each organism using each grouping method. The most significant new motif was found upstream of genes in the methane-metabolism functional group in Methanobacterium thermoautotrophicum. We found that at least 27% of the known E. coli DNA-regulatory motifs are conserved in one or more distantly related eubacteria. We also observed significant motifs that differed from the E. coli motif in other organisms upstream of sets of genes homologous to known E. coli regulons, including Crp, LexA, and ArcA in Bacillus subtilis; four anaerobic regulons in Archaeoglobus fulgidus (NarL, NarP, Fnr, and ModE); and the PhoB, PurR, RpoH, and FhlA regulons in other archaebacterial species. We also used motif conservation to aid in finding new motifs by grouping upstream regions from closely related bacteria, thus increasing the number of instances of the motif in the sequence to be aligned. For example, by grouping upstream sequences from three archaebacterial species, we found a conserved motif that may regulate ferrous ion transport that was not found in individual genomes. Discovery of conserved motifs becomes easier as the number of closely related genome sequences increases.

Computational Biology↗

[Whole microbial genome shotgun sequencing].

Two strategies introduced for whole genome sequencing, one is clone by clone method,the other is whole genome shotgun sequencing,for microbes which are very important to us,whole genome shotgun sequencing method is very convenient. In this article we discussed the library construction, long-to-short-ratio of insert, total number of reads should be sequenced, assembly and gap filling technologies of the whole microbial genome shotgun sequencing method while some examples presented.

English Abstract↗

On the nature of gene innovation: duplication patterns in microbial genomes.

Gene duplication is considered a major force in gene family expansion and gene innovation. As gene copies assume novel functions, they must avoid periods of neutrality or be deleted from the genome. Current opinions state that copies avoid neutrality through gene dosage effects. These copies are therefore selected from an early stage. This study concentrates on the flow of copies from recent duplication to gene innovation. We have studied 21 microbial genomes using amino acid divergence to describe paralog evolution in the long-term perspective. Five of these were studied in closer detail using nucleotide divergence for a shorter perspective. It was found that rates of duplication and deletion are high, with only a small fraction of duplications retained and apparently selected. This leads to a steady accumulation of paralogs, which seems to be of a similar magnitude in most of the genomes. Furthermore, it is found that genes of high expression level, as measured by their codon bias, are strongly underrepresented among the most recent duplications. Based on these and other observations, it is suggested that gene innovation is driven by amplification of weak, ancillary functions rather than strong, established functions.

DNA Transposable Elements↗

On the convergence of a clustering algorithm for protein-coding regions in microbial genomes.

MOTIVATION: As the number of fully sequenced prokaryotic genomes continues to grow rapidly, computational methods for reliably detecting protein-coding regions become even more important. Audic and Claverie (1998) Proc. Natl Acad. Sci. USA, 95, 10026-10031, have proposed a clustering algorithm for protein-coding regions in microbial genomes. The algorithm is based on three Markov models of order k associated with subsequences extracted from a given genome. The parameters of the three Markov models are recursively updated by the algorithm which, in simulations, always appear to converge to a unique stable partition of the genome. The partition corresponds to three kinds of regions: (1) coding on the direct strand, (2) coding on the complementary strand, (3) non-coding. RESULTS: Here we provide an explanation for the convergence of the algorithm by observing that it is essentially a form of the expectation maximization (EM) algorithm applied to the corresponding mixture model. We also provide a partial justification for the uniqueness of the partition based on identifiability. Other possible variations and improvements are briefly discussed.

Algorithms↗

Analysis of singleton ORFans in fully sequenced microbial genomes.

Singleton sequence ORFans are orphan ORFs (open reading frames) that have no detectable sequence similarity to any other sequence in the databases. ORFans are of particular interest not only as evolutionary puzzles but also because we can learn little about them using bioinformatics tools. Here, we present a first systematic analysis of singleton ORFans in the first 60 fully sequenced microbial genomes. We show that although ORFans have been underemphasized, the number of ORFans is steadily growing, currently accounting for 23,634 sequences. At the same time, the percentage of ORFans as a fraction of all sequences is slowly diminishing, and is currently about 14%. Short ORFans comprise about 61% of all ORFans. The abundance of short ORFans may be due to a yet unexplained artifact. The data also suggest that the number of longer ORFans may soon diminish as more genomes of closely related organisms become available. To better address the questions about the functions and origins of ORFans, we propose to focus further studies on the longer ORFans, with emphasis on three new types of ORFans: ORFan modules, paralogous ORFans, and orthologous ORFans. We conclude that the large number of ORFans reflects an intrinsic property of the genetic material not yet fully understood. Further computational and experimental studies aimed at understanding Nature's protein diversity should also include ORFans.

Genome↗

Codon-anticodon assignment and detection of codon usage trends in seven microbial genomes.

We have assigned codon-anticodon recognition patterns for the whole set of transfer RNAs of Haemophilus influenzae Rd, Methanococcus jannaschii, and Synechocystis sp. PCC6803 using sequence information derived from the complete genome sequence of these organisms and have tabulated them along with those previously reported for Escherichia coli, Mycoplasma genitalium, Mycoplasma pneumoniae, and Saccharomyces cerevisiae. Using the resulting codon-anticodon tables, the bias in codon usage of genes encoding the entire protein and ribosomal protein complement of each of the seven microbial genomes was analyzed. Then, the codon adaptation index (CAIrp) for each protein gene was calculated using the codon usage preference of the ribosomal protein genes of the corresponding organism. Of the seven genomes examined, six showed CAIrp scores that roughly coincided with the expected level of gene expression. The result demonstrates that CAIrp analysis may be useful for prediction of the expression level of unknown genes when all or at least considerable portions of the genome sequence are available.

Codon↗

The Enhanced Microbial Genomes Library.

Since the obtention of the complete sequence of Haemophilus influenzae Rd in 1995, the number of bacterial genomes entirely sequenced has regularly increased. A problem is that the quality of the annotations of these very large sequences is usually lower than those of the shorter entries encountered in the repository collections. Moreover, classical sequence database management systems have difficulties in handling entries of that size. In this context, we have decided to build the Enhanced Microbial Genomes Library (EMGLib) in which these two problems are alleviated. This library contains all the complete genomes from bacteria already sequenced and the yeast genome in GenBank format. The annotations are improved by the introduction of data on codon usage, gene orientation on the chromosome and gene families. It is possible to access EMGLib through two database systems set up on World Wide Web servers: the PBIL server at http://pbil.univ-lyon1.fr/emglib/emglib. html and the MICADO server at http://locus.jouy.inra.fr/micado

Base Sequence↗

Determination of microbial genome sizes by two-dimensional denaturing gradient gel electrophoresis.

In two-dimensional denaturing gradient gel electrophoresis, DNA is digested with a restriction endonuclease and the resulting DNA fragments are separated as a function of size by conventional agarose gel electrophoresis. Following this first dimension electrophoresis, the fragment distribution is placed at the top of a denaturing gradient slab gel and electrophoresis is carried out parallel to the gradient direction. This second dimension separation is a complex function of the base sequence of each fragment. Analysis of the DNA fragment distribution as a function of fragment size allows the DNA size to be calculated. This method has been applied to calculate three microbial genome sizes: Mycoplasma capricolum, 724 kb; Acholeplasma laidlawii, 1646 kb; and Hemophilus influenzae, 1833 kb.

Acholeplasma↗

A Sanger/pyrosequencing hybrid approach for the generation of high-quality draft assemblies of marine microbial genomes.

Since its introduction a decade ago, whole-genome shotgun sequencing (WGS) has been the main approach for producing cost-effective and high-quality genome sequence data. Until now, the Sanger sequencing technology that has served as a platform for WGS has not been truly challenged by emerging technologies. The recent introduction of the pyrosequencing-based 454 sequencing platform (454 Life Sciences, Branford, CT) offers a very promising sequencing technology alternative for incorporation in WGS. In this study, we evaluated the utility and cost-effectiveness of a hybrid sequencing approach using 3730xl Sanger data and 454 data to generate higher-quality lower-cost assemblies of microbial genomes compared to current Sanger sequencing strategies alone.

Biotechnology↗

Promoter prediction and annotation of microbial genomes based on DNA sequence and structural responses to superhelical stress.

BACKGROUND: In our previous studies, we found that the sites in prokaryotic genomes which are most susceptible to duplex destabilization under the negative superhelical stresses that occur in vivo are statistically highly significantly associated with intergenic regions that are known or inferred to contain promoters. In this report we investigate how this structural property, either alone or together with other structural and sequence attributes, may be used to search prokaryotic genomes for promoters. RESULTS: We show that the propensity for stress-induced DNA duplex destabilization (SIDD) is closely associated with specific promoter regions. The extent of destabilization in promoter-containing regions is found to be bimodally distributed. When compared with DNA curvature, deformability, thermostability or sequence motif scores within the -10 region, SIDD is found to be the most informative DNA property regarding promoter locations in the E. coli K12 genome. SIDD properties alone perform better at detecting promoter regions than other programs trained on this genome. Because this approach has a very low false positive rate, it can be used to predict with high confidence the subset of promoters that are strongly destabilized. When SIDD properties are combined with -10 motif scores in a linear classification function, they predict promoter regions with better than 80% accuracy. When these methods were tested with promoter and non-promoter sequences from Bacillus subtilis, they achieved similar or higher accuracies. We also present a strictly SIDD-based predictor for annotating promoter sequences in complete microbial genomes. CONCLUSION: In this report we show that the propensity to undergo stress-induced duplex destabilization (SIDD) is a distinctive structural attribute of many prokaryotic promoter sequences. We have developed methods to identify promoter sequences in prokaryotic genomes that use SIDD either as a sole predictor or in combination with other DNA structural and sequence properties. Although these methods cannot predict all the promoter-containing regions in a genome, they do find large sets of potential regions that have high probabilities of being true positives. This approach could be especially valuable for annotating those genomes about which there is limited experimental data.

Bacillus subtilis↗

A hybrid clustering approach to recognition of protein families in 114 microbial genomes.

BACKGROUND: Grouping proteins into sequence-based clusters is a fundamental step in many bioinformatic analyses (e.g., homology-based prediction of structure or function). Standard clustering methods such as single-linkage clustering capture a history of cluster topologies as a function of threshold, but in practice their usefulness is limited because unrelated sequences join clusters before biologically meaningful families are fully constituted, e.g. as the result of matches to so-called promiscuous domains. Use of the Markov Cluster algorithm avoids this non-specificity, but does not preserve topological or threshold information about protein families. RESULTS: We describe a hybrid approach to sequence-based clustering of proteins that combines the advantages of standard and Markov clustering. We have implemented this hybrid approach over a relational database environment, and describe its application to clustering a large subset of PDB, and to 328577 proteins from 114 fully sequenced microbial genomes. To demonstrate utility with difficult problems, we show that hybrid clustering allows us to constitute the paralogous family of ATP synthase F1 rotary motor subunits into a single, biologically interpretable hierarchical grouping that was not accessible using either single-linkage or Markov clustering alone. We describe validation of this method by hybrid clustering of PDB and mapping SCOP families and domains onto the resulting clusters. CONCLUSION: Hybrid (Markov followed by single-linkage) clustering combines the advantages of the Markov Cluster algorithm (avoidance of non-specific clusters resulting from matches to promiscuous domains) and single-linkage clustering (preservation of topological information as a function of threshold). Within the individual Markov clusters, single-linkage clustering is a more-precise instrument, discerning sub-clusters of biological relevance. Our hybrid approach thus provides a computationally efficient approach to the automated recognition of protein families for phylogenomic analysis.

Algorithms↗

Microbial genome sequencing.

Complete genome sequences of 30 microbial species have been determined during the past five years, and work in progress indicates that the complete sequences of more than 100 further microbial species will be available in the next two to four years. These results have revealed a tremendous amount of information on the physiology and evolution of microbial species, and should provide novel approaches to the diagnosis and treatment of infectious disease.

Anti-Infective Agents↗

Gut microbial genomes with paired isolates from China illustrate probiotic and cardiometabolic effects.

The gut microbiome displays genetic differences among populations, and characterization of the genomic landscape of the gut microbiome in China remains limited. Here, we present the Chinese Gut Microbial Reference (CGMR) set, comprising 101,060 high-quality metagenomic assembled genomes (MAGs) of 3,707 nonredundant species from 3,234 fecal samples across primarily rural Chinese locations, 1,376 live isolates mainly from lactic acid bacteria, and 987 novel species relative to worldwide databases. We observed region-specific coexisting MAGs and MAGs with probiotic and cardiometabolic functionalities. Preliminary mouse experiments suggest a probiotic effect of two Faecalibacillus intestinalis isolates in alleviating constipation, cardiometabolic influences of three Bacteroides fragilis_A isolates in obesity, and isolates from the genera Parabacteroides and Lactobacillus in host lipid metabolism. Our study expands the current microbial genomes with paired isolates and demonstrates potential host effects, contributing to the mechanistic understanding of host-microbe interactions.

Probiotics↗