Dealing with database explosion: a cautionary note.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
In order to extract the maximum amount of information from the rapidly accumulating genome sequences, all conserved genes need to be classified according to their homologous relationships. Comparison of proteins encoded in seven complete genomes from five major phylogenetic lineages and elucidation of consistent patterns of sequence similarities allowed the delineation of 720 clusters of orthologous groups (COGs). Each COG consists of individual orthologous proteins or orthologous sets of paralogs from at least three lineages. Orthologs typically have the same function, allowing transfer of functional information from one member to an entire COG. This relation automatically yields a number of functional predictions for poorly characterized genomes. The COGs comprise a framework for functional and evolutionary genome analysis.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
We designed PCR primers by using the DNA sequences of the soluble methane monooxygenase gene clusters of Methylosinus trichosporium OB3b and Methylococcus capsulatus (Bath), and these primers were found to be specific for four of the five structural genes in the soluble methane monooxygenase gene clusters of several methanotrophs. We also designed primers for the gram-negative methylotroph-specific methanol dehydrogenase gene moxF. The specificity of these primers was confirmed by hybridizing and sequencing the PCR products obtained. The primers were then used to amplify methanotroph DNAs in samples obtained from various aquatic and terrestrial environments. Our sequencing data suggest that a large number of different methanotrophs are present in peat samples and also that there is a high level of variability in the mmoC gene, which codes for the reductase component of the soluble methane monooxygenase, while the mmoX gene, which codes for the alpha subunit of the hydroxylase component of this enzyme complex, appears to be highly conserved in methanotrophs.
Recent work has raised a question as to the involvement of erythrose-4-phosphate, a product of the pentose phosphate pathway, in the metabolism of the methanogenic archaea (R. H. White, Biochemistry 43:7618-7627, 2004). To address the possible absence of erythrose-4-phosphate in Methanocaldococcus jannaschii, we have assayed cell extracts of this methanogen for the presence of this and other intermediates in the pentose phosphate pathway and have determined and compared the labeling patterns of sugar phosphates derived metabolically from [6,6-2H2]- and [U-13C]-labeled glucose-6-phosphate incubated with cell extracts. The results of this work have established the absence of pentose phosphate pathway intermediates erythrose-4-phosphate, xylose-5-phosphate, and sedoheptulose-7-phosphate in these cells and the presence of D-arabino-3-hexulose-6-phosphate, an intermediate in the ribulose monophosphate pathway. The labeling of the D-ara-bino-3-hexulose-6-phosphate, as well as the other sugar-Ps, indicates that this hexose-6-phosphate was the precursor to ribulose-5-phosphate that in turn was converted into ribose-5-phosphate by ribose-5-phosphate isomerase. Additional work has demonstrated that ribulose-5-phosphate is derived by the loss of formaldehyde from D-arabino-3-hexulose-6-phosphate, catalyzed by the protein product of the MJ1447 gene.
Archaeal motility occurs through the rotation of flagella that are distinct from the flagella found on bacteria. The differences between the two structures include the multi-flagellin nature of the archaeal filament, the widespread posttranslational modification of the flagellins and the presence of a short signal peptide on each flagellin that is cleaved by a specific signal peptidase prior to the incorporation of the mature flagellin into the flagellar filament. Research has revealed similarities between the archaeal flagellum and the type IV pilus, including the presence of similar unusual signal peptides on the flagellins and pilins, similarities in the amino acid sequences of the major structural proteins themselves, as well as similarities between potential assembly and processing components. The recent suggestion that type IV pili are part of a family of cell surface complexes, coupled with the similarities between type IV pili and archaeal flagella, raise questions about the evolution of these systems and possible inclusion of archaeal flagella into this surface complex family.
BACKGROUND: Escherichia coli guanine-N2 (m2G) methyltransferases (MTases) RsmC and RsmD modify nucleosides G1207 and G966 of 16S rRNA. They possess a common MTase domain in the C-terminus and a variable region in the N-terminus. Their C-terminal domain is related to the YbiN family of hypothetical MTases, but nothing is known about the structure or function of the N-terminal domain. RESULTS: Using a combination of sequence database searches and fold recognition methods it has been demonstrated that the N-termini of RsmC and RsmD are related to each other and that they represent a "degenerated" version of the C-terminal MTase domain. Novel members of the YbiN family from Archaea and Eukaryota were also indentified. It is inferred that YbiN and both domains of RsmC and RsmD are closely related to a family of putative MTases from Gram-positive bacteria and Archaea, typified by the Mj0882 protein from M. jannaschii (1dus in PDB). Based on the results of sequence analysis and structure prediction, the residues involved in cofactor binding, target recognition and catalysis were identified, and the mechanism of the guanine-N2 methyltransfer reaction was proposed. CONCLUSIONS: Using the known Mj0882 structure, a comprehensive analysis of sequence-structure-function relationships in the family of genuine and putative m2G MTases was performed. The results provide novel insight into the mechanism of m2G methylation and will serve as a platform for experimental analysis of numerous uncharacterized N-MTases.
BACKGROUND: Many current gene prediction methods use only one model to represent protein-coding regions in a genome, and so are less likely to predict the location of genes that have an atypical sequence composition. It is likely that future improvements in gene finding will involve the development of methods that can adequately deal with intra-genomic compositional variation. RESULTS: This work explores a new approach to gene-prediction, based on the Self-Organizing Map, which has the ability to automatically identify multiple gene models within a genome. The current implementation, named RescueNet, uses relative synonymous codon usage as the indicator of protein-coding potential. CONCLUSIONS: While its raw accuracy rate can be less than other methods, RescueNet consistently identifies some genes that other methods do not, and should therefore be of interest to gene-prediction software developers and genome annotation teams alike. RescueNet is recommended for use in conjunction with, or as a complement to, other gene prediction methods.
BACKGROUND: Amino acids in proteins are not used equally. Some of the differences in the amino acid composition of proteins are between species (mainly due to nucleotide composition and lifestyle) and some are between proteins from the same species (related to protein function, expression or subcellular localization, for example). As several factors contribute to the different amino acid usage in proteins, it is difficult both to analyze these differences and to separate the contributions made by each factor. RESULTS: Using a multi-way method called Tucker3, we have analyzed the amino composition of a set of 64 orthologous groups of proteins present in 62 archaea and bacteria. This dataset corresponds to essential proteins such as ribosomal proteins, tRNA synthetases and translational initiation or elongation factors, which are common to all the species analyzed. The Tucker3 model can be used to study the amino acid variability within and between species by taking into consideration the tridimensionality of the data set. We found that the main factor behind the amino acid composition of proteins is independent of the organism or protein function analyzed. This factor must be related to the biochemical characteristics of each amino acid. The difference between the non-ribosomal proteins and the ribosomal proteins (which are rich in arginine and lysine) is the main factor behind the differences in amino acid composition within species, while G+C content and optimal growth temperature are the main factors behind the differences in amino acid usage between species. CONCLUSION: We show that a multi-way method is useful for comparing the amino acid composition of several groups of orthologous proteins from the same group of species. This kind of dataset is extremely useful for detecting differences between and within species.
BACKGROUND: DNA tracts composed of only two bases are possible in six combinations: A+G (purines, R), C+T (pyrimidines, Y), G+T (Keto, K), A+C (Imino, M), A+T (Weak, W) and G+C (Strong, S). It is long known that all-pyrimidine tracts, complemented by all-purines tracts ("R.Y tracts"), are excessively present in analyzed DNA. We have previously shown that R.Y tracts are in vast excess in yeast promoters, and brought evidence for their role in gene regulation. Here we report the systematic mapping of all six binary combinations on the level of complete sequenced chromosomes, as well as in their different subregions. RESULTS: DNA tracts composed of the above binary base combinations have been mapped in seven sequenced chromosomes: Human chromosomes 21 and 22 (the major contigs); Drosophila melanogaster chr. 2R; Caenorhabditis elegans chr. I; Arabidopsis thaliana chr. II; Saccharomyces cerevisiae chr. IV and M. jannaschii. A huge over-representation, reaching million-folds, has been found for very long tracts of all binary motifs except S, in each of the seven organisms. Long R.Y tracts are the most excessive, except in D. melanogaster, where the K.M motif predominates. S (G, C rich) tracts are in excess mainly in CpG islands; the W motif predominates in bacteria. Many excessively long W tracts are nevertheless found also in the archeon and in the eukaryotes. The survey of complete chromosomes enables us, for the first time, to map systematically the intergenic regions. In human and other chromosomes we find the highest over-representation of the binary DNA tracts in the intergenic regions. These over-representations are only partly explainable by the presence of interspersed elements. CONCLUSIONS: The over-representation of long DNA tracts composed of five of the above motifs is the largest deviation from randomness so far established for DNA, and this in a wide range of eukaryotic and archeal chromosomes. A propensity for ready DNA unwinding is proposed as the functional role, explaining the evolutionary conservation of the huge excesses observed.
BACKGROUND: Phylogenetic analysis of the Archaea has been mainly established by 16S rRNA sequence comparison. With the accumulation of completely sequenced genomes, it is now possible to test alternative approaches by using large sequence datasets. We analyzed archaeal phylogeny using two concatenated datasets consisting of 14 proteins involved in transcription and 53 ribosomal proteins (3,275 and 6,377 positions, respectively). RESULTS: Important relationships were confirmed, notably the dichotomy of the archaeal domain as represented by the Crenarchaeota and Euryarchaeota, the sister grouping of Sulfolobales and Aeropyrum pernix, and the monophyly of a large group comprising Thermoplasmatales, Archaeoglobus fulgidus, Methanosarcinales and Halobacteriales, with the latter two orders forming a robust cluster. The main difference concerned the position of Methanopyrus kandleri, which grouped with Methanococcales and Methanobacteriales in the translation tree, whereas it emerged at the base of the euryarchaeotes in the transcription tree. The incongruent placement of M. kandleri is likely to be the result of a reconstruction artifact due to the high evolutionary rates displayed by the components of its transcription apparatus. CONCLUSIONS: We show that two informational systems, transcription and translation, provide a largely congruent signal for archaeal phylogeny. In particular, our analyses support the appearance of methanogenesis after the divergence of the Thermococcales and a late emergence of aerobic respiration from within methanogenic ancestors. We discuss the possible link between the evolutionary acceleration of the transcription machinery in M. kandleri and several unique features of this archaeon, in particular the absence of the elongation transcription factor TFS.
Cells respond to a wide variety of mechanical stimuli, ranging from thermal molecular agitation to potentially destructive cell swelling caused by osmotic pressure gradients. The cell membrane presents a major target of the external mechanical forces that act upon a cell, and mechanosensitive (MS) ion channels play a crucial role in the physiology of mechanotransduction. These detect and transduce external mechanical forces into electrical and/or chemical intracellular signals. Recent work has increased our understanding of their gating mechanism, physiological functions and evolutionary origins. In particular, there has been major progress in research on microbial MS channels. Moreover, cloning and sequencing of MS channels from several species has provided insights into their evolution, their physiological functions in prokaryotes and eukaryotes, and their potential roles in the pathology of disease.
The ubiquity of mechanosensitive (MS) channels triggered a search for their functional homologues in Archaea, the third domain of the phylogenetic tree. Two types of MS channels have been identified in the cell membranes of Haloferax volcanii using the patch clamp technique. Recently MS channels were identified and cloned from two archaeal species occupying different environmental habitats. These studies demonstrate that archaeal MS channels share structural and functional homology with bacterial MS channels. The mechanical force transmitted via the lipid bilayer alone activates all to date known prokaryotic MS channels. This implies the existence of a common gating mechanism for bacterial as well as archaeal MS channels according to the bilayer model. Based on recent evidence that the bilayer model also applies to eukaryotic MS channels, mechanosensory transduction probably originated along with the appearance of the first life forms according to simple biophysical principles. In support of this hypothesis the phylogenetic analysis revealed that prokaryotic MS channels of large and small conductance originated from a common ancestral molecule resembling the bacterial MscL channel protein. Furthemore, bacterial and archaeal MS channels share common structural motifs with eukaryotic channels of diverse function indicating the importance of identified structures to the gating mechanism of this family of channels. The comparative approach used throughout this review should contribute towards understanding of the evolution and molecular basis of mechanosensory transduction in general.
One of the essential maturation steps to yield functional tRNA molecules is the removal of 3'-trailer sequences by RNase Z. After RNase Z cleavage the tRNA nucleotidyl transferase adds the CCA sequence to the tRNA 3'-terminus, thereby generating the mature tRNA. Here we investigated whether a terminal CCA triplet as 3'-trailer or embedded in a longer 3'-trailer influences cleavage site selection by RNase Z using three activities: a recombinant plant RNase Z, a recombinant archaeal RNase Z and an RNase Z active wheat extract. A trailer of only the CCA trinucleotide is left intact by the wheat extract RNase Z but is removed by the recombinant plant and archaeal enzymes. Thus the CCA triplet is not recognized by the RNase Z enzyme itself, but rather requires cofactors still present in the extract. In addition, we investigated the influence of acceptor stem length on cleavage by RNase Z using variants of wild-type tRNATyr. While the wild type and the variant with 8 base pairs in the acceptor stem were processed efficiently by all three activities, variants with shorter and longer acceptor stems were poor substrates or were not cleaved at all.
The translocon is a protein-conducting channel conserved over all domains of life that serves to translocate proteins across or into membranes. Although this channel has been well studied for many years, the recent discovery of a high-resolution crystal structure opens up new avenues of exploration. Taking advantage of this, we performed molecular dynamics simulations of the translocon in a fully solvated lipid bilayer, examining the translocation abilities of monomeric SecYEbeta by forcing two helices comprised of different amino acid sequences to cross the channel. The simulations revealed that the so-called plug of SecYEbeta swings open during translocation, closing thereafter. Likewise, it was established that the so-called pore ring region of SecYEbeta forms an elastic, yet tight, seal around the translocating oligopeptides. The closed state of the channel was found to block permeation of all ions and water molecules; in the open state, ions were blocked. Our results suggest that the SecYEbeta monomer is capable of forming an active channel.
Explore the source record for details and available documents.