Evolution of bacterial genomes.
This review examines evolution of bacterial genomes with an emphasis on RNA based life, the transition to functional DNA and small evolving genomes (possible plasmids) that led to larger, functional bacterial genomes.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
This review examines evolution of bacterial genomes with an emphasis on RNA based life, the transition to functional DNA and small evolving genomes (possible plasmids) that led to larger, functional bacterial genomes.
Genome evolution in eukaryotes is predominantly driven by the dynamics of repetitive sequences, which vary widely in both copy number and sequence composition. Rates of repeat evolution differ between and within species and are likely modulated by both genetics and environment. To uncover factors shaping the rate of genome content evolution, we analyzed 1,142 resequenced Arabidopsis thaliana genomes using a novel K-mer based approach to characterize genome content variation and identify hypervariable regions underlying differences in repeat abundance. We next treated repeat abundance as a quantitative trait and performed genome-wide association analyses across more than 400 repeat families to identify the genetic basis of copy number variation. Integrating these results through a meta-GWAS approach revealed both cis-acting variants and more than 50 trans-acting loci that regulate repeat abundance genome-wide. Cis-acting variation was predominantly localized to pericentromeric and centromeric regions, whereas trans-acting loci were enriched for candidate genes involved in DNA replication, DNA repair, DNA methylation regulation. Finally, we found evidence that purifying selection acts against mutations that accelerate genome content divergence, favoring alleles that constrain repeat expansion. Together, these findings provide new insights into the genetic architecture and evolutionary forces shaping genome evolution in A. thaliana and establish a framework for investigating these processes in other plant species.
Explore the source record for details and available documents.
Del(13)Svea36H (Del36H) is a deletion of approximately 20% of mouse chromosome 13 showing conserved synteny with human chromosome 6p22.1-6p22.3/6p25. The human region is lost in some deletion syndromes and is the site of several disease loci. Heterozygous Del36H mice show numerous phenotypes and may model aspects of human genetic disease. We describe 12.7 Mb of finished, annotated sequence from Del36H. Del36H has a higher gene density than the draft mouse genome, reflecting high local densities of three gene families (vomeronasal receptors, serpins, and prolactins) which are greatly expanded relative to human. Transposable elements are concentrated near these gene families. We therefore suggest that their neighborhoods are gene factories, regions of frequent recombination in which gene duplication is more frequent. The gene families show different proportions of pseudogenes, likely reflecting different strengths of purifying selection and/or gene conversion. They are also associated with relatively low simple sequence concentrations, which vary across the region with a periodicity of approximately 5 Mb. Del36H contains numerous evolutionarily conserved regions (ECRs). Many lie in noncoding regions, are detectable in species as distant as Ciona intestinalis, and therefore are candidate regulatory sequences. This analysis will facilitate functional genomic analysis of Del36H and provides insights into mouse genome evolution.
Guanine plus cytosine (GC) content ranges broadly among bacterial genomes. In this study, we explore the use of a Brownian-motion model for the evolution of GC content over time. This model assumes that GC content varies over time in a continuous and homogeneous manner. Using this model and a maximum-likelihood approach, we analyzed the evolution of GC content across several bacterial phylogenies. Using three independent tests, we found that the observed divergence in GC content was consistent with a homogeneous Brownian-motion model. For example, similar rates of GC content evolution were inferred in several different bacterial subclades, indicating that there is relatively little rate heterogeneity in GC content evolution over broad evolutionary time scales. We thus argue that the homogeneous Brownian-motion model provides a good working model for GC content evolution. We then use this model to determine the overall rate of GC content evolution among eubacteria. We also determine the time frame over which GC content remains similar in related taxa, using a flexible definition for "similarity" in GC content so that, depending on the context, more or less stringent criteria may be applied. Our results have implications for models of sequence evolution, including those used for phylogenetic reconstruction and for inferring unusual changes in GC content.
It is generally accepted that the wide variation in genome size observed among eukaryotic species is more closely correlated with the amount of repetitive DNA than with the number of coding genes. Major types of repetitive DNA include transposable elements, satellite DNAs, simple sequences and tandem repeats, but reliable estimates of the relative contributions of these various types to total genome size have been hard to obtain. With the advent of genome sequencing, such information is starting to become available, but no firm conclusions can yet be made from the limited data currently available. Here, the ways in which transposable elements contribute both directly and indirectly to genome size variation are explored. Limited evidence is provided to support the existence of an approximately linear relationship between total transposable element DNA and genome size. Copy numbers per family are low and globally constrained in small genomes, but vary widely in large genomes. Thus, the partial release of transposable element copy number constraints appears to be a major characteristic of large genomes.
In the gram-negative model organism Escherichia coli, the effector molecule of the stringent response, (p)ppGpp, is synthesized by two different enzymes, RelA and SpoT, whereas in the gram-positive model organism Bacillus subtilis only one enzyme named Rel is responsible for this activity. Rel and SpoT also possess (p)ppGpp hydrolase activity. BLAST searches were used to identify orthologous genes in databases. The construction and bootstrapping of phylogenetic trees allowed classification of these orthologs. Four groups could be distinguished: With the exception of Neisseria and Bordetella (beta subdivision), the RelA and SpoT groups are exclusively found in the gamma subdivision of proteobacteria. Two Rel groups representing the actinobacterial and the Bacillus/Clostridium group were also identified. The SpoT proteins are related to the gram positive Rel proteins. RelA proteins carry substitutions in the HD domain (Aravind and Koonin, 1998, TIBS 23: 469-472) responsible for ppGpp degradation. A theory for the evolution of the specialized, paralogous relA and spoT genes is presented: After gene duplication of an ancestral rellike gene, the spoT and relA genes evolved from the duplicated genes. The distribution pattern of the paralogous RelA and SpoT proteins supports a new model of linear bacterial evolution (Gupta, 2000, FEMS Microbiol. Rev. 24: 367-402). This model postulates that the gamma subdivision of proteobacteria represents the most recently evolved bacterial lineage. However, two paralogous, closely related genes of Porphyromonas gingivalis (Cytophaga-Flavobacterium-Bacteroides phylum) encoding proteins with functions probably identical to the RelA and SpoT proteins do not fit in this model. Completely sequenced genomes of several obligately parasitic organisms (Treponema pallidum, Chlamydia species, Rickettsia prowazekii) and the obligate aphid symbiont Buchnera sp. APS as well as archaea do not contain rel-like genes but they are present in the Arabidopsis genome. In crosslinking experiments using different analogs of ppGpp as crosslinking reagents and RNA polymerase preparations of Escherichia coli, binding of ppGpp to distinct regions at the C-terminus of the beta subunit (the RpoB gene product) and/or at the N-terminus of the beta subunit (the RpoC gene product) was observed previously. RpoB and RpoC sequences of the species which do not possess a rel like gene do not exhibit specific insertions or deletions in the ppGpp binding regions.
BACKGROUND: The 1.83 Megabase (Mb) sequence of the Haemophilus influenzae chromosome, the first completed genome sequence of a cellular life form, has been recently reported. Approximately 75 % of the 4.7 Mb genome sequence of Escherichia coli is also available. The life styles of the two bacteria are very different - H. influenzae is an obligate parasite that lives in human upper respiratory mucosa and can be cultivated only on rich media, whereas E. coli is a saprophyte that can grow on minimal media. A detailed comparison of the protein products encoded by these two genomes is expected to provide valuable insights into bacterial cell physiology and genome evolution. RESULTS: We describe the results of computer analysis of the amino-acid sequences of 1703 putative proteins encoded by the complete genome of H. influenzae. We detected sequence similarity to proteins in current databases for 92 % of the H. influenzae protein sequences, and at least a general functional prediction was possible for 83 %. A comparison of the H. influenzae protein sequences with those of 3010 proteins encoded by the sequenced 75 % of the E. coli genome revealed 1128 pairs of apparent orthologs, with an average of 59 % identity. In contrast to the high similarity between orthologs, the genome organization and the functional repertoire of genes in the two bacteria were remarkably different. The smaller genome size of H. influenzae is explained, to a large extent, by a reduction in the number of paralogous genes. There was no long range colinearity between the E. coli and H. influenzae gene orders, but over 70 % of the orthologous genes were found in short conserved strings, only about half of which were operons in E. coli. Superposition of the H. influenzae enzyme repertoire upon the known E. coli metabolic pathways allowed us to reconstruct similar and alternative pathways in H. influenzae and provides an explanation for the known nutritional requirements. CONCLUSIONS: By comparing proteins encoded by the two bacterial genomes, we have shown that extensive gene shuffling and variation in the extent of gene paralogy are major trends in bacterial evolution; this comparison has also allowed us to deduce crucial aspects of the largely uncharacterized metabolism of H. influenzae.
BACKGROUND: HSP90 proteins are essential molecular chaperones involved in signal transduction, cell cycle control, stress management, and folding, degradation, and transport of proteins. HSP90 proteins have been found in a variety of organisms suggesting that they are ancient and conserved. In this study we investigate the nuclear genomes of 32 species across all kingdoms of organisms, and all sequences available in GenBank, and address the diversity, evolution, gene structure, conservation and nomenclature of the HSP90 family of genes across all organisms. RESULTS: Twelve new genes and a new type HSP90C2 were identified. The chromosomal location, exon splicing, and prediction of whether they are functional copies were documented, as well as the amino acid length and molecular mass of their polypeptides. The conserved regions across all protein sequences, and signature sequences in each subfamily were determined, and a standardized nomenclature system for this gene family is presented. The proeukaryote HSP90 homologue, HTPG, exists in most Bacteria species but not in Archaea, and it evolved into three lineages (Groups A, B and C) via two gene duplication events. None of the organellar-localized HSP90s were derived from endosymbionts of early eukaryotes. Mitochondrial TRAP and endoplasmic reticulum HSP90B separately originated from the ancestors of HTPG Group A in Firmicutes-like organisms very early in the formation of the eukaryotic cell. TRAP is monophyletic and present in all Animalia and some Protista species, while HSP90B is paraphyletic and present in all eukaryotes with the exception of some Fungi species, which appear to have lost it. Both HSP90C (chloroplast HSP90C1 and location-undetermined SP90C2) and cytosolic HSP90A are monophyletic, and originated from HSP90B by independent gene duplications. HSP90C exists only in Plantae, and was duplicated into HSP90C1 and HSP90C2 isoforms in higher plants. HSP90A occurs across all eukaryotes, and duplicated into HSP90AA and HSP90AB in vertebrates. Diplomonadida was identified as the most basal organism in the eukaryote lineage. CONCLUSION: The present study presents the first comparative genomic study and evolutionary analysis of the HSP90 family of genes across all kingdoms of organisms. HSP90 family members underwent multiple duplications and also subsequent losses during their evolution. This study established an overall framework of information for the family of genes, which may facilitate and stimulate the study of this gene family across all organisms.
Despite the fact that natural transformation was described long ago in Streptococcus pneumoniae, only a limited number of recombination genes have been identified. Two of them have recently been characterized at the molecular level, recA which encodes a protein essential for homologous recombination and mmsA which encodes the homologue of the Escherichia coli RecG protein. After a survey of the available information regarding the function of RecA, RecG, and other proteins such as the mismatch repair proteins HexA and HexB that can affect the outcome of recombinants, the different levels at which horizontal genetic exchange can be controlled are discussed. It is shown that the specific induction of the recA gene which occurs in competent cells is required for full recombination proficiency. Results regarding the ability of the Hex generalized mismatch repair system to prevent recombination between partially divergent sequences during transformation are also summarized. A structural analysis of homeologous recombinants which suggests that formation of mosaic recombinants can occur independently of mismatch repair in a single-step transformation is also reported. Finally, arguments in favor of an evolutionary origin of transformation as a means of genome evolution are discussed and the different types of recombination events observed which could potentially contribute to S. pneumoniae genome evolution are listed.
Genome-level studies are contributing to a major renaissance in crop science. In wheat, there are now more than 500,000 expressed sequence tags, and these are being used in conjunction with specially designed deletion stocks to unravel patterns of genome evolution, recombination and polyploid genome behavior.
At a small number of mammalian loci, only one of the two copies of a gene is expressed. Just which copy is expressed depends on the sex of the parent from which that copy was inherited. Such genes are said to be imprinted. The functional haploidy implied by imprinting has a number of population genetic consequences. Moreover, since diploidy is widely believed to be advantageous, the evolution of this non-Mendelian form of expression requires an explanation. Here I examine some of the theoretical and mathematical models investigating these two aspects of imprinting. For instance, the dynamics and equilibrium properties of many models of natural selection at imprinted loci are formally equivalent to models without imprinting. And different approaches to modeling the problem of the evolution of imprinting reveal the weakness of several of the apparent predictions of various verbal hypotheses about why imprinting has evolved.
A general evolutionary trend is the generation of organisms of increasing complexity, notwithstanding that reduction and simplification phenomena do occur in the evolutionary process. This paper proposes an evolutionary model incorporating the mechanisms of gene amplification and deletion. The evolutionary process leading to genomic complexity and the coexistence of simpler organisms with complicated ones were both simulated using the proposed model. The model was also used to investigate the influence of various factors on the evolution of complexity. The simulations indicated that the evolution of complexity is largely influenced by adaptation to complicated environments. Nevertheless, complex organisms require relatively more resources for survival and replication, which limits the on going tendency towards complexity. Moreover, the analysis showed that if the environment varies rapidly and the profit obtained from complexity is greater than the resources consumed, selection will tend to favor complexity. However, high living cost will tend to limit the trend of complexity and if the environment is relatively stable, reduction and simplification will become the dominant trends.
Genome comparisons have demonstrated that dramatic genetic change often underlies the emergence of new bacterial pathogens. Evolutionary analysis of Escherichia coli O157:H7, a pathogen that has emerged as a worldwide public health threat in the past two decades, has posited that this toxin-producing pathogen evolved in a series of steps from O55:H7, a recent ancestor of a nontoxigenic pathogenic clone associated with infantile diarrhea. We used comparative genomic hybridization with 50-mer oligonucleotide microarrays containing probes from both pathogenic and nonpathogenic genomes to infer when genes were acquired and lost. Many ancillary virulence genes identified in the O157 genome were already present in an O55:H7-like progenitor, with 27 of 33 genomic islands of >5 kb and specific for O157:H7 (O islands) that were acquired intact before the split from this immediate ancestor. Most (85%) of variably absent or present genes are part of prophages or phage-like elements. Divergence in gene content among these closely related strains was approximately 140 times greater than divergence at the nucleotide sequence level. A >100-kb region around the O-antigen gene cluster contained highly divergent sequences and also appears to be duplicated in its entirety in one lineage, suggesting that the whole region was cotransferred in the antigenic shift from O55 to O157. The beta-glucuronidase-positive O157 variants, although phylogenetically closest to the Sakai strain, were divergent for multiple adherence factors. These observations suggest that, in addition to gains and losses of phage elements, O157:H7 genomes are rapidly diverging and radiating into new niches as the pathogen disseminates.
We present nine diallelic models of genetic conflict in which one allele is imprintable and the other is not to examine how genomic imprinting may have evolved. Imprinting is presumed to be either maternal (i.e., the maternally derived gene is inactivated) or paternal. Females are assumed to be either completely monogamous or always bigamous, so that we may see any effect of multiple paternity. In contrast to previous verbal and quantitative genetic models, we find that genetic conflicts need not lead to paternal imprinting of growth inhibitors and maternal imprinting of growth enhancers. Indeed, in some of our models--those with strict monogamy--the dynamics of maternal and paternal imprinting are identical. Multiple paternity is not necessary for the evolution of imprinting, and in our models of maternal imprinting, multiple paternity has no effect at all. Nevertheless, multiple paternity favors the evolution of paternal imprinting of growth inhibitors and hinders that of growth enhancers. Hence, any degree of multiple paternity means that growth inhibitors are more likely to be paternally imprinted, and growth enhancers maternally so. In all of our models, stable polymorphism of imprinting status is possible and mean fitness can decrease over time. Neither of these behaviors have been predicted by previous models.
Pyrococcus sp. KOD1 is a newly isolated hyperthermophilic archaeon from a solfatara at a wharf on Kodakara Island, Kagoshima, Japan. A physical map of the KOD1 chromosome was constructed using pulsed-field gel electrophoresis of restriction fragments generated by AscI, PacI and PmeI. The order of the AscI fragments was deduced from Southern hybridization using the AscI, PmeI and PacI fragments as a probe. The derived physical map indicates that KOD1 possesses a circular-form genome and its size was estimated to be 2036 kb. Several cloned genes were hybridized to restriction fragments to locate their positions on the physical map. Some genes involved in the central dogma were located on the restricted segment of the genome. Novel characteristics of KOD1 enzymes are also introduced in this article.
Evolution of fitness values upon replication of viral populations is strongly influenced by the size of the virus population that participates in the infections. While large population passages often result in fitness gains, repeated plaque-to-plaque transfers result in average fitness losses. Here we develop a numerical model that describes fitness evolution of viral clones subjected to serial bottleneck events. The model predicts a biphasic evolution of fitness values in that a period of exponential decrease is followed by a stationary state in which fitness values display large fluctuations around an average constant value. This biphasic evolution is in agreement with experimental results of serial plaque-to-plaque transfers carried out with foot-and-mouth disease virus (FMDV) in cell culture. The existence of a stationary phase of fitness values has been further documented by serial plaque-to-plaque transfers of FMDV clones that had reached very low relative fitness values. The statistical properties of the stationary state depend on several parameters of the model, such as the probability of advantageous versus deleterious mutations, initial fitness, and the number of replication rounds. In particular, the size of the bottleneck is critical for determining the trend of fitness evolution.
There are no doubts that transposable elements (TEs) have greatly influenced genomes evolution. They have, however, evolved in different ways throughout mammals, plants, and invertebrates. In mammals they have been shown to be widely present but with low transposition activity; in plants they are responsible for large increases in genome size. In Drosophila, despite their low amount, transposition seems to be higher. Therefore, to understand how these elements have evolved in different genomes and how host genomes have proposed to go around them, are major questions on genome evolution. We analyzed sequences of the retrotransposable elements 412 in natural populations of the Drosophila simulans and D. melanogaster species that greatly differ in their amount of TEs. We identified new subfamilies of this element that were the result of mutation or insertion-deletion process, but also of interfamily recombinations. These new elements were well conserved in the D. simulans natural populations. The new regulatory regions produced by recombination could give rise to new elements able to overcome host control of transposition and, thus, become potential genome invaders.