Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

[Compact genomes].

The number of procaryotic genomes (both Archaea and Bacteria) completely sequenced is rapidly increasing since the publication in 1995 of the first ever finished one, Haemophilis influenzae. The small size and "simplicity" of these genomes make them ideal models for training in genomic before attacking more complex genomes, but they have also great intrinsic interest. Preliminary analyses of these compact genomes have detected many orphan genes, even in organisms previously extensively studied, as well as many families of duplicated genes. A major task now is to identify the function of these orphans by a combination of in silico, biochemical and genetic analyses (examples will be presented). Several genomes of hyperthermophiles have been or will be completely sequenced soon. Many of their genes have commercial (stable proteins), as well as medical interest (crystallization of proteins with eucaryotic homologs involved in pathogenesis). However, further work with these genomes will require the development of genetic tools for these hyperthermophiles. The complete understanding of genome evolution, structure and function will require the sequencing of many genomes at the different levels of the evolutionary scale. Sequencing of genomes from closely related organisms can be relevant to study genome plasticity, whilst sequencing of genome from different domains (Archaea, Bacteria, Eucarya) can help to reconstruct the Last Universal Common Ancestor (LUCA). The latter is a difficult task and will require not only classical molecular phylogenetic studies (which can be sometimes greatly misleading) but also in depth comparative analyses of all central genetic mechanisms in the three domains to infer their respective evolution. The fundamental problem is to determine if the compact genome of procaryotes is indeed a primitive one (as suggested by the term procaryote itself) or if it has been compacted from a more complex one by evolutionary forces related to the procaryotic way of life. Finally, taking into account the extreme diversity of procaryotes and their metabolism, it should be kept in mind that beside a core of genes essential for cellular life, the myriad of procaryotic genomes contain a mine of non essential genes with potential commercial or medical application. The total number of these genes probably outnumber the total number of eucaryotic genes.

Evolution, Molecular↗

The C-value enigma in plants and animals: a review of parallels and an appeal for partnership.

AIMS: Plants and animals represent the first two kingdoms recognized, and remain the two best-studied groups in terms of nuclear DNA content variation. Unfortunately, the traditional chasm between botanists and zoologists has done much to prevent an integrated approach to resolving the C-value enigma, the long-standing puzzle surrounding the evolution of genome size. This grand division is both unnecessary and counterproductive, and the present review aims to illustrate the numerous links between the patterns and processes found in plants and animals so that a stronger unity can be developed in the future. SCOPE: This review discusses the numerous parallels that exist in genome size evolution between plants and animals, including (i) the construction of large databases, (ii) the patterns of DNA content variation among taxa, (iii) the cytological, morphological, physiological and evolutionary impacts of genome size, (iv) the mechanisms by which genomes change in size, and (v) the development of new methodologies for estimating DNA contents. CONCLUSIONS: The fundamental questions of the C-value enigma clearly transcend taxonomic boundaries, and increased communication is therefore urged among those who study genome size evolution, whether in plants, animals or other organisms.

Animals↗

[Alu repeats in the human genome].

Highly repetitive DNA sequences account for more than 50% of the human genome. The L1 and Alu families harbor the most common mammalian long (LINEs) and short (SINEs) interspersed elements. Alu elements are each a dimer of similar, but not identical, fragments of total size about 300 bp, and originate from the 7SL RNA gene. Each element contains a bipartite promoter for RNA polymerase III, a poly(A) tract located between the monomers, a 3'-terminal poly(A) tract, and numerous CpG islands, and is flanked by short direct repeats. Alu repeats comprise more than 10% of the human genome and are capable of retroposition. Possibly, these elements played an important part in genome evolution. Insertion of an Alu element into a functionally important genome region or other Alu-dependent alterations of gene functions cause various hereditary disorders and are probably associated with carcinogenesis. In total, 14 Alu families differing in diagnostic mutations are known. Some of these, which are present in the human genome, are polymorphic and relatively recently inserted into new loci. Alu copies transposed during ethnic divergence of the human population are useful markers for evolutionary genetic studies.

Alu Elements↗

Quasispecies diversity determines pathogenesis through cooperative interactions in a viral population.

An RNA virus population does not consist of a single genotype; rather, it is an ensemble of related sequences, termed quasispecies. Quasispecies arise from rapid genomic evolution powered by the high mutation rate of RNA viral replication. Although a high mutation rate is dangerous for a virus because it results in nonviable individuals, it has been hypothesized that high mutation rates create a 'cloud' of potentially beneficial mutations at the population level, which afford the viral quasispecies a greater probability to evolve and adapt to new environments and challenges during infection. Mathematical models predict that viral quasispecies are not simply a collection of diverse mutants but a group of interactive variants, which together contribute to the characteristics of the population. According to this view, viral populations, rather than individual variants, are the target of evolutionary selection. Here we test this hypothesis by examining the consequences of limiting genomic diversity on viral populations. We find that poliovirus carrying a high-fidelity polymerase replicates at wild-type levels but generates less genomic diversity and is unable to adapt to adverse growth conditions. In infected animals, the reduced viral diversity leads to loss of neurotropism and an attenuated pathogenic phenotype. Notably, using chemical mutagenesis to expand quasispecies diversity of the high-fidelity virus before infection restores neurotropism and pathogenesis. Analysis of viruses isolated from brain provides direct evidence for complementation between members in the quasispecies, indicating that selection indeed occurs at the population level rather than on individual variants. Our study provides direct evidence for a fundamental prediction of the quasispecies theory and establishes a link between mutation rate, population dynamics and pathogenesis.

Animals↗

Experimental evolution reveals contrasting adaptive landscapes in lab and field environments.

Experimental evolution is widely used to infer microbial responses to environmental change, yet most laboratory studies impose constant, well-mixed conditions that differ fundamentally from fluctuating, spatially structured field environments. We compared genomic evolution in the leaf litter-associated bacterium Curtobacterium strain MMLR14_002 under control and warming treatments in laboratory culture and in a complementary field experiment. Laboratory-derived isolates accumulated more mutations per genome and exhibited stronger locus-level parallelism, with mutations recurring in a small number of coding loci. Field-derived isolates accumulated fewer mutations per genome, and these mutations rarely occurred in the same coding loci across replicate populations. Instead, field isolates exhibited a higher proportion of intergenic mutations, with mutations recurring in the same intergenic regions across independent field deployments. When coding mutations were detected in the field, they were distributed across functionally diffuse targets and more often involved metabolic pathways than the core cellular processes repeatedly targeted during laboratory evolution. Warming itself did not consistently influence mutation accumulation or the genomic distribution of mutations; instead, laboratory and field contexts primarily shaped the accumulation, targets, and repeatability of genomic change. These results suggest that laboratory thermal evolution identifies adaptive routes favored under sustained selection but may overestimate coding-level parallelism under heterogeneous field conditions. Bridging laboratory and field evolution will likely require experimental designs that incorporate temporal variability and spatial heterogeneity characteristic of natural systems.IMPORTANCEA central goal of experimental evolution is to infer how microbes evolve in nature from laboratory studies. Here, we evaluate this assumption by comparing genomic evolution of a leaf litter-associated Curtobacterium strain in laboratory and field warming experiments to identify broad patterns rather than isolate the contribution of any single environmental factor. We find that the strong parallelism at coding loci observed under laboratory conditions is reduced in the field, while mutations recurring in the same intergenic regions across field deployments suggest that parallel evolution in nature may more often involve regulatory noncoding regions rather than coding targets. These results show that environmental context reshapes adaptive landscapes and may limit the parallelism of coding-level genomic responses inferred from homogeneous laboratory conditions.

experimental evolution↗

A study of the middle-scale nucleotide clustering in DNA sequences of various origin and functionality, by means of a method based on a modified standard deviation.

The deviation from randomness in the distribution of nucleotides in genomic sequences is quantified and studied, using a modified standard deviation (MSD). This method implies a "per block" computation of the standard deviation of the nucleotide frequencies of occurrence, using local means (means taken in a neighborhood of each block). This quantity may serve as a scale-dependent measure of the nucleotide clustering. In the present work, the meso-scale of tenths of nucleotides is principally explored, by means of suitably adjusted filter parameters. This length scale is of an order of magnitude not directly affected by the grammar and syntax rules of the protein-coding procedure, remaining shorter than the scale of appearance of large-scale characteristics of the genome. MSD has been found to distinguish systematically between the sequences of different origin and functionality. The most near-random are found to be coding sequences of prokaryotes, while in intronic and intergenic regions of eukaryotic genomes, extended clustering of similar nucleotides is observed. The distributions of MSD values of large collections of sequences are found to be in most cases characteristic of their biological role and origin. Protein- and non-coding, prokaryotic and eukaryotic DNA as well as promoter, rRNA, viral and organelle sequences have been examined. The presented results corroborate a recently proposed model for genome evolution. The method is also applied for an assessment of the annotation of ORFs taken from the complete genome of Saccharomyces cerevisiae.

Animals↗

Gene factories, microfunctionalization and the evolution of gene families.

Gene duplication has long been considered an important force in genome evolution. In this article, I consider families of tandemly duplicated genes that show 'microfunctionalization' - genes encoding similar proteins with subtly different functions, such as olfactory receptors. I discuss the genomic processes giving rise to such microfunctionalized gene families and suggest that, like sites of chromosomal rearrangement and breakage, they are associated with relatively high concentrations of repetitive elements. I suggest that microfunctionalized gene families arise within gene factories: genomic regions rich in repetitive elements that undergo increased levels of unequal crossing-over.

Animals↗

The genome sequence of the filamentous fungus Neurospora crassa.

Neurospora crassa is a central organism in the history of twentieth-century genetics, biochemistry and molecular biology. Here, we report a high-quality draft sequence of the N. crassa genome. The approximately 40-megabase genome encodes about 10,000 protein-coding genes--more than twice as many as in the fission yeast Schizosaccharomyces pombe and only about 25% fewer than in the fruitfly Drosophila melanogaster. Analysis of the gene set yields insights into unexpected aspects of Neurospora biology including the identification of genes potentially associated with red light photobiology, genes implicated in secondary metabolism, and important differences in Ca2+ signalling as compared with plants and animals. Neurospora possesses the widest array of genome defence mechanisms known for any eukaryotic organism, including a process unique to fungi called repeat-induced point mutation (RIP). Genome analysis suggests that RIP has had a profound impact on genome evolution, greatly slowing the creation of new genes through genomic duplication and resulting in a genome with an unusually low proportion of closely related genes.

Calcium Signaling↗

Bacteriophage T4 genome.

Phage T4 has provided countless contributions to the paradigms of genetics and biochemistry. Its complete genome sequence of 168,903 bp encodes about 300 gene products. T4 biology and its genomic sequence provide the best-understood model for modern functional genomics and proteomics. Variations on gene expression, including overlapping genes, internal translation initiation, spliced genes, translational bypassing, and RNA processing, alert us to the caveats of purely computational methods. The T4 transcriptional pattern reflects its dependence on the host RNA polymerase and the use of phage-encoded proteins that sequentially modify RNA polymerase; transcriptional activator proteins, a phage sigma factor, anti-sigma, and sigma decoy proteins also act to specify early, middle, and late promoter recognition. Posttranscriptional controls by T4 provide excellent systems for the study of RNA-dependent processes, particularly at the structural level. The redundancy of DNA replication and recombination systems of T4 reveals how phage and other genomes are stably replicated and repaired in different environments, providing insight into genome evolution and adaptations to new hosts and growth environments. Moreover, genomic sequence analysis has provided new insights into tail fiber variation, lysis, gene duplications, and membrane localization of proteins, while high-resolution structural determination of the "cell-puncturing device," combined with the three-dimensional image reconstruction of the baseplate, has revealed the mechanism of penetration during infection. Despite these advances, nearly 130 potential T4 genes remain uncharacterized. Current phage-sequencing initiatives are now revealing the similarities and differences among members of the T4 family, including those that infect bacteria other than Escherichia coli. T4 functional genomics will aid in the interpretation of these newly sequenced T4-related genomes and in broadening our understanding of the complex evolution and ecology of phages-the most abundant and among the most ancient biological entities on Earth.

Bacteriophage T4↗

The next generation of microarray research: applications in evolutionary and ecological genomics.

Microarray technology is one of the key developments in recent years that has propelled biological research into the post-genomic era. With the ability to assay thousands to millions of features at the same time, microarray technology has fundamentally changed how biological questions are addressed, from examining one or a few genes to a collection of genes or the whole genome. This technology has much to offer in the study of genome evolution. After a brief introduction on the technology itself, we then focus on the use of microarrays to examine genome dynamics, to uncover novel functional elements in genomes, to unravel the evolution of regulatory networks, to identify genes important for behavioral and phenotypic plasticity, and to determine microbial community diversity in environmental samples. Although there are still practical issues in using microarrays, they will be alleviated by rapid advances in array technology and analysis methods, the availability of many genome sequences of closely related species and flexibility in array design. It is anticipated that the application of microarray technology will continue to better our understanding of evolution and ecology through the examination of individuals, populations, closely related species or whole microbial communities.

Ecology↗

Mammalian BEX, WEX and GASP genes: coding and non-coding chimaerism sustained by gene conversion events.

BACKGROUND: The identification of sequence innovations in the genomes of mammals facilitates understanding of human gene function, as well as sheds light on the molecular mechanisms which underlie these changes. Although gene duplication plays a major role in genome evolution, studies regarding concerted evolution events among gene family members have been limited in scope and restricted to protein-coding regions, where high sequence similarity is easily detectable. RESULTS: We describe a mammalian-specific expansion of more than 20 rapidly-evolving genes on human chromosome Xq22.1. Many of these are highly divergent in their protein-coding regions yet contain a conserved sequence motif in their 5' UTRs which appears to have been maintained by multiple events of concerted evolution. These events have led to the generation of chimaeric genes, each with a 5' UTR and a protein-coding region that possess independent evolutionary histories. We suggest that concerted evolution has occurred via gene conversion independently in different mammalian lineages, and these events have resulted in elevated G+C levels in the encompassing genomic regions. These concerted evolution events occurred within and between genes from three separate protein families ('brain-expressed X-linked' [BEX], WWbp5-like X-linked [WEX] and G-protein-coupled receptor-associated sorting protein [GASP]), which often are expressed in mammalian brains and associated with receptor mediated signalling and apoptosis. CONCLUSION: Despite high protein-coding divergence among mammalian-specific genes, we identified a DNA motif common to these genes' 5' UTR exons. The motif has undergone concerted evolution events independently of its neighbouring protein-coding regions, leading to formation of evolutionary chimaeric genes. These findings have implications for the identification of non protein-coding regulatory elements and their lineage-specific evolution in mammals.

5' Untranslated Regions↗

Mutualists and parasites: how to paint yourself into a (metabolic) corner.

Eukaryotes have developed an elaborate series of interactions with bacteria that enter their bodies and/or cells. Genome evolution of symbiotic and parasitic bacteria multiplying inside eukaryotic cells results in both convergent and divergent changes. The genome sequences of the symbiotic bacteria of aphids, Buchnera aphidicola, and the parasitic bacteria of body louse and humans, Rickettsia prowazekii, provide insights into these processes. Convergent genome characteristics include reduction in genome sizes and lowered G+C content values. Divergent evolution was recorded for amino acid and cell wall biosynthetic genes. The presence of pseudogenes in both genomes provides examples of recent gene inactivation events and offers clues to the process of genome deterioration and host-cell adaptation.

Adaptation, Physiological↗

Genome histories clarify evolution of the expansin superfamily: new insights from the poplar genome and pine ESTs.

Expansins comprise a superfamily of plant cell wall-loosening proteins that has been divided into four distinct families, EXPA, EXPB, EXLA and EXLB. In a recent analysis of Arabidopsis thaliana and Oryza sativa expansins, we proposed a further subdivision of the families into 17 clades, representing independent lineages in the last common ancestor of monocots and eudicots. This division was based on both traditional sequence-based phylogenetic trees and on position-based trees, in which genomic locations and dated segmental duplications were used to reconstruct gene phylogeny. In this article we review recent work concerning the patterns of expansin evolution in angiosperms and include additional insights gained from the genome of a second eudicot species, Populus trichocarpa, which includes at least 36 expansin genes. All of the previously proposed monocot-eudicot orthologous groups, but no additional ones, are represented in this species. The results also confirm that all of these clades are truly independent lineages. Furthermore, we have used position-based phylogeny to clarify the history of clades EXPA-II and EXPA-IV. Most of the growth of the expansin superfamily in the poplar lineage is likely due to a recent polyploidy event. Finally, some monocot-eudicot clades are shown to have diverged before the separation of the angiosperm and gymnosperm lineages.

Arabidopsis↗

Evolution of dnmt-2 and mbd-2-like genes in the free-living nematodes Pristionchus pacificus, Caenorhabditis elegans and Caenorhabditis briggsae.

Whole genome sequencing of several metazoan model organisms provides a platform for studying genome evolution. How representative are the genomes of these model organisms for their respective phyla? Within nematodes, for example, the free-living soil nematode Caenorhabditis elegans is a highly derived species with unusual genomic characters, such as a reduced Hox cluster (Curr. Biol., 13, 37-40) and the absence of a Hedgehog signaling system. Here, we describe the recent loss of a DNA methyltransferase-2 gene (dnmt-2) in C.elegans. A dnmt-2-like gene is present in the satellite model organism Pristionchus pacificus, another free-living nematode that diverged from C.elegans 200-300 million years ago. In contrast, C.elegans, Caenorhabditis briggsae and P.pacificus all contain an mbd-2-like gene, which encodes another essential component of the methylation system of higher animals and fungi. Cel-mbd-2 is expressed throughout development and RNA interference (RNAi) experiments result in variable phenotypes. In contrast, Cbr-mbd-2 RNAi results in paralyzed larval or adult worms suggesting recent changes of gene function within the genus Caenorhabditis. We speculate that both genes were part of an ancestral DNA methylation system in nematodes and that gene loss and sequence divergence have abolished DNA methylation in C.elegans.

Amino Acid Sequence↗

Characterisation of the chloroplast genome of Macrotyloma species: comparative analysis and phylogenomic insights.

Macrotyloma is an underutilised legume genus within the tribe Phaseoleae (Fabaceae) that includes nutritionally and agronomically important crops such as horse gram (Macrotyloma uniflorum) and Kersting's groundnut (Macrotyloma geocarpum). Despite their importance, knowledge of the chloroplast (cp.) genome of this genus remains limited. In this study, we assembled and analysed the complete chloroplast genomes of three Macrotyloma species: M. uniflorum, M. geocarpum, and M. axillare. The chloroplast genomes were assembled into two isoforms that differ in the orientation of the small single-copy (SSC) region. Genome sizes ranged from 150,811 to 151,013 bp and exhibited the canonical quadripartite structure, comprising a pair of inverted repeats (IRa and IRb; 26,416-26,436 bp each), a large single-copy region (LSC; 80,229-80,446 bp), and a small single-copy region (SSC; 17,710-17,711 bp). Each genome encoded 110 unique genes, including 4 rRNA genes, 30 tRNA genes, and 76 protein-coding genes. All three species also possessed the ~ 50 kb inversion in the LSC region, a synapomorphy shared among a large clade within the Papilionoideae subfamily of Fabaceae. Although overall chloroplast genome structure and organisation were highly conserved among Macrotyloma species, gene-wise nucleotide diversity analysis identified seven relatively variable genes: rps18, rps15, ccsA, ndhA, ycf1, ycf4, and psaI. Phylogenomic analysis based on complete chloroplast genomes robustly resolved Macrotyloma as a monophyletic group within the Phaseolinae clade of the Papilionoideae subfamily. Within the genus, M. uniflorum and M. axillare formed a strongly supported sister pair, with M. geocarpum sister to this clade. Overall, this study provides valuable insights into chloroplast genome evolution in Macrotyloma and enhances understanding of its phylogenetic placement within Phaseoleae, offering genomic resources for future evolutionary, taxonomic, and conservation studies of this underutilised legume genus.

Genome, Chloroplast↗

Sequence of the tomato chloroplast DNA and evolutionary comparison of solanaceous plastid genomes.

Tomato, Solanum lycopersicum (formerly Lycopersicon esculentum), has long been one of the classical model species of plant genetics. More recently, solanaceous species have become a model of evolutionary genomics, with several EST projects and a tomato genome project having been initiated. As a first contribution toward deciphering the genetic information of tomato, we present here the complete sequence of the tomato chloroplast genome (plastome). The size of this circular genome is 155,461 base pairs (bp), with an average AT content of 62.14%. It contains 114 genes and conserved open reading frames (ycfs). Comparison with the previously sequenced plastid DNAs of Nicotiana tabacum and Atropa belladonna reveals patterns of plastid genome evolution in the Solanaceae family and identifies varying degrees of conservation of individual plastid genes. In addition, we discovered several new sites of RNA editing by cytidine-to-uridine conversion. A detailed comparison of editing patterns in the three solanaceous species highlights the dynamics of RNA editing site evolution in chloroplasts. To assess the level of intraspecific plastome variation in tomato, the plastome of a second tomato cultivar was sequenced. Comparison of the two genotypes (IPA-6, bred in South America, and Ailsa Craig, bred in Europe) revealed no nucleotide differences, suggesting that the plastomes of modern tomato cultivars display very little, if any, sequence variation.

Amino Acid Sequence↗

Evolutionary genomics: new genes for new jobs.

Whole genome sequence analyses have confirmed that gene duplication and divergence play major roles in genome evolution. But the details of how young, functionally redundant gene duplicates escape mutational degradation have remained elusive. Several recent studies show that new genes survive because they evolve new, and sometimes essential, functions.

Animals↗

Papillomaviruses: different genes have different histories.

Papillomaviruses (PVs) infect stratified squamous epithelia in vertebrates. Some PVs are associated with different types of cancer and with certain benign lesions. It has been assumed that PVs coevolved with their hosts. However, recently it has been shown that different regions of the genome have different evolutionary histories. The PV genome has a modular nature and appeared after the addition of pre-existent blocks. This order of appearance in the PV genome is evident today in the different evolutionary rates of the different genes, with new genes--E5, E6 and E7--diverging faster than old genes--E1, E2, L2 and L1. Here, we propose an evolutionary framework aiming to integrate genome evolution, PV biology and epidemiology of PV infections.

Evolution, Molecular↗