Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

Comparison of multiple vertebrate genomes reveals the birth and evolution of human exons.

Orthologous gene structures in eight vertebrate species were compared on a genomic scale to detect the birth and maturation of new internal exons during the course of evolution. We found that 40% of new human exons are alternatively spliced, and most of these are cassette exons (exons that are either included or skipped in their entirety) with low inclusion rates. This proportion decreases steadily as older and older exons are examined, even as splicing efficiency increases. Remarkably, the great majority of new cassette exons are composed of highly repeated sequences, especially Alu. Many new cassette exons are 5' untranslated exons; the proportion that code for protein increases steadily with age. New protein-coding exons evolve at a high rate, as evidenced by the initially high substitution rates (K(s) and K(a)), as well as the SNP density compared with older exons. This dynamic picture suggests that de novo recruitment rather than shuffling is the major route by which exons are added to genes, and that species-specific repeats could play a significant role in recent evolution.

Alternative Splicing↗

Two patterns of genome organization in mammals: the chromosomal distribution of duplicate genes in human and mouse.

Gene duplication occurs repeatedly in the evolution of genomes, and the rearrangement of genomic segments has also occurred repeatedly over the evolution of eukaryotes. We studied the interaction of these two factors in mammalian evolution by comparing the chromosomal distribution of multigene families in human and mouse. In both species, gene families tended to be confined to a single chromosome to a greater extent than expected by chance. The average number of families shared between chromosomes was nearly 60% higher in mouse than in human, and human chromosomes rarely shared large numbers of gene families with more than one or two other chromosomes, whereas mouse chromosomes frequently did so. A higher proportion of duplicate gene pairs on the same chromosome originated from recent duplications in human than in mouse, whereas a higher proportion of duplicate gene pairs on separate chromosomes arose from ancient duplications in human than in mouse. These observations are most easily explained by the hypotheses that (1) most gene duplications arise in tandem and are subsequently separated by segmental rearrangement events, and (2) that the process of segmental rearrangement has occurred at a higher rate in the lineage of mouse than in that of human.

Animals↗

Did genomic imprinting and X chromosome inactivation arise from stochastic expression?

Both X chromosome inactivation and autosomal genomic imprinting generate a functional hemizygosity. Here we consider models that explain the evolution of genomic imprinting and X chromosome inactivation from novel perspectives. Specifically, we suggest that random (in)activation events are common in genes and gene clusters with a low probability of transcription. These generate variability that natural selection has acted on to evolve stable monoallelic expression. Possible selection forces might include a need for dosage compensation and the prevention of biallelic silencing where a total switch off would be lethal. Two different mechanisms can accomplish regular monoallelic expression - genomic imprinting and gene counting.

Animals↗

Microbial genome sequencing 2000: new insights into physiology, evolution and expression analysis.

The complete genome sequence has been reported for 24 microbial organisms. The genome organization and gene content of these organisms has revealed an incredible diversity. Nearly half of the open reading frames identified by these sequencing projects are for potential genes with no known biological function. Efforts to make evolutionary sense and biological sense of the gene content of these organisms have been initiated. The greatest future challenge of genomics will be to determine function for the unknown genes.

Bacteria↗

Archaeology and evolution of transfer RNA genes in the Escherichia coli genome.

Transfer RNA genes tend to be presented in multiple copies in the genomes of most organisms, from bacteria to eukaryotes. The evolution and genomic structure of tRNA genes has been a somewhat neglected area of molecular evolution. Escherichia coli, the first phylogenetic species for which more than two different strains have been sequenced, provides an invaluable framework to study the evolution of tRNA genes. In this work, a detailed analysis of the tRNA structure of the genomes of Escherichia coli strains K12, CFT073, and O157:H7, Shigella flexneri 2a 301, and Salmonella typhimurium LT2 was carried out. A phylogenetic analysis of these organisms was completed, and an archaeological map depicting the main events in the evolution of tRNA genes was drawn. It is shown that duplications, deletions, and horizontal gene transfers are the main factors driving tRNA evolution in these genomes. On average, 0.64 tRNA insertions/duplications occur every million years (Myr) per genome per lineage, while deletions occur at the slower rate of 0.30 per million years per genome per lineage. This work provides a first genomic glance at the problem of tRNA evolution as a repetitive process, and the relationship of this mechanism to genome evolution and codon usage is discussed.

Codon↗

Genome-wide molecular clock and horizontal gene transfer in bacterial evolution.

We describe a simple theoretical framework for identifying orthologous sets of genes that deviate from a clock-like model of evolution. The approach used is based on comparing the evolutionary distances within a set of orthologs to a standard intergenomic distance, which was defined as the median of the distribution of the distances between all one-to-one orthologs. Under the clock-like model, the points on a plot of intergenic distances versus intergenomic distances are expected to fit a straight line. A statistical technique to identify significant deviations from the clock-like behavior is described. For several hundred analyzed orthologous sets representing three well-defined bacterial lineages, the alpha-Proteobacteria, the gamma-Proteobacteria, and the Bacillus-Clostridium group, the clock-like null hypothesis could not be rejected for approximately 70% of the sets, whereas the rest showed substantial anomalies. Subsequent detailed phylogenetic analysis of the genes with the strongest deviations indicated that over one-half of these genes probably underwent a distinct form of horizontal gene transfer, xenologous gene displacement, in which a gene is displaced by an ortholog from a different lineage. The remaining deviations from the clock-like model could be explained by lineage-specific acceleration of evolution. The results indicate that although xenologous gene displacement is a major force in bacterial evolution, a significant majority of orthologous gene sets in three major bacterial lineages evolved in accordance with the clock-like model. The approach described here allows rapid detection of deviations from this mode of evolution on the genome scale.

Evolution, Molecular↗

Molecular evolution of an imprinted gene: repeatability of patterns of evolution within the mammalian insulin-like growth factor type II receptor.

The repeatability of patterns of variation in Ka/Ks and Ks is expected if such patterns are the result of deterministic forces. We have contrasted the molecular evolution of the mammalian insulin-like growth factor type II receptor (Igf2r) in the mouse-rat comparison with that in the human-cow comparison. In so doing, we investigate explanations for both the evolution of genomic imprinting and for Ks variation (and hence putatively for mutation rate evolution). Previous analysis of Igf2r, in the mouse-rat comparison, found Ka/Ks patterns that were suggested to be contrary to those expected under the conflict theory of imprinting. We find that Ka/Ks variation is repeatable and hence confirm these patterns. However, we also find that the molecular evolution of Igf2r signal sequences suggests that positive selection, and hence conflict, may be affecting this region. The variation in Ks across Igf2r is also repeatable. To the best of our knowledge this is the first demonstration of such repeatability. We consider three explanations for the variation in Ks across the gene: (1) that it is the result of mutational biases, (2) that it is the result of selection on the mutation rate, and (3) that it is the product of selection on codon usage. Explanations 2 and 3 predict a Ka-Ks correlation, which is not found. Explanation 3 also predicts a negative correlation between codon bias and Ks, which is also not found. However, in support of explanation 1 we do find that in rodents the rate of silent C --> T mutations at CpG sites does covary with Ks, suggesting that methylation-induced mutational patterns can explain some of the variation in Ks. We find evidence to suggest that this CpG effect is due to both variation in CpG density, and to variation in the frequency with which CpGs mutate. Interestingly, however, a GC4 analysis shows no covariance with Ks, suggesting that to eliminate methyl-associated effects CpG rates themselves must be analyzed. These results suggest that, in contrast to previous studies of intragenic variation, Ks patterns are not simply caused by the same forces responsible for Ka/Ks correlations.

Animals↗

Naturally occurring antisense: transcriptional leakage or real overlap?

Naturally occurring antisense transcription is associated with the regulation of gene expression through a variety of biological mechanisms. Several recent genome-wide studies reported the identification of potential antisense transcripts for thousands of mammalian genes, many of them resulting from alternatively polyadenylated transcripts or heterogeneous transcription start sites. However, it is not clear whether this transcriptional plasticity is intentional, leading to regulated overlap between the transcripts, or, alternatively, represents a "leakage" of the RNA transcription machinery. To address this question through an evolutionary approach, we compared the genomic organization of genes, with or without antisense, between human, mouse, and the pufferfish Fugu rubripes. Our hypothesis was that if two neighboring genes overlap and have a sense-antisense relationship, we would expect negative selection acting on the evolutionary separation between them. We found that antisense gene pairs are twice as likely to preserve their genomic organization throughout vertebrates' evolution compared to nonantisense pairs, implying an overlap existence in the ancestral genome. In addition, we show that increasing the genomic distance between pairs of genes having a sense-antisense relationship is selected against. These findings indicate that, at least in part, the abundance of antisense transcripts observed in expressed data represents real overlap rather than transcriptional leakage. Moreover, our results imply that natural antisense transcription has considerably affected vertebrate genome evolution.

Animals↗

Identification of intergenomic recombinations in unisexual salamanders of the genus Ambystoma by genomic in situ hybridization (GISH).

Unisexual salamanders in the genus Ambystoma (Amphibia, Caudata) are endemic to eastern North America and are mostly all-female polyploids. Two to four of the bisexual species, A. laterale, A. jeffersonianum, A. texanum and A. tigrinum, contribute to the nuclear genome of unisexuals and more than 20 combinations that range from diploid to pentaploid have been identified in this complex. Because the karyotypes of the four bisexual species are similar, homologous and homoeologous chromosomes in the unisexuals can not be distinguished by conventional or banded karyotypes. We chose two widespread unisexual genomic combinations (A.laterale-2 jeffersonianum [or LJJ] and A. 2 laterale-jeffersonianum [or LLJ]) and employed genomic in situ hybridization (GISH) to identify the genomes in these unisexuals. Under optimum conditions, GISH reliably distinguishes the respective chromosomes attributed to both A.laterale and A. jeffersonianum. Of four populations examined, two were found to have independently evolved homoeologous recombinants that persist in both LJJ and LLJ individuals. Our results refute the previous hypothesis of clonal integrity and independent evolution of the genome combinations in these unisexuals. Our data provide evidence for intergenomic interactions between maternal chromosomes during meiosis in unisexuals and help to explain previously observed non-homologous bivalents and/or quadrivalents among lampbrush chromosomes that were possibly initiated by partial homosequential pairing among the homo(eo)logues. To explore the utility of GISH in other members of the complex, probes developed from A. laterale were also applied to unisexuals that contained A. tigrinum and A. texanum genomes. GISH is an effective tool that can be used to identify and to quantify genomic constituents and to investigate intergenomic interactions in unisexual salamanders. GISH also has potential application to examine possible genomic evolution in other unisexuals.

Ambystoma↗

A cascade of complex subtelomeric duplications during the evolution of the hominoid and Old World monkey genomes.

Subtelomeric duplications of an obscure tubulin "genic" segment located near the telomere of human chromosome 4q35 have occurred at different evolutionary time points within the last 25 million years of the catarrhine (i.e., hominoid and Old World monkey) evolution. The analyses of these segments reported here indicate an exceptional level of evolutionary instability. Substantial intra- and interspecific differences in copy number and distribution are observed among cercopithecoid (Old World monkey) and hominoid genomes. Characterization of the hominoid duplicated segments reveals a strong positional bias within pericentromeric and subtelomeric regions of the genome. On the basis of phylogenetic analysis from predicted proteins and comparisons of nucleotide-substitution rates, we present evidence of a conserved b-tubulin gene among the duplications. Remarkably, the evolutionary conservation has occurred in a nonorthologous fashion, such that the functional copy has shifted its positional context between hominoids and cercopithecoids. We propose that, in a chimpanzee-human common ancestor, one of the paralogous copies assumed the original function, whereas the ancestral copy acquired mutations and eventually became silenced. Our analysis emphasizes the dynamic nature of duplication-mediated genome evolution and the delicate balance between gene acquisition and silencing.

Animals↗

The complete chloroplast genome sequence of Pelargonium x hortorum: organization and evolution of the largest and most highly rearranged chloroplast genome of land plants.

The chloroplast genome of Pelargonium x hortorum has been completely sequenced. It maps as a circular molecule of 217,942 bp and is both the largest and most rearranged land plant chloroplast genome yet sequenced. It features 2 copies of a greatly expanded inverted repeat (IR) of 75,741 bp each and, consequently, diminished single-copy regions of 59,710 and 6,750 bp. Despite the increase in size and complexity of the genome, the gene content is similar to that of other angiosperms, with the exceptions of a large number of pseudogenes, the recognition of 2 open reading frames (ORF56 and ORF42) in the trnA intron with similarities to previously identified mitochondrial products (ACRS and pvs-trnA), the losses of accD and trnT-ggu and, in particular, the presence of a highly divergent set of rpoA-like ORFs rather than a single, easily recognized gene for rpoA. The 3-fold expansion of the IR (relative to most angiosperms) accounts for most of the size increase of the genome, but an additional 10% of the size increase is related to the large number of repeats found. The Pelargonium genome contains 35 times as many 31 bp or larger repeats than the unrearranged genome of Spinacia. Most of these repeats occur near the rearrangement hotspots, and 2 different associations of repeats are localized in these regions. These associations are characterized by full or partial duplications of several genes, most of which appear to be nonfunctional copies or pseudogenes. These duplications may also be linked to the disruption of at least 1 but possibly 2 or 3 operons. We propose simple models that account for the major rearrangements with a minimum of 8 IR boundary changes and 12 inversions in addition to several insertions of duplicated sequence.

Chloroplasts↗

Echoes from the past--are we still in an RNP world?

Availability of the human genome sequence and those of other species is unmeasured in their value for a comprehensive understanding of the architecture, function and evolution of genomes and cells. Various mechanisms keep genomes in flux and generate intra- and interspecies variation. The conversion of RNA modules into DNA and their more or less random integration into chromosomes (retroposition) is in many lineages including our own the most pervasive and perhaps the most enigmatic. The proclivity of such events in extant multicellular eukaryotes, even in more recent evolutionary times, gives the impression that the transition period from the RNP (ribonucleoprotein) world to the emergence of modern cells, where DNA became the predominant carrier of genetic information, has lasted billions of years and is an endlessly drawn-out process rather than the punctuated event one might expect. Apart from the impact of such RNA-mediated processes as retroposition, the role of RNA in a wide variety of cellular functions has only recently become more widely appreciated.

Animals↗

A computational prediction of isochores based on hidden Markov models.

Mammalian genomes are organised into a mosaic of regions (in general more than 300 kb in length), with differing, relatively homogeneous G+C contents. The G+C content is the basic characteristic of isochores, but they have also been associated with many other biological properties. For instance, the genes are more compact and their density is highest in G+C rich isochores. Various ways of locating isochores in the human genome have been developed, but such methods use only the base composition of the DNA sequences. The present paper proposes a new method, based on a hidden Markov model, which takes into account several of the biological properties associated with the isochore structure of a genome. This method leads to good segmentation of the human genome into isochores, and also permits a new analysis of the known heterogeneity of G+C rich isochores: most (60%) of the G+C poor genes embedded in G+C rich isochores have UTR sequences characteristic of G+C rich genes. This genomic feature is discussed in the context of both evolution and genome function.

5' Untranslated Regions↗

[An introduction of several programs used in genomic analysis].

Genomics is a novel subject that has been developed accompanying with the progress of human genome project. Genomics deals with the chemistry component, structure organization and evolution of genome at global level. As genomics associated with huge data, bioinformatics plays an important role in these processes of data production, data management and data mining. At present, many reliable programs have been used in genomic research successfully, which are usually accessible and downloaded freely. We address here the principles of some programs used wildly in genomics such as sequence alignment, sequence assembly, repeat identification and gene prediction, which are exemplified with typical programs respectively.

English Abstract↗

A detailed RFLP map of cotton, Gossypium hirsutum x Gossypium barbadense: chromosome organization and evolution in a disomic polyploid genome.

We employ a detailed restriction fragment length polymorphism (RFLP) map to investigate chromosome organization and evolution in cotton, a disomic polyploid. About 46.2% of nuclear DNA probes detect RFLPs distinguishing Gossypium hirsutum and Gossypium barbadense; and 705 RFLP loci are assembled into 41 linkage groups and 4675 cM. The subgenomic origin (A vs. D) of most, and chromosomal identity of 14 (of 26), linkage groups is shown. The A and D subgenomes show similar recombinational length, suggesting that repetitive DNA in the physically larger A subgenome is recombinationally inert. RFLPs are somewhat more abundant in the D subgenome. Linkage among duplicated RFLPs reveals 11 pairs of homoelogous chromosomal regions-two appear homosequential, most differ by inversions, and at least one differs by a translocation. Most homoeologies involve chromosomes from different subgenomes, putatively reflecting the n = 13 to n = 26 polyploidization event of 1.1-1.9 million years ago. Several observations suggest that another, earlier, polyploidization event spawned n = 13 cottons, at least 25 million years ago. The cotton genome contains about 400-kb DNA per cM, hence map-based gene cloning is feasible. The cotton map affords new opportunities to study chromosome evolution, and to exploit Gossypium genetic resources for improvement of the world's leading natural fiber.

Biological Evolution↗

The use of comparative genomic hybridization to characterize genome dynamics and diversity among the serotypes of Shigella.

BACKGROUND: Compelling evidence indicates that Shigella species, the etiologic agents of bacillary dysentery, as well as enteroinvasive Escherichia coli, are derived from multiple origins of Escherichia coli and form a single pathovar. To further understand the genome diversity and virulence evolution of Shigella, comparative genomic hybridization microarray analysis was employed to compare the gene content of E. coli K-12 with those of 43 Shigella strains from all lineages. RESULTS: For the 43 strains subjected to CGH microarray analyses, the common backbone of the Shigella genome was estimated to contain more than 1,900 open reading frames (ORFs), with a mean number of 726 undetectable ORFs. The mosaic distribution of absent regions indicated that insertions and/or deletions have led to the highly diversified genomes of pathogenic strains. CONCLUSION: These results support the hypothesis that by gain and loss of functions, Shigella species became successful human pathogens through convergent evolution from diverse genomic backgrounds. Moreover, we also found many specific differences between different lineages, providing a window into understanding bacterial speciation and taxonomic relationships.

DNA, Bacterial↗