Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Accelerated evolution associated with genome reduction in a free-living prokaryote.

BACKGROUND: Three complete genomes of Prochlorococcus species, the smallest and most abundant photosynthetic organism in the ocean, have recently been published. Comparative genome analyses reveal that genome shrinkage has occurred within this genus, associated with a sharp reduction in G+C content. As all examples of genome reduction characterized so far have been restricted to endosymbionts or pathogens, with a host-dependent lifestyle, the observed genome reduction in Prochlorococcus is the first documented example of such a process in a free-living organism. RESULTS: Our results clearly indicate that genome reduction has been accompanied by an increased rate of protein evolution in P. marinus SS120 that is even more pronounced in P. marinus MED4. This acceleration has affected every functional category of protein-coding genes. In contrast, the 16S rRNA gene seems to have evolved clock-like in this genus. We observed that MED4 and SS120 have lost several DNA-repair genes, the absence of which could be related to the mutational bias and the acceleration of amino-acid substitution. CONCLUSIONS: We have examined the evolutionary mechanisms involved in this process, which are different from those known from host-dependent organisms. Indeed, most substitutions that have occurred in Prochlorococcus have to be selectively neutral, as the large size of populations imposes low genetic drift and strong purifying selection. We assume that the major driving force behind genome reduction within the Prochlorococcus radiation has been a selective process favoring the adaptation of this organism to its environment. A scenario is proposed for genome evolution in this genus.

Adaptation, Physiological↗

Functional genomics and enzyme evolution. Homologous and analogous enzymes encoded in microbial genomes.

Computational analysis of complete genomes, followed by experimental testing of emerging hypotheses--the area of research often referred to as 'functional genomics'--aims at deciphering the wealth of information contained in genome sequences and at using it to improve our understanding of the mechanisms of cell function. This review centers on the recent progress in the genome analysis with special emphasis on the new insights in enzyme evolution. Standard methods of predicting functions for new proteins are listed and the common errors in their application are discussed. A new method of improving the functional predictions is introduced, based on a phylogenetic approach to functional prediction, as implemented in the recently constructed Clusters of Orthologous Groups (COG) database (available at http:@www.ncbi.nlm.nih.gov/COG). This approach provides a convenient way to characterize the protein families (and metabolic pathways) that are present or absent in any given organism. Comparative analysis of microbial genomes based on this approach shows that metabolic diversity generally correlates with the genome size-parasitic bacteria code for fewer enzymes and lesser number of metabolic pathways than their free-living relatives. Comparison of different genomes reveals another evolutionary trend, the non-orthologous gene displacement of some enzymes by unrelated proteins with the same cellular function. An examination of the phylogenetic distribution of such cases provides new clues to the problems of biochemical evolution, including evolution of glycolysis and the TCA cycle.

Databases, Factual↗

Complex satellite DNA reshuffling in the polymorphic t(1;29) Robertsonian translocation and evolutionarily derived chromosomes in cattle.

We have analysed and mapped physically the satellite I, III (subunits pvu and sau) and IV DNA sequences in cattle using in-situ hybridization. Four breeds were analysed including individuals with a chromosome number of 2n = 60 and individuals with the widespread t(1;29) in the homozygous (2n = 58) and heterozygous state (2n = 59). All three satellite DNA families were present at the centromeres of the many but not all of the autosomal acrocentric chromosomes, and essentially absent from the sex chromosomes. In the translocated t(1;29) chromosome, the satellite DNA families showed a different pattern from that simply derived by fusion of the acrocentric autosomes and loss of satellite sequences, with no variation between breeds. A model of centromeric evolution is presented involving two independent events. Knowledge of mechanisms of translocation formation within cattle is important for a functional understanding of centromere and satellites, investigation of chromosomal abnormalities, and for understanding chromosomal fusion during evolution of other bovids and genome evolution in general.

Animals↗

Complete mitochondrial genome sequence of Urechis caupo, a representative of the phylum Echiura.

BACKGROUND: Mitochondria contain small genomes that are physically separate from those of nuclei. Their comparison serves as a model system for understanding the processes of genome evolution. Although hundreds of these genome sequences have been reported, the taxonomic sampling is highly biased toward vertebrates and arthropods, with many whole phyla remaining unstudied. This is the first description of a complete mitochondrial genome sequence of a representative of the phylum Echiura, that of the fat innkeeper worm, Urechis caupo. RESULTS: This mtDNA is 15,113 nts in length and 62% A+T. It contains the 37 genes that are typical for animal mtDNAs in an arrangement somewhat similar to that of annelid worms. All genes are encoded by the same DNA strand which is rich in A and C relative to the opposite strand. Codons ending with the dinucleotide GG are more frequent than would be expected from apparent mutational biases. The largest non-coding region is only 282 nts long, is 71% A+T, and has potential for secondary structures. CONCLUSIONS: Urechis caupo mtDNA shares many features with those of the few studied annelids, including the common usage of ATG start codons, unusual among animal mtDNAs, as well as gene arrangements, tRNA structures, and codon usage biases.

Amino Acid Sequence↗

Transposable elements drive evolution and perturb gene expression in Brassica rapa and B. oleracea.

Transposable elements (TEs) significantly influence genomic diversity and gene regulation in plants. Brassica rapa and B. oleracea, with their distinct domestication histories, offer excellent models to explore TE dynamics. Here, we developed a refined TE classification method and systematically analyzed TEs across 12 B. rapa and B. oleracea genomes, identifying 1878 TE families. Approximately half (49.5%) of these TE families were shared between the two species, reflecting a common evolutionary origin, whereas species-specific expansions, particularly among long-terminal repeat (LTR) retrotransposons, underscore their roles in genomic differentiation. We notably characterized a heat-responsive Ty1-copia family (Copia0035) in B. oleracea roots, distinguished by low GC content and the absence of CG and CHG methylation motifs, sharing regulatory similarities with the Arabidopsis heat-induced ONSEN element. Syntenic analyses of gene-TE associations highlighted significant intraspecies TE insertion variability, with more accession-specific insertions in B. rapa and more conserved insertions, often associated with distinct morphotypes in B. oleracea. Gene ontology enrichment indicated TE involvement in developmental, reproductive, and stress response pathways. Transcriptome analysis across diverse accessions revealed that genes proximal to TEs, particularly those regulating floral development and flowering time, exhibit increased expression variability. These findings advance our understanding of TE-mediated genome evolution in Brassica species and underscore their potential utility in breeding and genome engineering strategies for crop improvement.

DNA Transposable Elements↗

Molecular organization and evolution of mosquito genomes.

Given the importance of mosquitoes as disease vectors, relatively little is known about the molecular organization and evolution of mosquito genomes as compared to other insects such as fruit flies. The advances in recombinant DNA technology and the possibility that mosquito populations could be controlled and genetically manipulated using such technology has stimulated considerable research during the last few years in the areas of genome organization and evolution, genome mapping, endogenous transposable elements, and mapping and characterization of genes conferring susceptibility to different parasites and pathogens. This review summarizes research currently underway in our laboratory and elsewhere in these areas.

Animals↗

Gene finding with a hidden Markov model of genome structure and evolution.

MOTIVATION: A growing number of genomes are sequenced. The differences in evolutionary pattern between functional regions can thus be observed genome-wide in a whole set of organisms. The diverse evolutionary pattern of different functional regions can be exploited in the process of genomic annotation. The modelling of evolution by the existing comparative gene finders leaves room for improvement. RESULTS: A probabilistic model of both genome structure and evolution is designed. This type of model is called an Evolutionary Hidden Markov Model (EHMM), being composed of an HMM and a set of region-specific evolutionary models based on a phylogenetic tree. All parameters can be estimated by maximum likelihood, including the phylogenetic tree. It can handle any number of aligned genomes, using their phylogenetic tree to model the evolutionary correlations. The time complexity of all algorithms used for handling the model are linear in alignment length and genome number. The model is applied to the problem of gene finding. The benefit of modelling sequence evolution is demonstrated both in a range of simulations and on a set of orthologous human/mouse gene pairs. AVAILABILITY: Free availability over the Internet on www server: http://www.birc.dk/Software/evogene.

Algorithms↗

Genomic organization and evolution of the soybean SB92 satellite sequence.

Repetitive DNA sequences comprise a large percentage of plant genomes, and their characterization provides information about both species and genome evolution. We have isolated a recombinant clone containing a highly repeated DNA element (SB92) that is homologous to ca. 0.9% of the soybean genome or about 10(5) copies. This repeated sequence is tandemly arranged and is found in four or five major genomic locations. FISH analysis of metaphase chromosomes suggests that two of these locations are centromeric. We have determined the sequence of two cloned repeats and performed genomic sequencing to obtain a consensus sequence. The consensus repeat size was 92 bp and exhibited an average of 10% nucleotide substitution relative to the two cloned repeats. This high level of sequence diversity suggests an ancient origin but is inconsistent with the limited phylogenetic distribution of SB92, which is found at high copy number only in the annual soybeans. It therefore seems likely that this sequence is undergoing very rapid evolution.

Base Sequence↗

The organization and rate of evolution of wheat genomes are correlated with recombination rates along chromosome arms.

Genes detected by wheat expressed sequence tags (ESTs) were mapped into chromosome bins delineated by breakpoints of 159 overlapping deletions. These data were used to assess the organizational and evolutionary aspects of wheat genomes. Relative gene density and recombination rate increased with the relative distance of a bin from the centromere. Single-gene loci present once in the wheat genomes were found predominantly in the proximal, low-recombination regions, while multigene loci tended to be more frequent in distal, high-recombination regions. One-quarter of all gene motifs within wheat genomes were represented by two or more duplicated loci (paralogous sets). For 40 such sets, ancestral loci and loci derived from them by duplication were identified. Loci derived by duplication were most frequently located in distal, high-recombination chromosome regions whereas ancestral loci were most frequently located proximal to them. It is suggested that recombination has played a central role in the evolution of wheat genome structure and that gradients of recombination rates along chromosome arms promote more rapid rates of genome evolution in distal, high-recombination regions than in proximal, low-recombination regions.

Chromosome Mapping↗

[Structure and evolution of human sub-telomeric regions].

Recent progress in the field of human genome analysis has led to the development of new concepts in the definition of subtelomeric domains. Analysis of DNA sequences from human and yeast chromosome ends have shown that short stretches of degenerate TTAGGG are found at a distance from the telomeric repeats. These stretches define a boundary between two structurally different regions. The distal domain is characterised by numerous, short segments of interrupted homology to many other human telomeric regions and to a number of ESTs. The proximal domain shows much longer uninterrupted homology to a few chromosome ends. This domain evolved quickly within primates at least, as demonstrated by the detailed study of locus DNF92 which spread very recently in humans from 17 qter to at least ten other chromosome ends. At the different sites, presence-absence polymorphisms are observed within humans. The region remained single locus at the paralogous site in higher primates. Conversely, a human and orangutan single locus telomeric domain occupies multiple chromosome ends in chimpanzee. Balanced translocation is the likely mechanism through which the spreading occurred. Some members of the olfactory receptor gene family show a similar behaviour: multiple telomeric locations, and presence-absence polymorphism. Strikingly, the set of chromosome ends occupied by the two regions is identical, except for the two ancestral sites. Moreover, the relative frequency of detection of the region at the different sites indicates some kind of competition between the two regions. Consequently, these two regions represent major new tools to investigate recent human genome evolution and human genome diversity in different populations.

Animals↗

Molecular cytogenetic dissection of human chromosomes 3 and 21 evolution.

Chromosome painting in placental mammalians illustrates that genome evolution is marked by chromosomal synteny conservation and that the association of chromosomes 3 and 21 may be the largest widely conserved syntenic block known for mammals. We studied intrachromosomal rearrangements of the syntenic block 3/21 by using probes derived from chromosomal subregions with a resolution of up to 10-15 Mbp. We demonstrate that the rearrangements visualized by chromosome painting, mostly translocations, are only a fraction of the actual chromosomal changes that have occurred during evolution. The ancestral segment order for both primates and carnivores is still found in some species in both orders. From the ancestral primate/carnivore condition an inversion is needed to derive the pig homolog, and a fission of chromosome 21 and a pericentric inversion is needed to derive the Bornean orangutan condition. Two overlapping inversions in the chromosome 3 homolog then would lead to the chromosome form found in humans and African apes. This reconstruction of the origin of human chromosome 3 contrasts with the generally accepted scenario derived from chromosome banding in which it was proposed that only one pericentric inversion was needed. From the ancestral form for Old World primates (now found in the Bornean orangutan) a pericentric inversion and centromere shift leads to the chromosome ancestral for all Old World monkeys. Intrachromosomal rearrangements, as shown here, make up a set of potentially plentiful and informative markers that can be used for phylogenetic reconstruction and a more refined comparative mapping of the genome.

Animals↗

Strong regional heterogeneity in base composition evolution on the Drosophila X chromosome.

Fluctuations in base composition appear to be prevalent in Drosophila and mammal genome evolution, but their timescale, genomic breadth, and causes remain obscure. Here, we study base composition evolution within the X chromosomes of Drosophila melanogaster and five of its close relatives. Substitutions were inferred on six extant and two ancestral lineages for 14 near-telomeric and 9 nontelomeric genes. GC content evolution is highly variable both within the genome and within the phylogenetic tree. In the lineages leading to D. yakuba and D. orena, GC content at silent sites has increased rapidly near telomeres, but has decreased in more proximal (nontelomeric) regions. D. orena shows a 17-fold excess of GC-increasing vs. AT-increasing synonymous changes within a small (approximately 130-kb) region close to the telomeric end. Base composition changes within introns are consistent with changes in mutation patterns, but stronger GC elevation at synonymous sites suggests contributions of natural selection or biased gene conversion. The Drosophila yakuba lineage shows a less extreme elevation of GC content distributed over a wider genetic region (approximately 1.2 Mb). A lack of change in GC content for most introns within this region suggests a role of natural selection in localized base composition fluctuations.

Animals↗

Evolution of RNA genomes: does the high mutation rate necessitate high rate of evolution of viral proteins?

RNA genomes have been shown to mutate much more frequently than DNA genomes. It is generally assumed that this results in rapid evolution of RNA viral proteins. Here, an alternative hypothesis is proposed that close cooperation between positive-strand RNA viral proteins and those of the host cells required their coevolution, resulting in similar amino acid substitution rates. Constraints on compatibility with cellular proteins should determine, at any time, the covarion sets in RNA viral proteins. These ideas may be helpful in rationalizing the accumulating data on significant sequence similarities between proteins of positive-strand RNA viruses infecting evolutionarily distant hosts as well as between viral and cellular proteins.

Biological Evolution↗

An evolutionary model for the origin of non-randomness, long-range order and fractality in the genome.

We present a model for genome evolution, comprising biologically plausible events such as transpositions inside the genome and insertions of exogenous sequences. This model attempts to formulate a minimal proposition accounting for key statistical properties of genomes, avoiding, as far as possible, unsupportable hypotheses for the remote evolutionary past. The statistical properties that are observed in genomic sequences and are reproduced by the proposed model are: (i) deviations from randomness at different length scales, measured by suitable algorithms, (ii) a special form of size distribution (power law distribution) characterising different levels of genome organisation in the non-coding, and (iii) extensive resemblance in the alternation of coding and non-coding regions at several length scales (self-similarity) in long genomic sequences of higher eukaryotes.

DNA↗

Microevolutionary genomics of bacteria.

The availability of multiple complete genome sequences from the same species can facilitate attempts to systematically address basic questions in genome evolution. We refer to such efforts as "microevolutionary genomics". We report the results of comparative analyses of complete intraspecific genome (and proteome) sequences from four bacterial species--Chlamydophila pneumoniae, Escherichia coli, Helicobacter pylori and Neisseria meningitidis. Comparisons of average synonymous (K(s)) and nonsynonymous (K(a)) substitution rates were used to assess the influence of various biological factors on the rate of protein evolution. For example, E. coli experiences the most intense purifying selection of the species analyzed, and this may be due to the relatively larger population size of this species. In addition, essential genes were shown to be more evolutionarily conserved than nonessential genes in E. coli and duplicated genes have higher rates of evolution than unique genes for all species studied except C. pneumoniae. Different functional categories of genes were shown to evolve at significantly different rates emphasizing the role of category-specific functional constraints in determining evolutionary rates. Finally, functionally characterized genes tend to be conserved between strains, while uncharacterized genes are over-represented among the unique, strain-specific genes. This suggests the possibility that nonessential genes are responsible for driving the evolutionary diversification between strains.

Bacteria↗

Mutation pressure and the evolution of organelle genomic architecture.

The nuclear genomes of multicellular animals and plants contain large amounts of noncoding DNA, the disadvantages of which can be too weak to be effectively countered by selection in lineages with reduced effective population sizes. In contrast, the organelle genomes of these two lineages evolved to opposite ends of the spectrum of genomic complexity, despite similar effective population sizes. This pattern and other puzzling aspects of organelle evolution appear to be consequences of differences in organelle mutation rates. These observations provide support for the hypothesis that the fundamental features of genome evolution are largely defined by the relative power of two nonadaptive forces: random genetic drift and mutation pressure.

Animals↗

Neotelomeres and telomere-spanning chromosomal arm fusions in cancer genomes revealed by long-read sequencing.

Alterations in the structure and location of telomeres are pivotal in cancer genome evolution. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeats, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. These results provide a framework for the systematic study of telomeric repeats in cancer genomes, which could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Humans↗

Estimating the tempo and mode of gene family evolution from comparative genomic data.

Comparison of whole genomes has revealed that changes in the size of gene families among organisms is quite common. However, there are as yet no models of gene family evolution that make it possible to estimate ancestral states or to infer upon which lineages gene families have contracted or expanded. In addition, large differences in family size have generally been attributed to the effects of natural selection, without a strong statistical basis for these conclusions. Here we use a model of stochastic birth and death for gene family evolution and show that it can be efficiently applied to multispecies genome comparisons. This model takes into account the lengths of branches on phylogenetic trees, as well as duplication and deletion rates, and hence provides expectations for divergence in gene family size among lineages. The model offers both the opportunity to identify large-scale patterns in genome evolution and the ability to make stronger inferences regarding the role of natural selection in gene family expansion or contraction. We apply our method to data from the genomes of five yeast species to show its applicability.

Evolution, Molecular↗