Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Genomic characterization of recent human LINE-1 insertions: evidence supporting random insertion.

LINE-1 (L1) elements play an important creative role in genomic evolution by distributing both L1 and non-L1 DNA in a process called retrotransposition. A large percentage of the human genome consists of DNA that has been dispersed by the L1 transposition machinery. L1 elements are not randomly distributed in genomic DNA but are concentrated in regions with lower GC content. In an effort to understand the consequences of L1 insertions, we have begun an investigation of their genomic characteristics and the changes that occur to them over time. We compare human L1 insertions that were created either during recent human evolution or during the primate radiation. We report that L1 insertions are an important source for the creation of new microsatellites. We provide evidence that L1 first strand cDNA synthesis can occur from an internal priming event. We note that in contrast to older L1 insertions, recent L1s are distributed randomly in genomic DNA, and the shift in the L1 genomic distribution occurs relatively rapidly. Taken together, our data indicate that strong forces act on newly inserted L1 retrotransposons to alter their structure and distribution.

3' Flanking Region↗

Defining APOBEC-induced mutation signatures and modifying activities in yeast.

APOBEC cytidine deaminases guard cells in a variety of organisms from invading viruses and foreign nucleic acids. Recently, several human APOBECs have been implicated in mutating evolving cancer genomes. Expression of APOBEC3A and APOBEC3B in yeast allowed experimental derivation of the substitution patterns they cause in dividing cells, which provided critical links to these enzymes in the etiology of the COSMIC single base substitution (SBS) signatures 2 and 13 in human tumors. Additionally, the ability to scale yeast experiments to high-throughput screens allows use of this system to also investigate cellular pathways impacting the frequency of APOBEC-induced mutation. Here, we present validated methods utilizing yeast to determine APOBEC mutation signatures, genetic interactors, and chromosomal substrate preferences. These methods can be employed to assess the potential of other human APOBECs and APOBEC orthologs in different species to contribute to cancer genome evolution as well as define the pathways that protect the nuclear genome from inadvertent APOBEC activity during viral restriction.

Humans↗

Sequence analysis of an amphioxus cosmid containing a gene homologous to members of the aldo-keto reductase gene superfamily.

To gain an insight into vertebrate genome evolution, we have analysed the organization of an approximately 40-kb genomic clone of an amphioxus (Branchiostoma floridae) cosmid library. Amphioxus is considered as being the last non-vertebrate relative to vertebrates. Sequencing and analysis of the above clone using three different exon prediction programs (Grail, GenScan, Mzef) have led to the identification of a gene of the aldo-keto reductase family as well as further exons that gave a significant database match to known genes.

Alcohol Oxidoreductases↗

Genomes of Conopholis americana and Epifagus virginiana: two holoparasitic plants (Orobanchaceae).

Conopholis americana (American cancer-root) and Epifagus virginiana (beechdrops) are sister genera of holoparasitic plants (Orobanchaceae) native to eastern North America, parasitizing oaks and American beech, respectively. Both have served as models for plastid genome reduction, yet no nuclear genomes exist for either genus or any New World holoparasitic Orobanchaceae. Here we present the first nuclear genome assemblies for both species using PacBio HiFi sequencing. The C. americana assembly totals 1.82 Gb and E. virginiana totals 440 Mb, representing an approximately 4-fold difference in genome size between these sister genera. We observed a BUSCO completeness of 79% to 80% in both species, which is typical of holoparasites. While gene prediction identified 33,889 genes in C. americana and 21,031 in E. virginiana, repeat annotation revealed that LTR retrotransposons account for 78% of the genome size difference. These assemblies reveal contrasting mechanisms of genome evolution in sister holoparasitic genera and provide foundational resources for comparative genomics of parasitic plants.

Genome, Plant↗

Formation of solo-LTRs through unequal homologous recombination counterbalances amplifications of LTR retrotransposons in rice Oryza sativa L.

We studied the dynamics of hopi, Retrosat1, and RIRE3, three gypsy-like long terminal repeat (LTR) retrotransposons, in Oryza sativa L. genome. For each family, we assessed the phenetic relationships of the copies and estimated the date of insertion of the complete copies through the evaluation of their LTR divergence. We show that within each family, distinct phenetic groups have inserted at significantly different times, within the past 5 Myr and that two major amplification events may have occurred during this period. We show that solo-LTR formation through homologous unequal recombination has occurred in rice within the past 5 Myr for the three elements. We thus propose an increase/decrease model for rice genome evolution, in which both amplification and recombination processes drive variations in genome size.

Animals↗

Isolation and characterization of RNase LTR sequences of Ty1-copia retrotransposons in common bean (Phaseolus vulgaris L).

Retroelements have proved useful for molecular marker studies and play an important role in genome evolution. Ty1-copia retrotransposons are ubiquitous and heterogeneous in plant genomes, and although many elements have been isolated and characterized, almost no information about them is available in the literature for Phaseolus vulgaris L. We report here the isolation and characterization of new RNase long terminal repeat (LTR) sections of the Ty1-copia group for this crop plant. RNAse sections showed conserved amino acids with the downstream sections corresponding to the polypurine-tract and 5' sections of 3' LTRs. The RNase sections were aligned using ClustalX to find potential relationships between sequences. A comparison with this analysis was made using the partition analysis of quasispecies package (PAQ), which is specific for quasispecies-like populations. The analysis revealed eight distinct groups. To uncover LTR variability and potential conserved promoter motifs, we also designed new primers from the presumed polypurine-tract regions. A similarity search found short stretches similar to upstream and downstream regions of some genes. Conserved motifs, corresponding to transcription factor binding sites, were discovered through MatInspector software and two sequences characterized. From a putative LTR fragment, we then designed a new primer, which, through sequence-specific amplification polymorphism (SSAP), showed numerous polymorphic bands between two distinct P. vulgaris accessions.

Base Sequence↗

Diagnosing duplications--can it be done?

New genes arise through duplication and modification of DNA sequences on a range of scales: single gene duplication, duplication of large chromosomal fragments and whole-genome duplication. Each duplication mechanism has specific characteristics that influence the fate of the resulting duplicates, such as the size of the duplicated fragment, the potential for dosage imbalance, the preservation or disruption of regulatory control and genomic context. The ability to diagnose or identify the mechanism that produced a pair of paralogs has the potential to increase our ability to reconstruct evolutionary history, to understand the processes that govern genome evolution and to make functional predictions based on paralogy. The recent availability of large amounts of whole-genome sequence, often from several closely related species, has stimulated a wealth of new computational methods to diagnose gene duplications.

Animals↗

An integrative method for accurate comparative genome mapping.

We present MAGIC, an integrative and accurate method for comparative genome mapping. Our method consists of two phases: preprocessing for identifying "maximal similar segments," and mapping for clustering and classifying these segments. MAGIC's main novelty lies in its biologically intuitive clustering approach, which aims towards both calculating reorder-free segments and identifying orthologous segments. In the process, MAGIC efficiently handles ambiguities resulting from duplications that occurred before the speciation of the considered organisms from their most recent common ancestor. We demonstrate both MAGIC's robustness and scalability: the former is asserted with respect to its initial input and with respect to its parameters' values. The latter is asserted by applying MAGIC to distantly related organisms and to large genomes. We compare MAGIC to other comparative mapping methods and provide detailed analysis of the differences between them. Our improvements allow a comprehensive study of the diversity of genetic repertoires resulting from large-scale mutations, such as indels and duplications, including explicitly transposable and phagic elements. The strength of our method is demonstrated by detailed statistics computed for each type of these large-scale mutations. MAGIC enabled us to conduct a comprehensive analysis of the different forces shaping prokaryotic genomes from different clades, and to quantify the importance of novel gene content introduced by horizontal gene transfer relative to gene duplication in bacterial genome evolution. We use these results to investigate the breakpoint distribution in several prokaryotic genomes.

Algorithms↗

The large genome constraint hypothesis: evolution, ecology and phenotype.

BACKGROUND AND AIMS: If large genomes are truly saturated with unnecessary 'junk' DNA, it would seem natural that there would be costs associated ith accumulation and replication of this excess DNA. Here we examine the available evidence to support this hypothesis, which we term the 'large genome constraint'. We examine the large genome constraint at three scales: evolution, ecology, and the plant phenotype. SCOPE: In evolution, we tested the hypothesis that plant lineages with large genomes are diversifying more slowly. We found that genera with large genomes are less likely to be highly specious -- suggesting a large genome constraint on speciation. In ecology, we found that species with large genomes are under-represented in extreme environments -- again suggesting a large genome constraint for the distribution and abundance of species. Ultimately, if these ecological and evolutionary constraints are real, the genome size effect must be expressed in the phenotype and confer selective disadvantages. Therefore, in phenotype, we review data on the physiological correlates of genome size, and present new analyses involving maximum photosynthetic rate and specific leaf area. Most notably, we found that species with large genomes have reduced maximum photosynthetic rates - again suggesting a large genome constraint on plant performance. Finally, we discuss whether these phenotypic correlations may help explain why species with large genomes are trimmed from the evolutionary tree and have restricted ecological distributions. CONCLUSION: Our review tentatively supports the large genome constraint hypothesis.

Biological Evolution↗

Chance favors the prepared genome.

Most descriptions of mutation have emphasized its negative consequences, and randomness with respect to biological function. This book seeks to balance the discussion by emphasizing mechanisms that both diversify the genome and increase the probability that a genome's descendants will survive. This chapter provides a framework for, and overview of, the diverse contributions to this book; these contributions will be stimulating companions, well into the 21st Century, as we work to comprehend the information contained in genomic databases. Genomes that encode "better" amino acid sequences are at a selective advantage. Genomes that generate diversity also are at an advantage to the extent that they can navigate efficiently through the space of possible sequence changes. Biochemical systems that tend to increase the ratio of useful to destructive genetic change may harness preexisting information (horizontal gene transfer, DNA translocation and/or DNA duplication), focus the location, timing, and extent of genetic change, adjust the dynamic range of a gene's activity, and/or sample regulatory connections between sites distributed across the genome. Rejecting entirely random genetic variation as the substrate of genome evolution is not a refutation, but rather provides a deeper understanding, of the theory of natural selection of Darwin and Wallace. The fittest molecular strategies survive, along with descendants of the genomes that encode them.

Animals↗

Organization and structural evolution of four multigene families in Arabidopsis thaliana: AtLCAD, AtLGT, AtMYST and AtHD-GL2.

The Arabidopsis Genome Initiative has released up to now more than 80% of the genome sequence of Arabidopsis thaliana. About 70% of the identified genes have at least one paralogue. In order to understand the biological function of individual genes, it is essential to study the structure, expression and organization of the entire multigene family. A systematic analysis of multigene families, made possible by the amount of genomic sequence data available, provides important clues for the understanding of genome evolution and plasticity. In this paper, four multigene families of A. thaliana are characterized, namely LCAD, HD-GL2, LGT and MYST. Members of HD-GL2 and LCAD have already been reported in plants. The LGT genes specify proteins containing motifs of glycosyl transferase. No plant genes similar to the LGT genes have been reported to date. The novel MYST family, most likely plant-specific, encodes proteins with no identified function. Sequencing and in silico analysis led to the characterization of 29 novel genes belonging to these four gene families. The organization, structure and evolution of all the members of the four families are discussed, as well as their chromosome location. Expression data of some of the paralogues of each family are also presented.

Alcohol Oxidoreductases↗

X-linked genes evolve higher codon bias in Drosophila and Caenorhabditis.

Comparing patterns of molecular evolution between autosomes and sex chromosomes (such as X and W chromosomes) can provide insight into the forces underlying genome evolution. Here we investigate patterns of codon bias evolution on the X chromosome and autosomes in Drosophila and Caenorhabditis. We demonstrate that X-linked genes have significantly higher codon bias compared to autosomal genes in both Drosophila and Caenorhabditis. Furthermore, genes that become X-linked evolve higher codon bias gradually, over tens of millions of years. We provide several lines of evidence that this elevation in codon bias is due exclusively to their chromosomal location and not to any other property of X-linked genes. We present two possible explanations for these observations. One possibility is that natural selection is more efficient on the X chromosome due to effective haploidy of the X chromosomes in males and persistently low effective numbers of reproducing males compared to that of females. Alternatively, X-linked genes might experience stronger natural selection for higher codon bias as a result of maladaptive reduction of their dosage engendered by the loss of the Y-linked homologs.

Animals↗

The diversity of retrotransposons in the yeast Cryptococcus neoformans.

We have undertaken an analysis of the retrotransposons in the medically important basidiomycetous fungus Cryptococcus neoformans. Using the data generated by a C. neoformans genome sequencing project at the Stanford Genome Technology Center, 15 distinct families of LTR retrotransposons and several families of non-LTR retrotransposons were identified. Members of at least seven families have transposed recently and are probably still active. For several families, only partial elements could be identified and these are quite diverse in sequence, suggesting that they are ancient components of the C. neoformans genome. Most C. neoformans elements are not closely related to previously identified fungal retrotransposons, suggesting that the diversity of fungal retrotransposons has been only sparsely sampled to date. C. neoformans has fewer distinct retrotransposon families than Candida albicans (37 or more), in particular fewer families represented solely by ancient and inactive elements, but it has considerably more families than either Saccharomyces cerevisiae (five) or Schizosaccharomyces pombe (two). The findings suggest that elimination of retrotransposons is faster in C. neoformans than in C. albicans, but perhaps not as rapid as in S. cerevisiae or Sz. pombe. The identification of the retrotransposons of C. neoformans should assist in the molecular characterization of this important pathogen, and also further our understanding of the role played by retroelements in genome evolution.

Amino Acid Sequence↗

Genomic changes following host restriction in bacteria.

Many genomic sequences have been recently published for bacteria that can replicate only within eukaryotic hosts. Comparisons of genomic features with those of closely related bacteria retaining free-living stages indicate that rapid evolutionary change often occurs immediately after host restriction. Typical changes include a large increase in the frequency of mobile elements in the genome, chromosomal rearrangements mediated by recombination among these elements, pseudogene formation, and deletions of varying size. In anciently host-restricted lineages, the frequency of insertion sequence elements decreases as genomes become extremely small and strictly clonal. These changes represent a general syndrome of genome evolution, which is observed repeatedly in host-restricted lineages from numerous phylogenetic groups. Considerable variation also exists, however, in part reflecting unstudied aspects of the population structure and ecology of host-restricted bacterial lineages.

Bacteria↗

Modularity in the gain and loss of genes: applications for function prediction.

Genes that are clustered on multiple genomes and are likely to functionally interact tend to be gained or lost together during genome evolution. Here, we demonstrate that exceptions to this pattern indicate relatively distant functional interactions between the encoded proteins. Hence, this can be used to divide predicted clusters of functionally interacting proteins into sub-clusters, and as such, to refine the prediction of their function and functional interactions.

Bacterial Proteins↗