Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Discovery of human inversion polymorphisms by comparative analysis of human and chimpanzee DNA sequence assemblies.

With a draft genome-sequence assembly for the chimpanzee available, it is now possible to perform genome-wide analyses to identify, at a submicroscopic level, structural rearrangements that have occurred between chimpanzees and humans. The goal of this study was to investigate chromosomal regions that are inverted between the chimpanzee and human genomes. Using the net alignments for the builds of the human and chimpanzee genome assemblies, we identified a total of 1,576 putative regions of inverted orientation, covering more than 154 mega-bases of DNA. The DNA segments are distributed throughout the genome and range from 23 base pairs to 62 mega-bases in length. For the 66 inversions more than 25 kilobases (kb) in length, 75% were flanked on one or both sides by (often unrelated) segmental duplications. Using PCR and fluorescence in situ hybridization we experimentally validated 23 of 27 (85%) semi-randomly chosen regions; the largest novel inversion confirmed was 4.3 mega-bases at human Chromosome 7p14. Gorilla was used as an out-group to assign ancestral status to the variants. All experimentally validated inversion regions were then assayed against a panel of human samples and three of the 23 (13%) regions were found to be polymorphic in the human genome. These polymorphic inversions include 730 kb (at 7p22), 13 kb (at 7q11), and 1 kb (at 16q24) fragments with a 5%, 30%, and 48% minor allele frequency, respectively. Our results suggest that inversions are an important source of variation in primate genome evolution. The finding of at least three novel inversion polymorphisms in humans indicates this type of structural variation may be a more common feature of our genome than previously realized.

Animals↗

Paleogenomics or the search for remnant duplicated copies of the yeast DUP240 gene family in intergenic areas.

Duplication, resulting in gene redundancy, is well known to be a driving force of evolutionary change. Gene families are therefore useful targets for approaching genome evolution. To address the gene death process, we examined the fate of the 10-member-large S288C DUP240 family in 15 Saccharomyces cerevisiae strains. Using an original three-step method of analysis reported here, both slightly and highly degenerate DUP240 copies, called pseudo-open reading frames (ORFs) and relics, respectively, were detected in strain S288C. It was concluded that two previously annotated ORFs correspond, in fact, to pseudo-ORFs and three additional relics were identified in intergenic areas. Comparative intraspecies analysis of these degenerate DUP240 loci revealed that the two pseudo-ORFs are present in a nondegenerate state in some other strains. This suggests that within a given gene family different loci are the target of the gene erasure process, which is therefore strain dependent. Besides, the variable positions observed indicate that the relic sequence may diverge faster than the flanking regions. All in all, this study shows that short conserved protein motifs provide a useful tool for detecting and accurately mapping degenerate gene remnants. The present results also highlight the strong contribution of comparative genomics for gene relic detection because the possibility of finding short conserved protein motifs in intergenic regions (IRs) largely depends on the choice of the most closely related paralog or ortholog. By mapping new genetic components in previously annotated IRs, our study constitutes a further refinement step in the crucial stage of genome annotation and provides a strategy for retracing ancient chromosomal reshaping events and, hence, for deciphering genome history.

Amino Acid Sequence↗

The evolution of chronic infection strategies in the alpha-proteobacteria.

Many of the alpha-proteobacteria establish long-term, often chronic, interactions with higher eukaryotes. These interactions range from pericellular colonization through facultative intracellular multiplication to obligate intracellular lifestyles. A common feature in this wide range of interactions is modulation of host-cell proliferation, which sometimes leads to the formation of tumour-like structures in which the bacteria can grow. Comparative genome analyses reveal genome reduction by gene loss in the intracellular alpha-proteobacterial lineages, and genome expansion by gene duplication and horizontal gene transfer in the free-living species. In this review, we discuss alpha-proteobacterial genome evolution and highlight strategies and mechanisms used by these bacteria to infect and multiply in eukaryotic cells.

Alphaproteobacteria↗

De novo Genes in Plants: Origins, Mechanisms, and Functional Implications.

De novo genes originate from previously non-coding genomic regions. They provide an important source of lineage-specific innovation. In plants, these genes may contribute to adaptation, trait diversity and crop evolution. This review summarizes recent progress in plant de novo gene research. It first discusses major routes of gene birth, including transcription-first, open reading frame (ORF)-first and concurrent models. It also examines how nascent loci acquire regulatory control and enter existing biological networks. The review then summarizes their evolutionary features, including weak early constraint, rapid molecular change, restricted expression and structural refinement. It further discusses plant de novo genes involved in stress responses, seed germination, kernel dehydration, subspecies divergence, reproductive isolation and floral scent diversification. Current methods for identifying de novo genes remain limited by rapid sequence evolution, genome annotation quality, polyploidy and transposable elements. Whole-genome synteny alignment, multi-omics evidence and machine-learning approaches can improve candidate discovery. However, each method has important limitations. Finally, this review highlights key future questions in functional validation, latent coding potential in long non-coding RNAs, epigenetic activation, regulatory-network integration and crop improvement. These perspectives clarify how de novo genes shape plant adaptation and how they may be used in precision breeding and synthetic biology.

adaptive evolution↗

Comparative genome analyses of Arabidopsis spp.: inferring chromosomal rearrangement events in the evolutionary history of A. thaliana.

Comparative genome analysis is a powerful tool that can facilitate the reconstruction of the evolutionary history of the genomes of modern-day species. The model plant Arabidopsis thaliana with its n = 5 genome is thought to be derived from an ancestral n = 8 genome. Pairwise comparative genome analyses of A. thaliana with polyploid and diploid Brassicaceae species have suggested that rapid genome evolution, manifested by chromosomal rearrangements and duplications, characterizes the polyploid, but not the diploid, lineages of this family. In this study, we constructed a low-density genetic linkage map of Arabidopsis lyrata ssp. lyrata (A. l. lyrata; n = 8, diploid), the closest known relative of A. thaliana (MRCA approximately 5 Mya), using A. thaliana-specific markers that resolve into the expected eight linkage groups. We then performed comparative Bayesian analyses using raw mapping data from this study and from a Capsella study to infer the number and nature of rearrangements that distinguish the n = 8 genomes of A. l. lyrata and Capsella from the n = 5 genome of A. thaliana. We conclude that there is strong statistical support in favor of the parsimony scenarios of 10 major chromosomal rearrangements separating these n = 8 genomes from A. thaliana. These chromosomal rearrangement events contribute to a rate of chromosomal evolution higher than previously reported in this lineage. We infer that at least seven of these events, common to both sets of data, are responsible for the change in karyotype and underlie genome reduction in A. thaliana.

Arabidopsis↗

Repetitive extragenic palindromic sequences in the Pseudomonas syringae pv. tomato DC3000 genome: extragenic signals for genome reannotation.

Repetitive extragenic palindromic (REPs) sequences were first described in enterobacteriacea and later in Pseudomonas putida. We have detected a new variant (51 base pairs) of REP sequences that appears to be disseminated in more than 300 copies in the Pseudomonas syringae DC3000 genome. The finding of REP sequences in P. syringae confirms the broad presence of this type of repetitive sequence in bacteria. We analyzed the distribution of REP sequences and the structure of the clusters, and we show that palindromy is conserved. REP sequences appear to be allocated to the extragenic space, with a special preference for the intergenic spaces limited by convergent genes, while their presence is scarce between divergent genes. Using REP sequences as markers of extragenicity we re-annotated a set of genes of the P. syringae DC3000 genome demonstrating that REP sequences can be used for refinement of annotation of a genome. The similarity detected between virulence genes from evolutionarily distant pathogenic bacteria suggests the acquisition of clusters of virulence genes by horizontal gene transfer. We did not detect the presence of P. syringae REP elements in the principal pathogenicity gene clusters. This absence suggests that genome fragments lacking REP sequences could point to regions recently acquired from other organisms, and REP sequences might be new tracers for gaining insight into key aspects of bacterial genome evolution, especially when studying pathogenicity acquisition. In addition, as the P. syringae REP sequence is species-specific with respect to the sequenced genomes, it is an exceptional candidate for use as a fingerprint in precise genotyping and epidemiological studies.

Base Sequence↗

Rapid elimination of low-copy DNA sequences in polyploid wheat: a possible mechanism for differentiation of homoeologous chromosomes.

To study genome evolution in allopolyploid plants, we analyzed polyploid wheats and their diploid progenitors for the occurrence of 16 low-copy chromosome- or genome-specific sequences isolated from hexaploid wheat. Based on their occurrence in the diploid species, we classified the sequences into two groups: group I, found in only one of the three diploid progenitors of hexaploid wheat, and group II, found in all three diploid progenitors. The absence of group II sequences from one genome of tetraploid wheat and from two genomes of hexaploid wheat indicates their specific elimination from these genomes at the polyploid level. Analysis of a newly synthesized amphiploid, having a genomic constitution analogous to that of hexaploid wheat, revealed a pattern of sequence elimination similar to the one found in hexaploid wheat. Apparently, speciation through allopolyploidy is accompanied by a rapid, nonrandom elimination of specific, low-copy, probably noncoding DNA sequences at the early stages of allopolyploidization, resulting in further divergence of homoeologous chromosomes (partially homologous chromosomes of different genomes carrying the same order of gene loci). We suggest that such genomic changes may provide the physical basis for the diploid-like meiotic behavior of polyploid wheat.

Chromosomes↗

Accurate detection of tandem repeats exposes ubiquitous reuse of biological sequences.

Tandem repetition is one of the major processes underlying genome evolution and phenotypic diversification. While newly formed tandem repeats are often easy to identify, it is more challenging to detect repeat copies as they diverge over evolutionary timescales. Existing programs for finding tandem repeats return markedly different results, and it is unclear which predictions are more correct and how much room remains for improvement. Here, we introduce DetectRepeats, a new method that uses empirical information about structural repeats to improve the accuracy of repeat detection. We show that DetectRepeats advances the state-of-the-art by finding highly divergent repeats with relatively few false positive detections. We apply DetectRepeats to genomes across the tree of life to discover an enrichment of detectable tandem repeats within different genes, genome regions, and taxa. Furthermore, we use phylogenetic reconciliation to determine that some tandem repeats continue to evolve through intra-repeat unit replacement. In this manner, tandem repeats serve as a renewable genetic resource offering a bountiful source of alternative genetic material. Our work unlocks the confident detection of ancient tandem repeats, opening a doorway to future discoveries. DetectRepeats is part of the DECIPHER package for the R programming language and available via Bioconductor.

Tandem Repeat Sequences↗

Reconstruction of putative DNA virus from endogenous rice tungro bacilliform virus-like sequences in the rice genome: implications for integration and evolution.

BACKGROUND: Plant genomes contain various kinds of repetitive sequences such as transposable elements, microsatellites, tandem repeats and virus-like sequences. Most of them, with the exception of virus-like sequences, do not allow us to trace their origins nor to follow the process of their integration into the host genome. Recent discoveries of virus-like sequences in plant genomes led us to set the objective of elucidating the origin of the repetitive sequences. Endogenous rice tungro bacilliform virus (RTBV)-like sequences (ERTBVs) have been found throughout the rice genome. Here, we reconstructed putative virus structures from RTBV-like sequences in the rice genome and characterized to understand evolutionary implication, integration manner and involvements of endogenous virus segments in the corresponding disease response. RESULTS: We have collected ERTBVs from the rice genomes. They contain rearranged structures and no intact ORFs. The identified ERTBV segments were shown to be phylogenetically divided into three clusters. For each phylogenetic cluster, we were able to make a consensus alignment for a circular virus-like structure carrying two complete ORFs. Comparisons of DNA and amino acid sequences suggested the closely relationship between ERTBV and RTBV. The Oryza AA-genome species vary in the ERTBV copy number. The species carrying low-copy-number of ERTBV segments have been reported to be extremely susceptible to RTBV. The DNA methylation state of the ERTBV sequences was correlated with their copy number in the genome. CONCLUSIONS: These ERTBV segments are unlikely to have functional potential as a virus. However, these sequences facilitate to establish putative virus that provided information underlying virus integration and evolutionary relationship with existing virus. Comparison of ERTBV among the Oryza AA-genome species allowed us to speculate a possible role of endogenous virus segments against its related disease.

Amino Acid Sequence↗

Intron length evolution in Drosophila.

I present data on the evolution of intron lengths among 3 closely related Drosophila species, D. melanogaster, Drosophila simulans, and Drosophila yakuba. Using D. yakuba as an outgroup, I mapped insertion and deletion mutations in 148 introns (spanning approximately 30 kb) to the D. melanogaster and D. simulans lineages. Intron length evolution in the 2 sister species has been different: in D. melanogaster, X-linked introns have increased slightly in size, whereas autosomal ones have decreased slightly in size; in D. simulans, both X-linked and autosomal introns have decreased in size. To understand the possible evolutionary causes of these lineage- and chromosome-specific patterns of intron evolution, I studied insertion-deletion (indel) polymorphism and divergence in D. melanogaster. Small insertion mutations segregate at elevated frequencies and enjoy elevated probabilities of fixation, particularly on the X chromosome. In contrast, there is no detectable X chromosome effect on fixations in D. simulans. These findings suggest X chromosome-specific selection or biased gene conversion-gap repair favoring insertions in D. melanogaster but not in D. simulans. These chromosome- and lineage-specific patterns of indel substitution are not easily explained by existing general population genetic models of intron length evolution. Genomic data from D. melanogaster further suggest that the forces described here affect introns and intergenic regions similarly.

Animals↗

Use of long sequence alignments to study the evolution and regulation of mammalian globin gene clusters.

The determination of long segments of DNA sequences encompassing the beta- and alpha-globin gene clusters has provided an unprecedented data base for analysis of genome evolution and regulation of gene clusters. A newly developed computer tool kit generates local alignments between such long sequences in a space-efficient manner, helps the user analyze the alignments effectively, and finds consistently aligning blocks of sequences in multiple pairwise comparisons. Such sequence analyses among the beta-like globin gene clusters of human, galago, rabbit, and mouse have revealed the general patterns of evolution of this gene cluster. Alignments in the flanking regions are very useful in assigning orthologous relationships. Investigation of such matches between the mouse and human beta-like globin gene clusters has led to a reassessment of some orthologous assignments in mouse and to a revision of the proposed pathway for evolution of this gene cluster. In general, the interspersed repetitive elements have inserted independently, presumably via a retrotransposition mechanism, in the different mammalian lineages. However, some examples of ancient L1 repeats are found, including one between the epsilon- and gamma-globin genes that appears to have been in the ancestral eutherian gene cluster. Prominent matching sequences are found in a long region 5' to the epsilon-globin gene, the locus control region (LCR) that is a positive regulator of the entire gene cluster. Three-way alignments among the human, goat, and rabbit sequences can extend for > or = 3 kb in part of the LCR (DNase hypersensitive site 3), indicating that the cis-acting components of this complex regulatory region cover a long segment of DNA. In contrast to the beta-like globin gene clusters, the alpha-like globin gene clusters of many mammals occur in very G+C-rich isochores and contain prominent CpG islands. The regions between the alpha-like globin genes are evolving faster than the intergenic regions of the beta-like globin gene clusters. The contrasts between the two gene clusters can be attributed to differences in DNA metabolism in the isochore. The proximal control elements of the rabbit alpha-globin gene are located both 5' to and within the gene. All of this region is part of a prominent CpG island that may be acting as an extended, enhancer-independent promoter. One can hypothesize that the analogue to the LCR in the alpha-globin gene cluster may interface with the distinctive alpha-globin promoter in ways different from the interaction between the beta LCR and the promoters of beta-like globin genes.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals↗

Episode clustering in phylogenetic networks.

MOTIVATION: The classical duplication episode clustering (EC) model introduced by Guigó et al. in the 1990s provides a foundational approach for inferring genomic duplication events crucial to understanding genome evolution. This model clusters single gene duplications from a collection of gene trees at locations in the species tree to minimize the total number of such locations, called duplication episodes. However, it does not capture reticulate evolutionary histories. RESULTS: Here, we introduce NetEC, a novel extension of this problem to phylogenetic networks. To solve NetEC, we first develop a polynomial-time dynamic programming (DP) algorithm for testing whether a given set of network nodes can serve as episode locations. We then propose a main inference algorithm that utilizes this DP component to optimize the episode count; while the feasibility test runs in polynomial time, the full optimization has exponential worst-case complexity, and an optional heuristic mode is provided for larger instances. We also propose an extended episode analysis procedure that identifies additional genomic duplication candidates below reticulation nodes, complementing the main algorithm by resolving potential upward clustering of duplications induced by reticulation. We evaluate our method on simulated data and on an empirical Pandanales dataset comprising over 29 000 gene trees, demonstrating exact and accurate inference of genomic duplication events even in the presence of multiple reticulations. AVAILABILITY AND IMPLEMENTATION: All experiments were conducted using the NetEC tool (https://github.com/ppgorecki/netec), with all input data, scripts, and parameter settings for reproduction available in the same repository.

Phylogeny↗

A new family of chimeric retrotranscripts formed by a full copy of U6 small nuclear RNA fused to the 3' terminus of l1.

Long interspersed nuclear elements (LINE-1, L1) constitute a large family of mammalian retrotransposons that have been replicating and evolving in mammals for more than 100 million years and now compose 17% of the human genome. They have an important creative role in human genomic evolution through mechanisms such as new integrations, generation of processed pseudogenes, and transfer of non-L1 DNA flanking their 3' ends to new genomic locations. Here we present evidence that the L1 integration machinery was used for the creation of a new family of chimeric retrotranscripts, which contain a full copy of U6 small nuclear RNA and a 3' part of L1 at their 5' and 3' ends, respectively. There are at least 56 members of this family in the human genome. The integrations of such fused retrotranscripts into the human genome took place until recently. Here we report one U6-L1 insertion that is polymorphic in humans. We also propose a mechanism used to generate chimeric retrotranscripts.

Humans↗

Genomic analysis in the sting-2 quantitative trait locus for defensive behavior in the honey bee, Apis mellifera.

We have sequenced an 81-kb genomic region from the honey bee, Apis mellifera, associated with a quantitative trait locus (QTL) sting-2 for aggressive behavior. This sequence represents the first extensive study of the honey-bee genome structure encompassing putative genes in a QTL for a behavioral trait. Expression of 13 putative genes, as well as two transcripts that were present in a honey-bee EST database, was confirmed through reverse transcription analysis of mRNA from the honey-bee head. Whereas most transcripts exhibited little or no variation between European and Africanized honey-bee alleles, one transcript demonstrated significant nonsynonymous substitutions, deletions, and insertions. All 13 putative genes lacked similarity to known invertebrate or vertebrate proteins or transcripts. This observation may be reflective of the processes that determine the genomic evolution of an insect with social behavior and/or haplo-diploidy and are an indication of the unique nature of the honey-bee genome. These results make this sequence an invaluable research tool for the ongoing honey-bee whole-genome sequencing effort.

Animals↗

The Arabidopsis genome: a foundation for plant research.

The sequence of the first plant genome was completed and published at the end of 2000. This spawned a series of large-scale projects aimed at discovering the functions of the 25,000+ genes identified in Arabidopsis thaliana (Arabidopsis). This review summarizes progress made in the past five years and speculates about future developments in Arabidopsis research and its implications for crop science. The provision of large populations of gene disruption lines to the research community has greatly accelerated the impact of genomics on many areas of plant science. The tools and community organization required for plant integrative and systems biology approaches are now ready to accomplish the next big step in plant biology--the integration of knowledge and modeling of biological processes. In the future, plant science will continue to be enriched by the alignment of high-quality basic research (generally conducted in Arabidopsis), with strategic objectives in crop plants. The sequence and analysis of an increasing number of crop plant genomes enhance this alignment and provide new insights into genome evolution and crop plant domestication.

Arabidopsis↗

A novel function for spumaretrovirus integrase: an early requirement for integrase-mediated cleavage of 2 LTR circles.

Retroviral integration is central to viral persistence and pathogenesis, cancer as well as host genome evolution. However, it is unclear why integration appears essential for retrovirus production, especially given the abundance and transcriptional potential of non-integrated viral genomes. The involvement of retroviral endonuclease, also called integrase (IN), in replication steps apart from integration has been proposed, but is usually considered to be accessory. We observe here that integration of a retrovirus from the spumavirus family depends mainly on the quantity of viral DNA produced. Moreover, we found that IN directly participates to linear DNA production from 2-LTR circles by specifically cleaving the conserved palindromic sequence found at LTR-LTR junctions. These results challenge the prevailing view that integrase essential function is to catalyze retroviral DNA integration. Integrase activity upstream of this step, by controlling linear DNA production, is sufficient to explain the absolute requirement for this enzyme. The novel role of IN over 2-LTR circle junctions accounts for the pleiotropic effects observed in cells infected with IN mutants. It may explain why 1) 2-LTR circles accumulate in vivo in mutants carrying a defective IN while their linear and integrated DNA pools decrease; 2) why both LTRs are processed in a concerted manner. It also resolves the original puzzle concerning the integration of spumaretroviruses. More generally, it suggests to reassess 2-LTR circles as functional intermediates in the retrovirus cycle and to reconsider the idea that formation of the integrated provirus is an essential step of retrovirus production.

Animals↗

REvolutionH-tl 2.0: A fast and robust tool for decoding evolutionary gene histories.

REvolutionH-tl is a fast, scalable, and integrated software platform for inferring orthology relationships, gene trees, species trees, and reconciled evolutionary scenarios directly from sequence data. Built upon the formal framework of best match graphs (BMGs), REvolutionH-tl predicts orthogroups and orthologous gene pairs with high accuracy, requiring neither precomputed trees nor multiple external tools. The software reconstructs event-labeled gene and species trees, seamlessly integrating reconciliation to produce fast, accurate, and biologically insightful evolutionary scenarios. Through extensive benchmarking on synthetic datasets with known ground truth, REvolutionH-tl outperforms or matches the accuracy of established tools such as OrthoFinder, Proteinortho, RAxML, GeneRax, and RANGER-DTL, while achieving significantly lower runtimes. A key innovation of REvolutionH-tl is its built-in support for detailed, publication-ready visualizations, which allow users to explore genome evolution dynamics, orthogroup composition, and reconciliation results with clarity and ease. These visual features position REvolutionH-tl as the first platform of its kind to combine analytical precision with intuitive interpretability. The software is open-source, cross-platform, and freely available at https://pypi.org/project/revolutionhtl/, providing a robust solution for large-scale evolutionary analyses in comparative genomics.

Software↗

Protein evolution and codon usage bias on the neo-sex chromosomes of Drosophila miranda.

The neo-sex chromosomes of Drosophila miranda constitute an ideal system to study the effects of recombination on patterns of genome evolution. Due to a fusion of an autosome with the Y chromosome, one homolog is transmitted clonally. Here, I compare patterns of molecular evolution of 18 protein-coding genes located on the recombining neo-X and their homologs on the nonrecombining neo-Y chromosome. The rate of protein evolution has significantly increased on the neo-Y lineage since its formation. Amino acid substitutions are accumulating uniformly among neo-Y-linked genes, as expected if all loci on the neo-Y chromosome suffer from a reduced effectiveness of natural selection. In contrast, there is significant heterogeneity in the rate of protein evolution among neo-X-linked genes, with most loci being under strong purifying selection and two genes showing evidence for adaptive evolution. This observation agrees with theory predicting that linkage limits adaptive protein evolution. Both the neo-X and the neo-Y chromosome show an excess of unpreferred codon substitutions over preferred ones and no difference in this pattern was observed between the chromosomes. This suggests that there has been little or no selection maintaining codon bias in the D. miranda lineage. A change in mutational bias toward AT substitutions also contributes to the decline in codon bias. The contrast in patterns of molecular evolution between amino acid mutations and synonymous mutations on the neo-sex-linked genes can be understood in terms of chromosome-specific differences in effective population size and the distribution of selective effects of mutations.

Animals↗