Search PubMed⌕ Search

Biomedical subjects

Yves Van de Peer

Publications and source records attributed to Yves Van de Peer.

At least 19 recordsLinked to original sources

An improved chromosome-level genome resource for the sheep scab mite, Psoroptes ovis.

Sheep scab, caused by infestation with the ectoparasitic mite, Psoroptes ovis, represents a major welfare and economic challenge for the livestock industry. We present an updated 62.7 Mb genome assembly containing 10 chromosome-level scaffolds with 10,516 annotated protein-coding genes, which provides a useful resource for studies into resistance and control.

Psoroptes ovis↗

doubletrouble: an R/Bioconductor package for the identification, classification, and analysis of gene and genome duplications.

SUMMARY: Gene and genome duplications are major evolutionary forces that shape the diversity and complexity of life. However, different duplication modes have distinct impacts on gene function, expression, and regulation. Existing tools for identifying and classifying duplicated genes are either outdated or not user-friendly. Here, we present doubletrouble, an R/Bioconductor package that provides a comprehensive and robust framework for analyzing duplicated genes from genomic data. doubletrouble can detect and classify gene pairs as derived from six duplication modes (segmental, tandem, proximal, retrotransposon-derived, DNA transposon-derived, and dispersed duplications), calculate substitution rates, detect signatures of putative whole-genome duplication events, and visualize results as publication-ready figures. We applied doubletrouble to classify the duplicated gene repertoire in 822 eukaryotic genomes, and results were made available through a user-friendly web interface. AVAILABILITY AND IMPLEMENTATION: doubletrouble is available on Bioconductor (https://bioconductor.org/packages/doubletrouble), and the source code is available in a GitHub repository (https://github.com/almeidasilvaf/doubletrouble). doubletroubledb is available online at https://almeidasilvaf.github.io/doubletroubledb/.

Software↗

Interspecific transfer of genetic information through polyploid bridges.

Hybridization blurs species boundaries and leads to intertwined lineages resulting in reticulate evolution. Polyploidy, the outcome of whole genome duplication (WGD), has more recently been implicated in promoting and facilitating hybridization between polyploid species, potentially leading to adaptive introgression. However, because polyploid lineages are usually ephemeral states in the evolutionary history of life it is unclear whether WGD-potentiated hybridization has any appreciable effect on their diploid counterparts. Here, we develop a model of cytotype dynamics within mixed-ploidy populations to demonstrate that polyploidy can in fact serve as a bridge for gene flow between diploid lineages, where introgression is fully or partially hampered by the species barrier. Polyploid bridges emerge in the presence of triploid organisms, which despite critically low levels of fitness, can still allow the transfer of alleles between diploid states of independently evolving mixed-ploidy species. Notably, while marked genetic divergence prevents polyploid-mediated interspecific gene flow, we show that increased recombination rates can offset these evolutionary constraints, allowing a more efficient sorting of alleles at higher-ploidy levels before introgression into diploid gene pools. Additionally, we derive an analytical approximation for the rate of gene flow at the tetraploid level necessary to supersede introgression between diploids with nonzero introgression rates, which is especially relevant for plant species complexes, where interspecific gene flow is ubiquitous. Altogether, our results illustrate the potential impact of polyploid bridges on the (re)distribution of genetic material across ecological communities during evolution, representing a potential force behind reticulation.

Polyploidy↗

Microarray analysis of E2Fa-DPa-overexpressing plants uncovers a cross-talking genetic network between DNA replication and nitrogen assimilation.

Previously we have shown that overexpression of the heterodimeric E2Fa-DPa transcription factor in Arabidopsis thaliana results in ectopic cell division, increased endoreduplication, and an early arrest in development. To gain a better insight into the phenotypic behavior of E2Fa-DPa transgenic plants and to identify E2Fa-DPa target genes, a transcriptomic microarray analysis was performed. Out of 4,390 unique genes, a total of 188 had a twofold or more up- (84) or down-regulated (104) expression level in E2Fa-DPa transgenic plants compared to wild-type lines. Detailed promoter analysis allowed the identification of novel E2Fa-DPa target genes, mainly involved in DNA replication. Secondarily induced genes encoded proteins involved in cell wall biosynthesis, transcription and signal transduction or had an unknown function. A large number of metabolic genes were modified as well, among which, surprisingly, many genes were involved in nitrate assimilation. Our data suggest that the growth arrest observed upon E2Fa-DPa overexpression results at least partly from a nitrogen drain to the nucleotide synthesis pathway, causing decreased synthesis of other nitrogen compounds, such as amino acids and storage proteins.

Arabidopsis↗

Structural diversification and neo-functionalization during floral MADS-box gene evolution by C-terminal frameshift mutations.

Frameshift mutations generally result in loss-of-function changes since they drastically alter the protein sequence downstream of the frameshift site, besides creating premature stop codons. Here we present data suggesting that frameshift mutations in the C-terminal domain of specific ancestral MADS-box genes may have contributed to the structural and functional divergence of the MADS-box gene family. We have identified putative frameshift mutations in the conserved C-terminal motifs of the B-function DEF/AP3 subfamily, the A-function SQUA/AP1 subfamily and the E-function AGL2 subfamily, which are all involved in the specification of organ identity during flower development. The newly evolved C-terminal motifs are highly conserved, suggesting a de novo generation of functionality. Interestingly, since the new C-terminal motifs in the A- and B-function subfamilies are only found in higher eudicotyledonous flowering plants, the emergence of these two C-terminal changes coincides with the origin of a highly standardized floral structure. We speculate that the frameshift mutations described here are examples of co-evolution of the different components of a single transcription factor complex. 3' terminal frameshift mutations might provide an important but so far unrecognized mechanism to generate novel functional C-terminal motifs instrumental to the functional diversification of transcription factor families.

Amino Acid Motifs↗

Genomewide structural annotation and evolutionary analysis of the type I MADS-box genes in plants.

The type I MADS-box genes constitute a largely unexplored subfamily of the extensively studied MADS-box gene family, well known for its role in flower development. Genes of the type I MADS-box subfamily possess the characteristic MADS box but are distinguished from type II MADS-box genes by the absence of the keratin-like box. In this in silico study, we have structurally annotated all 47 members of the type I MADS-box gene family in Arabidopsis thaliana and exerted a thorough analysis of the C-terminal regions of the translated proteins. On the basis of conserved motifs in the C-terminal region, we could classify the gene family into three main groups, two of which could be further subdivided. Phylogenetic trees were inferred to study the evolutionary relationships within this large MADS-box gene subfamily. These suggest for plant type I genes a dynamic of evolution that is significantly different from the mode of both animal type I (SRF) and plant type II (MIKC-type) gene phylogeny. The presence of conserved motifs in the majority of these genes, the identification of Oryza sativa MADS-box type I homologues, and the detection of expressed sequence tags for Arabidopsis thaliana and other plant type I genes suggest that these genes are indeed of functional importance to plants. It is therefore even more intriguing that, from an experimental point of view, almost nothing is known about the function of these MADS-box type I genes.

Amino Acid Motifs↗

'Natural selection merely modified while redundancy created'--Susumu Ohno's idea of the evolutionary importance of gene and genome duplications.

Susumo Ohno's influential book Evolution by gene duplication dealt with the idea that gene and genome duplication events are the principal forces by which the genetic raw material is provided for increasing complexity during evolution. In 1970, the evidence for this hypothesis consisted mostly of karyotypic information, crude information by today's standard genetic data, DNA sequences. Nonetheless, although the type of data are outdated, the idea remained current and is still debated today in the age of complete genome sequences. Even more than thirty years after the initial publication more research than ever is being carried out on the evolutionary significance of gene and genome duplications and the contribution of these mechanisms to the advances in genomic and organismal evolution.

Evolution, Molecular↗

Genome duplication, a trait shared by 22000 species of ray-finned fish.

Through phylogeny reconstruction we identified 49 genes with a single copy in man, mouse, and chicken, one or two copies in the tetraploid frog Xenopus laevis, and two copies in zebrafish (Danio rerio). For 22 of these genes, both zebrafish duplicates had orthologs in the pufferfish (Takifugu rubripes). For another 20 of these genes, we found only one pufferfish ortholog but in each case it was more closely related to one of the zebrafish duplicates than to the other. Forty-three pairs of duplicated genes map to 24 of the 25 zebrafish linkage groups but they are not randomly distributed; we identified 10 duplicated regions of the zebrafish genome that each contain between two and five sets of paralogous genes. These phylogeny and synteny data suggest that the common ancestor of zebrafish and pufferfish, a fish that gave rise to approximately 22000 species, experienced a large-scale gene or complete genome duplication event and that the pufferfish has lost many duplicates that the zebrafish has retained.

Animals↗

Evidence that rice and other cereals are ancient aneuploids.

Detailed analyses of the genomes of several model organisms revealed that large-scale gene or even entire-genome duplications have played prominent roles in the evolutionary history of many eukaryotes. Recently, strong evidence has been presented that the genomic structure of the dicotyledonous model plant species Arabidopsis is the result of multiple rounds of entire-genome duplications. Here, we analyze the genome of the monocotyledonous model plant species rice, for which a draft of the genomic sequence was published recently. We show that a substantial fraction of all rice genes ( approximately 15%) are found in duplicated segments. Dating of these block duplications, their nonuniform distribution over the different rice chromosomes, and comparison with the duplication history of Arabidopsis suggest that rice is not an ancient polyploid, as suggested previously, but an ancient aneuploid that has experienced the duplication of one-or a large part of one-chromosome in its evolutionary past, approximately 70 million years ago. This date predates the divergence of most of the cereals, and relative dating by phylogenetic analysis shows that this duplication event is shared by most if not all of them.

Aneuploidy↗

Are all fishes ancient polyploids?

Euteleost fishes seem to have more copies of many genes than their tetrapod relatives. Three different mechanisms could explain the origin of these 'extra' fish genes. The duplicates may have been produced during a fish-specific genome duplication event. A second explanation is an increased rate of independent gene duplications in fish. A third possibility is that after gene or genome duplication events in the common ancestor of fish and tetrapods, the latter lost more genes. These three hypotheses have been tested by phylogenetic tree reconstruction. Phylogenetic analyses of sequences from human, mouse, chicken, frog (Xenopus laevis), zebrafish (Danio rerio) and pufferfish (Takifugu rubripes) suggest that ray-finned fishes are likely to have undergone a whole genome duplication event between 200 and 450 million years ago. We also comment here on the evolutionary consequences of this ancient genome duplication.

Animals↗

Investigating ancient duplication events in the Arabidopsis genome.

The complete genomic analysis of Arabidopsis thaliana has shown that a major fraction of the genome consists of paralogous genes that probably originated through one or more ancient large-scale gene or genome duplication events. However, the number and timing of these duplications still remains unclear, and several different hypotheses have been put forward recently. Here, we reanalyzed duplicated blocks found in the Arabidopsis genome described previously and determined their date of divergence based on silent substitution estimations between the paralogous genes and, where possible, by phylogenetic reconstruction. We show that methods based on averaging protein distances of heterogeneous classes of duplicated genes lead to unreliable conclusions and that a large fraction of blocks duplicated much more recently than assumed previously. We found clear evidence for one large-scale gene or even complete genome duplication event somewhere between 70 to 90 million years ago. Traces pointing to a much older (probably more than 200 million years) large-scale gene duplication event could be detected. However, for now it is impossible to conclude whether these old duplicates are the result of one or more large-scale gene duplication events.

Arabidopsis↗

The hidden duplication past of Arabidopsis thaliana.

Analysis of the genome sequence of Arabidopsis thaliana shows that this genome, like that of many other eukaryotic organisms, has undergone large-scale gene duplications or even duplications of the entire genome. However, the high frequency of gene loss after duplication events reduces colinearity and therefore the chance of finding duplicated regions that, at the extreme, no longer share homologous genes. In this study we show that heavily degenerated block duplications that can no longer be recognized by directly comparing two segments because of differential gene loss, can still be detected through indirect comparison with other segments. When these so-called hidden duplications in Arabidopsis are taken into account, many homologous genomic regions can be found in five to eight copies. This finding strongly implies that Arabidopsis has undergone three, but probably no more, rounds of genome duplications. Therefore, adding such hidden blocks to the duplication landscape of Arabidopsis sheds light on the number of polyploidy events that this model plant genome has undergone in its evolutionary past.

Arabidopsis↗

Dealing with saturation at the amino acid level: a case study based on anciently duplicated zebrafish genes.

The ray-finned fishes (Actinopterygii) seem to have two copies of many tetrapod (Sarcopterygii) genes. The origin of these duplicate fish genes is the subject of some controversy. One explanation for the existence of these extra fish genes could be an increase in the rate of independent gene duplications in fishes. Alternatively, gene duplicates in fish may have been formed in the ancestor of all or most Actinopterygii during a complete genome duplication event. A third possibility is that tetrapods have lost more genes than fish after gene or genome duplication events in the common ancestor of both lineages. These three hypotheses can be tested by phylogenetic reconstruction. Previously, we found that a large number of anciently duplicated genes of zebrafish are sister sequences in evolutionary trees suggesting that they were produced in Actinopterygii after the divergence of Sarcopterygii [Phil. Trans. R. Soc. Lond. B 356 (2001) 119]. On the other hand, several well-supported trees showed one of the two fish genes as the sister sequence to a monophyletic clade that included the second fish gene and genes from frog, chicken, mouse and human. These so-called outgroup topologies suggest that the origin of many fish duplicates predates the divergence of the Sarcopterygii and Actinopterygii and support the hypothesis that tetrapods have lost duplicates that have been retained in fish. Here we show that many of these 'outgroup' tree topologies are erroneous and can be corrected when mutational saturation is taken into account. To this end, a Java-based application has been developed to visualize the amount of saturation in amino acid sequences. The program graphically displays the number of observed frequent and rare amino acid replacements between pairs of sequences against their overall evolutionary distance. Discrimination between frequent and rare amino acid replacements is based on substitution probability matrices (e.g. PAM and BLOSUM). Evolutionary distances between sequences can be computed from the fraction of unsaturated sites only and evolutionary trees inferred by pairwise distance methods. When trees are computed by omitting the saturated fraction of sites, most fish duplicates are sister sequences.

Amino Acids↗

Phylogenetic analyses suggest lateral gene transfer from the mitochondrion to the apicoplast.

Apicomplexan protozoa contain a single mitochondrion and a multimembranous plastid-like organelle termed apicoplast. The size of the apicomplexan plastid genome is extremely small (35 kb) thus offering a limited number of genes for phylogenetic analysis. Moreover, the sequences of apicoplast genes are highly adenosine+thymidine-rich and rapidly evolving. Due to these facts, phylogenetic analyses based on different genes or the structure of the ribosomal operon show conflicting results and the evolutionary history of this exciting organelle remains unclear. Although it is evident that the apicoplast and its genome is plastid-derived, our detailed phylogenetic analysis of amino acid and nucleotide sequences of selected apicoplast ribosomal protein genes rpl2, rpl14 and rps12 show their possible mitochondrial origin. The affinity of apicoplast ribosomal proteins to their mitochondrial homologs is very stable and well supported. Based on our results we propose that apicoplasts might contain both plastid and mitochondrial genes, thus constituting a hybrid assembly.

Animals↗

Wanda: a database of duplicated fish genes.

Comparative genomics has shown that ray-finned fish (Actinopterygii) contain more copies of many genes than other vertebrates. A large number of these additional genes appear to have been produced during a genome duplication event that occurred early during the evolution of Actinopterygii (i.e. before the teleost radiation). In addition to this ancient genome duplication event, many lineages within Actinopterygii have experienced more recent genome duplications. Here we introduce a curated database named Wanda that lists groups of orthologous genes with one copy from man, mouse and chicken, one or two from tetraploid Xenopus and two or more ancient copies (i.e. paralogs) from ray-finned fish. The database also contains the sequence alignments and phylogenetic trees that were necessary for determining the correct orthologous and paralogous relationships among genes. Where available, map positions and functional data are also reported. The Wanda database should be of particular use to evolutionary and developmental biologists who are interested in the evolutionary and functional divergence of genes after duplication. Wanda is available at http://www.evolutionsbiologie.uni-konstanz.de/Wanda/.

Animals↗

The European database on small subunit ribosomal RNA.

The European database on SSU rRNA can be consulted via the World WideWeb at http://rrna.uia.ac.be/ssu/ and compiles all complete or nearly complete small subunit ribosomal RNA sequences. Sequences are provided in aligned format. The alignment takes into account the secondary structure information derived by comparative sequence analysis of thousands of sequences. Additional information such as literature references, taxonomy, secondary structure models and nucleotide variability maps, is also available.

Animals↗

PlantCARE, a database of plant cis-acting regulatory elements and a portal to tools for in silico analysis of promoter sequences.

PlantCARE is a database of plant cis-acting regulatory elements, enhancers and repressors. Regulatory elements are represented by positional matrices, consensus sequences and individual sites on particular promoter sequences. Links to the EMBL, TRANSFAC and MEDLINE databases are provided when available. Data about the transcription sites are extracted mainly from the literature, supplemented with an increasing number of in silico predicted data. Apart from a general description for specific transcription factor sites, levels of confidence for the experimental evidence, functional information and the position on the promoter are given as well. New features have been implemented to search for plant cis-acting regulatory elements in a query sequence. Furthermore, links are now provided to a new clustering and motif search method to investigate clusters of co-expressed genes. New regulatory elements can be sent automatically and will be added to the database after curation. The PlantCARE relational database is available via the World Wide Web at http://sphinx.rug.ac.be:8080/PlantCARE/.

Consensus Sequence↗

Detecting the undetectable: uncovering duplicated segments in Arabidopsis by comparison with rice.

Genome analysis shows that large-scale gene duplications have occurred in fungi, animals and plants, creating genomic regions that show similarity in gene content and order. However, the high frequency of gene loss reduces colinearity resulting in duplicated regions that, in the extreme, no longer share homologous genes. Here, we show that by comparison with an appropriate second genome, such paralogous regions can still be identified.

Arabidopsis↗