Search PubMed⌕ Search

Biomedical subjects

Jeroen Raes

Publications and source records attributed to Jeroen Raes.

18 recordsLinked to original sources

Prediction of effective genome size in metagenomic samples.

We introduce a novel computational approach to predict effective genome size (EGS; a measure that includes multiple plasmid copies, inserted sequences, and associated phages and viruses) from short sequencing reads of environmental genomics (or metagenomics) projects. We observe considerable EGS differences between environments and link this with ecologic complexity as well as species composition (for instance, the presence of eukaryotes). For example, we estimate EGS in a complex, organism-dense farm soil sample at about 6.3 megabases (Mb) whereas that of the bacteria therein is only 4.7 Mb; for bacteria in a nutrient-poor, organism-sparse ocean surface water sample, EGS is as low as 1.6 Mb. The method also permits evaluation of completion status and assembly bias in single-genome sequencing projects.

Artifacts↗

Nonsense-mediated mRNA decay: Target genes and functional diversification of effectors.

Recent genome-wide identification of nonsense-mediated mRNA decay (NMD) targets in yeast, fruitfly and human cells has provided insight into the biological functions and evolution of this mRNA quality control mechanism, revealing that NMD post-transcriptionally regulates an important fraction of the transcriptome. NMD targets are associated with a broad range of biological processes, but most of these targets are not encoded by orthologous genes across different species. Yeast and fruitfly NMD effectors regulate common targets in concert, but parallel pathways have evolved in humans, whereby NMD effectors have acquired additional functions. Thus, the phenotypic differences observed across species after inhibition of NMD are driven not only by the functional diversification of NMD effectors but also by changes in the repertoire of regulated genes.

Animals↗

An improved statistical method for detecting heterotachy in nucleotide sequences.

The principle of heterotachy states that the substitution rate of sites in a gene can change through time. In this article, we propose a powerful statistical test to detect sites that evolve according to the process of heterotachy. We apply this test to an alignment of 1289 eukaryotic rRNA molecules to 1) determine how widespread the phenomenon of heterotachy is in ribosomal RNA, 2) to test whether these heterotachous sites are nonrandomly distributed, that is, linked to secondary structure features of ribosomal RNA, and 3) to determine the impact of heterotachous sites on the bootstrap support of monophyletic groupings. Our study revealed that with 21 monophyletic taxa, approximately two-thirds of the sites in the considered set of sequences is heterotachous. Although the detected heterotachous sites do not appear bound to specific structural features of the small subunit rRNA, their presence is shown to have a large beneficial influence on the bootstrap support of monophyletic groups. Using extensive testing, we show that this may not be due to heterotachy itself but merely due to the increased substitution rate at the detected heterotachous sites.

Chi-Square Distribution↗

The TORNADO1 and TORNADO2 genes function in several patterning processes during early leaf development in Arabidopsis thaliana.

In multicellular organisms, patterning is a process that generates axes in the primary body plan, creates domains upon organ formation, and finally leads to differentiation into tissues and cell types. We identified the Arabidopsis thaliana TORNADO1 (TRN1) and TRN2 genes and their role in leaf patterning processes such as lamina venation, symmetry, and lateral growth. In trn mutants, the leaf venation network had a severely reduced complexity: incomplete loops, no tertiary or quaternary veins, and vascular islands. The leaf laminas were asymmetric and narrow because of a severely reduced cell number. We postulate that the imbalance between cell proliferation and cell differentiation and the altered auxin distribution in both trn mutants cause asymmetric leaf growth and aberrant venation patterning. TRN1 and TRN2 were epistatic to ASYMMETRIC LEAVES1 with respect to leaf asymmetry, consistent with their expression in the shoot apical meristem and leaf primordia. TRN1 codes for a large plant-specific protein with conserved domains also found in a variety of signaling proteins, whereas TRN2 encodes a transmembrane protein of the tetraspanin family whose phylogenetic tree is presented. Double mutant analysis showed that TRN1 and TRN2 act in the same pathway.

Arabidopsis↗

Nonrandom divergence of gene expression following gene and genome duplications in the flowering plant Arabidopsis thaliana.

BACKGROUND: Genome analyses have revealed that gene duplication in plants is rampant. Furthermore, many of the duplicated genes seem to have been created through ancient genome-wide duplication events. Recently, we have shown that gene loss is strikingly different for large- and small-scale duplication events and highly biased towards the functional class to which a gene belongs. Here, we study the expression divergence of genes that were created during large- and small-scale gene duplication events by means of microarray data and investigate both the influence of the origin (mode of duplication) and the function of the duplicated genes on expression divergence. RESULTS: Duplicates that have been created by large-scale duplication events and that can still be found in duplicated segments have expression patterns that are more correlated than those that were created by small-scale duplications or those that no longer lie in duplicated segments. Moreover, the former tend to have highly redundant or overlapping expression patterns and are mostly expressed in the same tissues, while the latter show asymmetric divergence. In addition, a strong bias in divergence of gene expression was observed towards gene function and the biological process genes are involved in. CONCLUSION: By using microarray expression data for Arabidopsis thaliana, we show that the mode of duplication, the function of the genes involved, and the time since duplication play important roles in the divergence of gene expression and, therefore, in the functional divergence of genes after duplication.

Amino Acid Substitution↗

Modeling gene and genome duplications in eukaryotes.

Recent analysis of complete eukaryotic genome sequences has revealed that gene duplication has been rampant. Moreover, next to a continuous mode of gene duplication, in many eukaryotic organisms the complete genome has been duplicated in their evolutionary past. Such large-scale gene duplication events have been associated with important evolutionary transitions or major leaps in development and adaptive radiations of species. Here, we present an evolutionary model that simulates the duplication dynamics of genes, considering genome-wide duplication events and a continuous mode of gene duplication. Modeling the evolution of the different functional categories of genes assesses the importance of different duplication events for gene families involved in specific functions or processes. By applying our model to the Arabidopsis genome, for which there is compelling evidence for three whole-genome duplications, we show that gene loss is strikingly different for large-scale and small-scale duplication events and highly biased toward certain functional classes. We provide evidence that some categories of genes were almost exclusively expanded through large-scale gene duplication events. In particular, we show that the three whole-genome duplications in Arabidopsis have been directly responsible for >90% of the increase in transcription factors, signal transducers, and developmental genes in the last 350 million years. Our evolutionary model is widely applicable and can be used to evaluate different assumptions regarding small- or large-scale gene duplication events in eukaryotic genomes.

Arabidopsis↗

GeneFarm, structural and functional annotation of Arabidopsis gene and protein families by a network of experts.

Genomic projects heavily depend on genome annotations and are limited by the current deficiencies in the published predictions of gene structure and function. It follows that, improved annotation will allow better data mining of genomes, and more secure planning and design of experiments. The purpose of the GeneFarm project is to obtain homogeneous, reliable, documented and traceable annotations for Arabidopsis nuclear genes and gene products, and to enter them into an added-value database. This re-annotation project is being performed exhaustively on every member of each gene family. Performing a family-wide annotation makes the task easier and more efficient than a gene-by-gene approach since many features obtained for one gene can be extrapolated to some or all the other genes of a family. A complete annotation procedure based on the most efficient prediction tools available is being used by 16 partner laboratories, each contributing annotated families from its field of expertise. A database, named GeneFarm, and an associated user-friendly interface to query the annotations have been developed. More than 3000 genes distributed over 300 families have been annotated and are available at http://genoplante-info.infobiogen.fr/Genefarm/. Furthermore, collaboration with the Swiss Institute of Bioinformatics is underway to integrate the GeneFarm data into the protein knowledgebase Swiss-Prot.

Arabidopsis↗

Functional divergence of proteins through frameshift mutations.

Frameshift mutations are generally considered to be deleterious and of little importance for the evolution of novel gene functions. However, by screening an exhaustive set of vertebrate gene families, we found that, when a second transcript encoding the original gene product compensates for this mutation, frameshift mutations can be retained for millions of years and enable new gene functions to be acquired.

Amino Acid Sequence↗

Nonsense-mediated mRNA decay factors act in concert to regulate common mRNA targets.

Nonsense-mediated mRNA decay (NMD) is a surveillance pathway that degrades mRNAs containing nonsense codons, and regulates the expression of naturally occurring transcripts. While NMD is not essential in yeast or nematodes, UPF1, a key NMD effector, is essential in mice. Here we show that NMD components are required for cell proliferation in Drosophila. This raises the question of whether NMD effectors diverged functionally during evolution. To address this question, we examined expression profiles in Drosophila cells depleted of all known metazoan NMD components. We show that UPF1, UPF2, UPF3, SMG1, SMG5, and SMG6 regulate in concert the expression of a cohort of genes with functions in a wide range of cellular activities, including cell cycle progression. Only a few transcripts were regulated exclusively by individual factors, suggesting that these proteins act mainly in the NMD pathway and their role in mRNA decay has not diverged substantially. Finally, the vast majority of NMD targets in Drosophila are not orthologs of targets previously identified in yeast or human cells. Thus phenotypic differences observed across species following inhibition of NMD can be largely attributed to changes in the repertoire of regulated genes.

Animals↗

Duplication and divergence: the evolution of new genes and old ideas.

Over 35 years ago, Susumu Ohno stated that gene duplication was the single most important factor in evolution. He reiterated this point a few years later in proposing that without duplicated genes the creation of metazoans, vertebrates, and mammals from unicellular organisms would have been impossible. Such big leaps in evolution, he argued, required the creation of new gene loci with previously nonexistent functions. Bold statements such as these, combined with his proposal that at least one whole-genome duplication event facilitated the evolution of vertebrates, have made Ohno an icon in the literature on genome evolution. However, discussion on the occurrence and consequences of gene and genome duplication events has a much longer, and often neglected, history. Here we review literature dealing with the occurrence and consequences of gene duplication, beginning in 1911. We document conceptual and technological advances in gene duplication research from this early research in comparative cytology up to recent research on whole genomes, "transcriptomes," and "interactomes."

Animals↗

Genomewide structural annotation and evolutionary analysis of the type I MADS-box genes in plants.

The type I MADS-box genes constitute a largely unexplored subfamily of the extensively studied MADS-box gene family, well known for its role in flower development. Genes of the type I MADS-box subfamily possess the characteristic MADS box but are distinguished from type II MADS-box genes by the absence of the keratin-like box. In this in silico study, we have structurally annotated all 47 members of the type I MADS-box gene family in Arabidopsis thaliana and exerted a thorough analysis of the C-terminal regions of the translated proteins. On the basis of conserved motifs in the C-terminal region, we could classify the gene family into three main groups, two of which could be further subdivided. Phylogenetic trees were inferred to study the evolutionary relationships within this large MADS-box gene subfamily. These suggest for plant type I genes a dynamic of evolution that is significantly different from the mode of both animal type I (SRF) and plant type II (MIKC-type) gene phylogeny. The presence of conserved motifs in the majority of these genes, the identification of Oryza sativa MADS-box type I homologues, and the detection of expressed sequence tags for Arabidopsis thaliana and other plant type I genes suggest that these genes are indeed of functional importance to plants. It is therefore even more intriguing that, from an experimental point of view, almost nothing is known about the function of these MADS-box type I genes.

Amino Acid Motifs↗

And then there were many: MADS goes genomic.

During the past decade, MADS-box genes have become known as key regulators in both reproductive and vegetative plant development. Traditional genetics and functional genomics tools are now available to elucidate the expression and function of this complex gene family on a much larger scale. Moreover, comparative analysis of the MADS-box genes in diverse flowering and non-flowering plants, boosted by bioinformatics, contributes to our understanding of how this important gene family has expanded during the evolution of land plants. Therefore, the recent advances in comparative and functional genomics should enable researchers to identify the full range of MADS-box gene functions, which should help us significantly in developing a better understanding of plant development and evolution.

Evolution, Molecular↗

Genome-wide characterization of the lignification toolbox in Arabidopsis.

Lignin, one of the most abundant terrestrial biopolymers, is indispensable for plant structure and defense. With the availability of the full genome sequence, large collections of insertion mutants, and functional genomics tools, Arabidopsis constitutes an excellent model system to profoundly unravel the monolignol biosynthetic pathway. In a genome-wide bioinformatics survey of the Arabidopsis genome, 34 candidate genes were annotated that encode genes homologous to the 10 presently known enzymes of the monolignol biosynthesis pathway, nine of which have not been described before. By combining evolutionary analysis of these 10 gene families with in silico promoter analysis and expression data (from a reverse transcription-polymerase chain reaction analysis on an extensive tissue panel, mining of expressed sequence tags from publicly available resources, and assembling expression data from literature), 12 genes could be pinpointed as the most likely candidates for a role in vascular lignification. Furthermore, a possible novel link was detected between the presence of the AC regulatory promoter element and the biosynthesis of G lignin during vascular development. Together, these data describe the full complement of monolignol biosynthesis genes in Arabidopsis, provide a unified nomenclature, and serve as a basis for further functional studies.

Alcohol Oxidoreductases↗

Investigating ancient duplication events in the Arabidopsis genome.

The complete genomic analysis of Arabidopsis thaliana has shown that a major fraction of the genome consists of paralogous genes that probably originated through one or more ancient large-scale gene or genome duplication events. However, the number and timing of these duplications still remains unclear, and several different hypotheses have been put forward recently. Here, we reanalyzed duplicated blocks found in the Arabidopsis genome described previously and determined their date of divergence based on silent substitution estimations between the paralogous genes and, where possible, by phylogenetic reconstruction. We show that methods based on averaging protein distances of heterogeneous classes of duplicated genes lead to unreliable conclusions and that a large fraction of blocks duplicated much more recently than assumed previously. We found clear evidence for one large-scale gene or even complete genome duplication event somewhere between 70 to 90 million years ago. Traces pointing to a much older (probably more than 200 million years) large-scale gene duplication event could be detected. However, for now it is impossible to conclude whether these old duplicates are the result of one or more large-scale gene duplication events.

Arabidopsis↗

Gene duplication, the evolution of novel gene functions, and detecting functional divergence of duplicates in silico.

Duplication of genes increases the amount of genetic material on which evolution can work and has been considered of major importance for the development of biological novelties or to explain important transitions that have occurred during biological evolution. Recently, much research has been devoted to the study of the evolutionary and functional divergence of duplicated genes. Since the majority of genes are part of gene families, there is considerable interest in predicting differences in function between duplicates and assessing the functional redundancy of genes within gene families. In this review, we discuss the strengths and limitations of both older and novel approaches to investigate the evolution of duplicated genes in silico.

Algorithms↗

Transcriptome analysis during cell division in plants.

Using synchronized tobacco Bright Yellow-2 cells and cDNA-amplified fragment length polymorphism-based genomewide expression analysis, we built a comprehensive collection of plant cell cycle-modulated genes. Approximately 1,340 periodically expressed genes were identified, including known cell cycle control genes as well as numerous unique candidate regulatory genes. A number of plant-specific genes were found to be cell cycle modulated. Other transcript tags were derived from unknown plant genes showing homology to cell cycle-regulatory genes of other organisms. Many of the genes encode novel or uncharacterized proteins, indicating that several processes underlying cell division are still largely unknown.

Cell Cycle↗

The automatic detection of homologous regions (ADHoRe) and its application to microcolinearity between Arabidopsis and rice.

It is expected that one of the merits of comparative genomics lies in the transfer of structural and functional information from one genome to another. This is based on the observation that, although the number of chromosomal rearrangements that occur in genomes is extensive, different species still exhibit a certain degree of conservation regarding gene content and gene order. It is in this respect that we have developed a new software tool for the Automatic Detection of Homologous Regions (ADHoRe). ADHoRe was primarily developed to find large regions of microcolinearity, taking into account different types of microrearrangements such as tandem duplications, gene loss and translocations, and inversions. Such rearrangements often complicate the detection of colinearity, in particular when comparing more anciently diverged species. Application of ADHoRe to the complete genome of Arabidopsis and a large collection of concatenated rice BACs yields more than 20 regions showing statistically significant microcolinearity between both plant species. These regions comprise from 4 up to 11 conserved homologous gene pairs. We predict the number of homologous regions and the extent of microcolinearity to increase significantly once better annotations of the rice genome become available.

Arabidopsis↗

Genome-wide analysis of core cell cycle genes in Arabidopsis.

Cyclin-dependent kinases and cyclins regulate with the help of different interacting proteins the progression through the eukaryotic cell cycle. A high-quality, homology-based annotation protocol was applied to determine the core cell cycle genes in the recently completed Arabidopsis genome sequence. In total, 61 genes were identified belonging to seven selected families of cell cycle regulators, for which 30 are new or corrections of the existing annotation. A new class of putative cell cycle regulators was found that probably are competitors of E2F/DP transcription factors, which mediate the G1-to-S progression. In addition, the existing nomenclature for cell cycle genes of Arabidopsis was updated, and the physical positions of all genes were compared with segmentally duplicated blocks in the genome, showing that 22 core cell cycle genes emerged through block duplications. This genome-wide analysis illustrates the complexity of the plant cell cycle machinery and provides a tool for elucidating the function of new family members in the future.

Amino Acid Sequence↗