Search PubMed⌕ Search

Biomedical subjects

Laurent Duret

Publications and source records attributed to Laurent Duret.

At least 19 recordsLinked to original sources

High prevalence of PRDM9-independent recombination hotspots in placental mammals.

In many mammals, recombination events are concentrated in hotspots directed by a sequence-specific DNA-binding protein named PRDM9. Intriguingly, PRDM9 has been lost several times in vertebrates, and notably among mammals, it has been pseudogenized in the ancestor of canids. In the absence of PRDM9, recombination hotspots tend to occur in promoter-like features such as CpG islands. It has thus been proposed that one role of PRDM9 could be to direct recombination away from PRDM9-independent hotspots. However, the ability of PRDM9 to direct recombination hotspots has been assessed in only a handful of species, and a clear picture of how much recombination occurs outside of PRDM9-directed hotspots in mammals is still lacking. In this study, we derived an estimator of past recombination activity based on signatures of GC-biased gene conversion in substitution patterns. We quantified recombination activity in PRDM9-independent hotspots in 52 species of boreoeutherian mammals. We observe a wide range of recombination rates at these loci: several species (such as mice, humans, some felids, or cetaceans) show a deficit of recombination, while a majority of mammals display a clear peak of recombination. Our results demonstrate that PRDM9-directed and PRDM9-independent hotspots can coexist in mammals and that their coexistence appears to be the rule rather than the exception. Additionally, we show that the location of PRDM9-independent hotspots is relatively more stable than that of PRDM9-directed hotspots, but that PRDM9-independent hotspots nevertheless evolve slowly in concert with DNA hypomethylation.

Animals↗

The Xist RNA gene evolved in eutherians by pseudogenization of a protein-coding gene.

The Xist noncoding RNA is the key initiator of the process of X chromosome inactivation in eutherian mammals, but its precise function and origin remain unknown. Although Xist is well conserved among eutherians, until now, no homolog has been identified in other mammals. We show here that Xist evolved, at least partly, from a protein-coding gene and that the loss of protein-coding function of the proto-Xist coincides with the four flanking protein genes becoming pseudogenes. This event occurred after the divergence between eutherians and marsupials, which suggests that mechanisms of dosage compensation have evolved independently in both lineages.

Animals↗

Evolutionary origin and maintenance of coexpressed gene clusters in mammals.

Gene order is not random with regard to gene expression in mammals: coexpressed genes, and in particular housekeeping genes, are clustered along chromosomes more often than expected by chance. To understand the origin of these clusters and to quantify the impact of this phenomenon on genome organization, we analyzed clusters of coexpressed genes in the human and mouse genomes. We show that neighboring genes experience continuous concerted expression changes during evolution, which leads to the formation of coexpressed gene clusters. The pattern of expression within these clusters evolves more slowly than the genomic average. Moreover, by studying gene order evolution, we show that some clusters are maintained by natural selection and, therefore, have a functional significance. However, we also demonstrate that some coexpressed gene clusters are the result of neutral coevolution effects, as illustrated by the clustering of genes escaping inactivation on the X chromosome. Moreover, we show that, although statistically significant, constraints on gene orders have a limited impact on mammalian genome organization, affecting only 3-5% of the pool of human and murine genes. It had been hypothesized that coexpressed gene clusters might correspond to large chromatin domains. In contradiction, we find that most of these clusters contain only 2 genes whose coexpression may be due to transcriptional read-through or the activity of bidirectional promoters.

Animals↗

GC content evolution of the human and mouse genomes: insights from the study of processed pseudogenes in regions of different recombination rates.

Processed pseudogenes are generated by reverse transcription of a functional gene. They are generally nonfunctional after their insertion and, as a consequence, are no longer subjected to the selective constraints associated with functional genes. Because of this property they can be used as neutral markers in molecular evolution. In this work, we investigated the relationship between the evolution of GC content in recently inserted processed pseudogenes and the local recombination pattern in two mammalian genomes (human and mouse). We confirmed, using original markers, that recombination drives GC content in the human genome and we demonstrated that this is also true for the mouse genome despite lower recombination rates. Finally, we discussed the consequences on isochores evolution and the contrast between the human and the mouse pattern.

Animals↗

No evidence for tissue-specific adaptation of synonymous codon usage in humans.

It has been proposed that the synonymous codon usage of human tissue-specific genes was under selective pressure to modulate the expression of proteins by codon-mediated translational control (Plotkin, J. B., H. Robins, and A. J. Levine. 2004. Tissue-specific codon usage and the expression of human genes. Proc. Natl. Acad. Sci. USA 101:12588-12591.) To test this model, we analyzed by internal correspondence analysis the codon usage of 2,126 human tissue-specific genes expressed in 18 different tissues. We confirm that synonymous codon usage differs significantly between the tissues. However, the effect is very weak: the variability of synonymous codon usage between tissues represents only 2.3% of the total codon usage variability. Moreover, this variability is directly linked to isochore-scale (>100 kb) variability of GC-content that affect both coding and introns or intergenic regions. This demonstrates that variations of synonymous codon usage between tissue-specific genes expressed in different tissues are due to regional variations of substitution patterns and not to translational selection.

Base Composition↗

Natural history of the ERVWE1 endogenous retroviral locus.

BACKGROUND: The human HERV-W multicopy family includes a unique proviral locus, termed ERVWE1, whose full-length envelope ORF was preserved through evolution by the action of a selective pressure. The encoded Env protein (Syncytin) is involved in hominoid placental physiology. RESULTS: In order to infer the natural history of this domestication process, a comparative genomic analysis of the human 7q21.2 syntenic regions in eutherians was performed. In primates, this region was progressively colonized by LTR-elements, leading to two different evolutionary pathways in Cercopithecidae and Hominidae, a genetic drift versus a domestication, respectively. CONCLUSION: The preservation in Hominoids of a genomic structure consisting in the juxtaposition of a retrotransposon-derived MaLR LTR and the ERVWE1 provirus suggests a functional link between both elements.

Animals↗

Homology-dependent methylation in primate repetitive DNA.

In mammals, several studies have suggested that levels of methylation are higher in repetitive DNA than in nonrepetitive DNA, possibly reflecting a genome-wide defense mechanism against deleterious effects associated with transposable elements (TEs). To analyze the determinants of methylation patterns in primate repetitive DNA, we took advantage of the fact that the methylation rate in the germ line is reflected by the transition rate at CpG sites. We assessed the variability of CpG substitution rates in nonrepetitive DNA and in various TE and retropseudogene families. We show that, unlike other substitution rates, the rate of transition at CpG sites is significantly (37%) higher in repetitive DNA than in nonrepetitive DNA. Moreover, this rate of CpG transition varies according to the number of repeats, their length, and their level of divergence from the ancestral sequence (up to 2.7 times higher in long, lowly divergent TEs compared with unique sequences). This observation strongly suggests the existence of a homology-dependent methylation (HDM) mechanism in mammalian genomes. We propose that HDM is a direct consequence of interfering RNA-induced transcriptional gene silencing.

Animals↗

Tree pattern matching in phylogenetic trees: automatic search for orthologs or paralogs in homologous gene sequence databases.

MOTIVATION: Comparative sequence analysis is widely used to study genome function and evolution. This approach first requires the identification of homologous genes and then the interpretation of their homology relationships (orthology or paralogy). To provide help in this complex task, we developed three databases of homologous genes containing sequences, multiple alignments and phylogenetic trees: HOBACGEN, HOVERGEN and HOGENOM. In this paper, we present two new tools for automating the search for orthologs or paralogs in these databases. RESULTS: First, we have developed and implemented an algorithm to infer speciation and duplication events by comparison of gene and species trees (tree reconciliation). Second, we have developed a general method to search in our databases the gene families for which the tree topology matches a peculiar tree pattern. This algorithm of unordered tree pattern matching has been implemented in the FamFetch graphical interface. With the help of a graphical editor, the user can specify the topology of the tree pattern, and set constraints on its nodes and leaves. Then, this pattern is compared with all the phylogenetic trees of the database, to retrieve the families in which one or several occurrences of this pattern are found. By specifying ad hoc patterns, it is therefore possible to identify orthologs in our databases.

Algorithms↗

Integr8 and Genome Reviews: integrated views of complete genomes and proteomes.

Integr8 is a new web portal for exploring the biology of organisms with completely deciphered genomes. For over 190 species, Integr8 provides access to general information, recent publications, and a detailed statistical overview of the genome and proteome of the organism. The preparation of this analysis is supported through Genome Reviews, a new database of bacterial and archaeal DNA sequences in which annotation has been upgraded (compared to the original submission) through the integration of data from many sources, including the EMBL Nucleotide Sequence Database, the UniProt Knowledgebase, InterPro, CluSTr, GOA and HOGENOM. Integr8 also allows the users to customize their own interactive analysis, and to download both customized and prepared datasets for their own use. Integr8 is available at http://www.ebi.ac.uk/integr8.

DNA, Archaeal↗

Polymorphix: a sequence polymorphism database.

Within-species sequence variation data are of special interest since they contain information about recent population/species history, and the molecular evolutionary forces currently in action in natural populations. These data, however, are presently dispersed within generalist databases, and are difficult to access. To solve this problem, we have developed Polymorphix, a database dedicated to sequence polymorphism. It contains within-species homologous sequence families built using EMBL/GenBank under suitable similarity and bibliographic criteria. Polymorphix is an ACNUC structured database allowing both simple and complex queries for population genomic studies. Alignments within families as well as phylogenetic trees can be download. When available, outgroups are included in the alignment. Polymorphix contains sequences from the nuclear, mitochondrial and chloroplastic genomes of every eukaryote species represented in EMBL. It can be accessed by a web interface (http://pbil.univ-lyon1.fr/polymorphix/query.php).

Animals↗

HOPPSIGEN: a database of human and mouse processed pseudogenes.

Processed pseudogenes result from reverse transcribed mRNAs. In general, because processed pseudogenes lack promoters, they are no longer functional from the moment they are inserted into the genome. Subsequently, they freely accumulate substitutions, insertions and deletions. Moreover, the ancestral structure of processed pseudogenes could be easily inferred using the sequence of their functional homologous genes. Owing to these characteristics, processed pseudogenes represent good neutral markers for studying genome evolution. Recently, there is an increasing interest for these markers, particularly to help gene prediction in the field of genome annotation, functional genomics and genome evolution analysis (patterns of substitution). For these reasons, we have developed a method to annotate processed pseudogenes in complete genomes. To make them useful to different fields of research, we stored them in a nucleic acid database after having annotated them. In this work, we screened both mouse and human complete genomes from ENSEMBL to find processed pseudogenes generated from functional genes with introns. We used a conservative method to detect processed pseudogenes in order to minimize the rate of false positive sequences. Within processed pseudogenes, some are still having a conserved open reading frame and some have overlapping gene locations. We designated as retroelements all reverse transcribed sequences and more strictly, we designated as processed pseudogenes, all retroelements not falling in the two former categories (having a conserved open reading or overlapping gene locations). We annotated 5823 retroelements (5206 processed pseudogenes) in the human genome and 3934 (3428 processed pseudogenes) in the mouse genome. Compared to previous estimations, the total number of processed pseudogenes was underestimated but the aim of this procedure was to generate a high-quality dataset. To facilitate the use of processed pseudogenes in studying genome structure and evolution, DNA sequences from processed pseudogenes, and their functional reverse transcribed homologs, are now stored in a nucleic acid database, HOPPSIGEN. HOPPSIGEN can be browsed on the PBIL (Pole Bioinformatique Lyonnais) World Wide Web server (http://pbil.univ-lyon1.fr/) or fully downloaded for local installation.

Animals↗

Relationship between gene expression and GC-content in mammals: statistical significance and biological relevance.

Mammalian chromosomes are characterized by large-scale variations of DNA base composition (the so-called isochores). In contradiction with previous studies, Lercher et al. (Hum. Mol. Genet., 12, 2411, 2003) recently reported a strong correlation between gene expression breadth and GC-content, suggesting that there might be a selective pressure favoring the concentration of housekeeping genes in GC-rich isochores. We reassessed this issue by examining in human and mouse the correlation between gene expression and GC-content, using different measures of gene expression (EST, SAGE and microarray) and different measures of GC-content. We show that correlations between GC-content and expression are very weak, and may vary according to the method used to measure expression. Such weak correlations have a very low predictive value. The strong correlations reported by Lercher et al. (2003) are because of the fact that they measured variables over neighboring genes windows. We show here that using gene windows artificially enhances the correlation. The assertion that the expression of a given gene depends on the GC-content of the region where it is located is therefore not supported by the data.

Animals↗

Identitag, a relational database for SAGE tag identification and interspecies comparison of SAGE libraries.

BACKGROUND: Serial Analysis of Gene Expression (SAGE) is a method of large-scale gene expression analysis that has the potential to generate the full list of mRNAs present within a cell population at a given time and their frequency. An essential step in SAGE library analysis is the unambiguous assignment of each 14 bp tag to the transcript from which it was derived. This process, called tag-to-gene mapping, represents a step that has to be improved in the analysis of SAGE libraries. Indeed, the existing web sites providing correspondence between tags and transcripts do not concern all species for which numerous EST and cDNA have already been sequenced. RESULTS: This is the reason why we designed and implemented a freely available tool called Identitag for tag identification that can be used in any species for which transcript sequences are available. Identitag is based on a relational database structure in order to allow rapid and easy storage and updating of data and, most importantly, in order to be able to precisely define identification parameters. This structure can be seen like three interconnected modules : the first one stores virtual tags extracted from a given list of transcript sequences, the second stores experimental tags observed in SAGE experiments, and the third allows the annotation of the transcript sequences used for virtual tag extraction. It therefore connects an observed tag to a virtual tag and to the sequence it comes from, and then to its functional annotation when available. Databases made from different species can be connected according to orthology relationship thus allowing the comparison of SAGE libraries between species. We successfully used Identitag to identify tags from our chicken SAGE libraries and for chicken to human SAGE tags interspecies comparison. Identitag sources are freely available on http://pbil.univ-lyon1.fr/software/identitag/ web site. CONCLUSIONS: Identitag is a flexible and powerful tool for tag identification in any single species and for interspecies comparison of SAGE libraries. It opens the way to comparative transcriptomic analysis, an emerging branch of biology.

Animals↗

Evidence of selection on the domesticated ERVWE1 env retroviral element involved in placentation.

The human endogenous retrovirus HERV-W multicopy family includes a unique proviral locus, termed ERVWE1, which contains gag and pol pseudogenes and has retained a full-length envelope open reading frame (ORF). This Env protein (syncytin) is a highly fusogenic membrane glycoprotein and has been proposed to be involved in hominoid placental physiology. To track the hallmarks of natural selection acting on the ERVWE1 env gene, the pattern of substitutions and indels was analyzed within all human HERV-W elements and along the ERVWE1 orthologous loci in chimpanzee, gorilla, orangutan, and gibbon. The comparison of ERVWE1 and paralogous HERV-W copies revealed an ERVWE1-specific signature consisting of a four amino acid deletion in the intracytoplasmic tail of the glycoprotein. We show that this deletion is crucial for the envelope fusogenic activity. The comparison of the human ERVWE1 locus with its orthologs demonstrates the existence of a selective pressure to maintain the env reading frame open. Notably, the 3' part of the env gene, encoding regions required for the fusion process, is under purifying selection. The identification of selective constraints on env ERVWE1 confirms that this retroviral locus has been recruited in the hominoid lineage to become a bona fide gene.

Amino Acid Sequence↗

Recombination drives the evolution of GC-content in the human genome.

Unraveling the evolutionary forces responsible for variations of neutral substitution patterns among taxa or along genomes is a major issue in the identification of functional sequence features. Mammalian genomes show large-scale regional variations of GC-content (the isochores), but the substitution processes at the origin of this structure are poorly understood. We have analyzed the pattern of neutral substitutions in 14.3 Mb of primate noncoding regions. We show that the GC-content toward which sequences are evolving is strongly correlated (r(2) = 0.61, P </= 2 10(-16)) with the rate of crossovers (notably in females). This demonstrates that recombination drives the evolution of base composition in human (probably via the process of biased gene conversion). The present substitution patterns are very different from what they had been in the past, resulting in a major modification of the isochore structure of our genome. This non-equilibrium situation suggests that changes of recombination rates occur relatively frequently during evolution, possibly as a consequence of karyotype rearrangements. These results have important implications for understanding the spatial and temporal variations of substitution processes in a broad range of sexual organisms, and for detecting the hallmarks of natural selection in DNA sequences.

Animals↗

The endogenous retroviral locus ERVWE1 is a bona fide gene involved in hominoid placental physiology.

The definitive demonstration of a role for a recently acquired gene is a difficult task, requiring exhaustive genetic investigations and functional analysis. The situation is indeed much more complicated when facing multicopy gene families, because most or portions of the gene are conserved among the hundred copies of the family. This is the case for the ERVWE1 locus of the human endogenous retrovirus W family (HERV-W), which encodes an envelope glycoprotein (syncytin) likely involved in trophoblast differentiation. Here we describe, in 155 individuals, the positional conservation of this locus and the preservation of the envelope ORF. Sequencing of the critical elements of the ERVWE1 provirus showed a striking conservation among the 48 alleles of 24 individuals, including the LTR elements involved in the transcriptional machinery, the splice sites involved in the maturation of subgenomic Env mRNA, and the Env ORF. The functionality and tissue specificity of the 5' LTR were demonstrated, as well as the fusogenic activity of the envelope polymorphic variants. Such functions were also shown to be preserved in the orthologous loci isolated from chimpanzee, gorilla, orangutan, and gibbon. This functional preservation among humans and during evolution strongly argued for the involvement of this recently acquired retroviral envelope glycoprotein in hominoid placental physiology.

Amino Acid Sequence↗