Search PubMedSearch

Biomedical subjects

D Mouchiroud

Publications and source records attributed to D Mouchiroud.

At least 19 recordsLinked to original sources

Impact of changes in GC content on the silent molecular clock in murids.

Murid nuclear genomes are more homogeneous in GC content than those of most mammals, which leads to the question of how such important compositional changes have accumulated. This paper reports on relationships between frequencies of synonymous differences and GC change, in the lineages leading to human and murids. For this, we used the four-species approach: GC changes between human and murids were compared to the frequencies of synonymous differences, measured between two independent species without GC change (bovine and pig), by using orthologous genes common to all four species. We report three conclusions: (1) Among genes with little GC change, 60% of the variability of synonymous substitution frequencies is explained by the gene-specific rate component. (2) GC changes in murid genomes are independent of the gene-specific rate component. Slowly evolving genes in pig bovine comparison can show strong GC change in murids. (3) By using a GC-independent estimate of the substitution rate, we show that GC changes in murid genomes increase synonymous substitution frequencies. The GC homogenization considerably weakens the gene-specific conservation of substitution rates in murids, and could explain part of the increase of evolutionary rates observed in this group. We present a mechanism that can account for the evolution of the GC homogenization in murids.

Animals

Molecular phylogeny of rodents, with special emphasis on murids: evidence from nuclear gene LCAT.

Phylogenetic relationships among 19 extant species of rodents, with special emphasis on rats, mice, and allied Muroidea, were studied using sequences of the nuclear protein-coding gene LCAT (lecithin:cholesterol acyltransferase), an enzyme of cholesterol metabolism. Analysis of 705 base pairs from the exonic regions of LCAT confirmed known groupings in and around Muroidea. Strong support was found for the families Sciuridae (squirrel and marmot) and Gliridae (dormice) and for suprafamilial taxa Muroidea and Caviomorpha (guinea pig and allies). Within Muroidea, the first branching leads to the fossorial mole rats Spalacinae and bamboo rats Rhizomyinae. The other Muroidea appear as a polytomy from which are issued Gerbillinae (gerbils), Murinae (rats and mice), Sigmodontinae (New World cricetids), Cricetinae (hamsters), and Arvicolinae (voles). Evidence from LCAT sequences agrees with that from a number of previous molecular and morphological studies, both concerning branching orders inside Muroidea and the bush-like radiation of rodent suprafamilial taxa (caviomorphs, sciurids, glirids, muroids), thus suggesting that this nuclear gene is an appropriate candidate for addressing questions of rodents relationships.

Animals

The major compositional transitions in the vertebrate genome.

The vertebrate genome underwent two major compositional transitions, between therapsids and mammals and between dinosaurs and birds. These transitions concerned a sizable part (roughly one-third) of the genome, the gene-richest part of it, and consisted in an increase in GC levels (GC is the molar fraction of guanine + cytosine in DNA) which affected both coding sequences (especially third codon positions) and noncoding sequences. These major transitions were studied here by comparing GC3 levels (GC3 is the GC of third codon positions) of orthologous genes from Xenopus, chicken, calf, and man.

Animals

Evolution of isochores in rodents.

The most deviant isochore pattern within mammals was found in rat and mouse; most other mammals possess a different kind of isochore organization called the "general pattern." However, isochore patterns remain largely unknown in rodents other than mouse and rat. To investigate the taxonomic distribution of isochore patterns in rodents, we sequenced the nuclear gene LCAT (lecithin:cholesterol acyltransferase) from 17 rodents species (bringing the total of LCAT sequences in rodent to 19) and compared their GC contents at third codon positions and in introns. We also analyzed an extensive sequence database from rodents other than rat and mouse. All murid LCAT sequences are much poorer in GC than all nonrodent LCAT sequences, and the hamster sequence database shows exactly the same isochore pattern as rat and mouse. Thus, all murids share the same special isochore pattern--GC homogenization. LCAT sequences are GC-poor in hystricomorphs too, but the guinea pig sequence database indicates that large changes in GC content occur without an overall modification of the isochore pattern. This novel mode of isochore evolution is called GC reordering. LCAT sequences also show that the evolution of isochores in sciurids and glirids is nonconservative in comparison with that in nonrodents. Thus, at least two novel patterns of isochore evolution were found. No rodent investigated to date shared the general mammalian pattern.

Animals

Human coding and noncoding DNA: compositional correlations.

As the correlations between GC levels in third codon positions (GC3) and intergenic sequence GC levels can be used to assess the distribution of genes in the human genome, they were studied in detail. Previous work from our laboratory has demonstrated the existence of linear correlations between GC levels of exons, introns, third codon positions, 5' flanking regions of genes, and long genomic DNA sequences (> or = 10 kb) or DNA molecules (50-100 kb) in which the genes are embedded. The present study confirms and extends the previous results using a larger set of data. Furthermore, an analysis of 4270 human genomic DNA and cDNA sequences has allowed us to confirm a correlation of GC3 against GC1+2. Recent additions to the sequence database have also allowed separate analyses of the 5' flanking regions of CpG island and non-CpG island genes as well as analyses of 3' flanking regions, which suggest that the GC levels of 3' flanking regions are closer to those of intergenic DNA than are those of other regions of genes.

Base Composition

Statistical analysis of vertebrate sequences reveals that long genes are scarce in GC-rich isochores.

We compared the exon/intron organization of vertebrate genes belonging to different isochore classes, as predicted by their GC content at third codon position. Two main features have emerged from the analysis of sequences published in GenBank: (1) genes coding for long proteins (i.e., > or = 500 aa) are almost two times more frequent in GC-poor than in GC-rich isochores; (2) intervening sequences (= sum of introns) are on average three times longer in GC-poor than in GC-rich isochores. These patterns are observed among human, mouse, rat, cow, and even chicken genes and are therefore likely to be common to all warm-blooded vertebrates. Analysis of Xenopus sequences suggests that the same patterns exist in cold-blooded vertebrates. It could be argued that such results do not reflect the reality because sequence databases are not representative of entire genomes. However, analysis of biases in GenBank revealed that the observed discrepancies between GC-rich and GC-poor isochores are not artifactual, and are probably largely underestimated. We investigated the distribution of microsatellites and interspersed repeats in introns of human and mouse genes from different isochores. This analysis confirmed previous studies showing that L1 repeats are almost absent from GC-rich isochores. Microsatellites and SINES (Alu, B1, B2) are found at roughly equal frequencies in introns from all isochore classes. Globally, the presence of repeated sequences does not account for the increased intron length in GC-poor isochores. The relationships between gene structure and global genome organization and evolution are discussed.

Animals

Frequencies of synonymous substitutions in mammals are gene-specific and correlated with frequencies of nonsynonymous substitutions.

The frequencies of synonymous substitutions of mammalian genes cover a much wider range than previously thought. We report here that the different frequencies found in homologous genes from a given mammalian pair are correlated with those in the same homologous genes from a different mammalian pair. This indicates that the frequencies of synonymous substitutions are gene-specific (as are the frequencies of nonsynonymous substitutions), or, in other words, that "fast" and "slow" genes in one mammal are fast and slow, respectively, in any other one. Moreover, the frequencies of synonymous substitutions are correlated with the frequencies of nonsynonymous substitution in the same genes.

Animals

HOVERGEN: a database of homologous vertebrate genes.

Comparison of homologous genes is a major step for many studies related to genome structure, function or evolution. Similarity search programs easily find genes homologous to a given sequence. However, only very tedious manual procedures allow the retrieval of all sets of homologous genes sequenced for a given set of species. Moreover, this search often generates errors due to the complexity of data to be managed simultaneously: phylogenetic trees, alignments, taxonomy, sequences and related information. HOVERGEN helps to solve these problems by integrating all this information. HOVERGEN corresponds to GenBank sequences from all vertebrate species, with some data corrected, clarified, or completed, notably to address the problem of redundancy. Coding sequences have been classified in gene families. Protein multiple alignments and phylogenetic trees have been calculated for each family. Sequences and related information have been structured in an ACNUC database which permits complex selections. A graphical interface has been developed to visualize and edit trees. Genes are displayed in color, according to their taxonomy. Users have directly access to all information attached to sequences and to multiple alignments simply by clicking on genes. This graphical tool gives thus a rapid and simple access to all data necessary to interpret homology relationships between genes. HOVERGEN allows the user to easily select sets of homologous vertebrate genes, and thus is particularly useful for comparative sequence analysis, or molecular evolution studies.

Animals

Silent substitutions in mammalian genomes and their evolutionary implications.

An analysis of silent substitutions in pairwise comparisons of homologous genes from different mammals has shown that, in spite of individual fluctuations, their frequencies (which are very strongly correlated with the frequency of substitutions per synonymous site calculated according to Li et al. 1985) do not vary, on the average, with the GC levels of silent positions. This holds in the general case, in which silent positions of pairs of homologous genes share the same composition, namely in the human/other primates, human/artiodactyls, and in the mouse/rat pairs, as well as in the special cases in which the composition of silent positions are different, namely in the human/rabbit and the human/rat (or human/mouse) pairs. A slightly lower frequency found for low GC values in the human/bovine and human/pig pairs seems to be due to the specific gene samples used. These results contradict the previously claimed existence of differences in mutation rates and of mutational biases in third codon positions of coding sequences located in different isochores of mammalian genomes. They also imply that the variations in nucleotide precursor pools through the cell cycle and the differences in replication timing, or in repair efficiency, which were reported for different isochores, do not lead, as claimed, to differences in mutation rates, not in mutational biases in mammals. The differences claimed appear to be due to using small gene samples when individual fluctuations from gene to gene are relatively large.

Animals

Compositional properties of coding sequences and mammalian phylogeny.

The compositional distributions of large DNA fragments reflect those of the isochores that make up vertebrate genomes and can provide novel phylogenetic insights in the case of mammalian genomes (see Sabeur et al. 1993). This approach has been complemented here by an analysis of the compositional patterns of coding sequences and their codon positions (which also reflect the isochore pattern) and by a comparison of the base compositions of codon positions from homologous genes in a number of pairs of species. The results obtained using these two approaches support the existence of a general compositional pattern for mammalian genomes and of a distinct pattern for Myomorpha. The other two "special" patterns identified in a megachiropteran and in pangolin could not be tested here.

Animals

Insect muscle actins differ distinctly from invertebrate and vertebrate cytoplasmic actins.

Invertebrate actins resemble vertebrate cytoplasmic actins, and the distinction between muscle and cytoplasmic actins in invertebrates is not well established as for vertebrate actins. However, Bombyx and Drosophila have actin genes specifically expressed in muscles. To investigate if the distinction between muscle and cytoplasmic actins evidenced by gene expression analysis is related to the sequence of corresponding genes, we compare the sequences of actin genes of these two insect species and of other Metazoa. We find that insect muscle actins form a family of related proteins characterized by about 10 muscle-specific amino acids. Insect muscle actins have clearly diverged from cytoplasmic actins and form a monophyletic group emerging from a cluster of closely related proteins including insect and vertebrate cytoplasmic actins and actins of mollusc, cestode, and nematode. We propose that muscle-specific actin genes have appeared independently at least twice during the evolution of animals: insect muscle actin genes have emerged from an ancestral cytoplasmic actin gene within the arthropod phylum, whereas vertebrate muscle actin genes evolved within the chordate lineage as previously described.

Actins

The compositional properties of human genes.

The present work represents the first attempt to study in greater detail previously proposed compositional correlations in genomes, based on a body of additional data relating to gene localizations as well as to extended flanking sequences extracted from gene banks. We have investigated the correlations that exist between (1) the GC levels of exons of human genes, and (2) the GC levels of either intergenic sequences or introns associated with the genes under consideration. In both cases, linear relationships with slopes close to unity were found. The similarity of the linear relationships indicates similar GC levels in intergenic sequences and introns located in the same isochores. Moreover, both intergenic sequences and introns showed GC levels 5-10% lower than the corresponding exons. The above findings considerably strengthen the previously drawn conclusion that coding and noncoding sequences (both inter- and intragenic) from the same isochores of the human genome are compositionally correlated. In addition, we find linear correlations between the GC levels of codon positions and of the intergenic sequences or introns associated with the corresponding genes, as well as among the GC levels of codon positions of genes.

Base Composition

Correlations between the compositional properties of human genes, codon usage, and amino acid composition of proteins.

We have analyzed the correlation that exists between the GC levels of third and first or second codon position for about 1400 human coding sequences. The linear relationship that was found indicates that the large differences in GC level of third codon positions of human genes are paralleled by smaller differences in GC levels of first and second codon positions. Whereas third codon position differences correspond to very large differences in codon usage within the human genome, the first and second codon position differences correspond to smaller, yet very remarkable, differences in the amino acid composition of encoded proteins. Because GC levels of codon positions are linearly correlated with the GC levels of the isochores harboring the corresponding genes, both codon usage and amino acid composition are different for proteins encoded by genes located in isochores of different GC levels. Furthermore, we have also shown that a linear relationship with a unit slope and a correlation coefficient of 0.77 exists between GC levels of introns and exons from the 238 human genes currently available for this analysis. Introns are, however, about 5% lower in GC, on average, than exons from the same genes.

Amino Acid Sequence

The distribution of genes in the human genome.

Previous investigations on the human genome determined: (i) the base compositions (GC levels) and the relative amounts of its isochore families; (ii) the compositional correlations (i.e., the correlations between GC levels) between third codon positions of a set of genes and the DNA fractions in which the genes were localized; and (iii) the compositional correlations between (a) third and first + second codon positions, as well as that between (b) introns and exons from the set of 'localized genes' and from all the coding sequences and genes (genomic sequences of exons + introns) available in gene banks. Here, we have shown that the correlations (iii, a and b) for 'localized genes' and genes from the bank are in full agreement, indicating that the former set is representative of the latter. We have then used the data (i) and the correlation (ii) to estimate the distribution of genes in isochore families. We have found that 34% of the genes are located in the GC-poor isochores (which represent 62% of the genome), 38% in the GC-rich isochores (31% of the genome) and 28% in the GC-richest isochores (3% of the genome). There is, therefore, a compositional gradient of gene concentration in the human genome. The gene density in the GC-richest 3% of the genome is about eight times higher than in the GC-rich 31%, and about 16 times higher than in the GC-poorest 62%.

Base Composition

Codon usage changes and sequence dissimilarity between human and rat.

This paper reports on the relationship between the number of silent differences and the codon usage changes in the lineages leading to human and rat. Examination of 102 pairs of homologous genes gives rise to four main conclusions: (1) We have previously demonstrated the existence of a codon usage change (called the minor shift) between human and rat; this was confirmed here with a larger sample. For genes with extreme C & G frequencies, the C & G level in the third codon position is less extreme in rat than in human. (2) Protein similarity and percentage of positive differences are the two main factors that discriminate homologous genes when characterized by differences between rat and human. By definition, positive differences result from silent changes between A or T and C or G with a direction implying a C & G content variation in the same direction as the overall gene variation. (3) For genes showing both codon usage change and low protein similarity, a majority of amino acid replacements contributes to C & G level variation in positions I and II in the same direction as the variation in position III. This is thus a new example of protein evolution due to constraints acting at the DNA level. (4) In heavy isochores (high C & G content) no direct correlation exists between codon usage change (measured by the dissymmetry of differences) and silent dissimilarity. In light isochores the opposite situation is observed: modification of codon usage is associated with a high synonymous dissimilarity. This result shows that, in some cases, modification of constrains acting at the DNA level could accelerate divergence between genomes.

Animals

The compositional distribution of coding sequences and DNA molecules in humans and murids.

The compositional distributions of coding sequences and DNA molecules (in the 50-100-kb range) are remarkably narrower in murids (rat and mouse) compared to humans (as well as to all other mammals explored so far). In murids, both distributions begin at higher and end at lower GC values. A comparison of homologous coding sequences from murids and humans revealed that their different compositional distributions are due to differences in GC levels in all three codon positions, particularly of genes located at both ends of the distribution. In turn, these differences are responsible for differences in both codon usage and amino acids. When GC levels at first + second codon positions and third codon positions, respectively, of murid genes are plotted against corresponding GC levels of homologous human genes, linear relationships (with very high correlation coefficients and slopes of about 0.78 and 0.60, respectively) are found. This indicates a conservation of the order of GC levels in homologous genes from humans and murids. (The same comparison for mouse and rat genes indicates a conservation of GC levels of homologous genes.) A similar linear relationship was observed when plotting GC levels of corresponding DNA fractions (as obtained by density gradient centrifugation in the presence of a sequence-specific ligand) from mouse and human. These findings indicate that orderly compositional changes affecting not only coding sequences but also noncoding sequences took place since the divergence of murids. Such directional fixations of mutations point to the existence of selective pressures affecting the genome as a whole.

Amino Acid Sequence

Compositional compartmentalization and gene composition in the genome of vertebrates.

The compositional distribution of coding sequences from five vertebrates (Xenopus, chicken, mouse, rat, and human) is shifted toward higher GC values compared to that of the DNA molecules (in the 35-85-kb size range) isolated from the corresponding genomes. This shift is due to the lower GC levels of intergenic sequences compared to coding sequences. In the cold-blooded vertebrate, the two distributions are similar in that GC-poor genes and GC-poor DNA molecules are largely predominant. In contrast, in the warm-blooded vertebrates, GC-rich genes are largely predominant over GC-poor genes, whereas GC-poor DNA molecules are largely predominant over GC-rich DNA molecules. As a consequence, the genomes of warm-blooded vertebrates show a compositional gradient of gene concentration. The compositional distributions of coding sequences (as well as of DNA molecules) showed remarkable differences between chicken and mammals, and between mouse (or rat) and human. Differences were also detected in the compositional distribution of housekeeping and tissue-specific genes, the former being more abundant among GC-rich genes.

Animals