Search PubMedSearch

Biomedical subjects

M Gouy

Publications and source records attributed to M Gouy.

At least 19 recordsLinked to original sources

A nonhyperthermophilic common ancestor to extant life forms.

The G+C nucleotide content of ribosomal RNA (rRNA) sequences is strongly correlated with the optimal growth temperature of prokaryotes. This property allows inference of the environmental temperature of the common ancestor to all life forms from knowledge of the G+C content of its rRNA sequences. A model of sequence evolution, assuming varying G+C content among lineages and unequal substitution rates among sites, was devised to estimate ancestral base compositions. This method was applied to rRNA sequences of various species representing the major lineages of life. The inferred G+C content of the common ancestor to extant life forms appears incompatible with survival at high temperature. This finding challenges a widely accepted hypothesis about the origin of life.

Animals

Microsporidian Encephalitozoon cuniculi, a unicellular eukaryote with an unusual chromosomal dispersion of ribosomal genes and a LSU rRNA reduced to the universal core.

Microsporidia are eukaryotic parasites lacking mitochondria, the ribosomes of which present prokaryote-like features. In order to better understand the structural evolution of rRNA molecules in microsporidia, the 5S and rDNA genes were investigated in Encephalitozoon cuniculi . The genes are not in close proximity. Non-tandemly arranged rDNA units are on every one of the 11 chromosomes. Such a dispersion is also shown in two other Encephalitozoon species. Sequencing of the 5S rRNA coding region reveals a 120 nt long RNA which folds according to the eukaryotic consensus structural shape. In contrast, the LSU rRNA molecule is greatly reduced in length (2487 nt). This dramatic shortening is essentially due to truncation of divergent domains, most of them being removed. Most variable stems of the conserved core are also deleted, reducing the LSU rRNA to only those structural features preserved in all living cells. This suggests that the E.cuniculi LSU rRNA performs only the basic mechanisms of translation. LSU rRNA phylogenetic analysis with the BASEML program favours a relatively recent origin of the fast evolving microsporidian lineage. Therefore, the prokaryote-like ribosomal features, such as the absence of ITS2, may be derived rather than primitive characters.

Animals

The non-redundant Bacillus subtilis (NRSub) database: update 1998.

The non-redundant Bacillus subtilis database (NRSub) has been developed in the context of the sequencing project devoted to this bacterium. As this project has reached completion, the whole genome is now available as a single contig. Thanks to the ACNUC database management system and its associated retrieval system Query_win, each functional region of the genome can be accessed individually. Extra annotations have been added such as accession numbers for the genes, locations on the genetic map, codon adaptation index values, as well as cross-references with other collections. NRSub is distributed through anonymous FTP as a text file in EMBL format and as an ACNUC database. It is also possible to access NRSub through two dedicated World Wide Web servers located in France (http://acnuc. univ-lyon1.fr/nrsub/nrsub.html ) and in Japan (http://ddbjs4h.genes. nig.ac.jp/ ).

Bacillus subtilis

Microsporidia, amitochondrial protists, possess a 70-kDa heat shock protein gene of mitochondrial evolutionary origin.

An intronless gene encoding a protein of 592 amino acid residues with similarity to 70-kDa heat shock proteins (HSP70s) has been cloned and sequenced from the amitochondrial protist Encephalitozoon cuniculi (phylum Microsporidia). Southern blot analyses show the presence of a single gene copy located on chromosome XI. The encoded protein exhibits an N-terminal hydrophobic leader sequence and two motifs shared by proteobacterial and mitochondrially expressed HSP70 homologs. Phylogenetic analysis using maximum likelihood and evolutionary distances place the E. cuniculi sequence in the cluster of mitochondrially expressed HSP70s, with a higher evolutionary rate than those of homologous sequences. Similar results were obtained after cloning a fragment of the homologous gene in the closely related species E. hellem. The presence of a nuclear targeting signal-like sequence supports a role of the Encephalitozoon HSP70 as a molecular chaperone of nuclear proteins. No evidence for cytosolic or endoplasmic reticulum forms of HSP70 was obtained through PCR amplification. These data suggest that Encephalitozoon species have evolved from an ancestor bearing mitochondria, which is in disagreement with the postulated presymbiotic origin of Microsporidia. The specific role and intracellular localization of the mitochondrial HSP70-like protein remain to be elucidated.

Amino Acid Sequence

Inferring pattern and process: maximum-likelihood implementation of a nonhomogeneous model of DNA sequence evolution for phylogenetic analysis.

A nonhomogeneous, nonstationary stochastic model of DNA sequence evolution allowing varying equilibrium G + C contents among lineages is devised in order to deal with sequences of unequal base compositions. A maximum-likelihood implementation of this model for phylogenetic analyses allows handling of a reasonable number of sequences. The relevance of the model and the accuracy of parameter estimates are theoretically and empirically assessed, using real or simulated data sets. Overall, a significant amount of information about past evolutionary modes can be extracted from DNA sequences, suggesting that process (rates of distinct kinds of nucleotide substitutions) and pattern (the evolutionary tree) can be simultaneously inferred. G + C contents at ancestral nodes are quite accurately estimated. The new method appears to be useful for phylogenetic reconstruction when base composition varies among compared sequences. It may also be suitable for molecular evolution studies.

Algorithms

Sensitivity of the relative-rate test to taxonomic sampling.

Relative-rate tests may be used to compare substitution rates between more than two sequences, which yields two main questions: What influence does the number of sequences have on relative-rate tests and what is the influence of the sampling strategy as characterized by the phylogenetic relationships between sequences? Using both simulations and analysis of real data from murids (APRT and LCAT nuclear genes), we show that comparing large numbers of species significantly improves the power of the test. This effect is stronger if species are more distantly related. On the other hand, it appears to be less rewarding to increase outgroup sampling than to use the single nearest outgroup sequence. Rates may be compared between paraphyletic ingroups and using paraphyletic outgroups, but unbalanced taxonomic sampling can bias the test. We present a simple phylogenetic weighting scheme which takes taxonomic sampling into account and significantly improves the relative-rate test in cases of unbalanced sampling. The answers are thus: (1) large taxonomic sampling of compared groups improves relative-rate tests, (2) sampling many outgroups does not bring significant improvement, (3) the only constraint on sampling strategy is that the outgroup be valid, and (4) results are more accurate when phylogenetic relationships between the investigated sequences are taken into account. Given current limitations of the maximum-likelihood and nonparametric approaches, the relative-rate test generalized to any number of species with phylogenetic weighting appears to be the most general test available to compare rates between lineages.

Adenine Phosphoribosyltransferase

Evolutionary affinities of the order Perissodactyla and the phylogenetic status of the superordinal taxa Ungulata and Altungulata.

Contrary to morphological claims, molecular data indicate that the order Perissodactyla (e.g., horses, rhinoceroses, and tapirs) is neither part of the superordinal taxon Paenungulata (Sirenia, Proboscidea, and Hyracoidea) nor an immediate outgroup of the paenungulates. Rather, Perissodactyla is closer to Carnivora and Cetartiodactyla (Cetacea+Artiodactyla) than it is to the paenungulates. Therefore, two morphologically defined superordinal taxa, Altungulata (Proboscidea, Sirenia, Hyracoidea, and Perissodactyla) and Ungulata (Altungulata and Cetartiodactyla), are invalidated. Perissodactyla, Carnivora, and Cetartiodactyla are shown to constitute a rather tight trichotomy. However, a molecular analysis of 36 protein sequences with a total concatenated length of 7885 aligned amino acids indicates that Perissodactyla is closer to Cetartiodactyla than either taxa is to Carnivora. The relationships among Paenungulata, Primates, and the clade consisting of Perissodactyla, Carnivora, and Cetartiodactylaa could not be resolved on the basis of the available data.

Amino Acid Sequence

Evolutionary distances between nucleotide sequences based on the distribution of substitution rates among sites as estimated by parsimony.

The rate of evolution of macromolecules such as ribosomal RNAs and proteins varies along the molecule because structural and functional constraints differ between sites. Many studies have shown that ignoring this variation in computing evolutionary distances leads to severe underestimation of sequence divergences, and thus can lead to misleading evolutionary tree inferences. We propose here a new parsimony-based method for computing evolutionary distances between pairs of sequences that takes into account this variation and estimates it from the data. This method applies to the number of substitutions per site in ribosomal RNA genes as well as to the number of nonsynonymous substitutions per codon for protein-coding genes and is especially suitable when large data sets (> or = 100 sequences) are analyzed. First, starting from a phylogeny constructed with usual distances, the maximum-parsimony method is used to infer the distribution of the number of substitutions that have occurred at each site (or codon) along this tree. This distribution is then fitted to an "invariant + truncated negative binomial" distribution that allows for invariant sites. Maximum-likelihood fitting of this distribution to different data sets showed that it agreed very well with real data. Noticeably, allowing for invariant sites seemed to be very important. Finally, two distance estimates were developed by introducing the distribution of site variability into the substitution models of Jukes and Cantor and of Kimura. The use of different numbers of aligned sequences (up to 1,000 rRNA sequences) showed that the parameters of the model are very sensitive to the number of sequences used to estimate them. However, if at least 100 sequences are considered, the two new distance estimates are quite stable with respect to the number of sequences used to fit the distribution. This stability is true for low as well as for high evolutionary distances. These new distances appeared to be much better estimates of the number of substitutions per site than the classical distances of Jukes and Cantor and of Kimura, which both greatly underestimate this number, so that they can serve as indexes to detect saturation. We conclude that the new distances are particularly suitable for phylogenetic analysis when very distantly related species and relatively large data sets are considered. Trees reconstructed using these distances are generally different from those constructed by means of the classical estimates. Using this new method, we showed that the mean evolutionary distance between Prokaryotes and Eukaryotes is substantially higher for the small-subunit than for the large-subunit rRNAs. This suggests than the former might have experienced a drastic change during the early evolution of Eukaryotes.

Algorithms

Extreme differences in rates of molecular evolution of foraminifera revealed by comparison of ribosomal DNA sequences and the fossil record.

Foraminifera have one of the best known fossil records among the unicellular eukaryotes. However, the origin and phylogenetic relationships of the extant foraminiferal lineages are poorly understood. To test the current paleontological hypotheses on evolution of foraminifera, we sequenced about 1,000 base pairs from the 3' end of the small subunit rRNA gene (SSU rDNA) in 22 species representing all major taxonomic groups. Phylogenies were derived using neighbor-joining, maximum-parsimony, and maximum-likelihood methods. All analyses confirm the monophyletic origin of foraminifera. Evolutionary relationships within foraminifera inferred from rDNA sequences, however, depend on the method of tree building and on the choice of analyzed sites. In particular, the position of planktonic foraminifera shows important variations. We have shown that these changes result from the extremely high rate of rDNA evolution in this group. By comparing the number of substitutions with the divergence times inferred from the fossil record, we have estimated that the rate of rDNA evolution in planktonic foraminifera is 50 to 100 times faster than in some benthic foraminifera. The use of the maximum-likelihood method and limitation of analyzed sites to the most conserved parts of the SSU rRNA molecule render molecular and paleontological data generally congruent.

Animals

Phylogenetic position of the order Lagomorpha (rabbits, hares and allies)

Ever since they have been classified as ruminants in the Old Testament (Leviticus 11:6, Deuteronomy 14:7) and equated with hyraxes in the vulgate Latin translation, rabbits and their relatives (order Lagomorpha) have frequently experienced radical changes in taxonomic rank. By using 91 orthologous protein sequences, we have attempted to answer the classical question "What, if anything, is a rabbit?". Here we show that Lagomorpha is significantly more closely related to Primates and Scandentia (tree shrews) than it is to rodents. This newly determined phylogenetic position invalidates the superordinal taxon Glires (Lagomorpha + Rodentia), and indicates that the morphological 'synapomorphies' previously used to cluster rodents and lagomorphs into Glires, may actually represent symplesiomorphies or homoplasies that are of no phylogenetic value. This raises the possibility that the ancestral eutherian morphotype may have possessed many rodent-like morphological characters.

Animals

WWW-query: an on-line retrieval system for biological sequence banks.

We have developed a World Wide Web (WWW) version of the sequence retrieval system Query: WWW-Query. This server allows to query nucleotide sequence banks in the EMBL/GenBank/DDBJ formats and protein sequence banks in the NBRF/PIR format. WWW-Query includes all the features of the on-line sequences browsers already available: possibility to build complex queries, integration of cross-references with different data banks, and access to the functional zones of biological interest. It also provides original services not available elsewhere: introduction of the notion of re-usable sequence lists, integration of dedicated helper applications for visualizing alignments and phylogenetic trees and links with multivariate methods for studying codon usage or for complementing phylogenies.

Amino Acid Sequence

SEAVIEW and PHYLO_WIN: two graphic tools for sequence alignment and molecular phylogeny.

SEAVIEW and PHYLO_WIN are two graphic tools for X Windows-Unix computers dedicated to sequence alignment and molecular phylogenetics. SEAVIEW is a sequence alignment editor allowing manual or automatic alignment through an interface with CLUSTALW program. Alignment of large sequences with extensive length differences is made easier by a dot-plot-based routine. The PHYLO_WIN program allows phylogenetic tree building according to most usual methods (neighbor joining with numerous distance estimates, maximum parsimony, maximum likelihood), and a bootstrap analysis with any of them. Reconstructed trees can be drawn, edited, printed, stored, evaluated according to numerous criteria. Taxonomic species groups and sets of conserved regions can be defined by mouse and stored into sequence files, thus avoiding multiple data files. Both tools are entirely mouse driven. On-line help makes them easy to use. They are freely available by anonymous ftp at biom3.univ-lyon1.fr/pub/ mol_phylogeny or http:@acnuc.univ-lyon1.fr/, or by e-mail to galtier@biomserv.univ-lyon1.fr.

Animals

Early origin of foraminifera suggested by SSU rRNA gene sequences.

Foraminifera are one of the largest groups of unicellular eukaryotes with probably the best known fossil record. However, the origin of foraminifera and their phylogenetic relationships with other eukaryotes are not well established. In particular, two recent reports, based on ribosomal RNA gene sequences, have reached strikingly different conclusions about foraminifera's evolutionary position within eukaryotes. Here, we present the complete small subunit (SSU) rRNA gene sequences of three species of foraminifera. Phylogenetic analysis of these sequences indicates that they branch very deeply in the eukaryotic evolutionary tree: later than those of the amitochondrial Archezoa, but earlier than those of the Euglenozoa and other mitochondria-bearing phyla. Foraminifera are clearly among the earliest eukaryotes with mitochondria, but because of the peculiar nature of their SSU genes we cannot be certain that they diverged first, as our data suggest.

Animals

Inferring phylogenies from DNA sequences of unequal base compositions.

A new method for computing evolutionary distances between DNA sequences is proposed. Contrasting with classical methods, the underlying model does not assume that sequence base compositions (A, C, G, and T contents) are at equilibrium, thus allowing unequal base compositions among compared sequences. This makes the method more efficient than the usual ones in recovering phylogenetic trees from sequence data when base composition is heterogeneous within the data set, as we show by using both simulated and empirical data. When applied to small-subunit ribosomal RNA sequences from several prokaryotic or eukaryotic organisms, this method provides evidence for an early divergence of the microsporidian Vairimorpha necatrix in the eukaryotic lineage.

Algorithms

Isolation and characterization of a cDNA encoding a chicken actin-like protein.

We report the isolation and characterization of a chicken cDNA which putatively encodes an actin-like protein (chACTL). This 394-amino-acid (aa) polypeptide shares sequence homology (81, 70 and 67% identical aa, respectively) with three actin-related proteins (ARP) described for Drosophila melanogaster (ARP14D), Caenorhabditis elegans (ACTL) and Saccharomyces cerevisiae (ACT2). At least six chACTL transcripts were detected in different tissues during chick embryogenesis. Sequence analysis suggests that at least three groups of ARP have been evolutionarily conserved.

Actins

NRSub: a non-redundant data base for the Bacillus subtilis genome.

We have organized the DNA sequences of Bacillus subtillis from the EMBL collection to build the NRSub data base. This data base is free from duplications and all detected overlapping sequences are merged into contigs. Data on gene mapping and codon usage are also included. NRSub is publically available through anonymous FTP in flat file format or structured on the form of an ACNUC data base. Under this format, it is possible to use NRSub with the retrieval program Query--win. This program integrates a graphical interface and may be installed on any kind of UNX computer under X Window and on which the Vibrant and Motif libraries are available.

Bacillus subtilis

HOVERGEN: a database of homologous vertebrate genes.

Comparison of homologous genes is a major step for many studies related to genome structure, function or evolution. Similarity search programs easily find genes homologous to a given sequence. However, only very tedious manual procedures allow the retrieval of all sets of homologous genes sequenced for a given set of species. Moreover, this search often generates errors due to the complexity of data to be managed simultaneously: phylogenetic trees, alignments, taxonomy, sequences and related information. HOVERGEN helps to solve these problems by integrating all this information. HOVERGEN corresponds to GenBank sequences from all vertebrate species, with some data corrected, clarified, or completed, notably to address the problem of redundancy. Coding sequences have been classified in gene families. Protein multiple alignments and phylogenetic trees have been calculated for each family. Sequences and related information have been structured in an ACNUC database which permits complex selections. A graphical interface has been developed to visualize and edit trees. Genes are displayed in color, according to their taxonomy. Users have directly access to all information attached to sequences and to multiple alignments simply by clicking on genes. This graphical tool gives thus a rapid and simple access to all data necessary to interpret homology relationships between genes. HOVERGEN allows the user to easily select sets of homologous vertebrate genes, and thus is particularly useful for comparative sequence analysis, or molecular evolution studies.

Animals