Search PubMed⌕ Search

Biomedical subjects

Manolo Gouy

Publications and source records attributed to Manolo Gouy.

18 recordsLinked to original sources

HoSeqI: automated homologous sequence identification in gene family databases.

UNLABELLED: We present a web service allowing to automatically assign sequences to homologous gene families from a set of databases. After identification of the most similar gene family to the query sequence, this sequence is added to the whole alignment and the phylogenetic tree of the family is rebuilt. Thus, the phylogenetic position of the query sequence in its gene family can be easily identified. AVAILABILITY: http://pbil.univ-lyon1.fr/software/HoSeqI/.

Algorithms↗

Origin and molecular evolution of receptor tyrosine kinases with immunoglobulin-like domains.

Receptor tyrosine kinases (RTKs) are involved in the control of fundamental cellular processes in metazoans. In vertebrates, RTK could be grouped in distinct classes based on the nature of their cognate ligand and modular composition of their extracellular domain. RTK with immunoglobulin-like domains (IG-like RTK) encompass several RTK classes and have been found in early metazoans, including sponges. Evolution of IG-like RTK is characterized by extended molecular and functional diversification, which prompted us to study their evolutionary history. For that purpose, a nonredundant data set including annotated protein sequences of IG-like RTK (n = 85) was built, representing 19 species ranging from sponges to humans. Phylogenetic trees were generated from alignment of conserved regions using maximum likelihood approach. Molecular phylogeny strongly suggests that IG-like RTK diversification occurred according to a complex scenario. In particular, we propose that specific cis duplications of a common ancestor to both platelet-derived growth factor receptor (class III) and vascular endothelial growth factor receptor (class V) families preceded two trans duplications. In contrast, other IG-like RTK genes, like Musk and PTK7, apparently did not evolve by duplications, whereas fibroblast growth factor receptors (class IV) evolved through two rounds of trans duplications. The proposed model of IG-like RTK evolution is supported by high bootstrap values and by the clustering of genes encoding class III and class V RTKs at specific chromosomal locations in mouse and human genomes.

Animals↗

Efficient likelihood computations with nonreversible models of evolution.

Recent advances in heuristics have made maximum likelihood phylogenetic tree estimation tractable for hundreds of sequences. Noticeably, these algorithms are currently limited to reversible models of evolution, in which Felsenstein's pulley principle applies. In this paper we show that by reorganizing the way likelihood is computed, one can efficiently compute the likelihood of a tree from any of its nodes with a nonreversible model of DNA sequence evolution, and hence benefit from cutting-edge heuristics. This computational trick can be used with reversible models of evolution without any extra cost. We then introduce nhPhyML, the adaptation of the nonhomogeneous nonstationary model of Galtier and Gouy (1998; Mol. Biol. Evol. 15:871-879) to the structure of PhyML, as well as an approximation of the model in which the set of equilibrium frequencies is limited. This new version shows good results both in terms of exploration of the space of tree topologies and ancestral G+C content estimation. We eventually apply it to rRNA sequences slowly evolving sites and conclude that the model and a wider taxonomic sampling still do not plead for a hyperthermophilic last universal common ancestor.

Algorithms↗

In silico whole-genome scanning of cancer-associated nonsynonymous SNPs and molecular characterization of a dynein light chain tumour variant.

Last decade has led to the accumulation of large amounts of data on cancer genetics, opening an unprecedented access to the mapping of cancer genes in the human genome. Single-nucleotide polymorphisms (SNPs), the most common form of DNA variation in humans, emerge as an invaluable tool for cancer association studies. These genotypic markers can be used to assay how alleles of candidate genes correlate with the malignant phenotype, and may provide new clues into the genetic modifications that characterize cancer onset. In this cancer-oriented study, we detail an SNP mining strategy based on the analysis of expressed sequence tags among publicly available databases. Our whole-genome approach provides a comprehensive and unbiased description of nonsynonymous SNPs (nsSNPs) in tumoral versus normal tissues. To gain further insights into the possible relationships between genetic variation and altered phenotype, locations of a subset of nsSNPs were mapped onto protein domains known to be critical for protein function. Computational methods were also used to predict the potential impact of these cancer-associated nsSNPs on protein structure and function. We illustrate our approach through the detailed biochemical and structural characterization of a previously unknown cancer-associated mutation (G79C) affecting the 8 kDa dynein light chain (DNCL1).

Computational Biology↗

Phylogenomics of life-or-death switches in multicellular animals: Bcl-2, BH3-Only, and BNip families of apoptotic regulators.

In this report, we conducted a comprehensive survey of Bcl-2 family members, a divergent group of proteins that regulate programmed cell death by an evolutionarily conserved mechanism. Using comparative sequence analysis, we found novel sequences in mammals, nonmammalian vertebrates, and in a number of invertebrates. We then asked what conclusions could be drawn from phyletic distribution, intron/exon structures, sequence/structure relationships, and phylogenetic analyses within the updated Bcl-2 family. First, multidomain members having a sequence pattern consistent with the conservation of the Bcl-X(L)/Bax/Bid topology appear to be restricted to multicellular animals and may share a common ancestry. Next, BNip proteins, which were originally identified based on their ability to bind to E1B 19K/Bcl-2 proteins, form three independent monophyletic branches with different evolutionary history. Lastly, a set of Bcl-2 homology 3-only proteins with unrelated secondary structures seems to have evolved after the origin of Metazoa and exhibits diverse expansion after speciation during vertebrate evolution.

Amino Acid Sequence↗

Tree pattern matching in phylogenetic trees: automatic search for orthologs or paralogs in homologous gene sequence databases.

MOTIVATION: Comparative sequence analysis is widely used to study genome function and evolution. This approach first requires the identification of homologous genes and then the interpretation of their homology relationships (orthology or paralogy). To provide help in this complex task, we developed three databases of homologous genes containing sequences, multiple alignments and phylogenetic trees: HOBACGEN, HOVERGEN and HOGENOM. In this paper, we present two new tools for automating the search for orthologs or paralogs in these databases. RESULTS: First, we have developed and implemented an algorithm to infer speciation and duplication events by comparison of gene and species trees (tree reconciliation). Second, we have developed a general method to search in our databases the gene families for which the tree topology matches a peculiar tree pattern. This algorithm of unordered tree pattern matching has been implemented in the FamFetch graphical interface. With the help of a graphical editor, the user can specify the topology of the tree pattern, and set constraints on its nodes and leaves. Then, this pattern is compared with all the phylogenetic trees of the database, to retrieve the families in which one or several occurrences of this pattern are found. By specifying ad hoc patterns, it is therefore possible to identify orthologs in our databases.

Algorithms↗

Molecular phylogeny of Myricaceae: a reexamination of host-symbiont specificity.

The phylogeny of 13 species of Myricaceae, the most ancient actinorhizal family involved in a nitrogen-fixing symbiosis with the actinomycete Frankia, was established by the analysis of their rbcL gene and 18S-26S ITS. The phylogenetic position of those species was then compared to their specificity of association with Frankia in their natural habitat and to their nodulation potential determined on greenhouse-grown seedlings. The results showed that Genus Myrica, including M. gale and M. hartwegii, and Genus Comptonia, including C. peregrina, belong to a phylogenetic cluster distinct from the other Myrica species transferred in a new genus, Morella. This grouping parallels the natural specificity of each cluster with Comptonia-Myrica and Morella being nodulated by two phylogenetically divergent clusters of Frankia strains, the Alnus and Elaeagnaceae-infective strains clusters, respectively. Under laboratory conditions, Comptonia and Morella had a nodulation potential larger than under natural conditions. From this study it appears that the Myricaceae are split into two different specificity groups. It can be hypothesized that the early divergence of the genera led to the selection of genetically diverse Frankia strains which is contradictory to the earlier proposal that evolution has proceeded toward narrower promiscuity within the family.

Evolution, Molecular↗

Horizontal transfer of two operons coding for hydrogenases between bacteria and archaea.

Using a phylogenetic approach, we discovered three putative horizontal transfers between bacterial and archaeal species involving large clusters of genes. One transfer involves an operon of 13 genes, called mbx, which probably was transferred into the genome of Thermotoga maritima from a species belonging or close to the Pyrococcus genus. The two others implied an operon of six genes, called ech, transferred independently to the genomes of Thermoanaerobacter tengcongensis and Desulfovibrio gigas, from a species belonging or close to the Methanosarcina genus. All these transfers affected operons coding for multisubunit membrane-bound (NiFe) hydrogenases involved in the energy metabolism of the donor genomes. The functionality of the transferred operons has not been experimentally demonstrated for T. maritima, whereas in D. gigas and T. tengcongensis the encoded multisubunit hydrogenase could have a role in energy conservation. This report adds several cases of horizontal gene transfers among hydrogenases already described.

Archaea↗

Clinical and environmental isolates of Legionella pneumophila serogroup 1 cannot be distinguished by sequence analysis of two surface protein genes and three housekeeping genes.

We used gene sequencing to determine whether clinical (sporadic, epidemic, and endemic) and environmental isolates of Legionella pneumophila serogroup (sg) 1 belong to specific lineages. A total of 178 clinical and environmental L. pneumophila sg 1 isolates, defined by pulsed-field gel electrophoresis and epidemiological data as sporadic, epidemic, or endemic, were analyzed for polymorphisms in five gene fragments. The fragments belonged to three housekeeping genes (coding for aconitase [acn], aspartate-beta-semialdehyde dehydrogenase [asd], and RNA polymerase beta subunit [rpoB]) and two surface protein genes (coding for the macrophage infectivity potentiator [mip] and the major outer membrane protein [mompS]). The phylogenetic tree inferred from sequence polymorphisms of the five genes identified two large clusters, one consisting of 133 poorly differentiated strains and containing two smaller clusters (10 and 2 strains) unrelated to each other and the other consisting of 42 strains. Clinical and environmental isolates could not be distinguished on this basis, and no link between genetic background and epidemiological type was found, suggesting that other factors are responsible for differences in pathogenicity.

Bacterial Proteins↗

Invertebrate data predict an early emergence of vertebrate fibrillar collagen clades and an anti-incest model.

Fibrillar collagens are involved in the formation of striated fibrils and are present from the first multicellular animals, sponges, to humans. Recently, a new evolutionary model for fibrillar collagens has been suggested (Boot-Handford, R. P., Tuckwell, D. S., Plumb, D. A., Farrington Rock, C., and Poulsom, R. (2003) J. Biol. Chem. 278, 31067-31077). In this model, a rare genomic event leads to the formation of the founder vertebrate fibrillar collagen gene prior to the early vertebrate genome duplications and the radiation of the vertebrate fibrillar collagen clades (A, B, and C). Here, we present the modular structure of the fibrillar collagen chains present in different invertebrates from the protostome Anopheles gambiae to the chordate Ciona intestinalis. From their modular structure and the use of a triple helix instead of C-propeptide sequences in phylogenetic analyses, we were able to show that the divergence of A and B clades arose early during evolution because alpha chains related to these clades are present in protostomes. Moreover, the event leading to the divergence of B and C clades from a founder gene arose before the appearance of vertebrates; altogether these data contradict the Boot-Handford model. Moreover, they indicate that all the key steps required for the formation of fibrils of variable structure and functionality arose step by step during invertebrate evolution.

Amino Acid Sequence↗

Phylogenetic study and identification of Vibrio splendidus-related strains based on gyrB gene sequences.

Different strains related to Vibrio splendidus have been associated with infection of aquatic animals. An epidemiological study of V. splendidus strains associated with Crassostrea gigas mortalities demonstrated genetic diversity within this group and suggested its polyphyletic nature. Recently 4 species, V. lentus, V. chagasii, V. pomeroyi and V. kanaloae, phenotypically related to V. splendidus, have been described, although biochemical methods do not clearly discriminate species within this group. Here, we propose a polyphasic approach to investigate their taxonomic relationships. Phylogenetic analysis of V. splendidus-related strains was carried out using the nucleotide sequences of 16S ribosomal DNA (16S rDNA) and gyrase B subunit (gyrB) genes. Species delineation based on 16S rDNA-sequencing is limited because of divergence between cistrons, roughly equivalent to divergence between strains. Despite a high level of sequence similarity, strains were separated into 2 clades. In the phylogenetic tree constructed on the basis of gyrB gene sequences, strains were separated into 5 independent clusters containing V. splendidus, V. lentus, V. chagasii-type strains and a putative new genomic species. This phylogenetic grouping was almost congruent with that based on DNA-DNA hybridisation analysis. V. pomeroyi, V. kanaloae and V. tasmaniensis-type strains clustered together in a fifth clade. The gyrB gene-sequencing approach is discussed as an alternative for investigating the taxonomy of Vibrio species.

Animals↗

Phylogenetic analysis of the complete genome sequence of Encephalitozoon cuniculi supports the fungal origin of microsporidia and reveals a high frequency of fast-evolving genes.

Microsporidia are unicellular eukaryotes living as obligate intracellular parasites. Lacking mitochondria, they were initially considered as having diverged before the endosymbiosis at the origin of mitochondria. That microsporidia were primitively amitochondriate was first questioned by the discovery of microsporidial sequences homologous to genes encoding mitochondrial proteins and then refuted by the identification of remnants of mitochondria in their cytoplasm. Various molecular phylogenies also cast doubt on the early divergence of microsporidia, these organisms forming a monophyletic group with or within the fungi. The 2001 proteins putatively encoded by the complete genome of Encephalitozoon cuniculi provided powerful data to test this hypothesis. Phylogenetic analysis of 99 proteins selected as adequate phylogenetic markers indicated that the E. cuniculi sequences having the lowest evolutionary rates preferentially clustered with fungal sequences or, more rarely, with both animal and fungal sequences. Because sequences with low evolutionary rates are less sensitive to the long-branch attraction artifact, we concluded that microsporidia are evolutionarily related to fungi. This analysis also allowed comparing the accuracy of several phylogenetic algorithms for a fast-evolving lineage with real rather than simulated sequences.

Animals↗

Integrated databanks access and sequence/structure analysis services at the PBIL.

The World Wide Web server of the PBIL (Pôle Bioinformatique Lyonnais) provides on-line access to sequence databanks and to many tools of nucleic acid and protein sequence analyses. This server allows to query nucleotide sequence banks in the EMBL and GenBank formats and protein sequence banks in the SWISS-PROT and PIR formats. The query engine on which our data bank access is based is the ACNUC system. It allows the possibility to build complex queries to access functional zones of biological interest and to retrieve large sequence sets. Of special interest are the unique features provided by this system to query the data banks of gene families developed at the PBIL. The server also provides access to a wide range of sequence analysis methods: similarity search programs, multiple alignments, protein structure prediction and multivariate statistics. An originality of this server is the integration of these two aspects: sequence retrieval and sequence analysis. Indeed, thanks to the introduction of re-usable lists, it is possible to perform treatments on large sets of data. The PBIL server can be reached at: http://pbil.univ-lyon1.fr.

Databases, Genetic↗

Automatic RNA secondary structure prediction with a comparative approach.

This paper presents an algorithm, DCFold, that automatically predicts the common secondary structure of a set of aligned homologous RNA sequences. It is based on the comparative approach. Helices are searched in one of the sequences, called the 'target sequence', and compared to the helices in the other sequences, called the 'test sequences'. Our algorithm searches in the target sequence for palindromes that have a high probability to define helices that are conserved in the test sequences. This selection of significant palindromes is based on criteria that take into account their length and their mutation rate. A recursive search of helices, starting from these likely ones, is implemented using the 'divide and conquer' approach. Indeed, as pseudo-knots are not searched by DCFold, a selected palindrome (p, p') makes possible to divide the initial sequence into two sequences, the internal one and the one resulting from the concatenation of the two external ones. New palindromes can be searched independently in these subsequences. This algorithm was run on ribosomal RNA sequences and recovered very efficiently their common secondary structures.

Algorithms↗

Functional and evolutionary analysis of a eukaryotic parasitic genome.

The DNA sequences of the 11 linear chromosomes of the approximately 2.9 Mbp genome of Encephalitozoon cuniculi, an obligate intracellular parasite of mammals, include approximately 2000 putative protein-coding genes. The compactness of this genome is associated with the length reduction of various genes. Essential functions are dependent on a minimal set of genes. Phylogenetic analysis supports the hypotheses that microsporidia are related to fungi and have retained a mitochondrion-derived organelle, the mitosome.

Animals↗

A phylogenomic approach to bacterial phylogeny: evidence of a core of genes sharing a common history.

It has been claimed that complete genome sequences would clarify phylogenetic relationships between organisms, but up to now, no satisfying approach has been proposed to use efficiently these data. For instance, if the coding of presence or absence of genes in complete genomes gives interesting results, it does not take into account the phylogenetic information contained in sequences and ignores hidden paralogies by using a BLAST reciprocal best hit definition of orthology. In addition, concatenation of sequences of different genes as well as building of consensus trees only consider the few genes that are shared among all organisms. Here we present an attempt to use a supertree method to build the phylogenetic tree of 45 organisms, with special focus on bacterial phylogeny. This led us to perform a phylogenetic study of congruence of tree topologies, which allows the identification of a core of genes supporting similar species phylogeny. We then used this core of genes to infer a tree. This phylogeny presents several differences with the rRNA phylogeny, notably for the position of hyperthermophilic bacteria.

Computational Biology↗

Recombination rate and the distribution of transposable elements in the Drosophila melanogaster genome.

We analyzed the distribution of 54 families of transposable elements (TEs; transposons, LTR retrotransposons, and non-LTR retrotransposons) in the chromosomes of Drosophila melanogaster, using data from the sequenced genome. The density of LTR and non-LTR retrotransposons (RNA-based elements) was high in regions with low recombination rates, but there was no clear tendency to parallel the recombination rate. However, the density of transposons (DNA-based elements) was significantly negatively correlated with recombination rate. The accumulation of TEs in regions of reduced recombination rate is compatible with selection acting against TEs, as selection is expected to be weaker in regions with lower recombination. The differences in the relationship between recombination rate and TE density that exist between chromosome arms suggest that TE distribution depends on specific characteristics of the chromosomes (chromatin structure, distribution of other sequences), the TEs themselves (transposition mechanism), and the species (reproductive system, effective population size, etc.), that have differing influences on the effect of natural selection acting against the TE insertions.

Animals↗

Sequence of the Small Subunit Ribosomal RNA Gene of Perkinsus atlanticus-like Isolated from Carpet Shell Clam in Galicia, Spain.

Parasites identified as Perkinsus atlanticus have been reported infecting carpet shell clams in Galicia (northwest Spain). We have sequenced the 18S ribosomal RNA gene of in vitro cultured Perkinsus atlanticus-like or hypnospores from diseased clams, and compared it with the same genomic region from P. marinus and Perkinsus sp. We have also compared the sequence of internal transcribed spacer (ITS) 1, ITS 2, and 5.8S rRNA from our isolate with the P. atlanticus GenBank sequence. The phylogenetic analysis of our cultured parasite based on the 18S gene led us to conclude that this isolate is not related to the genus Perkinsus but to the protists Anurofeca, Ichthyophonus, and Psorospermium, located near the animal-fungal divergence. These last two genera have been included, together with Dermocystidium, in the newly described DRIPs (Dermocystidium, rossete agent, Ichthyophonus, and Psorospermium) clade, recently named Mesomycetozoa.

Journal Article↗