Search PubMed⌕ Search

Biomedical subjects

Guy Perrière

Publications and source records attributed to Guy Perrière.

At least 19 recordsLinked to original sources

Integrating transcription factor binding site information with gene expression datasets.

MOTIVATION: Microarrays are widely used to measure gene expression differences between sets of biological samples. Many of these differences will be due to differences in the activities of transcription factors. In principle, these differences can be detected by associating motifs in promoters with differences in gene expression levels between the groups. In practice, this is hard to do. RESULTS: We combine correspondence analysis, between group analysis and co-inertia analysis to determine which motifs, from a database of promoter motifs, are strongly associated with differences in gene expression levels. Given a database of motifs and gene expression levels from a set of arrays, the method produces a ranked list of motifs associated with any specified split in the arrays. We give an example using the Gene Atlas compendium of gene expression levels for human tissues where we search for motifs that are associated with expression in central nervous system (CNS) or muscle tissues. Most of the motifs that we find are known from previous work to be strongly associated with expression in CNS or muscle. We give a second example using a published prostate cancer dataset where we can simply and clearly find which transcriptional pathways are associated with differences between benign and metastatic samples. AVAILABILITY: The source code is freely available upon request from the authors.

Algorithms↗

HoSeqI: automated homologous sequence identification in gene family databases.

UNLABELLED: We present a web service allowing to automatically assign sequences to homologous gene families from a set of databases. After identification of the most similar gene family to the query sequence, this sequence is added to the whole alignment and the phylogenetic tree of the family is rebuilt. Thus, the phylogenetic position of the query sequence in its gene family can be easily identified. AVAILABILITY: http://pbil.univ-lyon1.fr/software/HoSeqI/.

Algorithms↗

Physiological oxygenation status is required for fully differentiated phenotype in kidney cortex proximal tubules.

Hypoxia has been suspected to trigger transdifferentiation of renal tubular cells into myofibroblasts in an epithelial-to-mesenchymal transition (EMT) process. To determine the functional networks potentially altered by hypoxia, rat renal tubule suspensions were incubated under three conditions of oxygenation ranging from normoxia (lactate uptake) to severe hypoxia (lactate production). Transcriptome changes after 4 h were analyzed on a high scale by restriction fragment differential display. Among 1,533 transcripts found, 42% were maximally expressed under severe hypoxia and 8% under mild hypoxia (Po(2) = 48 mmHg), suggesting two different levels of oxygen sensing. Normoxia was required for full expression of the proximal tubule-specific transcripts 25-hydroxyvitamin D 1-hydroxylase (Cyp27b1) and l-pyruvate kinase (Pklr), transcripts involved in tissue cohesion such as fibronectin (Fn1) and N-cadherin (Cdh2), and non-muscle-type myosin transcripts. Mild hypoxia increased myogenin transcript level. Conversely, severe hypoxia increased transcripts involved in extracellular matrix remodeling, those of muscle-type myosins, and others involved in creatine phosphate synthesis and lactate transport (Slc16a7). Accordingly, microscopy showed loss of tubule aggregation under hypoxia, without tubular disruption. Hypoxia also increased the levels of kidney-specific transcripts normally restricted to the less oxygenated medullary zone and others specific for the distal part of the nephron. We conclude that extensive oxygen supply to the kidney tubule favors expression of its differentiated functions specifically in the proximal tubule, whose embryonic origin is mesenchymal. The phenotype changes could potentially permit transient adaptation to hypoxia but also favor pathological processes such as tissue invasion.

Animals↗

Origin and molecular evolution of receptor tyrosine kinases with immunoglobulin-like domains.

Receptor tyrosine kinases (RTKs) are involved in the control of fundamental cellular processes in metazoans. In vertebrates, RTK could be grouped in distinct classes based on the nature of their cognate ligand and modular composition of their extracellular domain. RTK with immunoglobulin-like domains (IG-like RTK) encompass several RTK classes and have been found in early metazoans, including sponges. Evolution of IG-like RTK is characterized by extended molecular and functional diversification, which prompted us to study their evolutionary history. For that purpose, a nonredundant data set including annotated protein sequences of IG-like RTK (n = 85) was built, representing 19 species ranging from sponges to humans. Phylogenetic trees were generated from alignment of conserved regions using maximum likelihood approach. Molecular phylogeny strongly suggests that IG-like RTK diversification occurred according to a complex scenario. In particular, we propose that specific cis duplications of a common ancestor to both platelet-derived growth factor receptor (class III) and vascular endothelial growth factor receptor (class V) families preceded two trans duplications. In contrast, other IG-like RTK genes, like Musk and PTK7, apparently did not evolve by duplications, whereas fibroblast growth factor receptors (class IV) evolved through two rounds of trans duplications. The proposed model of IG-like RTK evolution is supported by high bootstrap values and by the clustering of genes encoding class III and class V RTKs at specific chromosomal locations in mouse and human genomes.

Animals↗

Optimized between-group classification: a new jackknife-based gene selection procedure for genome-wide expression data.

BACKGROUND: A recent publication described a supervised classification method for microarray data: Between Group Analysis (BGA). This method which is based on performing multivariate ordination of groups proved to be very efficient for both classification of samples into pre-defined groups and disease class prediction of new unknown samples. Classification and prediction with BGA are classically performed using the whole set of genes and no variable selection is required. We hypothesize that an optimized selection of highly discriminating genes might improve the prediction power of BGA. RESULTS: We propose an optimized between-group classification (OBC) which uses a jackknife-based gene selection procedure. OBC emphasizes classification accuracy rather than feature selection. OBC is a backward optimization procedure that maximizes the percentage of between group inertia by removing the least influential genes one by one from the analysis. This selects a subset of highly discriminative genes which optimize disease class prediction. We apply OBC to four datasets and compared it to other classification methods. CONCLUSION: OBC considerably improved the classification and predictive accuracy of BGA, when assessed using independent data sets and leave-one-out cross-validation. AVAILABILITY: The R code is freely available [see Additional file 1] as well as supplementary information [see Additional file 2].

Algorithms↗

Type I polyketide synthases may have evolved through horizontal gene transfer.

Type I polyketide synthases (PKSI) are modular multidomain enzymes involved in the biosynthesis of many natural products of industrial interest. PKSI modules are minimally organized in three domains: ketosynthase (KS), acyltransferase (AT), and acyl carrier protein. The KS domain phylogeny of 23 PKSI clusters was determined. The results obtained suggest that many horizontal transfers of PKSI genes have occurred between actinomycetales species. Such gene transfers may explain the homogeneity and the robustness of the actinomycetales group since gene transfers between closely related species could mimic patterns generated by vertical inheritance. We suggest that the linearity and instability of actinomycetales chromosomes associated with their large quantity of genetic mobile elements have favored such horizontal gene transfers.

Acyltransferases↗

MADE4: an R package for multivariate analysis of gene expression data.

SUMMARY: MADE4, microarray ade4, is a software package that facilitates multivariate analysis of microarray gene-expression data. MADE4 accepts a wide variety of gene-expression data formats. MADE4 takes advantage of the extensive multivariate statistical and graphical functions in the R package ade4, extending these for application to microarray data. In addition, MADE4 provides new graphical and visualization tools that aid in interpretation of multivariate analysis of microarray data.

Algorithms↗

Tree pattern matching in phylogenetic trees: automatic search for orthologs or paralogs in homologous gene sequence databases.

MOTIVATION: Comparative sequence analysis is widely used to study genome function and evolution. This approach first requires the identification of homologous genes and then the interpretation of their homology relationships (orthology or paralogy). To provide help in this complex task, we developed three databases of homologous genes containing sequences, multiple alignments and phylogenetic trees: HOBACGEN, HOVERGEN and HOGENOM. In this paper, we present two new tools for automating the search for orthologs or paralogs in these databases. RESULTS: First, we have developed and implemented an algorithm to infer speciation and duplication events by comparison of gene and species trees (tree reconciliation). Second, we have developed a general method to search in our databases the gene families for which the tree topology matches a peculiar tree pattern. This algorithm of unordered tree pattern matching has been implemented in the FamFetch graphical interface. With the help of a graphical editor, the user can specify the topology of the tree pattern, and set constraints on its nodes and leaves. Then, this pattern is compared with all the phylogenetic trees of the database, to retrieve the families in which one or several occurrences of this pattern are found. By specifying ad hoc patterns, it is therefore possible to identify orthologs in our databases.

Algorithms↗

Horizontal transfer of two operons coding for hydrogenases between bacteria and archaea.

Using a phylogenetic approach, we discovered three putative horizontal transfers between bacterial and archaeal species involving large clusters of genes. One transfer involves an operon of 13 genes, called mbx, which probably was transferred into the genome of Thermotoga maritima from a species belonging or close to the Pyrococcus genus. The two others implied an operon of six genes, called ech, transferred independently to the genomes of Thermoanaerobacter tengcongensis and Desulfovibrio gigas, from a species belonging or close to the Methanosarcina genus. All these transfers affected operons coding for multisubunit membrane-bound (NiFe) hydrogenases involved in the energy metabolism of the donor genomes. The functionality of the transferred operons has not been experimentally demonstrated for T. maritima, whereas in D. gigas and T. tengcongensis the encoded multisubunit hydrogenase could have a role in energy conservation. This report adds several cases of horizontal gene transfers among hydrogenases already described.

Archaea↗

Update of NUREBASE: nuclear hormone receptor functional genomics.

Nuclear hormone receptors are an abundant class of ligand-activated transcriptional regulators, found in varying numbers in all animals. Based on our experience of managing the official nomenclature of nuclear receptors, we have developed NUREBASE, a database containing protein and DNA sequences, reviewed protein alignments and phylogenies, taxonomy and annotations for all nuclear receptors. New developments in NUREBASE include explicit declaration of alternative transcripts of each gene, and expression data for human and mouse nuclear receptors. The core of NUREBASE is reviewed, and it is completed by NUREBASE_DAILY, automatically updated every 24 h. All information on accessing and installing NUREBASE may be found at http://www. ens-lyon.fr/LBMC/laudet/nurebase/nurebase.html.

Alternative Splicing↗

The European ribosomal RNA database.

The European ribosomal RNA database aims to compile all complete or nearly complete ribosomal RNA sequences from both the small (SSU) and large (LSU) ribosomal subunits. All sequences are available in aligned format. Sequence alignment is based on the secondary structure of the molecules, as determined by comparative sequence analysis. Additional information about the sequences, such as taxonomic classification of the organism from which they have been obtained, and literature references are also provided. In order to identify the closest relatives to newly determined sequences, BLAST searches can be performed, after which the best matching sequences are aligned and a phylogenetic tree is inferred. As of 2003, the European ribosomal RNA database is maintained at Ghent University (Belgium). The database can be consulted at http://www.psb.ugent.be/rRNA/.

Animals↗

Phylogenetic analysis of polyketide synthase I domains from soil metagenomic libraries allows selection of promising clones.

The metagenomic approach provides direct access to diverse unexplored genomes, especially from uncultivated bacteria in a given environment. This diversity can conceal many new biosynthetic pathways. Type I polyketide synthases (PKSI) are modular enzymes involved in the biosynthesis of many natural products of industrial interest. Among the PKSI domains, the ketosynthase domain (KS) was used to screen a large soil metagenomic library containing more than 100,000 clones to detect those containing PKS genes. Over 60,000 clones were screened, and 139 clones containing KS domains were detected. A 700-bp fragment of the KS domain was sequenced for 40 of 139 randomly chosen clones. None of the 40 protein sequences were identical to those found in public databases, and nucleic sequences were not redundant. Phylogenetic analyses were performed on the protein sequences of three metagenomic clones to select the clones which one can predict to produce new compounds. Two PKS-positive clones do not belong to any of the 23 published PKSI included in the analysis, encouraging further analyses on these two clones identified by the selection process.

Amino Acid Sequence↗

Cross-platform comparison and visualisation of gene expression data using co-inertia analysis.

BACKGROUND: Rapid development of DNA microarray technology has resulted in different laboratories adopting numerous different protocols and technological platforms, which has severely impacted on the comparability of array data. Current cross-platform comparison of microarray gene expression data are usually based on cross-referencing the annotation of each gene transcript represented on the arrays, extracting a list of genes common to all arrays and comparing expression data of this gene subset. Unfortunately, filtering of genes to a subset represented across all arrays often excludes many thousands of genes, because different subsets of genes from the genome are represented on different arrays. We wish to describe the application of a powerful yet simple method for cross-platform comparison of gene expression data. Co-inertia analysis (CIA) is a multivariate method that identifies trends or co-relationships in multiple datasets which contain the same samples. CIA simultaneously finds ordinations (dimension reduction diagrams) from the datasets that are most similar. It does this by finding successive axes from the two datasets with maximum covariance. CIA can be applied to datasets where the number of variables (genes) far exceeds the number of samples (arrays) such is the case with microarray analyses. RESULTS: We illustrate the power of CIA for cross-platform analysis of gene expression data by using it to identify the main common relationships in expression profiles on a panel of 60 tumour cell lines from the National Cancer Institute (NCI) which have been subjected to microarray studies using both Affymetrix and spotted cDNA array technology. The co-ordinates of the CIA projections of the cell lines from each dataset are graphed in a bi-plot and are connected by a line, the length of which indicates the divergence between the two datasets. Thus, CIA provides graphical representation of consensus and divergence between the gene expression profiles from different microarray platforms. Secondly, the genes that define the main trends in the analysis can be easily identified. CONCLUSIONS: CIA is a robust, efficient approach to coupling of gene expression datasets. CIA provides simple graphical representations of the results making it a particularly attractive method for the identification of relationships between large datasets.

Breast Neoplasms↗

The source of laterally transferred genes in bacterial genomes.

BACKGROUND: Laterally transferred genes have often been identified on the basis of compositional features that distinguish them from ancestral genes in the genome. These genes are usually A+T-rich, arguing either that there is a bias towards acquiring genes from donor organisms having low G+C contents or that genes acquired from organisms of similar genomic base compositions go undetected in these analyses. RESULTS: By examining the genome contents of closely related, fully sequenced bacteria, we uncovered genes confined to a single genome and examined the sequence features of these acquired genes. The analysis shows that few transfer events are overlooked by compositional analyses. Most observed lateral gene transfers do not correspond to free exchange of regular genes among bacterial genomes, but more probably represent the constituents of phages or other selfish elements. CONCLUSIONS: Although bacteria tend to acquire large amounts of DNA, the origin of these genes remains obscure. We have shown that contrary to what is often supposed, their composition cannot be explained by a previous genomic context. In contrast, these genes fit the description of recently described genes in lambdoid phages, named 'morons'. Therefore, results from genome content and compositional approaches to detect lateral transfers should not be cited as evidence for genetic exchange between distantly related bacteria.

Arginine↗

Integrated databanks access and sequence/structure analysis services at the PBIL.

The World Wide Web server of the PBIL (Pôle Bioinformatique Lyonnais) provides on-line access to sequence databanks and to many tools of nucleic acid and protein sequence analyses. This server allows to query nucleotide sequence banks in the EMBL and GenBank formats and protein sequence banks in the SWISS-PROT and PIR formats. The query engine on which our data bank access is based is the ACNUC system. It allows the possibility to build complex queries to access functional zones of biological interest and to retrieve large sequence sets. Of special interest are the unique features provided by this system to query the data banks of gene families developed at the PBIL. The server also provides access to a wide range of sequence analysis methods: similarity search programs, multiple alignments, protein structure prediction and multivariate statistics. An originality of this server is the integration of these two aspects: sequence retrieval and sequence analysis. Indeed, thanks to the introduction of re-usable lists, it is possible to perform treatments on large sets of data. The PBIL server can be reached at: http://pbil.univ-lyon1.fr.

Databases, Genetic↗

G+C3 structuring along the genome: a common feature in prokaryotes.

The heterogeneity of gene nucleotide content in prokaryotic genomes is commonly interpreted as the result of three main phenomena: (1) genes undergo different selection pressures both during and after translation (affecting codon and amino acid choice); (2) genes undergo different mutational pressure whether they are on the leading or lagging strand; and (3) genes may have different phylogenetic origins as a result of lateral transfers. However, this view neglects the necessity of organizing genetic information on a chromosome that needs to be replicated and folded, which may add constraints to single gene evolution. As a consequence, genes are potentially subjected to different mutation and selection pressures, depending on their position in the genome. In this paper, we analyze the structuring of different codon usage measures along completely sequenced bacterial genomes. We show that most of them are highly structured, suggesting that genes have different base content, depending on their location on the chromosome. A peculiar pattern of genome structure, with a tendency toward an A+T-enrichment near the replication terminus, is found in most bacterial phyla and may reflect common chromosome constraints. Several species may have lost this pattern, probably because of genome rearrangements or integration of foreign DNA. We show that in several species, this enrichment is associated with an increase of evolutionary rate and we discuss the evolutionary implications of these results. We argue that structural constraints acting on the circular chromosome are not negligible and that this natural structuring of bacterial genomes may be a cause of overestimation in lateral gene transfer predictions using codon composition indices.

Base Composition↗

RTKdb: database of Receptor Tyrosine Kinase.

Receptor Tyrosine Kinases (RTK) are transmembrane receptors specifically found in metazoans. They represent an excellent model for studying evolution of cellular processes in metazoans because they encompass large families of modular proteins and belong to a major family of contingency generating molecules in eukaryotic cells: the protein kinases. Because tyrosine kinases have been under close scrutiny for many years in various species, they are associated with a wealth of information, mainly in mammals. Presently, most categories of RTK were identified in mammals, but in a near future other model species will be sequenced, and will bring us RTKs from other metazoan clades. Thus, collecting RTK sequences would provide a good starting point as a new model for comparative and evolutionary studies applying to multigene families. In this context, we are developing the Receptor Tyrosine Kinase database (RTKdb), which is the only database on tyrosine kinase receptors presently available. In this database, protein sequences from eight model metazoan species are organized under the format previously used for the HOVERGEN, HOBACGEN and NUREBASE systems. RTKdb can be accessed through the PBIL (Pôle Bioinformatique Lyonnais) World Wide Web server at http://pbil.univ-lyon1.fr/RTKdb/, or through the FamFetch graphical user interface available at the same address.

Animals↗

Use of correspondence discriminant analysis to predict the subcellular location of bacterial proteins.

Correspondence discriminant analysis (CDA) is a multivariate statistical method derived from discriminant analysis which can be used on contingency tables. We have used CDA to separate Gram negative bacteria proteins according to their subcellular location. The high resolution of the discrimination obtained makes this method a good tool to predict subcellular location when this information is not known. The main advantage of this technique is its simplicity. Indeed, by computing two linear formulae on amino acid composition, it is possible to classify a protein into one of the three classes of subcellular location we have defined. The CDA itself can be computed with the ADE-4 software package that can be downloaded, as well as the data set used in this study, from the Pôle Bio-Informatique Lyonnais (PBIL) server at http://pbil.univ-lyon1.fr.

Bacterial Proteins↗