Search PubMed⌕ Search

Biomedical subjects

P Rouzé

Publications and source records attributed to P Rouzé.

At least 19 recordsLinked to original sources

The genome of black cottonwood, Populus trichocarpa (Torr. & Gray).

We report the draft genome of the black cottonwood tree, Populus trichocarpa. Integration of shotgun sequence assembly with genetic mapping enabled chromosome-scale reconstruction of the genome. More than 45,000 putative protein-coding genes were identified. Analysis of the assembled genome revealed a whole-genome duplication event; about 8000 pairs of duplicated genes from that event survived in the Populus genome. A second, older duplication event is indistinguishably coincident with the divergence of the Populus and Arabidopsis lineages. Nucleotide substitution, tandem gene duplication, and gross chromosomal rearrangement appear to proceed substantially more slowly in Populus than in Arabidopsis. Populus has more protein-coding genes than Arabidopsis, ranging on average from 1.4 to 1.6 putative Populus homologs for each Arabidopsis gene. However, the relative frequency of protein domains in the two genomes is similar. Overrepresented exceptions in Populus include genes associated with lignocellulosic wall biosynthesis, meristem development, disease resistance, and metabolite transport.

Arabidopsis↗

Annotation of a 95-kb Populus deltoides genomic sequence reveals a disease resistance gene cluster and novel class I and class II transposable elements.

Poplar has become a model system for functional genomics in woody plants. Here, we report the sequencing and annotation of the first large contiguous stretch of genomic sequence (95 kb) of poplar, corresponding to a bacterial artificial chromosome clone mapped 0.6 centiMorgan from the Melampsora larici-populina resistance locus. The annotation revealed 15 putative genetic objects, of which five were classified as hypothetical genes that were similar only with expressed sequence tags from poplar. Ten putative objects showed similarity with known genes, of which one was similar to a kinase. Three other objects corresponded to the toll/interleukin-1 receptor/nucleotide-binding site/leucine-rich repeat class of plant disease resistance genes, of which two were predicted to encode an amino terminal nuclear localization signal. Four objects were homologous to the Ty1/ copia family of class I transposable elements, one of which was designated Retropop and interrupted one of the disease resistance genes. Two other objects constituted a novel Spm-like class II transposable element, which we designated Magali.

Amino Acid Sequence↗

A higher-order background model improves the detection of promoter regulatory elements by Gibbs sampling.

MOTIVATION: Transcriptome analysis allows detection and clustering of genes that are coexpressed under various biological circumstances. Under the assumption that coregulated genes share cis-acting regulatory elements, it is important to investigate the upstream sequences controlling the transcription of these genes. To improve the robustness of the Gibbs sampling algorithm to noisy data sets we propose an extension of this algorithm for motif finding with a higher-order background model. RESULTS: Simulated data and real biological data sets with well-described regulatory elements are used to test the influence of the different background models on the performance of the motif detection algorithm. We show that the use of a higher-order model considerably enhances the performance of our motif finding algorithm in the presence of noisy data. For Arabidopsis thaliana, a reliable background model based on a set of carefully selected intergenic sequences was constructed. AVAILABILITY: Our implementation of the Gibbs sampler called the Motif Sampler can be used through a web interface: http://www.esat.kuleuven.ac.be/~thijs/Work/MotifSampler.html. CONTACT: gert.thijs@esat.kuleuven.ac.be; yves.moreau@esat.kuleuven.ac.be

Algorithms↗

Gene prediction and gene classes in Arabidopsis thaliana.

Gene prediction methods for eukaryotic genomes still are not fully satisfying. One way to improve gene prediction accuracy, proven to be relevant for prokaryotes, is to consider more than one model of genes. Thus, we used our classification of Arabidopsis thaliana genes in two classes (CU(1) and CU(2)), previously delineated according to statistical features, in the GeneMark gene identification program. For each gene class, as well as for the two classes combined, a Markov model was developed (respectively, GM-CU(1), GM-CU(2) and GM-all) and then used on a test set of 168 genes to compare their respective efficiency. We concluded from this analysis that GM-CU(1) is more sensitive than GM-CU(2) which seems to be more specific to a gene type. Besides, GM-all does not give better results than GM-CU(1) and combining results from GM-CU(1) and GM-CU(2) greatly improve prediction efficiency in comparison with predictions made with GM-all only. Thus, this work confirms the necessity to consider more than one gene model for gene prediction in eukaryotic genomes, and to look for gene classes in order to build these models.

Arabidopsis↗

PPMdb: a plant plasma membrane database.

PPMdb is a proteome database dedicated to proteins from plant plasma membranes. It provides comprehensive two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) maps, partial amino acid sequences and expression data. All this information is gathered and structured in a relational database, after being analyzed and annotated. PPMdb includes active links to related biological databases (EMBL, GenBank, GenPep, and SWISS-PROT and TrEMBL) as well as to MEDLINE abstracts. Information on specific protein spots can be displayed by clicking on the 2-D maps. In addition, users can query the database by accession number, protein name, pI and MW, and cellular location. Access to PPMdb is available at the following URL: http://sphinx.rug. ac.be:8080.

Amino Acid Sequence↗

The sense of naturally transcribed antisense RNAs in plants.

Naturally occurring antisense transcripts are well documented in mammals and prokaryotes but little is known about their existence and effects in plants. Generally, antisense RNAs are believed to control gene expression negatively by annealing to the complementary sequences of the sense transcript. The resulting double-stranded RNAs are thought either to affect RNA stability, transcription and/or translation directly, or to generate a signal for gene silencing and defense against viruses.

Gene Expression Regulation, Plant↗

Plant genomics.

The rapidity with which genomic sequences of the model plant Arabidopsis thaliana and soon of rice are becoming available has strongly boosted plant molecular biology research. Here, two main genomic fields will be discussed: the progress in different structural genome projects, such as mapping, sequencing, genome organization and comparative genomics, and the so-called functional genomics approaches to analyze the genome using such molecular tools as transcript profiling, micro-arrays, and insertional mutagenesis. In addition a section on bioinformatics is included.

DNA, Plant↗

Evidence for an ancient chromosomal duplication in Arabidopsis thaliana by sequencing and analyzing a 400-kb contig at the APETALA2 locus on chromosome 4.

As part of the European Scientists Sequencing Arabidopsis program, a contiguous region (396607 bp) located on chromosome 4 around the APETALA2 gene was sequenced. Analysis of the sequence and comparison to public databases predicts 103 genes in this area, which represents a gene density of one gene per 3.85 kb. Almost half of the genes show no significant homology to known database entries. In addition, the first 45 kb of the contig, which covers 11 genes, is similar to a region on chromosome 2, as far as coding sequences are concerned. This observation indicates that ancient duplications of large pieces of DNA have occurred in Arabidopsis.

Arabidopsis↗

Classification of Arabidopsis thaliana gene sequences: clustering of coding sequences into two groups according to codon usage improves gene prediction.

While genomic sequences are accumulating, finding the location of the genes remains a major issue that can be solved only for about a half of them by homology searches. Prediction methods are thus required, but unfortunately are not fully satisfying. Most prediction methods implicitly assume a unique model for genes. This is an oversimplification as demonstrated by the possibility to group coding sequences into several classes in Escherichia coli and other genomes. As no classification existed for Arabidopsis thaliana, we classified genes according to the statistical features of their coding sequences. A clustering algorithm using a codon usage model was developed and applied to coding sequences from A. thaliana, E. coli, and a mixture of both. By using it, Arabidopsis sequences were clustered into two classes. The CU1 and CU2 classes differed essentially by the choice of pyrimidine bases at the codon silent sites: CU2 genes often use C whereas CU1 genes prefer T. This classification discriminated the Arabidopsis genes according to their expressiveness, highly expressed genes being clustered in CU2 and genes expected to have a lower expression, such as the regulatory genes, in CU1. The algorithm separated the sequences of the Escherichia-Arabidopsis mixed data set into five classes according to the species, except for one class. This mixed class contained 89 % Arabidopsis genes from CU1 and 11 % E. coli genes, mostly horizontally transferred. Interestingly, most genes encoding organelle-targeted proteins, except the photosynthetic and photoassimilatory ones, were clustered in CU1. By tailoring the GeneMark CDS prediction algorithm to the observed coding sequence classes, its quality of prediction was greatly improved. Similar improvement can be expected with other prediction systems.

Algorithms↗

PlantCARE, a plant cis-acting regulatory element database.

PlantCARE is a database of plant cis- acting regulatory elements, enhancers and repressors. Besides the transcription motifs found on a sequence, it also offers a link to the EMBL entry that contains the full gene sequence as well as a description of the conditions in which a motif becomes functional. The information on these sites is given by matrices, consensus and individual site sequences on particular genes, depending on the available information. PlantCARE is a relational database available via the web at the URL: http://sphinx.rug.ac.be:8080/PlantCARE/

Arabidopsis↗

Genome annotation: which tools do we have for it?

Genome data have to be converted into knowledge to be useful to biologists. Many valuable computational tools have already been developed to help annotation of plant genome sequences, and these may be improved further, for example by identification of more gene regulatory elements. The lack of a standard computer-assisted annotation platform for eukaryotic genomes remains major bottle-neck.

Arabidopsis↗

Evaluation of gene prediction software using a genomic data set: application to Arabidopsis thaliana sequences.

MOTIVATION: The annotation of the Arabidopsis thaliana genome remains a problem in terms of time and quality. To improve the annotation process, we want to choose the most appropriate tools to use inside a computer-assisted annotation platform. We therefore need evaluation of prediction programs with Arabidopsis sequences containing multiple genes. RESULTS: We have developed AraSet, a data set of contigs of validated genes, enabling the evaluation of multi-gene models for the Arabidopsis genome. Besides conventional metrics to evaluate gene prediction at the site and the exon levels, new measures were introduced for the prediction at the protein sequence level as well as for the evaluation of gene models. This evaluation method is of general interest and could apply to any new gene prediction software and to any eukaryotic genome. The GeneMark.hmm program appears to be the most accurate software at all three levels for the Arabidopsis genomic sequences. Gene modeling could be further improved by combination of prediction software. AVAILABILITY: The AraSet sequence set, the Perl programs and complementary results and notes are available at http://sphinx.rug.ac.be:8080/biocomp/napav/. CONTACT: Pierre.Rouze@gengenp.rug.ac.be.

Alternative Splicing↗

Sequence analysis of a 40-kb Arabidopsis thaliana genomic region located at the top of chromosome 1.

As a contribution to the European Scientists Sequencing Arabidopsis (BIOTECH ESSA) project, a contig of almost 40kb has been sequenced at the extreme top of chromosome 1, around the Arabidopsis thaliana gene coding for a member of the 1-aminocyclopropane-1-carboxylate synthesis gene family. The region contains, besides the ACS1 gene itself, 10 putative genes, all new for Arabidopsis. Among these are three genes encoding kinases, a late embryogenesis-abundant protein, a MADS box-containing protein, a dehydrogenase, and a Myb-related transcription factor. In addition, six cDNAs have been sequenced that correspond to this region.

Arabidopsis↗

Analysis of T-DNA-mediated translational beta-glucuronidase gene fusions.

Three random translational beta-glucuronidase (gus) gene fusions were previously obtained in Arabidopsis thaliana, using Agrobacterium-mediated transfer of a gus coding sequence without promoter and ATG initiation site. These were analysed by IPCR amplification of the sequence upstream of gus and nucleotide sequence analysis. In one instance, the gus sequence was fused, in inverse orientation, to the nos promoter sequence of a truncated tandem T-DNA copy and translated from a spurious ATG in this sequence. In the second transgenic line, the gus gene was fused to A. thaliana DNA, 27 bp downstream an ATG. In this line, a large deletion occurred at the target site of the T-DNA. In the third line, gus is fused in frame to a plant DNA sequence after the eighth codon of an open reading frame encoding a protein of 619 amino acids. This protein has significant homology with animal and plant (receptor) serine/threonine protein kinases. The twelve subdomains essential for kinase activity are conserved. The presence of a potential signal peptide and a membrane-spanning domain suggests that it may be a receptor kinase. These data confirm that plant genes can be tagged as functional translational gene fusions.

Amino Acid Sequence↗

Use of a proteome strategy for tagging proteins present at the plasma membrane.

A plasma membrane (PM) fraction was purified from Arabidopsis thaliana using a standard procedure and analyzed by two-dimensional (2D) gel electrophoresis. The proteins were classified according to their relative abundance in PM or cell membrane supernatant fractions. Eighty-two of the 700 spots detected on the PM 2D gels were microsequenced. More than half showed sequence similarity to proteins of known function. Of these, all the spots in the PM-specific and PM-enriched fractions, together with half of the spots with similar abundance in PM fraction and supernatant, have previously been found at the PM, supporting the validity of this approach. Extrapolation from this analysis indicates that (i) approximately 550 polypeptides found at the PM could be resolved on 2D gels; (ii) that numerous proteins with multiple locations are found at the PM; and (iii) that approximately 80% of PM-specific spots correspond to proteins with unknown function. Among the later, half are represented by ESTs or cDNAs in databases. In this way, several unknown gene products were potentially localized to the PM. These data are discussed with respect to the efficiency of organelle proteome approaches to link systematically genomic data to genome expression. It is concluded that generalized proteomes can constitute a powerful resource, with future completion of Arabidopsis genome sequencing, for genome-wide exploration of plant function.

Amino Acid Sequence↗

Sequence analysis of a 24-kb contiguous genomic region at the Arabidopsis thaliana PFL locus on chromosome 1.

As part of the European Union program of European Scientist Sequencing Arabidopsis (ESSA), the DNA sequence of a 24.053-bp insert of cosmid clone CC17J13 was determined. The cosmid is located on chromosome 1 at the PFL locus (position 30 cM). Analysis of the sequence and comparison to public databases predicts seven genes in this area, thus approximately one gene every 3.3 kb. Three cDNAs corresponding to genes in this region were also sequenced. The homologies and/or possible functions of the (putative) genes are discussed. Proteins encoded by genes in this region include a polyadenylate-binding protein (PAB-3) and a GTP-binding protein (Rab7) as well as a novel protein, possibly involved in double-stranded RNA unwinding and apoptosis. Intriguingly, the gene encoding the PAB-3 protein, which is very specifically expressed, is flanked by putative matrix attachment regions.

Arabidopsis↗

A branch point consensus from Arabidopsis found by non-circular analysis allows for better prediction of acceptor sites.

Little knowledge exists about branch points in plants; it has even been claimed that plant introns lack conserved branch point sequences similar to those found in vertebrate introns. A putative branch point consensus sequence for Arabidopsis thaliana resembling the well known metazoan consensus sequence has been proposed, but this is based on search of sequences similar to those in yeast and metazoa. Here we present a novel consensus sequence found by a non-circular approach. A hidden Markov model with a fixed A nucleotide was trained on sequences upstream of the acceptor site. The consensus found by the Markov model shares features with the metazoan consensus, but differs in its details from the consensus proposed earlier. Despite the fact that branch point consensus sequences in plants are weak, we show that a prediction scheme incorporating them leads to a substantial improvement in the recognition of true acceptor sites; the false positive rate being reduced by a factor of 2. We take this as an indication that the consensus found here is the genuine one and that the branch point does play a role in the proper recognition of the acceptor site in plants.

Arabidopsis↗

Splice site prediction in Arabidopsis thaliana pre-mRNA by combining local and global sequence information.

Artificial neural networks have been combined with a rule based system to predict intron splice sites in the dicot plant Arabidopsis thaliana. A two step prediction scheme, where a global prediction of the coding potential regulates a cutoff level for a local prediction of splice sites, is refined by rules based on splice site confidence values, prediction scores, coding context and distances between potential splice sites. In this approach, the prediction of splice sites mutually affect each other in a non-local manner. The combined approach drastically reduces the large amount of false positive splice sites normally haunting splice site prediction. An analysis of the errors made by the networks in the first step of the method revealed a previously unknown feature, a frequent T-tract prolongation containing cryptic acceptor sites in the 5' end of exons. The method presented here has been compared with three other approaches, GeneFinder, Gene-Mark and Grail. Overall the method presented here is an order of magnitude better. We show that the new method is able to find a donor site in the coding sequence for the jelly fish Green Fluorescent Protein, exactly at the position that was experimentally observed in A.thaliana transformants. Predictions for alternatively spliced genes are also presented, together with examples of genes from other dicots, monocots and algae. The method has been made available through electronic mail (NetPlantGene@cbs.dtu.dk), or the WWW at http://www.cbs.dtu.dk/NetPlantGene.html

Algorithms↗