Search PubMed⌕ Search

Biomedical subjects

P Déhais

Publications and source records attributed to P Déhais.

9 recordsLinked to original sources

Gene prediction and gene classes in Arabidopsis thaliana.

Gene prediction methods for eukaryotic genomes still are not fully satisfying. One way to improve gene prediction accuracy, proven to be relevant for prokaryotes, is to consider more than one model of genes. Thus, we used our classification of Arabidopsis thaliana genes in two classes (CU(1) and CU(2)), previously delineated according to statistical features, in the GeneMark gene identification program. For each gene class, as well as for the two classes combined, a Markov model was developed (respectively, GM-CU(1), GM-CU(2) and GM-all) and then used on a test set of 168 genes to compare their respective efficiency. We concluded from this analysis that GM-CU(1) is more sensitive than GM-CU(2) which seems to be more specific to a gene type. Besides, GM-all does not give better results than GM-CU(1) and combining results from GM-CU(1) and GM-CU(2) greatly improve prediction efficiency in comparison with predictions made with GM-all only. Thus, this work confirms the necessity to consider more than one gene model for gene prediction in eukaryotic genomes, and to look for gene classes in order to build these models.

Arabidopsis↗

PPMdb: a plant plasma membrane database.

PPMdb is a proteome database dedicated to proteins from plant plasma membranes. It provides comprehensive two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) maps, partial amino acid sequences and expression data. All this information is gathered and structured in a relational database, after being analyzed and annotated. PPMdb includes active links to related biological databases (EMBL, GenBank, GenPep, and SWISS-PROT and TrEMBL) as well as to MEDLINE abstracts. Information on specific protein spots can be displayed by clicking on the 2-D maps. In addition, users can query the database by accession number, protein name, pI and MW, and cellular location. Access to PPMdb is available at the following URL: http://sphinx.rug. ac.be:8080.

Amino Acid Sequence↗

Evidence for an ancient chromosomal duplication in Arabidopsis thaliana by sequencing and analyzing a 400-kb contig at the APETALA2 locus on chromosome 4.

As part of the European Scientists Sequencing Arabidopsis program, a contiguous region (396607 bp) located on chromosome 4 around the APETALA2 gene was sequenced. Analysis of the sequence and comparison to public databases predicts 103 genes in this area, which represents a gene density of one gene per 3.85 kb. Almost half of the genes show no significant homology to known database entries. In addition, the first 45 kb of the contig, which covers 11 genes, is similar to a region on chromosome 2, as far as coding sequences are concerned. This observation indicates that ancient duplications of large pieces of DNA have occurred in Arabidopsis.

Arabidopsis↗

Classification of Arabidopsis thaliana gene sequences: clustering of coding sequences into two groups according to codon usage improves gene prediction.

While genomic sequences are accumulating, finding the location of the genes remains a major issue that can be solved only for about a half of them by homology searches. Prediction methods are thus required, but unfortunately are not fully satisfying. Most prediction methods implicitly assume a unique model for genes. This is an oversimplification as demonstrated by the possibility to group coding sequences into several classes in Escherichia coli and other genomes. As no classification existed for Arabidopsis thaliana, we classified genes according to the statistical features of their coding sequences. A clustering algorithm using a codon usage model was developed and applied to coding sequences from A. thaliana, E. coli, and a mixture of both. By using it, Arabidopsis sequences were clustered into two classes. The CU1 and CU2 classes differed essentially by the choice of pyrimidine bases at the codon silent sites: CU2 genes often use C whereas CU1 genes prefer T. This classification discriminated the Arabidopsis genes according to their expressiveness, highly expressed genes being clustered in CU2 and genes expected to have a lower expression, such as the regulatory genes, in CU1. The algorithm separated the sequences of the Escherichia-Arabidopsis mixed data set into five classes according to the species, except for one class. This mixed class contained 89 % Arabidopsis genes from CU1 and 11 % E. coli genes, mostly horizontally transferred. Interestingly, most genes encoding organelle-targeted proteins, except the photosynthetic and photoassimilatory ones, were clustered in CU1. By tailoring the GeneMark CDS prediction algorithm to the observed coding sequence classes, its quality of prediction was greatly improved. Similar improvement can be expected with other prediction systems.

Algorithms↗

PlantCARE, a plant cis-acting regulatory element database.

PlantCARE is a database of plant cis- acting regulatory elements, enhancers and repressors. Besides the transcription motifs found on a sequence, it also offers a link to the EMBL entry that contains the full gene sequence as well as a description of the conditions in which a motif becomes functional. The information on these sites is given by matrices, consensus and individual site sequences on particular genes, depending on the available information. PlantCARE is a relational database available via the web at the URL: http://sphinx.rug.ac.be:8080/PlantCARE/

Arabidopsis↗

Evaluation of gene prediction software using a genomic data set: application to Arabidopsis thaliana sequences.

MOTIVATION: The annotation of the Arabidopsis thaliana genome remains a problem in terms of time and quality. To improve the annotation process, we want to choose the most appropriate tools to use inside a computer-assisted annotation platform. We therefore need evaluation of prediction programs with Arabidopsis sequences containing multiple genes. RESULTS: We have developed AraSet, a data set of contigs of validated genes, enabling the evaluation of multi-gene models for the Arabidopsis genome. Besides conventional metrics to evaluate gene prediction at the site and the exon levels, new measures were introduced for the prediction at the protein sequence level as well as for the evaluation of gene models. This evaluation method is of general interest and could apply to any new gene prediction software and to any eukaryotic genome. The GeneMark.hmm program appears to be the most accurate software at all three levels for the Arabidopsis genomic sequences. Gene modeling could be further improved by combination of prediction software. AVAILABILITY: The AraSet sequence set, the Perl programs and complementary results and notes are available at http://sphinx.rug.ac.be:8080/biocomp/napav/. CONTACT: Pierre.Rouze@gengenp.rug.ac.be.

Alternative Splicing↗

Sequence analysis of a 40-kb Arabidopsis thaliana genomic region located at the top of chromosome 1.

As a contribution to the European Scientists Sequencing Arabidopsis (BIOTECH ESSA) project, a contig of almost 40kb has been sequenced at the extreme top of chromosome 1, around the Arabidopsis thaliana gene coding for a member of the 1-aminocyclopropane-1-carboxylate synthesis gene family. The region contains, besides the ACS1 gene itself, 10 putative genes, all new for Arabidopsis. Among these are three genes encoding kinases, a late embryogenesis-abundant protein, a MADS box-containing protein, a dehydrogenase, and a Myb-related transcription factor. In addition, six cDNAs have been sequenced that correspond to this region.

Arabidopsis↗

Sequence analysis of a 24-kb contiguous genomic region at the Arabidopsis thaliana PFL locus on chromosome 1.

As part of the European Union program of European Scientist Sequencing Arabidopsis (ESSA), the DNA sequence of a 24.053-bp insert of cosmid clone CC17J13 was determined. The cosmid is located on chromosome 1 at the PFL locus (position 30 cM). Analysis of the sequence and comparison to public databases predicts seven genes in this area, thus approximately one gene every 3.3 kb. Three cDNAs corresponding to genes in this region were also sequenced. The homologies and/or possible functions of the (putative) genes are discussed. Proteins encoded by genes in this region include a polyadenylate-binding protein (PAB-3) and a GTP-binding protein (Rab7) as well as a novel protein, possibly involved in double-stranded RNA unwinding and apoptosis. Intriguingly, the gene encoding the PAB-3 protein, which is very specifically expressed, is flanked by putative matrix attachment regions.

Arabidopsis↗