Search PubMed⌕ Search

Biomedical subjects

Simon Penel

Publications and source records attributed to Simon Penel.

6 recordsLinked to original sources

Tree pattern matching in phylogenetic trees: automatic search for orthologs or paralogs in homologous gene sequence databases.

MOTIVATION: Comparative sequence analysis is widely used to study genome function and evolution. This approach first requires the identification of homologous genes and then the interpretation of their homology relationships (orthology or paralogy). To provide help in this complex task, we developed three databases of homologous genes containing sequences, multiple alignments and phylogenetic trees: HOBACGEN, HOVERGEN and HOGENOM. In this paper, we present two new tools for automating the search for orthologs or paralogs in these databases. RESULTS: First, we have developed and implemented an algorithm to infer speciation and duplication events by comparison of gene and species trees (tree reconciliation). Second, we have developed a general method to search in our databases the gene families for which the tree topology matches a peculiar tree pattern. This algorithm of unordered tree pattern matching has been implemented in the FamFetch graphical interface. With the help of a graphical editor, the user can specify the topology of the tree pattern, and set constraints on its nodes and leaves. Then, this pattern is compared with all the phylogenetic trees of the database, to retrieve the families in which one or several occurrences of this pattern are found. By specifying ad hoc patterns, it is therefore possible to identify orthologs in our databases.

Algorithms↗

Integr8 and Genome Reviews: integrated views of complete genomes and proteomes.

Integr8 is a new web portal for exploring the biology of organisms with completely deciphered genomes. For over 190 species, Integr8 provides access to general information, recent publications, and a detailed statistical overview of the genome and proteome of the organism. The preparation of this analysis is supported through Genome Reviews, a new database of bacterial and archaeal DNA sequences in which annotation has been upgraded (compared to the original submission) through the integration of data from many sources, including the EMBL Nucleotide Sequence Database, the UniProt Knowledgebase, InterPro, CluSTr, GOA and HOGENOM. Integr8 also allows the users to customize their own interactive analysis, and to download both customized and prepared datasets for their own use. Integr8 is available at http://www.ebi.ac.uk/integr8.

DNA, Archaeal↗

Polymorphix: a sequence polymorphism database.

Within-species sequence variation data are of special interest since they contain information about recent population/species history, and the molecular evolutionary forces currently in action in natural populations. These data, however, are presently dispersed within generalist databases, and are difficult to access. To solve this problem, we have developed Polymorphix, a database dedicated to sequence polymorphism. It contains within-species homologous sequence families built using EMBL/GenBank under suitable similarity and bibliographic criteria. Polymorphix is an ACNUC structured database allowing both simple and complex queries for population genomic studies. Alignments within families as well as phylogenetic trees can be download. When available, outgroups are included in the alignment. Polymorphix contains sequences from the nuclear, mitochondrial and chloroplastic genomes of every eukaryote species represented in EMBL. It can be accessed by a web interface (http://pbil.univ-lyon1.fr/polymorphix/query.php).

Animals↗

YodA from Escherichia coli is a metal-binding, lipocalin-like protein.

We have determined the crystal structure of YodA, an Escherichia coli protein of unknown function. YodA had been identified under conditions of cadmium stress, and we confirm that it binds metals such as cadmium and zinc. We have also found nickel bound in one of the crystal forms. YodA is composed of two domains: a main lipocalin/calycin-like domain and a helical domain. The principal metal-binding site lies on one side of the calycin domain, thus making YodA the first metal-binding lipocalin known. Our experiments suggest that YodA expression may be part of a more general stress response. From sequence analogy with the C-terminal domain of a metal-binding receptor of a member of bacterial ATP-binding cassette transporters, we propose a three-dimensional model for this receptor and suggest that YodA may have a receptor-type partner in E. coli.

Adenosine Triphosphate↗

Integrated databanks access and sequence/structure analysis services at the PBIL.

The World Wide Web server of the PBIL (Pôle Bioinformatique Lyonnais) provides on-line access to sequence databanks and to many tools of nucleic acid and protein sequence analyses. This server allows to query nucleotide sequence banks in the EMBL and GenBank formats and protein sequence banks in the SWISS-PROT and PIR formats. The query engine on which our data bank access is based is the ACNUC system. It allows the possibility to build complex queries to access functional zones of biological interest and to retrieve large sequence sets. Of special interest are the unique features provided by this system to query the data banks of gene families developed at the PBIL. The server also provides access to a wide range of sequence analysis methods: similarity search programs, multiple alignments, protein structure prediction and multivariate statistics. An originality of this server is the integration of these two aspects: sequence retrieval and sequence analysis. Indeed, thanks to the introduction of re-usable lists, it is possible to perform treatments on large sets of data. The PBIL server can be reached at: http://pbil.univ-lyon1.fr.

Databases, Genetic↗

Length preferences and periodicity in beta-strands. Antiparallel edge beta-sheets are more likely to finish in non-hydrogen bonded rings.

We analysed the length distributions of different types of beta-strand in a high resolution, non-homologous set of 500 protein structures, finding differences in their mean lengths. Antiparallel edge strands in strand-turn-strand motifs show a preference for an even number of residues. This propensity is enhanced if the length is corrected for beta-bulges, which insert an extra residue into the strand. Residues in antiparallel edge beta-strands alternate between being in hydrogen bonded and non-hydrogen bonded rings. Antiparallel edges with an even number of residues are more likely to have their final beta residue in a non-hydrogen bonded ring. This suggests that non-hydrogen bonded rings are intrinsically more stable than hydrogen bonded rings, perhaps because its side chain packing is closer. Therefore, we suggest that a simple way to increase beta-hairpin stability, or the stability of an antiparallel edge strand, is to have a non-hydrogen bonded ring at the end of the strand.

Computational Biology↗