Search PubMed⌕ Search

Biomedical subjects

Irena I Artamonova

Publications and source records attributed to Irena I Artamonova.

4 recordsLinked to original sources

PEDANT genome database: 10 years online.

The PEDANT genome database provides exhaustive annotation of 468 genomes by a broad set of bioinformatics algorithms. We describe recent developments of the PEDANT Web server. The all-new Graphical User Interface (GUI) implemented in Javatrade mark allows for more efficient navigation of the genome data, extended search capabilities, user customization and export facilities. The DNA and Protein viewers have been made highly dynamic and customizable. We also provide Web Services to access the entire body of PEDANT data programmatically. Finally, we report on the application of association rule mining for automatic detection of potential annotation errors. PEDANT is freely accessible to academic users at http://pedant.gsf.de.

Computer Graphics↗

Mining sequence annotation databanks for association patterns.

MOTIVATION: Millions of protein sequences currently being deposited to sequence databanks will never be annotated manually. Similarity-based annotation generated by automatic software pipelines unavoidably contains spurious assignments due to the imperfection of bioinformatics methods. Examples of such annotation errors include over- and underpredictions caused by the use of fixed recognition thresholds and incorrect annotations caused by transitivity based information transfer to unrelated proteins or transfer of errors already accumulated in databases. One of the most difficult and timely challenges in bioinformatics is the development of intelligent systems aimed at improving the quality of automatically generated annotation. A possible approach to this problem is to detect anomalies in annotation items based on association rule mining. RESULTS: We present the first large-scale analysis of association rules derived from two large protein annotation databases-Swiss-Prot and PEDANT-and reveal novel, previously unknown tendencies of rule strength distributions. Most of the rules are either very strong or very weak, with rules in the medium strength range being relatively infrequent. Based on dynamics of error correction in subsequent Swiss-Prot releases and on our own manual analysis we demonstrate that exceptions from strong rules are, indeed, significantly enriched in annotation errors and can be used to automatically flag them. We identify different strength dependencies of rules derived from different fields in Swiss-Prot. A compositional breakdown of association rules generated from PEDANT in terms of their constituent items indicates that most of the errors that can be corrected are related to gene functional roles. Swiss-Prot errors are usually caused by under-annotation owing to its conservative approach, whereas automatically generated PEDANT annotation suffers from over-annotation. AVAILABILITY: All data generated in this study are available for download and browsing at http://pedant.gsf.de/ARIA/index.htm.

Conserved Sequence↗

Evolution of the exon-intron structure and alternative splicing of the MAGE-A family of cancer/testis antigens.

Cancer/testis antigens (CT-antigens) are proteins that are predominantly expressed in cancer and testis and thus are possible targets for immunotherapy. Most of them form large multigene families. The evolution of the MAGE-A family of CT-antigens is characterized by four processes: (1) gene duplications; (2) duplications of the initial exon; (3) point mutations and short insertions/deletions inactivating splicing sites or creating new sites; and (4) deletions removing sites and creating chimeric exons. All this concerns the genomic regions upstream of the coding region, creating a wide diversity of isoforms with different 5'-untranslated regions. Many of these isoforms are gene-specific and have emerged due to point mutations in alternative and constitutive splicing sites. There are also examples of chimeric mRNAs, likely produced by splicing of read-through transcripts. Since there is consistent use of homologous sites for different genes and no random, indiscriminant use of preexisting cryptic sites, it is likely that most observed isoforms are functional, and do not result from relaxed control in transformed cells.

Alternative Splicing↗

Low conservation of alternative splicing patterns in the human and mouse genomes.

Alternative splicing has recently emerged as a major mechanism of generating protein diversity in higher eukaryotes. We compared alternative splicing isoforms of 166 pairs of orthologous human and mouse genes. As the mRNA and EST libraries of human and mouse are not complete and thus cannot be compared directly, we instead analyzed whether known cassette exons or alternative splicing sites from one genome are conserved in the other genome. We demonstrate that about half of the analyzed genes have species-specific isoforms, and about a quarter of elementary alternatives are not conserved between the human and mouse genomes. The detailed results of this study are available at www.ig-msk.ru:8005/HMG_paper.

Alternative Splicing↗