Search PubMed⌕ Search

Biomedical subjects

Christina L Zheng

Publications and source records attributed to Christina L Zheng.

5 recordsLinked to original sources

MAASE: an alternative splicing database designed for supporting splicing microarray applications.

Alternative splicing is a prominent feature of higher eukaryotes. Understanding of the function of mRNA isoforms and the regulation of alternative splicing is a major challenge in the post-genomic era. The development of mRNA isoform sensitive microarrays, which requires precise splice-junction sequence information, is a promising approach. Despite the availability of a large number of mRNAs and ESTs in various databases and the efforts made to align transcript sequences to genomic sequences, existing alternative splicing databases do not offer adequate information in an appropriate format to aid in splicing array design. Here we describe our effort in constructing the Manually Annotated Alternatively Spliced Events (MAASE) database system, which is specifically designed to support splicing microarray applications. MAASE comprises two components: (1) a manual/computational annotation tool for the efficient extraction of critical sequence and functional information for alternative splicing events and (2) a user-friendly database of annotated events that allows convenient export of information to aid in microarray design and data analysis. We provide a detailed introduction and a step-by-step user guide to the MAASE database system to facilitate future large-scale annotation efforts, integration with other alternative splicing databases, and splicing array fabrication.

Alternative Splicing↗

Characteristics and regulatory elements defining constitutive splicing and different modes of alternative splicing in human and mouse.

Alternative splicing is a major contributor to genomic complexity, disease, and development. Previous studies have captured some of the characteristics that distinguish alternative splicing from constitutive splicing. However, most published work only focuses on skipped exons and/or a single species. Here we take advantage of the highly curated data in the MAASE database (see related paper in this issue) to analyze features that characterize different modes of splicing. Our analysis confirms previous observations about alternative splicing, including weaker splicing signals at alternative splice sites, higher sequence conservation surrounding orthologous alternative exons, shorter exon length, and more frequent reading frame maintenance in skipped exons. In addition, our study reveals potentially novel regulatory principles underlying distinct modes of alternative splicing and a role of a specific class of repeat elements (transposons) in the origin/evolution of alternative exons. These features suggest diverse regulatory mechanisms and evolutionary paths for different modes of alternative splicing.

Alternative Splicing↗

Rival penalized competitive learning (RPCL): a topology-determining algorithm for analyzing gene expression data.

DNA arrays have become the immediate choice in the analysis of large-scale expression measurements. Understanding the expression pattern of genes provide functional information on newly identified genes by computational approaches. Gene expression pattern is an indicator of the state of the cell, and abnormal cellular states can be inferred by comparing expression profiles. Since co-regulated genes, and genes involved in a particular pathway, tend to show similar expression patterns, clustering expression patterns has become the natural method of choice to differentiate groups. However, most methods based on cluster analysis suffer from the usual problems (i) dead units, and (ii) the problem of determining the correct number of clusters (k) needed to classify the data. Selecting the k has been an open problem of pattern recognition and statistics for decades. Since clustering reveals similar patterns present in the data, fixing this number strongly influences the quality of the result. While there is no theoretical solution to this problem, the number of clusters can be decided by a heuristic clustering algorithm called rival penalized competitive learning (RPCL). We present a novel implementation of RPCL that transforms the correct number of clusters problem to the tractable problem of clustering based on the degree of similarity. This is biologically significant since our implementation clusters functionally co-regulated genes and genes that present similar patterns of expression. This new approach reveals potential genes that are co-involved in a biological process. This implementation of the RPCL algorithm is useful in differentiating groups involved in concerted functional regulation and helps to progressively home into patterns, which are closely similar.

Algorithms↗

MODULEWRITER: a program for automatic generation of database interfaces.

MODULEWRITER is a PERL object relational mapping (ORM) tool that automatically generates database specific application programming interfaces (APIs) for SQL databases. The APIs consist of a package of modules providing access to each table row and column. Methods for retrieving, updating and saving entries are provided, as well as other generally useful methods (such as retrieval of the highest numbered entry in a table). MODULEWRITER provides for the inclusion of user-written code, which can be preserved across multiple runs of the MODULEWRITER program.

Computational Biology↗

On selecting features from splice junctions: an analysis using information theoretic and machine learning approaches.

The computational recognition of precise splice junctions is a challenge faced in the analysis of newly sequenced genomes. This is challenging due to the fact that the distribution of sequence patterns in these regions is not always distinct. Our objective is to understand the sequence signatures at the splice junctions, not simply to create an artificial recognition system. We use a combination of a neural network based calliper randomization approach and an information theoretic based feature selection approach for this purpose. This has been done in an effort to understand regions that harbor information content and to extract features relevant for the prediction of splice junctions. The analysis using the neural network based calliper randomization approach revealed regions important in the internal representation of the network model. The calliper approach captured both correlated as well as independently important features. The feature selection approach captures features that are independently informative. The two different methods can capture features with different properties. Comparative analysis of the results using both the methods help to infer about the kind of information present in the region.

Alternative Splicing↗