Search PubMed⌕ Search

Biomedical subjects

Kaizhong Zhang

Publications and source records attributed to Kaizhong Zhang.

12 recordsLinked to original sources

RNA-RNA interaction prediction and antisense RNA target search.

Recent studies demonstrating the existence of special noncoding "antisense" RNAs used in post transcriptional gene regulation have received considerable attention. These RNAs are synthesized naturally to control gene expression in C. elegans, Drosophila, and other organisms; they are known to regulate plasmid copy numbers in E. coli as well. Small RNAs have also been artificially constructed to knock out genes of interest in humans and other organisms for the purpose of finding out more about their functions. Although there are a number of algorithms for predicting the secondary structure of a single RNA molecule, no such algorithm exists for reliably predicting the joint secondary structure of two interacting RNA molecules or measuring the stability of such a joint structure. In this paper, we describe the RNA-RNA interaction prediction (RIP) problem between an antisense RNA and its target mRNA and develop efficient algorithms to solve it. Our algorithms minimize the joint free energy between the two RNA molecules under a number of energy models with growing complexity. Because the computational resources needed by our most accurate approach is prohibitive for long RNA molecules, we also describe how to speed up our techniques through a number of heuristic approaches while experimentally maintaining the original accuracy. Equipped with this fast approach, we apply our method to discover targets for any given antisense RNA in the associated genome sequence.

Adenosine Triphosphatases↗

Improving the sensitivity and specificity of protein homology search by incorporating predicted secondary structures.

In this paper, we improve the homology search performance by the combination of the predicted protein secondary structures and protein sequences. Previous research suggested that the straightforward combination of predicted secondary structures did not improve the homology search performance, mostly because of the errors in the structure prediction. We solved this problem by taking into account the confidence scores output by the prediction programs.

Computational Biology↗

MetricMap: an embedding technique for processing distance-based queries in metric spaces.

In this paper, we present an embedding technique, called MetricMap, which is capable of estimating distances in a pseudometric space. Given a database of objects and a distance function for the objects, which is a pseudometric, we map the objects to vectors in a pseudo-Euclidean space with a reasonably low dimension while preserving the distance between two objects approximately. Such an embedding technique can be used as an approximate oracle to process a broad class of distance-based queries. It is also adaptable to data mining applications such as data clustering and classification. We present the theory underlying MetricMap and conduct experiments to compare MetricMap with other methods including MVP-tree and M-tree in processing the distance-based queries. Experimental results on both protein and RNA data show the good performance and the superiority of MetricMap over the other methods.

Algorithms↗

SPIDER: software for protein identification from sequence tags with de novo sequencing error.

For the identification of novel proteins using MS/MS, de novo sequencing software computes one or several possible amino acid sequences (called sequence tags) for each MS/MS spectrum. Those tags are then used to match, accounting amino acid mutations, the sequences in a protein database. If the de novo sequencing gives correct tags, the homologs of the proteins can be identified by this approach and software such as MS-BLAST is available for the matching. However, de novo sequencing very often gives only partially correct tags. The most common error is that a segment of amino acids is replaced by another segment with approximately the same masses. We developed a new efficient algorithm to match sequence tags with errors to database sequences for the purpose of protein and peptide identification. A software package, SPIDER, was developed and made available on Internet for free public use. This paper describes the algorithms and features of the SPIDER software.

Algorithms↗

Multiple RNA structure alignment.

Ribonucleic Acid (RNA) structures can be viewed as a special kind of strings where characters in a string can bond with each other. The question of aligning two RNA structures has been studied for a while, and there are several successful algorithms that are based upon different models. In this paper, by adopting the model introduced in Wang and Zhang,(19) we propose two algorithms to attack the question of aligning multiple RNA structures. Our methods are to reduce the multiple RNA structure alignment problem to the problem of aligning two RNA structure alignments. Meanwhile, we will show that the framework of sequence center star alignment algorithm can be applied to the problem of multiple RNA structure alignment, and if the triangle inequality is met in the scoring matrix, the approximation ratio of the algorithm remains to be 2-2(over)n, where n is the total number of structures.

Algorithms↗

SPIDER: software for protein identification from sequence tags with de novo sequencing error.

For the identification of novel proteins using MS/MS, de novo sequencing software computes one or several possible amino acid sequences (called sequence tags) for each MS/MS spectrum. Those tags are then used to match, accounting amino acid mutations, the sequences in a protein database. If the de novo sequencing gives correct tags, the homologs of the proteins can be identified by this approach and software such as MS-BLAST is available for the matching. However, de novo sequencing very often gives only partially correct tags. The most common error is that a segment of amino acids is replaced by another segment with approximately the same masses. We developed a new efficient algorithm to match sequence tags with errors to database sequences for the purpose of protein and peptide identification. A software package, SPIDER, was developed and made available on Internet for free public use. This paper describes the algorithms and features of the SPIDER software.

Algorithms↗

Multiple RNA structure alignment.

RNA structures can be viewed as a kind of special strings with some characters bonded with each other. The question of aligning two RNA structures has been studied for a while, and there are several successful algorithms that are based upon different models. In this paper, by adopting the model introduced in [18], we propose two algorithms to attack the question of aligning multiple RNA structures. We reduce the multiple RNA structure alignment problem to the problem of aligning two RNA structure alignments.

Algorithms↗

An algorithm for detecting homologues of known structured RNAs in genomes.

Distinct RNA structures are frequently involved in a wide-range of functions in various biological mechanisms. The three dimensional RNA structures solved by X-ray crystallography and various well-established RNA phylogenetic structures indicate that functional RNAs have characteristic RNA structural motifs represented by specific combinations of base pairings and conserved nucleotides in the loop region. Discovery of well-ordered RNA structures and their homologues in genome-wide searches will enhance our ability to detect the RNA structural motifs and help us to highlight their association with functional and regulatory RNA elements. We present here a novel computer algorithm, HomoStRscan, that takes a single RNA sequence with its secondary structure to search for homologous-RNAs in complete genomes. This novel algorithm completely differs from other currently used search algorithms of homologous structures or structural motifs. For an arbitrary segment (or window) given in the target sequence, that has similar size to the query sequence, HomoStRscan finds the most similar structure to the input query structure and computes the maximal similarity score (MSS) between the two structures. The homologousRNA structures are then statistically inferred from the MSS distribution computed in the target genome. The method provides a flexible, robust and fine search tool for any homologous structural RNAs.

Algorithms↗

PEAKS: powerful software for peptide de novo sequencing by tandem mass spectrometry.

A number of different approaches have been described to identify proteins from tandem mass spectrometry (MS/MS) data. The most common approaches rely on the available databases to match experimental MS/MS data. These methods suffer from several drawbacks and cannot be used for the identification of proteins from unknown genomes. In this communication, we describe a new de novo sequencing software package, PEAKS, to extract amino acid sequence information without the use of databases. PEAKS uses a new model and a new algorithm to efficiently compute the best peptide sequences whose fragment ions can best interpret the peaks in the MS/MS spectrum. The output of the software gives amino acid sequences with confidence scores for the entire sequences, as well as an additional novel positional scoring scheme for portions of the sequences. The performance of PEAKS is compared with Lutefisk, a well-known de novo sequencing software, using quadrupole-time-of-flight (Q-TOF) data obtained for several tryptic peptides from standard proteins.

Amino Acid Sequence↗

RNA molecules with structure dependent functions are uniquely folded.

Cis-acting elements in post-transcriptional regulation of gene expression are often correlated with distinct local RNA secondary structure. These structures are expected to be significantly more ordered than those anticipated at random because of evolutionary constraints and intrinsic structural properties. In this study, we introduce a computing method to calculate two quantitative measures, NRd and Stscr, for estimating the uniqueness of an RNA secondary structure. NRd is a normalized score based on evaluating how different a natural RNA structure is from those predicted for its randomly shuffled variants. The lower the score NRd the more well ordered is the natural RNA structure. The statistical significance of NRd compared with that computed from structural comparisons among large numbers of randomly permuted sequences is represented by a standardized score, STSCR: We tested the method on the trans-activation response element and Rev response element of HIV-1 mRNA, internal ribosome entry sequence of hepatitis C virus, Tetrahymena thermophila rRNA intron, 100 tRNAs and 14 RNase P RNAs. Our data indicate that functional RNA structures have high Stscr, while other structures have low Stscr. We conclude that RNA functional molecules and/or cis-acting elements with structure dependent functions possess well ordered conformations and they are uniquely folded as measured by this technique.

Animals↗

A general edit distance between RNA structures.

Arc-annotated sequences are useful in representing the structural information of RNA sequences. In general, RNA secondary and tertiary structures can be represented as a set of nested arcs and a set of crossing arcs, respectively. Since RNA functions are largely determined by molecular confirmation and therefore secondary and tertiary structures, the comparison between RNA secondary and tertiary structures has received much attention recently. In this paper, we propose the notion of edit distance to measure the similarity between two RNA secondary and tertiary structures, by incorporating various edit operations performed on both bases and arcs (i.e., base-pairs). Several algorithms are presented to compute the edit distance between two RNA sequences with various arc structures and under various score schemes, either exactly or approximately, with provably good performance. Preliminary experimental tests confirm that our definition of edit distance and the computation model are among the most reasonable ones ever studied in the literature.

Algorithms↗