Search PubMed⌕ Search

PubMed · 12150718

Kangaroo--a pattern-matching program for biological sequences.

Abstract

BACKGROUND: Biologists are often interested in performing a simple database search to identify proteins or genes that contain a well-defined sequence pattern. Many databases do not provide straightforward or readily available query tools to perform simple searches, such as identifying transcription binding sites, protein motifs, or repetitive DNA sequences. However, in many cases simple pattern-matching searches can reveal a wealth of information. We present in this paper a regular expression pattern-matching tool that was used to identify short repetitive DNA sequences in human coding regions for the purpose of identifying potential mutation sites in mismatch repair deficient cells. RESULTS: Kangaroo is a web-based regular expression pattern-matching program that can search for patterns in DNA, protein, or coding region sequences in ten different organisms. The program is implemented to facilitate a wide range of queries with no restriction on the length or complexity of the query expression. The program is accessible on the web at http://bioinfo.mshri.on.ca/kangaroo/ and the source code is freely distributed at http://sourceforge.net/projects/slritools/. CONCLUSION: A low-level simple pattern-matching application can prove to be a useful tool in many research settings. For example, Kangaroo was used to identify potential genetic targets in a human colorectal cancer variant that is characterized by a high frequency of mutations in coding regions containing mononucleotide repeats.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Doron Betel, Christopher W V Hogue. 2002-07-31. Kangaroo--a pattern-matching program for biological sequences.. https://doi.org/10.1186/1471-2105-3-20

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Analysis of mutations made during active synthesis or extension of mismatched substrates further define the mechanism of HIV-RT mutagenesis.

The effect of reverse transcriptase (RT) catalyzed mutations on continued extension of the nascent DNA chain was investigated. A system using the alpha-lac gene of beta-galactosidase as template and two sets of conditions was used. In one, RT was allowed to reassociate with the primer-template after falling off, while in a second RT was sequestered after dissociating. In the first condition, subsequent extension of errors that may have initially caused enzyme dissociation can occur. In the second, such errors would not be extended. Fully extended products were assayed by alpha-complementation to assess mutation frequency. A lower frequency in the latter scenario implies that some errors caused the polymerase to dissociate. Allowing only a single binding event lowered the mutation frequency of the products by about 1/2 suggesting that approximately 1 in 2 errors terminated synthesis. In other experiments, when added to a primer-template with a terminal mismatch at the 3' end, RT dissociated from the template about 50-90% of the time (depending on mismatch type) rather than extending. Running start reactions indicated that extension was more likely if an actively synthesizing RT made the mutation. RT RNase H cleavage analysis showed that 3' mismatches weakened the association of RT with the primer-terminus. Taken together, these results suggest that an actively synthesizing RT enzyme that has just made a mistake is likely bound in a configuration that generally enhances extension of the mistake. This is in contrast to RTs that must bind to then extend mismatches. The importance of these findings with respect to the mechanism of mutagenesis is discussed.

Base Pair Mismatch↗

Solution structure of the ActD-5'-CCGTT3GTGG-3' complex: drug interaction with tandem G.T mismatches and hairpin loop backbone.

Binding of actinomycin D (ActD) to the seemingly single-stranded DNA (ssDNA) oligomer 5'-CCGTT3 GTGG-3' has been studied in solution using high-resolution nuclear magnetic resonance (NMR) techniques. A strong binding constant (8 x 10(6) M(-1)) and high quality NMR spectra have allowed us to determine the initial DNA structure using distance geometry as well as the final ActD-5'-CCGTT3 GTGG-3' complex structure using constrained molecular dynamics calculations. The DNA oligomer 5'-CCGTT3GTGG-3' in the complex forms a hairpin structure with tandem G.T mismatches at the stem region next to a loop of three stacked thymine bases pointing toward the major groove. Bipartite T2O-GH1 and T2O-G2NH2 hydrogen bonds were detected for the G.T mismatches that further stabilize this unusual DNA hairpin. The phenoxazone chromophore of ActD intercalates nicely between the tandem G.T mismatches in essentially one major orientation. Additional hydrophobic interactions between the ActD quinoid amino acid residues with the loop T5-T6-T7 backbone protons were also observed. The hydrophobic G-phenoxazone-G interaction in the ActD-5'-CCGTT3GTGG-3' complex is more robust than that of the classical ActD- 5'-CCGCT3GCGG-3' complex, consistent with the roughly 2-fold stronger binding of ActD to the 5'-CCGTT3GTGG-3' sequence than to its 5'-CCG CT3GCGG-3' counterpart. Stabilization by ActD of a hairpin containing non-canonical stem base pairs further strengthens the notion that ActD or other related compounds may serve as a sequence- specific ssDNA-binding agent that inhibits human immunodeficiency virus (HIV) and other retroviruses replicating through ssDNA intermediates.

Base Pair Mismatch↗

Unusual DNA duplex and hairpin motifs.

Single-stranded DNA or double-stranded DNA has the potential to adopt a wide variety of unusual duplex and hairpin motifs in the presence (trans) or absence (cis) of ligands. Several principles for the formation of those unusual structures have been established through the observation of a number of recurring structural motifs associated with different sequences. These include: (i) internal loops of consecutive mismatches can occur in a B-DNA duplex when sheared base pairs are adjacent to each other to confer extensive cross- and intra-strand base stacking; (ii) interdigitated (zipper-like) duplex structures form instead when sheared G*A base pairs are separated by one or two pairs of purine*purine mismatches; (iii) stacking is not restricted to base, deoxyribose also exhibits the potential to do so; (iv) canonical G*C or A.T base pairs are flexible enough to exhibit considerable changes from the regular H-bonded conformation. The paired bases become stacked when bracketed by sheared G.A base pairs, or become extruded out and perpendicular to their neighboring bases in the presence of interacting drugs; (v) the purine-rich and pyrimidine-rich loop structures are notably different in nature. The purine-rich loops form compact triloop structures closed by a sheared G*A, A*A, A*C or sheared-like G(anti)*C(syn) base pair that is stacked by a single residue. On the other hand, the pyrimidine-rich loops with a thymidine in the first position exhibit no base pairing but are characterized by the folding of the thymidine residue into the minor groove to form a compact loop structure. Identification of such diverse duplex or hairpin motifs greatly enlarges the repertoire for unusual DNA structural formation.

Base Pair Mismatch↗