Search PubMed⌕ Search

Biomedical subjects

H T Wareham

Publications and source records attributed to H T Wareham.

4 recordsLinked to original sources

Exact algorithms for computing pairwise alignments and 3-medians from structure-annotated sequences (extended abstract).

Given the problem of mutation saturation in ancient molecular sequences, there is great interest in inferring phylogenies from higher-order types of molecular data that change more slowly, such as genomic organization and the secondary and tertiary structures of ribosomal RNA and proteins. In this paper, we define edit distances based on two representations of RNA secondary structure, arc annotation and hierarchical string annotation, and give algorithms for computing these distances on pairs of annotated sequences, aligning pairs of annotated sequences, and computing 3-median annotated sequences from triples of annotated sequences. The 3-median algorithms can be used as part of a well-known iterative heuristic for inferring phylogenies. All given algorithms are adapted from algorithms for computing longest common annotated subsequences of pairs of annotated sequences.

Algorithms↗

Stochastic heuristic algorithms for target motif identification (extended abstract).

Target motifs are motifs that are "close" to one or more substrings in each sequence in one given set of sequences but are far from every substring in another given set of sequences. Target motifs have pharmaceutical applications; unfortunately, the problem of identifying target motifs is NP-hard and is thus unlikely to have efficient optimal solution algorithms. In this paper, we propose a set of simple modifications to the Gibbs Sampling heuristic for finding motifs which allows this heuristic to detect target motifs. We also present the results of several experiments relative to both simulated and real datasets which suggest that this modified heuristic is good at detecting target motifs under a variety of conditions.

Algorithms↗

A simplified proof of the NP- and MAX SNP-hardness of multiple sequence tree alignment.

We give a simple proof which shows that the multiple sequence tree alignment problem from molecular biology is both NP-complete and MAX SNP-hard. Our proof of MAX SNP-hardness is simpler than that given previously by Wang and Jiang. These results suggest that it is unlikely that the multiple sequence tree alignment problem has polynomial-time algorithms that produce either optimal solutions or approximate solutions whose cost may be arbitrarily close to optimal.

Algorithms↗

Parameterized complexity analysis in computational biology.

Many computational problems in biology involve parameters for which a small range of values cover important applications. We argue that for many problems in this setting, parameterized computational complexity rather than NP-completeness is the appropriate tool for studying apparent intractability. At issue in the theory of parameterized complexity is whether a problem can be solved in time O(n alpha) for each fixed parameter value, where alpha is a constant independent of the parameter. In addition to surveying this complexity framework, we describe a new result for the Longest Common Subsequence problem. In particular, we show that the problem is hard for W[t] for all t when parameterized by the number of strings and the size of the alphabet. Lower bounds on the complexity of this basic combinatorial problem imply lower bounds on more general sequence alignment and consensus discovery problems. We also describe a number of open problems pertaining to the parameterized complexity of problems in computational biology where small parameter values are important.

Algorithms↗