Search PubMed⌕ Search

Biomedical subjects

Ward C Wheeler

Publications and source records attributed to Ward C Wheeler.

5 recordsLinked to original sources

Iterative pass optimization of sequence data.

The problem of determining the minimum-cost hypothetical ancestral sequences for a given cladogram is known to be NP-complete. This "tree alignment" problem has motivated the considerable effort placed in multiple sequence alignment procedures. Wheeler in 1996 proposed a heuristic method, direct optimization, to calculate cladogram costs without the intervention of multiple sequence alignment. This method, though more efficient in time and more effective in cladogram length than many alignment-based procedures, greedily optimizes nodes based on descendent information only. In their proposal of an exact multiple alignment solution, Sankoff et al. in 1976 described a heuristic procedure--the iterative improvement method--to create alignments at internal nodes by solving a series of median problems. The combination of a three-sequence direct optimization with iterative improvement and a branch-length-based cladogram cost procedure, provides an algorithm that frequently results in superior (i.e., lower) cladogram costs. This iterative pass optimization is both computation and memory intensive, but economies can be made to reduce this burden. An example in arthropod systematics is discussed.

Algorithms↗

Implied alignment: a synapomorphy-based multiple-sequence alignment method and its use in cladogram search.

A method to align sequence data based on parsimonious synapomorphy schemes generated by direct optimization (DO; earlier termed optimization alignment) is proposed. DO directly diagnoses sequence data on cladograms without an intervening multiple-alignment step, thereby creating topology-specific, dynamic homology statements. Hence, no multiple-alignment is required to generate cladograms. Unlike general and globally optimal multiple-alignment procedures, the method described here, implied alignment (IA), takes these dynamic homologies and traces them back through a single cladogram, linking the unaligned sequence positions in the terminal taxa via DO transformation series. These "lines of correspondence" link ancestor-descendent states and, when displayed as linearly arrayed columns without hypothetical ancestors, are largely indistinguishable from standard multiple alignment. Since this method is based on synapomorphy, the treatment of certain classes of insertion-deletion (indel) events may be different from that of other alignment procedures. As with all alignment methods, results are dependent on parameter assumptions such as indel cost and transversion:transition ratios. Such an IA could be used as a basis for phylogenetic search, but this would be questionable since the homologies derived from the implied alignment depend on its natal cladogram and any variance, between DO and IA + Search, due to heuristic approach. The utility of this procedure in heuristic cladogram searches using DO and the improvement of heuristic cladogram cost calculations are discussed.

Animals↗

'Pluralism' and the aims of phylogenetic research.

In science, and particularly in the field of phylogenetic systematics, investigators may choose among different methods to analyze their data. These methods include neighbor-joining (or other genetic distance approaches), maximum-likelihood, and cladistic parsimony, among others. These distinct methods of analysis differ considerably in how they process information from the observed data. However, many published molecular analyses utilize trees generated under more than one of these methods, which we will call a 'pluralistic' approach. Here, we explore the statistical, philosophical and operational aspects of the pluralistic approach. We suggest that the pluralistic approach is misguided from all three perspectives and we propose an alternative, logically consistent, strategy as an aim of phylogenetic research.

Animals↗

DNA multiple sequence alignments.

In this chapter we examine the procedure of multiple sequence alignment. We first examine the heuristic procedures commonly used in multiple sequence alignment. Next we examine sources of ambiguity involved in the alignment procedure. We suggest that several alignment parameters be employed to examine alignment sensitivity. We end by presenting an experiment with humans showing the ambiguity involved in manual alignment.

Humans↗

Theory and practice of parallel direct optimization.

Our ability to collect and distribute genomic and other biological data is growing at a staggering rate (Pagel, 1999). However, the synthesis of these data into knowledge of evolution is incomplete. Phylogenetic systematics provides a unifying intellectual approach to understanding evolution but presents formidable computational challenges. A fundamental goal of systematics, the generation of evolutionary trees, is typically approached as two distinct NP-complete problems: multiple sequence alignment and phylogenetic tree search. The number of cells in a multiple alignment matrix are exponentially related to sequence length. In addition, the number of evolutionary trees expands combinatorially with respect to the number of organisms or sequences to be examined. Biologically interesting datasets are currently comprised of hundreds of taxa and thousands of nucleotides and morphological characters. This standard will continue to grow with the advent of highly automated sequencing and development of character databases. Three areas of innovation are changing how evolutionary computation can be addressed: (1) novel concepts for determination of sequence homology, (2) heuristics and shortcuts in tree-search algorithms, and (3) parallel computing. In this paper and the online software documentation we describe the basic usage of parallel direct optimization as implemented in the software POY (ftp://ftp.amnh.org/pub/molecular/poy).

Animals↗