Search PubMedSearch

Biomedical subjects

P O Lewis

Publications and source records attributed to P O Lewis.

4 recordsLinked to original sources

A genetic algorithm for maximum-likelihood phylogeny inference using nucleotide sequence data.

Phylogeny reconstruction is a difficult computational problem, because the number of possible solutions increases with the number of included taxa. For example, for only 14 taxa, there are more than seven trillion possible unrooted phylogenetic trees. For this reason, phylogenetic inference methods commonly use clustering algorithms (e.g., the neighbor-joining method) or heuristic search strategies to minimize the amount of time spent evaluating nonoptimal trees. Even heuristic searches can be painfully slow, especially when computationally intensive optimality criteria such as maximum likelihood are used. I describe here a different approach to heuristic searching (using a genetic algorithm) that can tremendously reduce the time required for maximum-likelihood phylogenetic inference, especially for data sets involving large numbers of taxa. Genetic algorithms are simulations of natural selection in which individuals are encoded solutions to the problem of interest. Here, labeled phylogenetic trees are the individuals, and differential reproduction is effected by allowing the number of offspring produced by each individual to be proportional to that individual's rank likelihood score. Natural selection increases the average likelihood in the evolving population of phylogenetic trees, and the genetic algorithm is allowed to proceed until the likelihood of the best individual ceases to improve over time. An example is presented involving rbcL sequence data for 55 taxa of green plants. The genetic algorithm described here required only 6% of the computational effort required by a conventional heuristic search using tree bisection/reconnection (TBR) branch swapping to obtain the same maximum-likelihood topology.

Algorithms

Success of maximum likelihood phylogeny inference in the four-taxon case.

We used simulated data to investigate a number of properties of maximum-likelihood (ML) phylogenetic tree estimation for the case of four taxa. Simulated data were generated under a broad range of conditions, including wide variation in branch lengths, differences in the ratio of transition and transversion substitutions, and the absence of presence of gamma-distributed site-to-site rate variation. Data were analyzed in the ML framework with two different substitution models, and we compared the ability of the two models to reconstruct the correct topology. Although both models were inconsistent for some branch-length combinations in the presence of site-to-site variation, the models were efficient predictors of topology under most simulation conditions. We also examined the performance of the likelihood ratio (LR) test for significant positive interior branch length. This test was found to be misleading under many simulation conditions, rejecting too often under some simulation conditions. Under the null hypothesis of zero length internal branch, LR statistics are assumed to be asymptotically distributed chi 2(1); with limited data, the distribution of LR statistics under the null hypothesis varies from chi 2(1).

Computer Simulation

Deterministic paternity exclusion using RAPD markers.

The Random Amplified Polymorphic DNA (RAPD) technique can potentially provide hundreds of polymorphic markers for use by ecologists studying mating systems in natural populations. We consider here the implications of the dominance displayed by RAPD markers for deterministic paternity assignment. Our goal was to provide a means for assessing the costs associated with such a study for ecologists who might be considering the use of RAPD markers for paternity analysis. The theoretical expected proportion of offspring for which all males except the true father can be exlucded (P(ET)) is calculated for both dominant and codominant marker systems. The ability to assign paternity unambiguously generally increases with the number of loci and the frequency of the recessive allele (but only up to a point), and decreases with increasing sample size (number of individuals surveyed). The gain in P(ET) with decreasing sample size is unexpectedly slight. Not surprisingly, the performance of dominant markers at paternity exclusion is, in general, greatly exceeded by codominant markers, with the exception of the case in which the frequency of the recessive allele is high at all loci. In this case, codominant markers perform only slightly better than do dominant markers. Thus, a researcher should expect to score more than 50 RAPD loci for each offspring for most applications of paternity exclusion analysis.(ABSTRACT TRUNCATED AT 250 WORDS)

Alleles