Search PubMed⌕ Search

Biomedical subjects

J P Huelsenbeck

Publications and source records attributed to J P Huelsenbeck.

At least 19 recordsLinked to original sources

Detecting positively selected amino acid sites using posterior predictive P-values.

Identifying positively selected amino acid sites is an important approach for making inference about the function of proteins; an amino acid site that is undergoing positive selection is likely to play a key role in the function of the protein. We present a new Bayesian method for identifying positively selected amino acid sites and apply the method to a data set of hemagglutinin sequences from the Influenza virus. We show that the results of the new methods are in accordance with results obtained using previous methods. More importantly, we also demonstrate how the method can be used for making further inferences about the evolutionary history of the sequences. For example, we demonstrate that sites that are positively selected tend to have a preponderance of conservative amino acid substitutions.

Amino Acid Sequence↗

Bayesian inference of phylogeny and its impact on evolutionary biology.

As a discipline, phylogenetics is becoming transformed by a flood of molecular data. These data allow broad questions to be asked about the history of life, but also present difficult statistical and computational problems. Bayesian inference of phylogeny brings a new perspective to a number of outstanding issues in evolutionary biology, including the analysis of large phylogenetic trees and complex evolutionary models and the detection of the footprint of natural selection in DNA sequences.

Algorithms↗

Phylogeny, genome evolution, and host specificity of single-stranded RNA bacteriophage (family Leviviridae).

Bacteriophage of the family Leviviridae have played an important role in molecular biology where representative species, such as Q beta and MS2, have been studied as model systems for replication, translation, and the role of secondary structure in gene regulation. Using nucleotide sequences from the coat and replicase genes we present the first statistical estimate of phylogeny for the family Leviviridae using maximum-likelihood and Bayesian estimation. Our analyses reveal that the coliphage species are a monophyletic group consisting of two clades representing the genera Levivirus and Allolevivirus. The Pseudomonas species PP7 diverged from its common ancestor with the coliphage prior to the ancient split between these genera and their subsequent diversification. Differences in genome size, gene composition, and gene expression are shown with a high probability to have changed along the lineage leading to the Allolevivirus through gene expansion. The change in genome size of the Allolevivirus ancestor may have catalyzed subsequent changes that led to their current genome organization and gene expression.

Allolevivirus↗

MRBAYES: Bayesian inference of phylogenetic trees.

SUMMARY: The program MRBAYES performs Bayesian inference of phylogeny using a variant of Markov chain Monte Carlo. AVAILABILITY: MRBAYES, including the source code, documentation, sample data files, and an executable, is available at http://brahms.biology.rochester.edu/software.html.

Algorithms↗

Empirical and hierarchical Bayesian estimation of ancestral states.

Several methods have been proposed to infer the states at the ancestral nodes on a phylogeny. These methods assume a specific tree and set of branch lengths when estimating the ancestral character state. Inferences of the ancestral states, then, are conditioned on the tree and branch lengths being true. We develop a hierarchical Bayes method for inferring the ancestral states on a tree. The method integrates over uncertainty in the tree, branch lengths, and substitution model parameters by using Markov chain Monte Carlo. We compare the hierarchical Bayes inferences of ancestral states with inferences of ancestral states made under the assumption that a specific tree is correct. We find that the methods are correlated, but that accommodating uncertainty in parameters of the phylogenetic model can make inferences of ancestral states even more uncertain than they would be in an empirical Bayes analysis.

Animals↗

Accommodating phylogenetic uncertainty in evolutionary studies.

Many evolutionary studies use comparisons across species to detect evidence of natural selection and to examine the rate of character evolution. Statistical analyses in these studies are usually performed by means of a species phylogeny to accommodate the effects of shared evolutionary history. The phylogeny is usually treated as known without error; this assumption is problematic because inferred phylogenies are subject to both stochastic and systematic errors. We describe methods for accommodating phylogenetic uncertainty in evolutionary studies by means of Bayesian inference. The methods are computationally intensive but general enough to be applied in most comparative evolutionary studies.

Animals↗

A compound poisson process for relaxing the molecular clock.

The molecular clock hypothesis remains an important conceptual and analytical tool in evolutionary biology despite the repeated observation that the clock hypothesis does not perfectly explain observed DNA sequence variation. We introduce a parametric model that relaxes the molecular clock by allowing rates to vary across lineages according to a compound Poisson process. Events of substitution rate change are placed onto a phylogenetic tree according to a Poisson process. When an event of substitution rate change occurs, the current rate of substitution is modified by a gamma-distributed random variable. Parameters of the model can be estimated using Bayesian inference. We use Markov chain Monte Carlo integration to evaluate the posterior probability distribution because the posterior probability involves high dimensional integrals and summations. Specifically, we use the Metropolis-Hastings-Green algorithm with 11 different move types to evaluate the posterior distribution. We demonstrate the method by analyzing a complete mtDNA sequence data set from 23 mammals. The model presented here has several potential advantages over other models that have been proposed to relax the clock because it is parametric and does not assume that rates change only at speciation events. This model should prove useful for estimating divergence times when substitution rates vary across lineages.

Bayes Theorem↗

A Bayesian framework for the analysis of cospeciation.

Information on the history of cospeciation and host switching for a group of host and parasite species is contained in the DNA sequences sampled from each. Here, we develop a Bayesian framework for the analysis of cospeciation. We suggest a simple model of host switching by a parasite on a host phylogeny in which host switching events are assumed to occur at a constant rate over the entire evolutionary history of associated hosts and parasites. The posterior probability density of the parameters of the model of host switching are evaluated numerically using Markov chain Monte Carlo. In particular, the method generates the probability density of the number of host switches and of the host switching rate. Moreover, the method provides information on the probability that an event of host switching is associated with a particular pair of branches. A Bayesian approach has several advantages over other methods for the analysis of cospeciation. In particular, it does not assume that the host or parasite phylogenies are known without error; many alternative phylogenies are sampled in proportion to their probability of being correct.

Animals↗

Variation in the pattern of nucleotide substitution across sites.

A model of nucleotide substitution that allows the transition/transversion rate bias to vary across sites was constructed. We examined the fit of this model using likelihood-ratio tests by analyzing 13 protein coding genes and 1 pseudogene. Likelihood-ratio testing indicated that a model that allows variation in the transition/transversion rate bias across sites provided a significant improvement in fit for most protein coding genes but not for the pseudogene. When the analysis was repeated with parameters estimated separately for first, second, and third codon positions, strong heterogeneity was uncovered for the first and second codon positions; the variation in the transition/transversion rate was generally weaker at the third codon position. The transition rate bias and branch lengths are underestimated when variation in the transition/transversion rate was not accommodated, suggesting that it may be important to accommodate variation in the pattern of nucleotide substitution for accurate estimation of evolutionary parameters.

Animals↗

Effect of nonindependent substitution on phylogenetic accuracy.

All current phylogenetic methods assume that DNA substitutions are independent among sites. However, ample empirical evidence suggests that the process of substitution is not independent but is, in fact, temporally and spatially correlated. The robustness of several commonly used phylogenetic methods to the assumption of independent substitution is examined. A compound Poisson process is used to model DNA substitution. This model assumes that substitution events are Poisson-distributed in time and that the number of substitutions associated with each event is geometrically distributed. The asymptotic properties of phylogenetic methods do not appear to change under a compound Poisson process of DNA substitution. Moreover, the rank order of the performance of different methods does not change. However, all phylogenetic methods become less efficient when substitution follows a compound Poisson process.

Base Composition↗

Base compositional bias and phylogenetic analyses: a test of the "flying DNA" hypothesis.

Phylogenetic methods can produce biased estimates of phylogeny when base composition varies along different lineages. Pettigrew (1994, Curr. Biol. 4:277-280) has suggested that base composition bias is responsible for the apparent support for the monophyly of bats (Chiroptera: megabats and microbats) from several different nuclear and mitochondrial genes. Pettigrew's "flying DNA" hypothesis makes several predictions: (1) that metabolic constraints associated with flying result in elevated levels of adenine and thymine throughout the genome of both megabats and microbats, (2) that the resulting base compositional bias in bats is sufficient to mislead phylogenetic methods and account for the support for bat monophyly from several nuclear and mitochondrial genes, and (3) that phylogenetic analysis using pairwise distances corrected for compositional bias should eliminate the support for bat monophyly. We tested these predictions by analyzing DNA sequences from two nuclear and three mitochondrial genes. The predicted base compositional bias does not appear to exist in some of the genes, and in other genes the differences in AT content are very small. Analyses under a wide diversity of criteria and models of evolution, including analyses that take base composition into account (using log-determinant distances), all strongly support bat monophyly. Moreover, simulation analyses indicate that even extreme bias toward AT-base composition in bats would be insufficient to explain the observed levels of support for bat monophyly. These analyses provide no support for the "flying DNA" hypothesis, whereas the monophyly of bats appears to be well supported by the DNA sequence data.

Animals↗

Phylogenetic methods come of age: testing hypotheses in an evolutionary context.

The use of molecular phylogenies to examine evolutionary questions has become commonplace with the automation of DNA sequencing and the availability of efficient computer programs to perform phylogenetic analyses. The application of computer simulation and likelihood ratio tests to evolutionary hypotheses represents a recent methodological development in this field. Likelihood ratio tests have enabled biologists to address many questions in evolutionary biology that have been difficult to resolve in the past, such as whether host-parasite systems are cospeciating and whether models of DNA substitution adequately explain observed sequences.

Animals↗

Exceptional convergent evolution in a virus.

Replicate lineages of the bacteriophage phiX 174 adapted to growth at high temperature on either of two hosts exhibited high rates of identical, independent substitutions. Typically, a dozen or more substitutions accumulated in the 5.4-kilobase genome during propagation. Across the entire data set of nine lineages, 119 independent substitutions occurred at 68 nucleotide sites. Over half of these substitutions, accounting for one third of the sites, were identical with substitutions in other lineages. Some convergent substitutions were specific to the host used for phage propagation, but others occurred across both hosts. Continued adaptation of an evolved phage at high temperature, but on the other host, led to additional changes that included reversions of previous substitutions. Phylogenetic reconstruction using the complete genome sequence not only failed to recover the correct evolutionary history because of these convergent changes, but the true history was rejected as being a significantly inferior fit to the data. Replicate lineages subjected to similar environmental challenges showed similar rates of substitution and similar rates of fitness improvement across corresponding times of adaptation. Substitution rates and fitness improvements were higher during the initial period of adaptation than during a later period, except when the host was changed.

Bacteriophage phi X 174↗

Is the Felsenstein zone a fly trap?

Although long-branch attraction, the incorrect grouping of long lineages in a phylogeny because of systematic error, has been identified as a potential source of error in phylogenetic analysis for almost two decades, no empirical examples of the phenomenon exist. Here, I outline several criteria for identifying long-branch attraction and apply these criteria to 18S ribosomal DNA (rDNA) sequence data for 13 insects. Parsimony and minimum evolution with p distances group the two longest branches together (those leading to Strepsiptera and Diptera). Simulation studies show that the long branches are long enough to attract. When a tree is assumed in which Strepsiptera and Diptera are separated and many data sets are simulated for that tree (using the parameter estimates for that tree for the original data), parsimony analysis of the simulated data consistently groups Strepsiptera and Diptera. Analyses of the 18S rDNA sequences using methods that are less sensitive to the problem of long-branch attraction estimate trees in which the long branches are separate.

Animals↗

The robustness of two phylogenetic methods: four-taxon simulations reveal a slight superiority of maximum likelihood over neighbor joining.

The robustness (sensitivity to violation of assumptions) of the maximum-likelihood and neighbor-joining methods was examined using simulation. Maximum likelihood and neighbor joining were implemented with Jukes-Cantor, Kimura, and gamma models of DNA substitution. Simulations were performed in which the assumptions of the methods were violated to varying degrees on three model four-taxon trees. The performance of the methods was evaluated with respect to ability to correctly estimate the unrooted four-taxon tree. Maximum likelihood outperformed neighbor joining in 29 of the 36 cases in which the assumptions of both methods were satisfied. In 133 of 180 of the simulations in which the assumptions of the maximum-likelihood and neighbor-joining methods were violated, maximum likelihood outperformed neighbor joining. These results are consistent with a general superiority of maximum likelihood over neighbor joining under comparable conditions. They extend and clarify an earlier study that found an advantage for neighbor joining over maximum likelihood for gamma-distributed mutation rates.

DNA↗