Biomedical subjects
A Rambaut
Publications and source records attributed to A Rambaut.
Estimating divergence dates from molecular sequences.
The ability to date the time of divergence between lineages using molecular data provides the opportunity to answer many important questions in evolutionary biology. However, molecular dating techniques have previously been criticized for failing to adequately account for variation in the rate of molecular evolution. We present a maximum-likelihood approach to estimating divergence times that deals explicitly with the problem of rate variation. This method has many advantages over previous approaches including the following: (1) a rate constancy test excludes data for which rate heterogeneity is detected; (2) date estimates are generated with confidence intervals that allow the explicit testing of hypotheses regarding divergence times; and (3) a range of sequences and fossil dates are used, removing the reliance on a single calculated calibration rate. We present tests of the accuracy of our method, which show it to be robust to the effects of some modes of rate variation. In addition, we test the effect of substitution model and length of sequence on the accuracy of the dating technique. We believe that the method presented here offers solutions to many of the problems facing molecular dating and provides a platform for future improvements to such analyses.
Elucidating the population histories and transmission dynamics of papillomaviruses using phylogenetic trees.
Using gene genealogies constructed from gene sequence data, we show that both the mucosal and cutaneous papillomaviruses (PV)-supergroups A and B-appear to have been transmitted through susceptible populations faster than exponentially. The data and methods involved (1) examining the PV database for phylogenetic signal in an L1 open reading frame (ORF) fragment and an E1 ORF segment, (2) demonstrating that the same two fragments have evolved in a way consistent with a molecular clock, and (3) applying methods of phylogenetic tree analysis that test different scenarios for the dynamics of viral transmission within populations. The results indicate increases in PV populations of both supergroups A and B in the recent past. This form of the increases, which fit a null model of population growth with an exponent increasing in time, is compatible with the fact that human populations have grown at a faster than exponential rate, thus increasing the numbers of susceptible hosts for HPVs. There are, however, indications that the population of supergroup A has now stopped increasing in size.
Seq-Gen: an application for the Monte Carlo simulation of DNA sequence evolution along phylogenetic trees.
MOTIVATION: Seq-Gen is a program that will simulate the evolution of nucleotide sequences along a phylogeny, using common models of the substitution process. A range of models of molecular evolution are implemented, including the general reversible model. Nucleotide frequencies and other parameters of the model may be given and site-specific rate heterogeneity can also be incorporated in a number of ways. Any number of trees may be read in and the program will produce any number of data sets for each tree. Thus, large sets of replicate simulations can be easily created. This can be used to test phylogenetic hypotheses using the parametric bootstrap. AVAILABILITY: Seq-Gen can be obtained by WWW from http:/(/)evolve.zoo.ox.ac.uk/Seq-Gen/seq-gen.html++ + or by FTP from ftp:/(/)evolve.zoo.ox.ac.uk/packages/Seq-Gen/. The package includes the source code, manual and example files. An Apple Macintosh version is available from the same sites.
End-Epi: an application for inferring phylogenetic and population dynamical processes from molecular sequences.
MOTIVATION: Phylogenetic trees constructed from molecular sequences contain information about the evolutionary or population dynamical processes that created them. Here we describe a computer package (End-Epi) that uses graphical methods to allow researchers to make inferences about these processes from their data. Statistical analyses can be performed to test the consistency of the data with various competing hypotheses. AVAILABILITY: End-Epi can be obtained by WWW from http://evolve.zoo.ox.ac.uk/ and by anonymous FTP from ftp://evolve.zoo.ox.ac.uk/packages/End-Epi10.hqx. This file contains the compiled application, the manual and a test tree.
PSeq-Gen: an application for the Monte Carlo simulation of protein sequence evolution along phylogenetic trees.
Explore the source record for details and available documents.
Inferring the population history of an epidemic from a phylogenetic tree.
Phylogenetic or family trees reconstructed from a sample of individuals belonging to a population or species can be used to infer population dynamic history. A method for making such inferences involves visual inspection of the graphs of the numbers of lineages in the phylogeny plotted against the times when they occur. One transformation of the lineages axis often produces a linear lineages-through-time plot if the population has been growing exponentially. However, that transformation is known to fail under certain circumstances. We report a series of simulation studies designed to determine the conditions under which the transformation fails. As long as the sample represents less than 0.5% of individuals in the total population, the Type 1 error rate is acceptable. This means that studies of virus populations are likely not to suffer from transformation failure, so long as a random sample from the population is represented. When the transformation does fail, the resultant lineages-through-time plot is very similar to a two-stage epidemic process.
Isolation and sequence analysis of a cDNA encoding the c subunit of a vacuolar-type H(+)-ATPase from the CAM plant Kalanchoë daigremontiana.
We report the sequence of a cDNA clone encoding the c ("16 kDa') subunit of a vacuolar-type H(+)-ATPase (V-ATPase) from Kalanchoë daigremontiana, a plant in which the cell vacuole plays a pivotal role in crassulacean acid metabolism. The clone, pKVA211, was isolated from a K. daigremontiana leaf cDNA library constructed in lambda ZAP II using a homologous PCR-generated cDNA probe for the V-ATPase c subunit. The KVA211 cDNA was 839 nucleotides long and included a 20 bp poly(A)+ tail together with a complete 495 bp coding region for a polypeptide with a predicted molecular mass of 16659 Da. The deduced amino acid sequence was highly conserved across the wide range of eukaryotes (vertebrates, invertebrates, fungi, plants and protozoa) in which this gene has now been identified. Sequence comparison of several PCR products and genomic Southern analysis indicated that the V-ATPase c subunit in K. daigremontiana is encoded by a small multi-gene family. Steady-state levels of the KVA211 mRNA were much higher in leaves than in roots or flowers, and expression of this transcript in leaves was shown to be strongly light-dependent.
Recombination between sequences of hepatitis B virus from different genotypes.
A comparison of 25 hepatitis B virus (HBV) isolates for which complete genome sequences are available revealed two that occupied different positions in phylogenetic trees reconstructed from different open reading frames. Further analysis indicated that this incongruence was the result of recombination between viruses of different genomic and antigenic types. Both putative recombinants originated from geographic regions where multiple genotypes are known to cocirculate. A search of the sequence databases showed evidence of similar intergenotypic recombinants. These observations indicate that recombination between divergent strains may represent an important source of genetic variation in HBV.
Determinants of rate variation in mammalian DNA sequence evolution.
Attempts to analyze variation in the rates of molecular evolution among mammalian lineages have been hampered by paucity of data and by nonindependent comparisons. Using phylogenetically independent comparisons, we test three explanations for rate variation which predict correlations between rate variation and generation time, metabolic rate, and body size. Mitochondrial and nuclear genes, protein coding, rRNA, and nontranslated sequences from 61 mammal species representing 14 orders are used to compare the relative rates of sequence evolution. Correlation analyses performed on differences in genetic distance since common origin of each pair against differences in body mass, generation time, and metabolic rate reveal that substitution rate at fourfold degenerate sites in two out of three protein sequences is negatively correlated with generation time. In addition, there is a relationship between the rate of molecular evolution and body size for two nuclear-encoded sequences. No evidence is found for an effect of metabolic rate on rate of sequence evolution. Possible causes of variation in substitution rate between species are discussed.
Bi-De: an application for simulating phylogenetic processes.
Birth-Death (Bi-De) is an application for the Apple Macintosh which simulates the growth of phylogenetic trees using various models of lineage birth and death. The trees produced are intended to be analogous to those reconstructed from molecular sequence data. The user may define a constant birth rate and death rate or a function describing how these rates vary by time or population size. Instantaneous mass extinctions can also be simulated. The package allows the tree produced to be used as a template for the simulated evolution of molecular sequence data under a range of different transition models.
Complete nucleotide sequence and transcriptional analysis of snakehead fish retrovirus.
The complete genome of the snakehead fish retrovirus has been cloned and sequenced, and its transcriptional profile in cell culture has been determined. The 11.2-kb provirus displays a complex expression pattern capable of encoding accessory proteins and is unique in the predicted location of the env initiation codon and signal peptide upstream of gag and the common splice donor site. The virus is distinguishable from all known retrovirus groups by the presence of an arginine tRNA primer binding site. The coding regions are highly divergent and show a number of unusual characteristics, including a large Gag coiled-coil region, a Pol domain of unknown function, and a long, lentiviral-like, Env cytoplasmic domain. Phylogenetic analysis of the Pol sequence emphasizes the divergent nature of the virus from the avian and mammalian retroviruses. The snakehead virus is also distinct from a previously characterized complex fish retrovirus, suggesting that discrete groups of these viruses have yet to be identified in the lower vertebrates.
Inferring population history from molecular phylogenies.
Variable molecular sequences sampled from a population can be used to infer its dynamic history. Graphical methods are developed and applied to real data, illustrating ways of navigating through hypothesis space with two landmarks for reference: constant population size and exponentially growing population size.
Revealing the history of infectious disease epidemics through phylogenetic trees.
Phylogenetic trees play an increasing role in molecular epidemiology, where they have been used to understand the forces that shape patterns of viral sequence diversity. Phylogenetic trees can also be used to trace the dynamics of viral transmission within populations. Case studies document the worldwide spread of Human Immunodeficiency Virus type 1 (HIV-1) and hepatitis C virus (HCV). Despite similarities between these viruses, especially in their transmission routes, they are shown to have very different epidemiological histories. A possible reason for the difference is that HCV has coexisted longer with human populations.
Comparative analysis by independent contrasts (CAIC): an Apple Macintosh application for analysing comparative data.
CAIC is an application for the Apple Macintosh which allows the valid analysis of comparative (multi-species) data sets that include continuous variables. Comparison among species is the most common technique for testing hypotheses of how organisms are adapted to their environments, but standard statistical tests like regression should not be used with species data. Such tests assume independence of data points, but related species often share traits by common descent rather than through independent adaptation. CAIC uses a phylogeny of the species in the data set to partition the variance among species into independent comparisons (technically, linear contrasts), each comparison being made at a different node in the phylogeny. There are two partitioning procedures--one used when all variables are continuous, the other when one variable is discrete. The resulting comparisons can be analysed validly in standard statistical packages to test hypotheses about correlated evolution among traits, to estimate parameters such as allometric exponents, and to compare rates of evolution. Previous versions of the package have already been used widely; this version is simpler to use and works on a wider range of machines. The package and manual are freely available by anonymous ftp or from the authors.