Search PubMed⌕ Search

Biomedical subjects

David Posada

Publications and source records attributed to David Posada.

At least 19 recordsLinked to original sources

GARD: a genetic algorithm for recombination detection.

MOTIVATION: Phylogenetic and evolutionary inference can be severely misled if recombination is not accounted for, hence screening for it should be an essential component of nearly every comparative study. The evolution of recombinant sequences can not be properly explained by a single phylogenetic tree, but several phylogenies may be used to correctly model the evolution of non-recombinant fragments. RESULTS: We developed a likelihood-based model selection procedure that uses a genetic algorithm to search multiple sequence alignments for evidence of recombination breakpoints and identify putative recombinant sequences. GARD is an extensible and intuitive method that can be run efficiently in parallel. Extensive simulation studies show that the method nearly always outperforms other available tools, both in terms of power and accuracy and that the use of GARD to screen sequences for recombination ensures good statistical properties for methods aimed at detecting positive selection. AVAILABILITY: Freely available http://www.datamonkey.org/GARD/

Algorithms↗

MtArt: a new model of amino acid replacement for Arthropoda.

A statistical approach was applied to select those models that best fit each individual mitochondrial (mt) protein at different taxonomic levels of metazoans. The existing mitochondrial replacement matrices, MtREV and MtMam, were found to be the best-fit models for the mt-proteins of vertebrates, with the exception of Nd6, at different taxonomic levels. Remarkably, existing mitochondrial matrices generally failed to best-fit invertebrate mt-proteins. In an attempt to better model the evolution of invertebrate mt-proteins, a new replacement matrix, named MtArt, was constructed based on arthropod mt-proteomes. The new model was found to best fit almost all analyzed invertebrate mt-protein data sets. The observed pattern of model fit across the different data sets indicates that no single replacement matrix is able to describe the general evolutionary properties of mt-proteins but rather that taxonomical biases and/or the existence of different mt-genetic codes have great influence on which model is selected.

Amino Acid Sequence↗

Automated phylogenetic detection of recombination using a genetic algorithm.

The evolution of homologous sequences affected by recombination or gene conversion cannot be adequately explained by a single phylogenetic tree. Many tree-based methods for sequence analysis, for example, those used for detecting sites evolving nonneutrally, have been shown to fail if such phylogenetic incongruity is ignored. However, it may be possible to propose several phylogenies that can correctly model the evolution of nonrecombinant fragments. We propose a model-based framework that uses a genetic algorithm to search a multiple-sequence alignment for putative recombination break points, quantifies the level of support for their locations, and identifies sequences or clades involved in putative recombination events. The software implementation can be run quickly and efficiently in a distributed computing environment, and various components of the methods can be chosen for computational expediency or statistical rigor. We evaluate the performance of the new method on simulated alignments and on an array of published benchmark data sets. Finally, we demonstrate that prescreening alignments with our method allows one to analyze recombinant sequences for positive selection.

Algorithms↗

ModelTest Server: a web-based tool for the statistical selection of models of nucleotide substitution online.

ModelTest server is a web-based application for the selection of models of nucleotide substitution using the program ModelTest. The server takes as input a text file with likelihood scores for the set of candidate models. Models can be selected with hierarchical likelihood ratio tests, or with the Akaike or Bayesian information criteria. The output includes several statistics for the assessment of model selection uncertainty, for model averaging or to estimate the relative importance of model parameters. The server can be accessed at http://darwin.uvigo.es/software/modeltest_server.html.

Base Composition↗

GenDecoder: genetic code prediction for metazoan mitochondria.

Although the majority of the organisms use the same genetic code to translate DNA, several variants have been described in a wide range of organisms, both in nuclear and organellar systems, many of them corresponding to metazoan mitochondria. These variants are usually found by comparative sequence analyses, either conducted manually or with the computer. Basically, when a particular codon in a query-species is linked to positions for which a specific amino acid is consistently found in other species, then that particular codon is expected to translate as that specific amino acid. Importantly, and despite the simplicity of this approach, there are no available tools to help predicting the genetic code of an organism. We present here GenDecoder, a web server for the characterization and prediction of mitochondrial genetic codes in animals. The analysis of automatic predictions for 681 metazoans aimed us to study some properties of the comparative method, in particular, the relationship among sequence conservation, taxonomic sampling and reliability of assignments. Overall, the method is highly precise (99%), although highly divergent organisms such as platyhelminths are more problematic. The GenDecoder web server is freely available from http://darwin.uvigo.es/software/gendecoder.html.

Amino Acid Sequence↗

Parallel evolution of the genetic code in arthropod mitochondrial genomes.

The genetic code provides the translation table necessary to transform the information contained in DNA into the language of proteins. In this table, a correspondence between each codon and each amino acid is established: tRNA is the main adaptor that links the two. Although the genetic code is nearly universal, several variants of this code have been described in a wide range of nuclear and organellar systems, especially in metazoan mitochondria. These variants are generally found by searching for conserved positions that consistently code for a specific alternative amino acid in a new species. We have devised an accurate computational method to automate these comparisons, and have tested it with 626 metazoan mitochondrial genomes. Our results indicate that several arthropods have a new genetic code and translate the codon AGG as lysine instead of serine (as in the invertebrate mitochondrial genetic code) or arginine (as in the standard genetic code). We have investigated the evolution of the genetic code in the arthropods and found several events of parallel evolution in which the AGG codon was reassigned between serine and lysine. Our analyses also revealed correlated evolution between the arthropod genetic codes and the tRNA-Lys/-Ser, which show specific point mutations at the anticodons. These rather simple mutations, together with a low usage of the AGG codon, might explain the recurrence of the AGG reassignments.

Animals↗

Recombination estimation under complex evolutionary models with the coalescent composite-likelihood method.

The composite-likelihood estimator (CLE) of the population recombination rate considers only sites with exactly two alleles under a finite-sites mutation model (McVean, G. A. T., P. Awadalla, and P. Fearnhead. 2002. A coalescent-based method for detecting and estimating recombination from gene sequences. Genetics 160:1231-1241). While in such a model the identity of alleles is not considered, the CLE has been shown to be robust to minor misspecification of the underlying mutational model. However, there are many situations where the putative mutation and demographic history can be quite complex. One good example is rapidly evolving pathogens, like HIV-1. First we evaluated the performance of the CLE and the likelihood permutation test (LPT) under more complex, realistic models, including a general time reversible (GTR) substitution model, rate heterogeneity among sites (Gamma), positive selection, population growth, population structure, and noncontemporaneous sampling. Second, we relaxed some of the assumptions of the CLE allowing for a four-allele, GTR + Gamma model in an attempt to use the data more efficiently. Through simulations and the analysis of real data, we concluded that the CLE is robust to severe misspecifications of the substitution model, but underestimates the recombination rate in the presence of exponential growth, population mixture, selection, or noncontemporaneous sampling. In such cases, the use of more complex models slightly increases performance in some occasions, especially in the case of the LPT. Thus, our results provide for a more robust application of the estimation of recombination rates.

Alleles↗

Longitudinal population analysis of dual infection with recombination in two strains of HIV type 1 subtype B in an individual from a Phase 3 HIV vaccine efficacy trial.

This study documents a case of coinfection (simultaneous infection of an individual with two or more strains) of two HIV-1 subtype B strains in an individual from a Phase 3 HIV-1 vaccine efficacy trial, conducted in North American and the Netherlands. We examined 86 full-length gp120 (env) gene sequences from this individual collected from nine different time points over a 20-month period. We estimated evolutionary relationships using maximum likelihood and Bayesian methods and inferred recombination breakpoints and recombinant sequences using phylogenetic and substitutional methods. These analyses identified two strongly supported monophyletic clades (clades A and B) of 14 and 69 sequences each and a small paraphyletic recombinant clade of three sequences. We then studied the genetic characteristics of these lineages by comparing estimates of genetic diversity generated by mutation and recombination and adaptive selection within a coalescent and maximum likelihood framework. Our results suggest significant differences on the evolutionary dynamics of these strains. We then discuss the implications of these results for vaccine development.

AIDS Vaccines↗

Perkinsoide chabelardi n. gen., a protozoan parasite with an intermediate evolutionary position: possible cause of the decrease of sardine fisheries?

Phenotypic scrutiny on the life cycle of Icthyodinium chabelardi (Perkinsoide chabelardi n. gen.) based on ultrastructural techniques, and molecular phylogenetic analysis of RNA gene sequences, were carried out in order to elucidate the taxonomic position of this parasite. The absence of plastid, presence of trichocysts, and chromosomes or chromatin condensed and low in number, suggested that this protozoan could be considered a dinoflagellate syndinial parasite. However, the life cycle, schizogonic divisions and structure of schizonts inside the host, the nuclei without the typical dinoflagellate appearance, presence of rhoptrias-like structures, a possible pseudo-conoid, and the biflagellated spore, resembled those of the genus Perkinsus. Phylogenetic analysis of genes transcribing for the RNA forming the small subunit and the large subunit suggests that this parasite has an ambiguous evolutionary position within the group formed by dinoflagellates, perkinsids and syndinials. Because of differences with dinoflagellates and similarities with perkinsids, we propose to change the generic name to P. chabelardi n. gen. High stationary infection prevalence on Sardina pilchardus eggs was observed. This protozoan parasite caused the death of all the infected sardine eggs, and therefore a high impact in the recruitment of this fishery in the Atlantic coast is expected.

Animals↗

Identification of a novel HIV-1 complex circulating recombinant form (CRF18_cpx) of Central African origin in Cuba.

BACKGROUND: Analysis of partial pol and env sequences have indicated a high diversity of HIV-1 genetic forms in Cuba, including two potential novel circulating recombinant forms (CRF): U/H and D/A. OBJECTIVES: To determine whether U/H recombinant viruses from Cuba, detected in 7% of samples, represent a novel HIV-1 CRF, and to identify non-Cuban viruses related to this recombinant form. METHODS: Near full-length genome amplification was carried out by nested polymerase chain reaction in four overlapping DNA segments of two epidemiologically unlinked viruses in uncultured peripheral blood mononuclear cells. The sequences were analysed phylogenetically. Recombinant structures and phylogenetic relationships were analysed by bootscanning and by maximum likelihood. Searches for related viruses in databases were initially based on sequence homology and sharing of signature nucleotides. RESULTS: Both Cuban viruses clustered uniformly in bootscans all along the genome with each other and with a virus from Cameroon, CM53379, indicating that all three represent the same recombinant form. Their genome comprised multiple segments clustering with subtypes A1, F, G, H and K, as well as segments failing to cluster with recognized subtypes. The newly defined CRF, designated CRF18_cpx, was phylogenetically related in partial segments to CRF13_cpx, CRF04_cpx and 36 additional viruses, most of them from Central Africa. One of the viruses from Cameroon, sequenced in the near full-length genome, was a CRF18_cpx/subtype G secondary recombinant. CONCLUSIONS: A novel HIV-1 complex circulating recombinant form (CRF18_cpx) has been identified that is circulating in Cuba and Central Africa.

Africa, Central↗

On the phylogenetic placement of human T cell leukemia virus type 1 sequences associated with an Andean mummy.

Recently, the putative finding of ancient human T cell leukemia virus type 1 (HTLV-1) long terminal repeat (LTR) DNA sequences in association with a 1500-year-old Chilean mummy has stirred vigorous debate. The debate is based partly on the inherent uncertainties associated with phylogenetic reconstruction when only short sequences of closely related genotypes are available. However, a full analysis of what phylogenetic information is present in the mummy data has not previously been published, leaving open the question of what precisely is the range of admissible interpretation. To fulfill this need, we re-analyzed the mummy data in a new way. We first performed phylogenetic analysis of 188 published LTR DNA sequences from extant strains belonging to the HTLV-1 Cosmopolitan clade, using the method of statistical parsimony which is designed both to optimize phylogenetic resolution among sequences with little evolutionary divergence, and to permit precise mapping of individual sequence mutations onto branches of a divergence network. We then deduced possible phylogenetic positions for the two main categories of published Chilean mummy sequences, based on their published 157-nucleotide LTR sequences. The possible phylogenetic placements for one of the mummy sequence categories are consistent with a modern origin. However, one of these placements for the other mummy sequence category falls very close to the root of the Cosmopolitan clade, consistent with an ancient origin for both this mummy sequence and the Cosmopolitan clade.

Asian People↗

TreeScan: a bioinformatic application to search for genotype/phenotype associations using haplotype trees.

SUMMARY: We present the software implementation of the tree scanning method to detect associations between genetic haplotypes and quantitative traits, utilizing the evolutionary history of the haplotypes, in samples of unrelated individuals. AVAILABILITY: The program is available free of charge, under the GNU General Public License. A package including C source code, a Makefile, and Windows (DOS) and Macintosh binaries, can be downloaded from http://darwin.uvigo.es

Chromosome Mapping↗

ProtTest: selection of best-fit models of protein evolution.

SUMMARY: Using an appropriate model of amino acid replacement is very important for the study of protein evolution and phylogenetic inference. We have built a tool for the selection of the best-fit model of evolution, among a set of candidate models, for a given protein sequence alignment. AVAILABILITY: ProtTest is available under the GNU license from http://darwin.uvigo.es

Algorithms↗

Using models of nucleotide evolution to build phylogenetic trees.

Molecular phylogenetics and its applications are popular and useful tools for making comparative investigations in genetics; however, estimating phylogenetic trees is not always straightforward. Some phylogenetic estimators use an explicit model of nucleotide evolution to estimate evolutionary parameters such as branch lengths and tree topology. There are many models to choose from, and use of the optimal model for a particular data set is important to avoid a loss of power and accuracy in phylogenetic estimations. Here, we review some molecular evolutionary forces and the parameters included in some common models of evolution used to interpret resulting patterns of molecular variation. We present some statistical methods of selecting a particular model of nucleotide evolution, and provide an empirical example of model selection. Statistical model selection strikes a balance between the bias introduced by some models and the increased variance of parameter estimates that results from using other models.

Animals↗

The evolutionary value of recombination is constrained by genome modularity.

Genetic recombination is a fundamental evolutionary mechanism promoting biological adaptation. Using engineered recombinants of the small single-stranded DNA plant virus, Maize streak virus (MSV), we experimentally demonstrate that fragments of genetic material only function optimally if they reside within genomes similar to those in which they evolved. The degree of similarity necessary for optimal functionality is correlated with the complexity of intragenomic interaction networks within which genome fragments must function. There is a striking correlation between our experimental results and the types of MSV recombinants that are detectable in nature, indicating that obligatory maintenance of intragenome interaction networks strongly constrains the evolutionary value of recombination for this virus and probably for genomes in general.

Evolution, Molecular↗

Tree scanning: a method for using haplotype trees in phenotype/genotype association studies.

We use evolutionary trees of haplotypes to study phenotypic associations by exhaustively examining all possible biallelic partitions of the tree, a technique we call tree scanning. If the first scan detects significant associations, additional rounds of tree scanning are used to partition the tree into three or more allelic classes. Two worked examples are presented. The first is a reanalysis of associations between haplotypes at the Alcohol Dehydrogenase locus in Drosophila melanogaster that was previously analyzed using a nested clade analysis, a more complicated technique for using haplotype trees to detect phenotypic associations. Tree scanning and the nested clade analysis yield the same inferences when permutation testing is used with both approaches. The second example is an analysis of associations between variation in various lipid traits and genetic variation at the Apolipoprotein E (APOE) gene in three human populations. Tree scanning successfully identified phenotypic associations expected from previous analyses. Tree scanning for the most part detected more associations and provided a better biological interpretative framework than single SNP analyses. We also show how prior information can be incorporated into the tree scan by starting with the traditional three electrophoretic alleles at APOE. Tree scanning detected genetically determined phenotypic heterogeneity within all three electrophoretic allelic classes. Overall, tree scanning is a simple, powerful, and flexible method for using haplotype trees to detect phenotype/genotype associations at candidate loci.

Alcohol Dehydrogenase↗

Pharmacogenetic study of statin therapy and cholesterol reduction.

CONTEXT: Polymorphisms in genes involved in cholesterol synthesis, absorption, and transport may affect statin efficacy. OBJECTIVE: To evaluate systematically whether genetic variation influences response to pravastatin therapy. DESIGN, SETTING, AND POPULATION: The DNA of 1536 individuals treated with pravastatin, 40 mg/d, was analyzed for 148 single-nucleotide polymorphisms (SNPs) within 10 candidate genes related to lipid metabolism. Variation within these genes was then examined for associations with changes in lipid levels observed with pravastatin therapy during a 24-week period. MAIN OUTCOME MEASURE: Changes in lipid levels in response to pravastatin therapy. RESULTS: Two common and tightly linked SNPs (linkage disequilibrium r2 = 0.90; heterozygote prevalence = 6.7% for both) were significantly associated with reduced efficacy of pravastatin therapy. Both of these SNPs were in the gene coding for 3-hydroxy-3-methylglutaryl-coenzyme A (HMG-CoA) reductase, the target enzyme that is inhibited by pravastatin. For example, compared with individuals homozygous for the major allele of one of the SNPs, individuals with a single copy of the minor allele had a 22% smaller reduction in total cholesterol (-32.8 vs -42.0 mg/dL [-0.85 vs -1.09 mmol/L]; P =.001; absolute difference, 9.2 mg/dL [95% confidence interval [CI], 3.8-14.6 mg/dL]) and a 19% smaller reduction in low-density lipoprotein (LDL) cholesterol (-27.7 vs -34.1 mg/dL [-0.72 vs -0.88 mmol/L]; P =.005; absolute difference, 6.4 mg/dL [95% CI, 2.2-10.6 mg/dL]). The association for total cholesterol reduction persisted even after adjusting for multiple tests on all 33 SNPs evaluated in the HMG-CoA reductase gene as well as for all 148 SNPs evaluated was similar in magnitude and direction among men and women and was present in the ethnically diverse total cohort as well as in the majority subgroup of white participants. No association for either SNP was observed for the change in high-density lipoprotein (HDL) cholesterol (P>.80) and neither was associated with baseline lipid levels among those actively treated or among those who did not receive the drug. Among the remaining genes, less robust associations were found for squalene synthase and change in total cholesterol, apolipoprotein E and change in LDL cholesterol, and cholesteryl ester transfer protein and change in HDL cholesterol, although none of these met our conservative criteria for purely pharmacogenetic effects. CONCLUSION: Individuals heterozygous for a genetic variant in the HMG-CoA reductase gene may experience significantly smaller reductions in cholesterol when treated with pravastatin.

Aged↗