Search PubMed⌕ Search

Biomedical subjects

Hervé Philippe

Publications and source records attributed to Hervé Philippe.

At least 19 recordsLinked to original sources

Phylogenetic analyses of nuclear, mitochondrial, and plastid multigene data sets support the placement of Mesostigma in the Streptophyta.

All extant green plants belong to 1 of 2 major lineages, commonly known as the Chlorophyta (most of the green algae) and the Streptophyta (land plants and their closest green algal relatives). The scaly green flagellate Mesostigma viride has an important place in the debate on the origin of green plants. However, there have been conflicting results from molecular systematics as to whether Mesostigma diverges before the Chlorophyta/Streptophyta split or is an early diverging flagellate member of the Streptophyta. Previous studies employed either a limited taxon sampling (plastid and mitochondrial genomes) or a small number of phylogenetically informative sites (single nuclear genes). Here, we use large data sets from the nuclear (125 proteins; 29,319 positions), mitochondrial (33 proteins; 6,622 positions), and plastid (50 proteins; 10,137 positions) genomes with an expanded taxon sampling (21, 13, and 28 species, respectively) to reevaluate the phylogenetic position of Mesostigma. Our study supports the placement of Mesostigma in the Streptophyta (as an early diverging lineage) and provides evidence that systematic biases have played a role in generating some of the previous conflicting results. Importantly, we demonstrate that using an increased taxon sampling as well as more realistic models of evolution allows increasing congruence among the nuclear, mitochondrial, and plastid data sets.

Base Sequence↗

Lack of resolution in the animal phylogeny: closely spaced cladogeneses or undetected systematic errors?

A recent phylogenomic study reported that the animal phylogeny was unresolved despite the use of 50 genes. This lack of resolution was interpreted as "a positive signature of closely spaced cladogenetic events." Here, we propose that this lack of resolution is rather due to the mutual cancellation of the phylogenetic signal (historical) and the nonphylogenetic signal (due to systematic errors) that results from inadequate taxon sampling and/or model of sequence evolution. Starting with a data set of comparable size, we use 3 different strategies to reduce the nonphylogenetic signal: 1) increasing the number of species; 2) replacing a fast-evolving species by a slowly evolving one; and 3) using a better model of sequence evolution. In all cases, the phylogenetic resolution is markedly improved, in agreement with our hypothesis that the originally reported lack of resolution was artifactual.

Algorithms↗

Large-scale sequencing and the new animal phylogeny.

Although comparisons of gene sequences have revolutionised our understanding of the animal phylogenetic tree, it has become clear that, to avoid errors in tree reconstruction, a large number of genes from many species must be considered: too few genes and stochastic errors predominate, too few taxa and systematic errors appear. We argue here that, to gather many sequences from many taxa, the best use of resources is to sequence a small number of expressed sequence tags (1000-5000 per species) from as many taxa as possible. This approach counters both sources of error, gives the best hope of a well-resolved phylogeny of the animals and will act as a central resource for a carefully targeted genome sequencing programme.

Animals↗

A maximum likelihood framework for protein design.

BACKGROUND: The aim of protein design is to predict amino-acid sequences compatible with a given target structure. Traditionally envisioned as a purely thermodynamic question, this problem can also be understood in a wider context, where additional constraints are captured by learning the sequence patterns displayed by natural proteins of known conformation. In this latter perspective, however, we still need a theoretical formalization of the question, leading to general and efficient learning methods, and allowing for the selection of fast and accurate objective functions quantifying sequence/structure compatibility. RESULTS: We propose a formulation of the protein design problem in terms of model-based statistical inference. Our framework uses the maximum likelihood principle to optimize the unknown parameters of a statistical potential, which we call an inverse potential to contrast with classical potentials used for structure prediction. We propose an implementation based on Markov chain Monte Carlo, in which the likelihood is maximized by gradient descent and is numerically estimated by thermodynamic integration. The fit of the models is evaluated by cross-validation. We apply this to a simple pairwise contact potential, supplemented with a solvent-accessibility term, and show that the resulting models have a better predictive power than currently available pairwise potentials. Furthermore, the model comparison method presented here allows one to measure the relative contribution of each component of the potential, and to choose the optimal number of accessibility classes, which turns out to be much higher than classically considered. CONCLUSION: Altogether, this reformulation makes it possible to test a wide diversity of models, using different forms of potentials, or accounting for other factors than just the constraint of thermodynamic stability. Ultimately, such model-based statistical analyses may help to understand the forces shaping protein sequences, and driving their evolution.

Amino Acid Sequence↗

Assessing site-interdependent phylogenetic models of sequence evolution.

In recent works, methods have been proposed for applying phylogenetic models that allow for a general interdependence between the amino acid positions of a protein. As of yet, such models have focused on site interdependencies resulting from sequence-structure compatibility constraints, using simplified structural representations in combination with a set of statistical potentials. This structural compatibility criterion is meant as a proxy for sequence fitness, and the methods developed thus far can incorporate different site-interdependent fitness proxies based on other measurements. However, no methods have been proposed for comparing and evaluating the adequacy of alternative fitness proxies in this context, or for more general comparisons with canonical models of protein evolution. In the present work, we apply Bayesian methods of model selection-based on numerical calculations of marginal likelihoods and posterior predictive checks-to evaluate models encompassing the site-interdependent framework. Our application of these methods indicates that considering site-interdependencies, as done here, leads to an improved model fit for all data sets studied. Yet, we find that the use of pairwise contact potentials alone does not suitably account for across-site rate heterogeneity or amino acid exchange propensities; for such complexities, site-independent treatments are still called for. The most favored models combine the use of statistical potentials with a suitably rich site-independent model. Altogether, the methodology employed here should allow for a more rigorous and systematic exploration of different ways of modeling explicit structural constraints, or any other site-interdependent criterion, while best exploiting the richness of previously proposed models.

Bayes Theorem↗

Tunicates and not cephalochordates are the closest living relatives of vertebrates.

Tunicates or urochordates (appendicularians, salps and sea squirts), cephalochordates (lancelets) and vertebrates (including lamprey and hagfish) constitute the three extant groups of chordate animals. Traditionally, cephalochordates are considered as the closest living relatives of vertebrates, with tunicates representing the earliest chordate lineage. This view is mainly justified by overall morphological similarities and an apparently increased complexity in cephalochordates and vertebrates relative to tunicates. Despite their critical importance for understanding the origins of vertebrates, phylogenetic studies of chordate relationships have provided equivocal results. Taking advantage of the genome sequencing of the appendicularian Oikopleura dioica, we assembled a phylogenomic data set of 146 nuclear genes (33,800 unambiguously aligned amino acids) from 14 deuterostomes and 24 other slowly evolving species as an outgroup. Here we show that phylogenetic analyses of this data set provide compelling evidence that tunicates, and not cephalochordates, represent the closest living relatives of vertebrates. Chordate monophyly remains uncertain because cephalochordates, albeit with a non-significant statistical support, surprisingly grouped with echinoderms, a hypothesis that needs to be tested with additional data. This new phylogenetic scheme prompts a reappraisal of both morphological and palaeontological data and has important implications for the interpretation of developmental and genomic studies in which tunicates and cephalochordates are used as model animals.

Animals↗

An unusual choanoflagellate protein released by Hedgehog autocatalytic processing.

Hedgehog proteins are important cell-cell signalling proteins utilized during the development of multicellular animals. Members of the hedgehog gene family have not been detected outside the Metazoa, raising unanswered questions about their evolutionary origin. Here we report a highly unusual hedgehog-related gene from a choanoflagellate, a close unicellular relative of the animals. The deduced C-terminal domain, Hoglet-C, is homologous to the autocatalytic domain of Hedgehog proteins and is predicted to function in autocatalytic cleavage of the precursor peptide. In contrast, the N-terminal Hoglet-N peptide has no similarity to the signalling peptide of Hedgehog (Hh-N). Instead, Hoglet-N is deduced to be a secreted protein with an enormous threonine-rich domain of unprecedented size and purity (over 200 threonine residues) and two polysaccharide-binding domains. Structural modelling reveals that these domains have a novel combination of features found in cellulose-binding domains (CBD) of types IIa and IIb, and are expected to bind cellulose. We propose that the two CBD domains enable Hoglet-N to bind to plant matter, tethering an amorphous nucleophilic anchor, facilitating transient adhesion of the choanoflagellate cell. Since Hh-C and Hoglet-C are homologous, but Hh-N and Hoglet-N are not, we argue that metazoan hedgehog genes evolved by fusion of two distinct genes.

Amino Acid Sequence↗

Phylogenomics: the beginning of incongruence?

Until recently, molecular phylogenies based on a single or few orthologous genes often yielded contradictory results. Using multiple genes in a large concatenation was proposed to end these incongruences. Here we show that single-gene phylogenies often produce incongruences, albeit ones lacking statistically significant support. By contrast, the use of different tree reconstruction methods on different partitions of the concatenated supergene leads to well-resolved, but real (i.e. statistically significant) incongruences. Gathering a large amount of data is not sufficient to produce reliable trees, given the current limitation of tree reconstruction methods, especially when the quality of data is poor. We propose that selecting only data that contain minimal nonphylogenetic signals takes full advantage of phylogenomics and markedly reduces incongruence.

Genome, Fungal↗

A simple and robust statistical test for detecting the presence of recombination.

Recombination is a powerful evolutionary force that merges historically distinct genotypes. But the extent of recombination within many organisms is unknown, and even determining its presence within a set of homologous sequences is a difficult question. Here we develop a new statistic, phi(w), that can be used to test for recombination. We show through simulation that our test can discriminate effectively between the presence and absence of recombination, even in diverse situations such as exponential growth (star-like topologies) and patterns of substitution rate correlation. A number of other tests, Max chi2, NSS, a coalescent-based likelihood permutation test (from LDHat), and correlation of linkage disequilibrium (both r2 and /D'/) with distance, all tend to underestimate the presence of recombination under strong population growth. Moreover, both Max chi2 and NSS falsely infer the presence of recombination under a simple model of mutation rate correlation. Results on empirical data show that our test can be used to detect recombination between closely as well as distantly related samples, regardless of the suspected rate of recombination. The results suggest that phi(w) is one of the best approaches to distinguish recurrent mutation from recombination in a wide variety of circumstances.

Computer Simulation↗

[Molecular dating in the genomic era].

The comparison of DNA and protein sequences of extant species might be informative for reconstructing the chronology of evolutionary events on Earth. A phylogenetic tree inferred from molecular data directly depicts the evolutionary affinities of species and indirectly allows estimating the age of their origin and diversification. Molecular dating is achieved by assuming the molecular clock hypothesis, i.e., that the rate of change of nucleotide and amino acid sequences is on average constant over geological time. If paleontological calibrations are available, then absolute divergence times of species can be estimated. However, three major difficulties potentially hamper molecular dating : (1) a limited sample of genes and organisms, (2) a limited number of fossil references, and (3) pervasive variations of molecular evolutionary rates among genomes and species. To circumvent these problems, different solutions have been recently proposed. Larger data sets are built with more genes and more species sampled through the mining of an increasing number of genomes. Moreover, independent key fossils are identified to calibrate molecular clocks, and the uncertainty on their age is integrated in subsequent analyses. Finally, models of molecular rate variations are constructed, and incorporated in the so-called relaxed molecular clock approaches. As an illustration of these improvements, we mention that the debated age of the animal (bilaterian metazoans) diversification may have occurred between 642-761 million years ago (Mya), roughly 100 Ma before the Cambrian explosion. Among mammals, the initial diversification of major placental groups may have taken place around 100 Mya, well before the Cretaceous/Tertiary boundary marking the extinction of dinosaurs.

Animals↗

Computing Bayes factors using thermodynamic integration.

In the Bayesian paradigm, a common method for comparing two models is to compute the Bayes factor, defined as the ratio of their respective marginal likelihoods. In recent phylogenetic works, the numerical evaluation of marginal likelihoods has often been performed using the harmonic mean estimation procedure. In the present article, we propose to employ another method, based on an analogy with statistical physics, called thermodynamic integration. We describe the method, propose an implementation, and show on two analytical examples that this numerical method yields reliable estimates. In contrast, the harmonic mean estimator leads to a strong overestimation of the marginal likelihood, which is all the more pronounced as the model is higher dimensional. As a result, the harmonic mean estimator systematically favors more parameter-rich models, an artefact that might explain some recent puzzling observations, based on harmonic mean estimates, suggesting that Bayes factors tend to overscore complex models. Finally, we apply our method to the comparison of several alternative models of amino-acid replacement. We confirm our previous observations, indicating that modeling pattern heterogeneity across sites tends to yield better models than standard empirical matrices.

Amino Acid Sequence↗

Taking the first steps towards a standard for reporting on phylogenies: Minimum Information About a Phylogenetic Analysis (MIAPA).

In the eight years since phylogenomics was introduced as the intersection of genomics and phylogenetics, the field has provided fundamental insights into gene function, genome history and organismal relationships. The utility of phylogenomics is growing with the increase in the number and diversity of taxa for which whole genome and large transcriptome sequence sets are being generated. We assert that the synergy between genomic and phylogenetic perspectives in comparative biology would be enhanced by the development and refinement of minimal reporting standards for phylogenetic analyses. Encouraged by the development of the Minimum Information About a Microarray Experiment (MIAME) standard, we propose a similar roadmap for the development of a Minimal Information About a Phylogenetic Analysis (MIAPA) standard. Key in the successful development and implementation of such a standard will be broad participation by developers of phylogenetic analysis software, phylogenetic database developers, practitioners of phylogenomics, and journal editors.

Genomics↗

Effects of laparoscopic pneumoperitoneum and changes in position on arterial pulse pressure wave-form: comparison between morbidly obese and normal-weight patients.

BACKGROUND: Laparoscopic adjustable gastric banding (LAGB) is commonly indicated in morbidly obese patients. There is controversy regarding the hemodynamic effects of pneumoperitoneum (PNP) in obese patients. PNP and changes in body posture have complex effects on venous return that may be detected by respiratory changes in the arterial pressure waveform. The aim of this study was to compare pneumoperitoneum-induced and reverse Trendelenburg (RT) changes in arterial pulse pressure in obese and normal-weight patients. METHODS: 15 morbidly obese patients undergoing LAGB were compared to 15 normal-weight patients undergoing laparoscopic surgery. Arterial pressure was non-invasively recorded using an arterial tonometer. Respiratory changes in pulse pressure (deltaPp) were recorded in the supine position without and with PNP, and in RT position with pneumoperitoneum. RESULTS: PNP increased deltaPp values in normal weight (P<0.001), but not in obese patients. RT position increased deltaPp values in obese patients, but did not cause additional changes in normal-weight patients. CONCLUSIONS: Unlike normal-weight patients, PNP in the supine position has minimal effect on the arterial pulse-pressure wave-form in obese patients. This observation may reflect physiological differences in total blood volume and loading conditions of the heart between morbidly obese and normal-weight patients, which affect venous return during PNP. Differences in abdominal vascular zone conditions between obese and normal weight-patients may explain these results.

Adult↗

Heterotachy and long-branch attraction in phylogenetics.

BACKGROUND: Probabilistic methods have progressively supplanted the Maximum Parsimony (MP) method for inferring phylogenetic trees. One of the major reasons for this shift was that MP is much more sensitive to the Long Branch Attraction (LBA) artefact than is Maximum Likelihood (ML). However, recent work by Kolaczkowski and Thornton suggested, on the basis of simulations, that MP is less sensitive than ML to tree reconstruction artefacts generated by heterotachy, a phenomenon that corresponds to shifts in site-specific evolutionary rates over time. These results led these authors to recommend that the results of ML and MP analyses should be both reported and interpreted with the same caution. This specific conclusion revived the debate on the choice of the most accurate phylogenetic method for analysing real data in which various types of heterogeneities occur. However, variation of evolutionary rates across species was not explicitly incorporated in the original study of Kolaczkowski and Thornton, and in most of the subsequent heterotachous simulations published to date, where all terminal branch lengths were kept equal, an assumption that is biologically unrealistic. RESULTS: In this report, we performed more realistic simulations to evaluate the relative performance of MP and ML methods when two kinds of heterogeneities are considered: (i) within-site rate variation (heterotachy), and (ii) rate variation across lineages. Using a similar protocol as Kolaczkowski and Thornton to generate heterotachous datasets, we found that heterotachy, which constitutes a serious violation of existing models, decreases the accuracy of ML whatever the level of rate variation across lineages. In contrast, the accuracy of MP can either increase or decrease when the level of heterotachy increases, depending on the relative branch lengths. This result demonstrates that MP is not insensitive to heterotachy, contrary to the report of Kolaczkowski and Thornton. Finally, in the case of LBA (i.e. when two non-sister lineages evolved faster than the others), ML outperforms MP over a wide range of conditions, except for unrealistic levels of heterotachy. CONCLUSION: For realistic combinations of both heterotachy and variation of evolutionary rates across lineages, ML is always more accurate than MP. Therefore, ML should be preferred over MP for analysing real data, all the more so since parametric methods also allow one to handle other types of biological heterogeneities much better, such as among sites rate variation. The confounding effects of heterotachy on tree reconstruction methods do exist, but can be eschewed by the development of mixture models in a probabilistic framework, as proposed by Kolaczkowski and Thornton themselves.

Animals↗

Monophyly of primary photosynthetic eukaryotes: green plants, red algae, and glaucophytes.

Between 1 and 1.5 billion years ago, eukaryotic organisms acquired the ability to convert light into chemical energy through endosymbiosis with a Cyanobacterium (e.g.,). This event gave rise to "primary" plastids, which are present in green plants, red algae, and glaucophytes ("Plantae" sensu Cavalier-Smith). The widely accepted view that primary plastids arose only once implies two predictions: (1) all plastids form a monophyletic group, as do (2) primary photosynthetic eukaryotes. Nonetheless, unequivocal support for both predictions is lacking (e.g.,). In this report, we present two phylogenomic analyses, with 50 genes from 16 plastid and 15 cyanobacterial genomes and with 143 nuclear genes from 34 eukaryotic species, respectively. The nuclear dataset includes new sequences from glaucophytes, the less-studied group of primary photosynthetic eukaryotes. We find significant support for both predictions. Taken together, our analyses provide the first strong support for a single endosymbiotic event that gave rise to primary photosynthetic eukaryotes, the Plantae. Because our dataset does not cover the entire eukaryotic diversity (but only four of six major groups in), further testing of the monophyly of Plantae should include representatives from eukaryotic lineages for which currently insufficient sequence information is available.

Bayes Theorem↗

Site interdependence attributed to tertiary structure in amino acid sequence evolution.

Standard likelihood-based frameworks in phylogenetics consider the process of evolution of a sequence site by site. Assuming that sites evolve independently greatly simplifies the required calculations. However, this simplification is known to be incorrect in many cases. Here, a computational method that allows for general dependence between sites of a sequence is investigated. Using this method, measures acting as sequence fitness proxies can be considered over a phylogenetic tree. In this work, a set of statistically derived amino acid pairwise potentials, developed in the context of protein threading, is used to account for what we call the structural fitness of a sequence. We describe a model combining statistical potentials with an empirical amino acid substitution matrix. We propose such a combination as a useful way of capturing the complexity of protein evolution. Finally, we outline features of the model using three datasets and show the approach's sensitivity to different tree topologies.

Amino Acid Sequence↗

Multigene analyses of bilaterian animals corroborate the monophyly of Ecdysozoa, Lophotrochozoa, and Protostomia.

Almost a decade ago, a new phylogeny of bilaterian animals was inferred from small-subunit ribosomal RNA (rRNA) that claimed the monophyly of two major groups of protostome animals: Ecdysozoa (e.g., arthropods, nematodes, onychophorans, and tardigrades) and Lophotrochozoa (e.g., annelids, molluscs, platyhelminths, brachiopods, and rotifers). However, it received little additional support. In fact, several multigene analyses strongly argued against this new phylogeny. These latter studies were based on a large amount of sequence data and therefore showed an apparently strong statistical support. Yet, they covered only a few taxa (those for which complete genomes were available), making systematic artifacts of tree reconstruction more probable. Here we expand this sparse taxonomic sampling and analyze a large data set (146 genes, 35,371 positions) from a diverse sample of animals (35 species). Our study demonstrates that the incongruences observed between rRNA and multigene analyses were indeed due to long-branch attraction artifacts, illustrating the enormous impact of systematic biases on phylogenomic studies. A refined analysis of our data set excluding the most biased genes provides strong support in favor of the new animal phylogeny and in addition suggests that urochordates are more closely related to vertebrates than are cephalochordates. These findings have important implications for the interpretation of morphological and genomic data.

Animals↗