Search PubMed⌕ Search

Biomedical subjects

Pierre Darlu

Publications and source records attributed to Pierre Darlu.

13 recordsLinked to original sources

Data-mining methods as useful tools for predicting individual drug response: application to CYP2D6 data.

OBJECTIVES: Selecting a maximally informative subset of polymorphisms to predict a clinical outcome, such as drug response, requires appropriate search methods due to the increased dimensionality associated with looking at multiple genotypes. In this study, we investigated the ability of several pattern recognition methods to identify the most informative markers in the CYP2D6 gene for the prediction of CYP2D6 metabolizer status. METHODS: Four data-mining tools were explored: decision trees, random forests, artificial neural networks, and the multifactor dimensionality reduction (MDR) method. Marker selection was performed separately in eight population samples of different ethnic origin to evaluate to what extent the most informative markers differ across ethnic groups. RESULTS: Our results show that the number of polymorphisms required to predict CYP2D6 metabolic phenotype with a high accuracy can be dramatically reduced owing to the strong haplotype block structure observed at CYP2D6. MDR and neural networks provided nearly identical results and performed the best. CONCLUSION: Data-mining methods, such as MDR and neural networks, appear as promising tools to improve the efficiency of genotyping tests in pharmacogenetics with the ultimate goal of pre-screening patients for individual therapy selection with minimum genotyping effort.

Cytochrome P-450 CYP2D6↗

Clustering of haplotypes based on phylogeny: how good a strategy for association testing?

Haplotypes are now widely used in association studies between markers and disease susceptibility locus. However, when a large number of markers are considered, the number of possible haplotypes increases leading to two problems: an increased number of degrees of freedom that may result in a lack of power and the existence of rare haplotypes that may be difficult to take into account in the statistical analysis. In a recent paper, Durrant et al proposed a method, CLADHC, to group haplotypes based on distance matrices and showed that this could considerably increase the power of the association test as compared to either single-locus analysis or haplotype analysis without prior grouping. Although the authors considered different one-disease-locus susceptibility models in their simulations, they did not study the impact of the linkage disequilibrium (LD) pattern and of the susceptibility allele frequency on their conclusions. Here, we show, using haplotype data from five regions of the genome of different lengths and with different LD patterns, that, when a single disease susceptibility locus is simulated, the prior grouping of haplotypes based on the algorithm of Durrant et al does not increase the power of association testing except in very particular situations of LD patterns and allele frequencies.

Computer Simulation↗

SNP selection at the NAT2 locus for an accurate prediction of the acetylation phenotype.

PURPOSE: Genetic polymorphisms in the N-acetyltransferase 2 gene determine the individual acetylator status, which influences both the toxicity and efficacy profile of acetylated drugs. Determination of an individual's acetylation phenotype prior to initiation of therapy, through DNA-based tests, should permit to improve therapy response and reduce adverse events. However, due to extensive linkage disequilibrium between markers within NAT2, the genotyping of closely spaced markers yields highly redundant data: testing them all is expensive and often unnecessary. The objective of this study is to establish the optimal strategy to define, in the genetic context of a given ethnic group, the most informative set of single-nucleotide polymorphisms that best enables accurate prediction of acetylation phenotype. METHODS: Three classification methods have been investigated (classification trees, artificial neural networks and multifactor dimensionality reduction method) in order to find the optimal set of single-nucleotide polymorphisms enabling the most efficient classification of individuals in rapid and slow acetylators. RESULTS: Our results show that, in almost all population samples, only one or two single-nucleotide polymorphisms would be enough to obtain a good predictive capacity with no or only a modest reduction in power relative to direct assays of all common markers. In contrast, in Black African populations, where lower levels of linkage disequilibrium are observed at NAT2, a larger number of single-nucleotide polymorphisms are required to predict acetylation phenotype. CONCLUSION: The results of this study will be helpful for the design of time- and cost-effective pharmacogenetic tests (adapted to specific populations) that could be used as routine tools in clinical practice.

Acetylation↗

Very virulent infectious bursal disease virus: reduced pathogenicity in a rare natural segment-B-reassorted isolate.

The purpose of this study was to compare the molecular epidemiology of infectious bursal disease virus (IBDV) segments A and B of 50 natural or vaccine IBDV strains that were isolated or produced between 1972 and 2002 in 17 countries from four continents, with phenotypes ranging from attenuated to very virulent (vv). These strains were subjected to sequence and phylogenetic analysis based on partial sequences of genome segments A and B. Although there is co-evolution of the two genome segments (70 % of strains kept the same genetic relatives in the segment A- and B-defined consensus trees), several strains (26 %) were identified with the incongruence length difference test as exhibiting a significantly different phylogenetic relationship depending on which segment was analysed. This suggested that natural reassortment could have occurred. One of the possible naturally occurring reassortant strains, which exhibited a segment A related to the vvIBDV cluster whereas its segment B was not, was thoroughly sequenced (coding sequence of both segments) and submitted to a standardized experimental characterization of its acute pathogenicity. This strain induced significantly less mortality than typical vvIBDVs; however, the mechanisms for this reduced pathogenicity remain unknown, as no significant difference in the bursal lesions, post-infectious antibody response or virus production in the bursa was observed in challenged chickens.

Animals↗

Inferring haplotypes at the NAT2 locus: the computational approach.

BACKGROUND: Numerous studies have attempted to relate genetic polymorphisms within the N-acetyltransferase 2 gene (NAT2) to interindividual differences in response to drugs or in disease susceptibility. However, genotyping of individuals single-nucleotide polymorphisms (SNPs) alone may not always provide enough information to reach these goals. It is important to link SNPs in terms of haplotypes which carry more information about the genotype-phenotype relationship. Special analytical techniques have been designed to unequivocally determine the allocation of mutations to either DNA strand. However, molecular haplotyping methods are labour-intensive and expensive and do not appear to be good candidates for routine clinical applications. A cheap and relatively straightforward alternative is the use of computational algorithms. The objective of this study was to assess the performance of the computational approach in NAT2 haplotype reconstruction from phase-unknown genotype data, for population samples of various ethnic origin. RESULTS: We empirically evaluated the effectiveness of four haplotyping algorithms in predicting haplotype phases at NAT2, by comparing the results with those directly obtained through molecular haplotyping. All computational methods provided remarkably accurate and reliable estimates for NAT2 haplotype frequencies and individual haplotype phases. The Bayesian algorithm implemented in the PHASE program performed the best. CONCLUSION: This investigation provides a solid basis for the confident and rational use of computational methods which appear to be a good alternative to infer haplotype phases in the particular case of the NAT2 gene, where there is near complete linkage disequilibrium between polymorphic markers.

Algorithms↗

On the use of haplotype phylogeny to detect disease susceptibility loci.

BACKGROUND: The cladistic approach proposed by Templeton has been presented as promising for the study of the genetic factors involved in common diseases. This approach allows the joint study of multiple markers within a gene by considering haplotypes and grouping them in nested clades. The idea is to search for clades with an excess of cases as compared to the whole sample and to identify the mutations defining these clades as potential candidate disease susceptibility sites. However, the performance of this approach for the study of the genetic factors involved in complex diseases has never been studied. RESULTS: In this paper, we propose a new method to perform such a cladistic analysis and we estimate its power through simulations. We show that under models where the susceptibility to the disease is caused by a single genetic variant, the cladistic test is neither really more powerful to detect an association nor really more efficient to localize the susceptibility site than an individual SNP testing. However, when two interacting sites are responsible for the disease, the cladistic analysis greatly improves the probability to find the two susceptibility sites. The impact of the linkage disequilibrium and of the tree characteristics on the efficiency of the cladistic analysis are also discussed. An application on a real data set concerning the CARD15 gene and Crohn disease shows that the method can successfully identify the three variant sites that are involved in the disease susceptibility. CONCLUSION: The use of phylogenies to group haplotypes is especially interesting to pinpoint the sites that are likely to be involved in disease susceptibility among the different markers identified within a gene.

Crohn Disease↗

Selection-driven transcriptome polymorphism in Escherichia coli/Shigella species.

To explore the role of transcriptome polymorphism in adaptation of organisms to their environment, we evaluated this parameter for the Escherichia coli/Shigella bacterial species, which is composed of well-characterized phylogenetic groups that exhibit characteristic life styles ranging from commensalism to intracellular pathogenicity. Both the genomic content and the transcriptome of 10 strains representative of the major E. coli/Shigella phylogenetic groups were evaluated using macroarrays displaying the 4290 K12-MG1655 open reading frames (ORFs). Although Shigella and enteroinvasive E. coli (EIEC) are not monophyletic, phylogenetic analysis of the binary coded (presence/absence) gene content data showed that these organisms group together due to similar patterns of undetectable K12-MG1655 genes. The variation in transcript abundance was then analyzed using a core genome of 2880 genes present in all strains, after adjusting RNA hybridization signals for DNA hybridization signals. Nonrandom changes in gene expression during the evolution of the E. coli/Shigella species were evidenced. Phylogenetic analysis of transcriptome data again showed that Shigella and EIEC strains group together in terms of gene expression, and this convergence involved groups of genes displaying biologically coherent patterns of functional divergence. Unlike the other E. coli strains evaluated, Shigella and EIEC are intracellular pathogens, and therefore face similar selective pressures. Thus, within the E. coli/Shigella species, strains exhibiting a particular life style have converged toward a specific gene expression pattern in a subset of genes common to the species, revealing the role of selection in shaping transcriptome polymorphism.

Caco-2 Cells↗

Decreasing the effects of horizontal gene transfer on bacterial phylogeny: the Escherichia coli case study.

Phylogenetic reconstructions of bacterial species from DNA sequences are hampered by the existence of horizontal gene transfer. One possible way to overcome the confounding influence of such movement of genes is to identify and remove sequences which are responsible for significant character incongruence when compared to a reference dataset free of horizontal transfer (e.g., multilocus enzyme electrophoresis, restriction fragment length polymorphism, or random amplified polymorphic DNA) using the incongruence length difference (ILD) test of Farris et al. [Cladistics 10 (1995) 315]. As obtaining this "whole genome dataset" prior to the reconstruction of a phylogeny is clearly troublesome, we have tested alternative approaches allowing the release from such reference dataset, designed for a species with modest level of horizontal gene transfer, i.e., Escherichia coli. Eleven different genes available or sequenced in this work were studied in a set of 30 E. coli reference (ECOR) strains. Either using ILD to test incongruence between each gene against the all remaining (in this case 10) genes in order to remove sequences responsible for significant incongruence, or using just a simultaneous analysis without removals, gave robust phylogenies with slight topological differences. The use of the ILD test remains a suitable method for estimating the level of horizontal gene transfer in bacterial species. Supertrees also had suitable properties to extract the phylogeny of strains, because the way they summarize taxonomic congruence clearly limits the impact of individual gene transfers on the global topology. Furthermore, this work allowed a significant improvement of the accuracy of the phylogeny within E. coli.

DNA, Bacterial↗

Analysis of the French National Registry of unrelated bone marrow donors, using surnames as a tool for improving geographical localisation of HLA haplotypes.

The first statistical analysis of the French National Registry of volunteer bone marrow donors estimated the probabilities of haplotype frequencies separately for each of the 20 administrative regions of France. Here we propose to use donors' surnames to increase the accuracy of location of the donor's geographical origin. This approach allows us to estimate haplotype frequencies for administrative entities (90 departments) smaller than regions and to correct for bias resulting from recent mobility. We analysed 30,777 donors typed for HLA-A,B and 17,745 donors typed for HLA-A,B,DR,DQ. By using the donors' surnames, we identified common and rare haplotypes (those found in only one department) and estimated the degree of HLA polymorphism at the department level. We also identified departments with a distinctive genetic structure (for example, Paris, Corsica, Pyrenees and Meurthe-et-Moselle). By providing a more accurate geographical distribution of HLA polymorphisms in France, this study will enable us to optimise policies for recruiting bone marrow donors and to improve the fit between the donor file and patients' needs.

Algorithms↗

When does the incongruence length difference test fail?

This paper examines the efficiency of the incongruence length difference test (ILD) proposed by Farris et al. (1994) for assessing the incongruence between sets of characters. DNA sequences were simulated under various evolutionary conditions: (1) following symmetric or asymmetric trees, (2) with various mutation rates, (3) with constant or variable evolutionary rates along the branches, and (4) with different among-site substitution rates. We first compared two sets of sequences generated along the same tree and under the same evolutionary conditions. The probability of a Type-I error (wrongly rejecting the true hypothesis of congruence) was substantially below the standard 5% level of significance given by the ILD test; this finding indicates that the choice of the 5% level is rather conservative in this case. We then compared two data sets, still generated along the same tree, but under different evolutionary conditions (constant vs. variable evolutionary rate, homogeneity vs. heterogeneity rate of substitution). Under these conditions, the probability of rejecting the true hypothesis of congruence was greater than the 5% given by the ILD test and increased with the number of sites and the degree to which the tree was asymmetric. Finally, the comparison of the two data sets, simulated under contrasting tree structures (symmetric vs. asymmetric) but under the same evolutionary conditions, led us to reject the hypothesis of congruence, albeit weakly, particularly when the number of informative sites was low and among-site substitution rate heterogeneous. We conclude that the ILD test has only limited power to detect incongruence caused by differences in the evolutionary conditions or in the tree topology, except when numerous characters are present and the substitution rate is homogeneous from site to site.

Artifacts↗

Genetic polymorphism of human herpesvirus-7 among human populations.

The analysis of three human herpesvirus-7 (HHV-7) genes encoding phosphoprotein p100, glycoprotein B and major capsid protein respectively had previously shown the existence of distinct gene alleles, leading to the concept of HHV-7 variants. We have analysed the distribution of HHV-7 variants among 297 distinct subjects who belonged to different human populations from Africa, Asia, Europe and America. Two variants, designated Co1 and Co2, were found in 52% and 20% of studied subjects. Ten other variants, designated Co3-Co12, were less frequent and classified into two groups related to Co1 and Co2 respectively. While the former group was ubiquitous and the most frequent in Africa and Asia, the latter one was predominantly found in European and Mongol populations. Despite the high stability of the HHV-7 genome, a few nucleotide substitutions at precise positions define distinct variants which, to some extent, behave as markers of human populations.

Africa↗

Genetic structure of Algerian populations.

Blood samples were collected in Algeria from 4,444 army recruits and tested for 10 genetic polymorphic systems. These samples were collected from territorial Wilayas (administrative units of Algeria) from which the young soldiers had originated. Based on similar geography and economic and political history, these Wilayas were clustered into 10 regions. These regions, not part of the governmental administrative units, were characterized by allelic frequencies, and analyzed using R-matrix principal components, Wright's F(ST), spatial autocorrelation, and Mantel tests. Hierarchical relationships between the culturally defined regions were examined using two different analytical methods of phylogenetic tree constructions: neighbor-joining, and unweighted pair group average arithmetic (UPGMA). These results indicated the predominance of genetic homogeneity due to the gene flow between regions, but with some migration emanating from sub-Saharan Africa and Mediterranean Europe. Wright's F(ST) value of 0.0063, based on 16 alleles, suggested a relatively small genetic microdifferentiation of the regions. In Algeria, gene flow apparently swamped most of the effects of stochastic processes and disrupted the relationship between geography and genetics, as characterized by the isolation-by-distance model. Some genetic differences and similarities were observed between regions or clusters of regions. The resulting genetic structure of the Algerian populations is best explained by a combination of gene flow, ecology, and history.

Algeria↗