Search PubMed⌕ Search

Biomedical subjects

L Excoffier

Publications and source records attributed to L Excoffier.

At least 19 recordsLinked to original sources

Multiple maternal origins and weak phylogeographic structure in domestic goats.

Domestic animals have played a key role in human history. Despite their importance, however, the origins of most domestic species remain poorly understood. We assessed the phylogenetic history and population structure of domestic goats by sequencing a hypervariable segment (481 bp) of the mtDNA control region from 406 goats representing 88 breeds distributed across the Old World. Phylogeographic analysis revealed three highly divergent goat lineages (estimated divergence >200,000 years ago), with one lineage occurring only in eastern and southern Asia. A remarkably similar pattern exists in cattle, sheep, and pigs. These results, combined with recent archaeological findings, suggest that goats and other farm animals have multiple maternal origins with a possible center of origin in Asia, as well as in the Fertile Crescent. The pattern of goat mtDNA diversity suggests that all three lineages have undergone population expansions, but that the expansion was relatively recent for two of the lineages (including the Asian lineage). Goat populations are surprisingly less genetically structured than cattle populations. In goats only approximately 10% of the mtDNA variation is partitioned among continents. In cattle the amount is >/=50%. This weak structuring suggests extensive intercontinental transportation of goats and has intriguing implications about the importance of goats in historical human migrations and commerce.

Animals↗

An extensive analysis of Y-chromosomal microsatellite haplotypes in globally dispersed human populations.

The genetic variance at seven Y-chromosomal microsatellite loci (or short tandem repeats [STRs]) was studied among 986 male individuals from 20 globally dispersed human populations. A total of 598 different haplotypes were observed, of which 437 (73.1%) were each found in a single male only. Population-specific haplotype-diversity values were.86-.99. Analyses of haplotype diversity and population-specific haplotypes revealed marked population-structure differences between more-isolated indigenous populations (e.g., Central African Pygmies or Greenland Inuit) and more-admixed populations (e.g., Europeans or Surinamese). Furthermore, male individuals from isolated indigenous populations shared haplotypes mainly with male individuals from their own population. By analysis of molecular variance, we found that 76.8% of the total genetic variance present among these male individuals could be attributed to genetic differences between male individuals who were members of the same population. Haplotype sharing between populations, phi(ST) statistics, and phylogenetic analysis identified close genetic affinities among European populations and among New Guinean populations. Our data illustrate that Y-chromosomal STR haplotypes are an ideal tool for the study of the genetic affinities between groups of male subjects and for detection of population structure.

Africa↗

A simple method of removing the effect of a bottleneck and unequal population sizes on pairwise genetic distances.

In this paper, we derive the expectation of two popular genetic distances under a model of pure population fission allowing for unequal population sizes. Under the model, we show that conventional genetic distances are not proportional to the divergence time and generally overestimate it due to unequal genetic drift and to a bottleneck effect at the divergence time. This bias cannot be totally removed even if the present population sizes are known. Instead, we present a method to estimate the divergence times between populations which is based on the average number of nucleotide differences within and between populations. The method simultaneously estimates the divergence time, the ancestral population size and the relative sizes of the derived populations. A simulation study revealed that this method is essentially unbiased and that it leads to better estimates than traditional approaches for a very wide range of parameter values. Simulations also indicated that moderate population growth after divergence has little effect on the estimates of all three estimated parameters. An application of our method to a comparison of humans and chimpanzee mitochondrial DNA diversity revealed that common chimpanzees have a significantly larger female population size than humans.

Animals↗

A linkage disequilibrium map of the MHC region based on the analysis of 14 loci haplotypes in 50 French families.

A sample of 100 individuals from 50 French families of known pedigrees were typed for 14 loci of the HLA region (DPB1, DQB1, DQA1, DRB1, DRB3, 4, 5, C4B, C4A, Bf, C2, TNFa, TNFb, B, Cw, A). Linkage disequilibrium in each pair of loci was investigated by an exact test using a Markov chain algorithm. The results indicate no disequilibrium between DPB1 and the other loci, whereas the other class II genes are all significantly linked to each other. Linkage disequilibrium is also detected between some pairs of class I and class II-class I loci despite the long physical distance separating the loci (e.g. A-B, Cw-DRB1). On the other hand, some contiguous loci of the class III region are found to be in equilibrium with each other. Several hypotheses including selection, but also unequal allelic diversity at different MHC loci are discussed to explain this complex pattern of linkage disequilibrium.

Chromosome Mapping↗

Maternal and paternal lineages in Albania and the genetic structure of Indo-European populations.

Mitochondrial DNA HV1 sequences and Y chromosome haplotypes (DYS19 STR and YAP) were characterised in an Albanian sample and compared with those of several other Indo-European populations from the European continent. No significant difference was observed between Albanians and most other Europeans, despite the fact that Albanians are clearly different from all other Indo-Europeans linguistically. We observe a general lack of genetic structure among Indo-European populations for both maternal and paternal polymorphisms, as well as low levels of correlation between linguistics and genetics, even though slightly more significant for the Y chromosome than for mtDNA. Altogether, our results show that the linguistic structure of continental Indo-European populations is not reflected in the variability of the mitochondrial and Y chromosome markers. This discrepancy could be due to very recent differentiation of Indo-European populations in Europe and/or substantial amounts of gene flow among these populations.

Albania↗

Inferring the impact of linguistic boundaries on population differentiation: application to the Afro-Asiatic-Indo-European case.

We present here a quantitative way to assess the impact of language-family boundaries on population differentiation and to evaluate the homogeneity of the genetic processes along these boundaries. Our estimator (delta a) of the impact of the boundary is based on an isolation by distance (IBD) model and measures the added genetic distance between populations located on different sides of the boundary. We compare this statistic with another estimator of group differentiation (F(CT)) computed under an analysis of variance framework that does not assume any particular spatial structure of the populations. Monte Carlo simulations are used to study the behaviour of these statistics under a two-dimensional stepping-stone model. Simulations show that F(CT) can suggest the existence of a frontier when populations only differ because of IBD. This spurious behaviour is much less frequent for the delta a statistic. However, the large variance associated with the delta a statistic, and the fact that it should only be computed in the presence of IBD, may limit the use of this statistic. Overall, the origin and the effect of the boundary is best understood by comparing different statistics and by testing for the presence of IBD on each side of the boundary as well as across the boundary. We illustrate our approach by examining the boundary between Afro-Asiatic and Indo-European populations. These populations are globally genetically differentiated, but the effect of the linguistic boundary on gene flow seems geographically very heterogeneous. This boundary appears to be the result of a secondary contact between two differentiation centres rather than an enhancer of population differentiation.

Africa↗

Is the Gibraltar strait a barrier to gene flow for the bat Myotis myotis (Chiroptera: Vespertilionidae)?

Because of their role in limiting gene flow, geographical barriers like mountains or seas often coincide with intraspecific genetic discontinuities. Although the Strait of Gibraltar represents such a potential barrier for both plants and animals, few studies have been conducted on its impact on gene flow. Here we test this effect on a bat species (Myotis myotis) which is apparently distributed on both sides of the strait. Six colonies of 20 Myotis myotis each were sampled in southern Spain and northern Morocco along a linear transect of 1350 km. Results based on six nuclear microsatellite loci reveal no significant population structure within regions, but a complete isolation between bats sampled on each side of the strait. Variability at 600 bp of a mitochondrial gene (cytochrome b) confirms the existence of two genetically distinct and perfectly segregating clades, which diverged several million years ago. Despite the narrowness of the Gibraltar Strait (14 km), these molecular data suggest that neither males, nor females from either region have ever reproduced on the opposite side of the strait. Comparisons of molecular divergence with bats from a closely related species (M. blythii) suggest that the North African clade is possibly a distinct taxon warranting full species rank. We provisionally refer to it as Myotis cf punicus Felten 1977, but a definitive systematic understanding of the whole Mouse-eared bat species complex awaits further genetic sampling, especially in the Eastern Mediterranean areas.

Animals↗

Why hunter-gatherer populations do not show signs of pleistocene demographic expansions.

The mitochondrial DNA diversity of 62 human population samples was examined for potential signals of population expansions. Stepwise expansion times were estimated by taking into account heterogeneity of mutation rates among sites. Assuming an mtDNA divergence rate of 33% per million years, most populations show signals of Pleistocene expansions at around 70,000 years (70 KY) ago in Africa and Asia, 55 KY ago in America, and 40 KY ago in Europe and the Middle East, whereas the traces of the oldest expansions are found in East Africa (110 KY ago for the Turkana). The genetic diversity of two groups of populations (most Amerindian populations and present-day hunter-gatherers) cannot be explained by a simple stepwise expansion model. A multivariate analysis of the genetic distances among 61 populations reveals that populations that did not undergo demographic expansions show increased genetic distances from other populations, confirming that the demography of the populations strongly affects observed genetic affinities. The absence of traces of Pleistocene expansions in present-day hunter-gatherers seems best explained by the occurrence of recent bottlenecks in those populations, implying a difference between Pleistocene (approximately 1,800 KY to 10 KY ago) and Holocene (10 KY to present) hunter-gatherers demographies, a difference that occurred after, and probably in response to, the Neolithic expansions of the other populations.

Anthropology↗

Minisatellite mutational processes reduce F(st) estimates.

We have used a new method for binning minisatellite alleles (semi-automated allele aggregation) and report the extent of population diversity detectable by eleven minisatellite loci in 2,689 individuals from 19 human populations distributed widely throughout the world. Whereas population relationships are consistent with those found in other studies, our estimate of genetic differentiation (F(st)) between populations is less than 8%, which is lower than comparative estimates of between 10%-15% obtained by using other sources of polymorphism data. We infer that mutational processes are involved in reducing F(st) estimates from minisatellite data because, first, the lowest F(st) estimates are found at loci showing autocorrelated frequencies among alleles of similar size and, second, F(st) declines with heterozygosity but by more than predicted assuming simple models of mutation. These conclusions are consistent with the view that minisatellites are subject to selective or mutational constraints in addition to those expected under simple step-wise mutation models.

Alleles↗

Estimation of past demographic parameters from the distribution of pairwise differences when the mutation rates vary among sites: application to human mitochondrial DNA.

Distributions of pairwise differences often called "mismatch distributions" have been extensively used to estimate the demographic parameters of past population expansions. However, these estimations relied on the assumption that all mutations occurring in the ancestry of a pair of genes lead to observable differences (the infinite-sites model). This mutation model may not be very realistic, especially in the case of the control region of mitochondrial DNA, where this methodology has been mostly applied. In this article, we show how to infer past demographic parameters by explicitly taking into account a finite-sites model with heterogeneity of mutation rates. We also propose an alternative way to derive confidence intervals around the estimated parameters, based on a bootstrap approach. By checking the validity of these confidence intervals by simulations, we find that only those associated with the timing of the expansion are approximately correctly estimated, while those around the population sizes are overly large. We also propose a test of the validity of the estimated demographic expansion scenario, whose proper behavior is verified by simulation. We illustrate our method with human mitochondrial DNA, where estimates of expansion times are found to be 10-20% larger when taking into account heterogeneity of mutation rates than under the infinite-sites model.

DNA, Mitochondrial↗

Substitution rate variation among sites in mitochondrial hypervariable region I of humans and chimpanzees.

Mitochondrial D-loop hypervariable region I (HVI) sequences are widely used in human molecular evolutionary studies, and therefore accurate assessment of rate heterogeneity among sites is essential. We used the maximum-likelihood method to estimate the gamma shape parameter alpha for variable substitution rates among sites for HVI from humans and chimpanzees to provide estimates for future studies. The complete data of 839 humans and 224 chimpanzees, as well as many subsets of these data, were analyzed to examine the effect of sequence sampling. The effects of the genealogical tree and the nucleotide substitution model were also examined. The transition/transversion rate ratio (kappa) is estimated to be about 25, although much larger and biased estimates were also obtained from small data sets at low divergences. Estimates of alpha were 0.28-0.39 for human data sets of different sizes and 0.20-0.39 for data sets including different chimpanzee subspecies. The combined data set of both species gave estimates of 0.42-0.45. While all those estimates suggest highly variable substitution rates among sites, smaller samples tend to give smaller estimates of alpha. Possible causes for this pattern were examined, such as biases in the estimation procedure and shifts in the rate distribution along certain lineages. Computer simulations suggest that the estimation procedure is quite reliable for large trees but can be biased for small samples at low divergences. Thus, an alpha of 0.4 appears suitable for both humans and chimpanzees. Estimates of alpha can be affected by the nucleotide sites included in the data, the overall tree length (the amount of sequence divergence), the number of rate classes used for the estimation, and to a lesser extent, the included sequences. The genealogical tree, the substitution model, and demographic processes such as population expansion do not have much effect.

Animals↗

Incorporating genotypes of relatives into a test of linkage disequilibrium.

Genetic data from autosomal loci in diploids generally consist of genotype data for which no phase information is available, making it difficult to implement a test of linkage disequilibrium. In this paper, we describe a test of linkage disequilibrium based on an empirical null distribution of the likelihood of a sample. Information on the genotypes of related individuals is explicitly used to help reconstruct the gametic phase of the independent individuals. Simulation studies show that the present approach improves on estimates of linkage disequilibrium gathered from samples of completely independent individuals but only if some offspring are sampled together with their parents. The failure to incorporate some parents sharply decreases the sensitivity and accuracy of the test. Simulations also show that for multiallelic data (more than two alleles) our testing procedure is not as powerful as an exact test based on known haplotype frequencies, owing to the interaction between departure from Hardy-Weinberg equilibrium and linkage disequilibrium.

Genetic Diseases, Inborn↗

Different genetic components in the Ethiopian population, identified by mtDNA and Y-chromosome polymorphisms.

Seventy-seven Ethiopians were investigated for mtDNA and Y chromosome-specific variations, in order to (1) define the different maternal and paternal components of the Ethiopian gene pool, (2) infer the origins of these maternal and paternal lineages and estimate their relative contributions, and (3) obtain information about ancient populations living in Ethiopia. The mtDNA was studied for the RFLPs relative to the six classical enzymes (HpaI, BamHI, HaeII, MspI, AvaII, and HincII) that identify the African haplogroup L and the Caucasoid haplogroups I and T. The sample was also examined at restriction sites that define the other Caucasoid haplogroups (H, U, V, W, X, J, and K) and for the simultaneous presence of the DdeI10394 and AluI10397 sites, which defines the Asian haplogroup M. Four polymorphic systems were examined on the Y chromosome: the TaqI/12f2 and the 49a,f RFLPs, the Y Alu polymorphic element (DYS287), and the sY81-A/G (DYS271) polymorphism. For comparison, the last two Y polymorphisms were also examined in 87 Senegalese previously classified for the two TaqI RFLPs. Results from these markers led to the hypothesis that the Ethiopian population (1) experienced Caucasoid gene flow mainly through males, (2) contains African components ascribable to Bantu migrations and to an in situ differentiation process from an ancestral African gene pool, and (3) exhibits some Y-chromosome affinities with the Tsumkwe San (a very ancient African group). Our finding of a high (20%) frequency of the "Asian" DdeI10394AluI10397 (++) mtDNA haplotype in Ethiopia is discussed in terms of the "out of Africa" model.

Black People↗

Inferring admixture proportions from molecular data.

We derive here two new estimators of admixture proportions based on a coalescent approach that explicitly takes into account molecular information as well as gene frequencies. These estimators can be applied to any type of molecular data (such as DNA sequences, restriction fragment length polymorphisms [RFLPs], or microsatellite data) for which the extent of molecular diversity is related to coalescent times. Monte Carlo simulation studies are used to analyze the behavior of our estimators. We show that one of them (mY) appears suitable for estimating admixture from molecular data because of its absence of bias and relatively low variance. We then compare it to two conventional estimators that are based on gene frequencies. mY proves to be less biased than conventional estimators over a wide range of situations and especially for microsatellite data. However, its variance is larger than that of conventional estimators when parental populations are not very differentiated. The variance of mY becomes smaller than that of conventional estimators only if parental populations have been kept separated for about N generations and if the mutation rate is high. Simulations also show that several loci should always be studied to achieve a drastic reduction of variance and that, for microsatellite data, the mean square error of mY rapidly becomes smaller than that of conventional estimators if enough loci are surveyed. We apply our new estimator to the case of admixed wolflike Canid populations tested for microsatellite data.

Animals↗

Human genetic affinities for Y-chromosome P49a,f/TaqI haplotypes show strong correspondence with linguistics.

Numerous population samples from around the world have been tested for Y chromosome-specific p49a,f/TaqI restriction polymorphisms. Here we review the literature as well as unpublished data on Y-chromosome p49a,f/TaqI haplotypes and provide a new nomenclature unifying the notations used by different laboratories. We use this large data set to study worldwide genetic variability of human populations for this paternally transmitted chromosome segment. We observe, for the Y chromosome, an important level of population genetics structure among human populations (FST = .230, P < .001), mainly due to genetic differences among distinct linguistic groups of populations (FCT = .246, P < .001). A multivariate analysis based on genetic distances between populations shows that human population structure inferred from the Y chromosome corresponds broadly to language families (r = .567, P < .001), in agreement with autosomal and mitochondrial data. Times of divergence of linguistic families, estimated from their internal level of genetic differentiation, are fairly concordant with current archaeological and linguistic hypotheses. Variability of the p49a,f/TaqI polymorphic marker is also significantly correlated with the geographic location of the populations (r = .613, P < .001), reflecting the fact that distinct linguistic groups generally also occupy distinct geographic areas. Comparison of Y-chromosome and mtDNA RFLPs in a restricted set of populations shows a globally high level of congruence, but it also allows identification of unequal maternal and paternal contributions to the gene pool of several populations.

DNA Probes↗

Testing for linkage disequilibrium in genotypic data using the Expectation-Maximization algorithm.

We generalize an approach suggested by Hill (Heredity, 33, 229-239, 1974) for testing for significant association among alleles at two loci when only genotype and not haplotype frequencies are available. The principle is to use the Expectation-Maximization (EM) algorithm to resolve double heterozygotes into haplotypes and then apply a likelihood ratio test in order to determine whether the resolutions of haplotypes are significantly nonrandom, which is equivalent to testing whether there is statistically significant linkage disequilibrium between loci. The EM algorithm in this case relies on the assumption that genotype frequencies at each locus are in Hardy-Weinberg proportions. This method can accommodate X-linked loci and samples from haplodiploid species. We use three methods for testing significance of the likelihood ratio: the empirical distribution in a large number of randomized data sets, the X2 approximation for the distribution of likelihood ratios, and the Z2 test. The performance of each method is evaluated by applying it to simulated data sets and comparing the tail probability with the tail probability from Fisher's exact test applied to the actual haplotype data. For realistic sample sizes (50-150 individuals) all three methods perform well with two or three alleles per locus, but only the empirical distribution is adequate when there are five to eight alleles per locus, as is typical of hypervariable loci such as microsatellites. The method is applied to a data set of 32 microsatellite loci in a Finnish population and the results confirm the theoretical predictions. We conclude that with highly polymorphic loci, the EM algorithm does lead to a useful test for linkage disequilibrium, but that it is necessary to find the empirical distribution of likelihood ratios in order to perform a test of significance correctly.

Algorithms↗

A generic estimation of population subdivision using distances between alleles with special reference for microsatellite loci.

Several estimators of population differentiation have been proposed in the recent past to deal with various types of genetic markers (i.e., allozymes, nucleotide sequences, restriction fragment length polymorphisms, or microsatellites). We discuss the relationships among these estimators and show how a single analysis of variance framework can accomodate these qualitatively different data types.

Alleles↗

The impact of population expansion and mutation rate heterogeneity on DNA sequence polymorphism.

In order to study the effect of mutation rate heterogeneity on patterns of DNA polymorphism, we simulated samples of DNA sequences with gamma-distributed nucleotide substitution rates in stationary and expanding populations. We find that recent population expansions and mutation rate heterogeneity have similar effects on several polymorphism indicators, like the shape and the mean of the observed pairwise difference distribution, or the number of segregating sites. The inferred size of population expansion thus appears overestimated if nucleotides have dissimilar substitution rates. Interestingly, population expansion and uneven mutation rates have contrasting effects on Tajima's D statistic when acting separately, and the consequence on the associated test of selective neutrality is investigated. The patterns of polymorphism of several human populations analyzed for the mitochondrial control region are examined, mainly showing the difficulty in quantifying the respective contribution of past demographic history and uneven mutation rates from a single sampled evolutionary process. However, substitution rates appear more heterogeneous in the second hypervariable segment of the control region than in the first segment.

Animals↗