Search PubMedSearch

Biomedical subjects

R R Sokal

Publications and source records attributed to R R Sokal.

At least 19 recordsLinked to original sources

Indo-European origins: a computer-simulation test of five hypotheses.

Allele frequency distributions were generated by computer simulation of five models of microevolution in European populations. Genetic distances calculated from these distributions were compared with observed genetic distances among Indo-European speakers. The simulated models differ in complexity, but all incorporate random genetic drift and short-range gene flow (isolation by distance). The best correlations between observed and simulated data were obtained for two models where dispersal of Neolithic farmers from the Near East depends only on population growth. More complex models, where the timing of the farmers' expansion is constrained by archaeological time data, fail to account for a larger fraction of the observed genetic variation; this is also the case for a model including late Neolithic migrations from the Pontic steppes. The genetic structure of current populations speaking Indo-European languages seems therefore to largely reflect a Neolithic expansion. This is consistent with the hypothesis of a parallel spread of farming technologies and a proto-Indo-European language in the Neolithic. Allele-frequency gradients among Indo-European speakers may be due either to incomplete admixture between dispersing farmers, who presumably spoke proto-Indo-European, and pre-existing hunters and gatherers (as in the traditional demic diffusion hypothesis), or to founder effects during the farmers' dispersal. By contrast, successive migrational waves from the East, if any, do not seem to have had genetic consequences detectable by the present comparison of observed and simulated allele frequencies.

Alleles

Origins of Indo-Europeans and the spread of agriculture in Europe: comparison of lexicostatistical and genetic evidence.

A series of tests was undertaken to relate lexicostatistical dissimilarities (LAN) among 48 Indo-European languages to distances representing various causal hypotheses. The comparison is limited to languages currently spoken in Europe. The putative causal distance matrices include (1) geographic (GEO) distances between the languages, (2) distances representing the origin of agriculture (OOA), (3) distances representing a model postulated by C. Renfrew (REN) concerning transformations that gave rise to the major Indo-European language families in Europe, and (4) distances representing a competing hypothesis by M. Gimbutas (GIM) concerning the origin and spread of Indo-European languages in Europe. Pairwise Mantel tests of the matrices show that OOA correlates better with LAN than does REN, supporting Renfrew's basic hypothesis of the dispersal of the Indo-European languages with the spread of agriculture but showing less effect for his postulated transformations. Partial correlation of LAN with OOA when GEO is held constant is significant at p = 0.004, whereas REN is no longer correlated with LAN when GEO is held constant. When repeated for only seven languages chosen to represent the seven major families of Indo-European languages currently spoken in Europe, the results differed appreciably, yielding a negative, albeit nonsignificant, partial correlation between OOA and LAN when GEO is held constant. This apparent contradiction led us to develop some new statistical approaches to examine, confirm, and explain the patterns. Decomposing the Mantel correlation coefficients for the 48 Indo-European languages into several additive correlation components showed that much of the positive component of the correlation coefficient was contributed by LAN, OOA correlation within language families, particularly within the Germanic family, covering up the negative contributions between language families. The differentiation of the seven major Indo-European language branches in Europe seems unrelated to the times of the origin of agriculture. This finding fails to support the fundamental assumption of Renfrew's hypothesis. There are also no significant correlations between LAN and REN or GIM. A series of Monte Carlo experiments confirmed these findings. Consideration of the accumulated evidence from genetics supports the model of demic diffusion during the origin of agriculture. However, published genetic studies and the present study lend no support to the notion that the early farmers were indeed the Indo-Europeans.

Agriculture

Worldwide analysis of genetic and linguistic relationships of human populations.

In this study we relate language differences on a global scale with genetic distances for the same populations. The analysis is carried out on more populations (130) but fewer genetic systems (11) than earlier studies. We constructed an overall genetic distance matrix that allowed for missing values. A separate genetic distance matrix was also computed for each genetic system, and matching matrices of linguistic and geographic distances were associated with each genetic distance matrix, because the number of populations used differed among the genetic systems studied. Significant matrix correlations between language and genetics were found for both overall genetic distances and a substantial number of genetic systems, even when the effects of geographic distances were held constant. This demonstrates a significant correspondence between genetics and language on a global scale. Genetic matrices were correlated with two different linguistic distance matrices: one with higher (supraphyletic) taxonomic structure, in which among other features sub-Saharan Africans separate from non-Africans in the basal split and the Eurasiatic superphylum is postulated; and one without such structure. The correlations yield no genetic evidence to support the proposed higher linguistic structure. UPGMA and neighbor-joining trees were constructed for linguistic and genetic data. The proposed African split pattern is not supported by these data. Both types of trees indicate a pattern of grouping of east Asians, Arctic populations, and Australian natives separating from Caucasoid and African populations.

Ethnicity

Origins of the Indo-Europeans: genetic evidence.

Two theories of the origins of the Indo-Europeans currently compete. M. Gimbutas believes that early Indo-Europeans entered southeastern Europe from the Pontic Steppes starting ca. 4500 B.C. and spread from there. C. Renfrew equates early Indo-Europeans with early farmers who entered southeastern Europe from Asia Minor ca. 7000 BC and spread through the continent. We tested genetic distance matrices for each of 25 systems in numerous Indo-European-speaking samples from Europe. To match each of these matrices, we created other distance matrices representing geography, language, time since origin of agriculture, Gimbutas' model, and Renfrew's model. The correlation between genetics and language is significant. Geography, when held constant, produces a markedly lower, yet still highly significant partial correlation between genetics and language, showing that more remains to be explained. However, none of the remaining three distances--time since origin of agriculture, Gimbutas' model, or Renfrew's model--reduces the partial correlation further. Thus, neither of the two theories appears able to explain the origin of the Indo-Europeans as gauged by the genetics-language correlation.

Blood Group Antigens

Genetic evidence for the spread of agriculture in Europe by demic diffusion.

European agriculture originated in the Near East about 9,000 years ago. The Neolithic reached almost all areas suitable for agriculture by 5,000 yr BP (before present). The routes and times of the spread of agriculture through Europe are relatively well established, but not its manner of spreading. This could have been by cultural diffusion with few genetic consequences. By contrast, Ammerman and Cavalli-Sforza proposed that the spread of farming increased local population densities, causing demic expansion into new territory and diffusive gene flow between the neolithic farmers and mesolithic groups. We have now tested observed genetic patterns against expectations derived from the demic expansion hypothesis. We found significant partial correlations of genetic distances with a distance matrix especially designed to represent the spread of agriculture on that continent, when geographic distances are held constant. These findings support the hypothesis of Ammerman and Cavalli-Sforza and invite further investigation into Renfrew's hypothesis on the origin of the Indo-European languages.

Agriculture

Spatial differentiation of RH and GM haplotype frequencies in Sub-Saharan Africa and its relation to linguistic affinities.

This study analyzes patterns of variation in eight GM and seven Rhesus (RH) haplotypes across sub-Saharan Africa. We examine the concordance with genetic patterns of both geographic and language-family relationships by spatial analysis and ordination techniques. The genetic variation has significant spatial structure, but positive autocorrelation declines neither asymptotically nor proportionally with increasing distance. Evidently, neither isolation by distance with increasing distance. Evidently, neither isolation by distance nor clinical migration-selection models account for the observed genetic structure. Language-family relationship is the best predictor of genetic relationship and may reflect historic migrations and expansions of ethnically different peoples within sub-Saharan Africa. Yet the greatest part of the genetic variance remains unexplained by the models we have tested.

Africa, Central

Ancient movement patterns determine modern genetic variances in Europe.

A summary ethnohistory database on population movements in Europe between 2000 B.C. and A.D. 1970 was related to genetic variances and distances based on 26 genetic systems. For the purposes of these analyses, Europe was divided into 85 terrestrial quadrats measuring 5 degrees x 5 degrees. Counts, stratified by time, were taken of the number of movements out of and into each quadrat (called source and target counts, respectively) and between each pair of quadrats. The source and target counts have distinct and different patterns in Europe and vary significantly over time. Central Europe and the Pontic area have the quadrats with the highest source counts, and the Balkans have the highest target counts. Modern genetic variances per quadrat are significantly correlated with source and target counts, somewhat more prominently with source counts. Genetic distances between pairs of quadrats are correlated strongly with geographic distances and moderately and negatively correlated with the total number of movements between these quadrats. Partial correlations of genetic distances with total number of movements, holding geographic distance constant, are small and mostly nonsignificant. These results are interpreted in light of our knowledge of the history and biology of the populations concerned.

Databases, Factual

Genetic population structure of Italy. II. Physical and cultural barriers to gene flow.

Three approaches were employed to evaluate the relative importance of geographic and linguistic factors in maintaining genetic differentiation of Italian populations as shown by blood groups and erythrocyte and serum markers. Genetic distances are closer to linguistic than to geographic distances. Gene-frequency change across 12 linguistic boundaries is significantly more rapid than at random locations. The zones of sharp genetic variation correspond to physical barriers to gene flow and to boundaries between dialect families, which overlap widely. However, two linguistically differentiated populations appear genetically differentiated despite the absence of physical obstacles to gene flow around them. The Po River is associated with abrupt genetic change only in the area where it corresponds to a dialect boundary. At most loci the genetic population structure seems affected by linguistic rather than geographic factors; exceptions are the systems that were subject to malarial selection in geographically close but linguistically heterogeneous localities. Gene flow appears to homogenize gene frequencies within regions corresponding to dialect families but not between them, leading to the patchy distributions of allele frequencies that were detected in an earlier study.

Alleles

Genetic population structure of Italy. I. Geographic patterns of gene frequencies.

The diversity of spatial patterns of 61 allele frequencies for 20 genetic systems (15 loci) in Italy is presented. Blood antigens, enzymes, and proteins were analyzed. The total number of data points over all systems and localities was 1119. We used homogeneity tests, one-dimensional and directional spatial correlograms, and SYMAP interpolated surfaces. The data matrices were reduced by clustering techniques to reveal the principal patterns. Only a few allele frequency surfaces are strongly correlated across loci. All systems but one (ADA) exhibit significant heterogeneity in allele frequencies among the localities. Significant spatial patterns are shown by 27 of the 61 surfaces. Only one pattern (cde; system 4.19) is clinal; another (PGM1) exhibits a pure isolation by distance pattern; the others show long-range differentiation in addition to the short-distance decline of autocorrelation expected under isolation by distance. There is a marked decline in overall genetic similarity with distance for most variables. The 27 spatially significant alleles in Italy are also significantly patterned in Europe, but in all but 2 cases the country-wide and continent-wide patterns differ. The Italian patterns are due to forces specific to Italy. Differential selection for alleles associated with malaria is still evident. Whereas short-range differentiation can with malaria is still evident. Whereas short-range differentiation can be explained by isolation by distance, long-range differentiation appears to be due to demographic changes in certain populations that may be maintained by physical and linguistic isolation.

Gene Frequency

Genetic affinities of Jewish populations.

Genetic relations between various Jewish (J) and non-Jewish (NJ) populations were assessed using two sets of data. The first set contained 12 pairs of matched J and NJ populations from Europe, the Middle East, and North Africa, for which 10 common polymorphic genetic systems (13 loci) were available. The second set included 22 polymorphic genetic systems (26 loci) with various numbers of populations (ranging from 21 to 51) for each system. Therefore, each system was studied separately. Nei's standard genetic distance (D) matrices obtained for these two sets of data were tested against design matrices specifying hypotheses concerning the affiliations of the tested populations. The tests against single designs were carried out by means of Mantel tests. Our results consistently show lower distances among J populations than with their NJ neighbors, most simply explained by the common origin of the former. Yet, there is evidence also of genetic similarity between J and corresponding NJ populations, suggesting reciprocal gene flow between these populations or convergent selection in a common environment. The results of our study also indicate that stochastic factors are likely to have played a role in masking the descent relationships of the J populations.

Gene Frequency

Zones of sharp genetic change in Europe are also linguistic boundaries.

A newly elaborated method, "Wombling," for detecting regions of abrupt change in biological variables was applied to 63 human allele frequencies in Europe. Of the 33 gene-frequency boundaries discovered in this way, 31 are coincident with linguistic boundaries marking contiguous regions of different language families, languages, or dialects. The remaining two boundaries (through Iceland and Greece) separate descendants of different ethnic or geographical provenance but lack modern linguistic correlates. These findings support a model of genetic differentiation in Europe in which the genetic structure of the population is determined mainly by gene flow and admixture, rather than by adaptation to varying environmental conditions. Of the 33 boundaries, 27 reflect diverse population origins at often distant locations. Language affiliation of European populations plays a major role in maintaining and probably causing genetic differences.

Alleles

Genetic differences among language families in Europe.

We investigated whether 59 allele frequencies and 10 cranial variables differed among speakers of the 12 modern language families in Europe. Although this is a classical analysis of variance design, special techniques had to be developed for the analysis because of spatial autocorrelation of both biological and language data. The method examines pooled sums of squares within language families. These are compared with the same quantities obtained by randomly partitioning the available data points in Europe into internally cohesive subsets representing the same sample sizes for each language family as in the originally observed data. Our results suggest that for numerous genetic systems, population samples differ more among language families than they do within families. These findings are considered in relation to two contrasting models: a model of random spatial differentiation of gene frequencies unrelated to language and a model of aboriginal genetic differences among speakers of different language groups. Our observed findings suggest partial validity of both models.

Alleles

Spatial patterns of human gene frequencies in Europe.

The aims of this study of spatial patterns of human gene frequencies in Europe are twofold. One is to present new methodology developed for the analysis of such data. The other is to report on the diversity of spatial patterns observed in Europe and their interpretation as evidence of population processes. Spatial variation in 59 allele and haplotype frequencies (26 genetic systems) for polymorphisms in blood antigens, enzymes, and proteins is analyzed for an aggregate of 3,384 localities, using homogeneity tests, one-dimensional and directional spatial correlograms, and SYMAP interpolated surfaces. The data matrices are reduced to reveal the principal patterns by clustering techniques. The findings of this study can be summarized as follows: 1) There is significant heterogeneity in allele frequencies among the localities for all but one genetic system. 2) There are significant spatial patterns for most allele frequencies. 3) There is a substantial minority of clinal patterns in these populations. Clinal trends are found more frequently in HLA alleles than for other variables. North-south and northwest-southwest gradients predominate. 4) There is a strong decline in overall genetic similarity with geographic distance for most variables. 5) There are few, if any, appreciable correlations in pairs of allele frequencies over the continent, and there is little interesting correlation structure in the resulting correlation matrix. 6) Few spatial correlograms are markedly similar to each other, yet they form well-defined clusters. Spatial variation patterns, therefore, differ among allele frequencies. Patterns of human gene frequencies in modern Europe are diverse and complex. No single model suffices for interpretation of the observed genetic structure. Some clinal patterns reported here support the Neolithic demic-expansion hypothesis, others suggest latitudinal selection. Most of the clinal patterns are in HLA alleles, but there is also evidence from ABO for east-west migration diffusion. The majority of patterns are patchy, consistent with hypotheses of isolation by distance or of settlement of genetically differing, subsequently expanding ethnic groups. While undoubtedly there has been an ongoing stochastic process of differentiation consistent with the isolation-by-distance model, this has not obscured the directional patterns caused by migration (demic diffusion), and has perhaps only reinforced the contribution from settlement of ethnic units to patterns of genetic variation. However, the impact of the latter is most difficult to discern and requires further methodological developments.

ABO Blood-Group System

Spatial autocorrelation analysis of migration and selection.

We test various assumptions necessary for the interpretation of spatial autocorrelation analysis of gene frequency surfaces, using simulations of Wright's isolation-by-distance model with migration or selection superimposed. Increasing neighborhood size enhances spatial autocorrelation, which is reduced again for the largest neighborhood sizes. Spatial correlograms are independent of the mean gene frequency of the surface. Migration affects surfaces and correlograms when immigrant gene frequency differentials are substantial. Multiple directions of migration are reflected in the correlograms. Selection gradients yield clinal correlograms; other selection patterns are less clearly reflected in their correlograms. Sequential migration from different directions and at different gene frequencies can be disaggregated into component migration vectors by means of principal components analysis. This encourages analysis by such methods of gene frequency surfaces in nature. The empirical results of these findings lend support to the inference structure developed earlier for spatial autocorrelation analysis.

Computer Simulation

Genetic changes across language boundaries in Europe.

By means of three different methods we investigated whether 59 allele frequencies and ten cranial variables show increased change at 29 language-family boundaries in Europe. The quadrat-variance method compares variances of map quadrats crossed by language-family boundaries to variances of quadrats that are not crossed. The rate-of-change method examines the directional derivative of surfaces of the variables perpendicular to a language-family boundary and compares these derivatives to the same quantities obtained by randomly placing the language boundaries on the map of Europe. The difference method tests whether these variables differ more across language-family boundaries than across randomly placed boundaries. These special data-analytic techniques had to be developed to avoid the problem of spatial autocorrelation of both language and biological data. All three methods indicate increased genetic change at language-family boundaries. Clearer and more pronounced results are obtained by the first two methods than by the difference method. Thirteen language-family boundaries show significant gene frequency change by at least one of the methods. Changes are more marked in gene frequencies than in cranial variables. Different allele frequencies mark the increased change at different language boundaries. A model, based on the known history of each language-family boundary, was constructed to predict whether given boundaries should exhibit increased genetic change. The model is in good agreement with the observed results.

Alleles

Classification of the European language families by genetic distance.

Genetic distances among speakers of the European language families were computed by using gene-frequency data for human blood group antigens, enzymes, and proteins of 26 genetic systems. Each system was represented by a different subset of 3369 localities across Europe. By subjecting the matrix of distances to numerical taxonomic procedures, we obtained a grouping of the language families of Europe by their genetic distances as contrasted with their linguistic relationships. The resulting classification largely reflects geographic propinquity rather than linguistic origins. This is evidence for the primary importance of short-range interdemic gene flow in shaping the modern gene pools of Europe. Yet, some language families--i.e., Basque, Finnic (including Lappish), and Semitic (Maltese)--have distant genetic relationships with their geographic neighbors. These results indicate that European gene pools still reflect the remote origins of some ethnic units subsumed by these major linguistic groups.

Ethnicity

Genetic, geographic, and linguistic distances in Europe.

Genetic and taxonomic distances were computed for 3466 samples of human populations in Europe based on 97 allele frequencies and 10 cranial variables. Since the actual samples employed differed among the genetic systems studied, the genetic distances were computed separately for each system, as were matrices of geographic distances and of linguistic distances based on membership in the same language family or phylum. Significant matrix correlations between genetics and geography were found for the majority of systems; somewhat less frequent are significant correlations between genetics and language. The effects of the two factors can be separated by means of partial matrix correlations. These show significant values for both genetics and geography, language kept constant, and genetics and language, geography kept constant, with a tendency for the former to be higher. These findings demonstrate that speakers of different language families in Europe differ genetically and that this difference remains even after geographic differentiation is allowed for. The greater effect of geography than of language may be due to the several factors that bring about spatial differentiation in human populations.

Europe