Search PubMedSearch

SEARCH · Search PubMed

Results for “population structure”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Association mapping in structured populations.

The use, in association studies, of the forthcoming dense genomewide collection of single-nucleotide polymorphisms (SNPs) has been heralded as a potential breakthrough in the study of the genetic basis of common complex disorders. A serious problem with association mapping is that population structure can lead to spurious associations between a candidate marker and a phenotype. One common solution has been to abandon case-control studies in favor of family-based tests of association, such as the transmission/disequilibrium test (TDT), but this comes at a considerable cost in the need to collect DNA from close relatives of affected individuals. In this article we describe a novel, statistically valid, method for case-control association studies in structured populations. Our method uses a set of unlinked genetic markers to infer details of population structure, and to estimate the ancestry of sampled individuals, before using this information to test for associations within subpopulations. It provides power comparable with the TDT in many settings and may substantially outperform it if there are conflicting associations in different subpopulations.

Alleles

Sex ratio theory in geographically structured populations.

Equilibrium sex ratios have been determined analytically under Wright's island model in order to determine the effect of population structure with limited dispersal. When mating occurs before dispersal, the dispersal rate has little or no effect, and the equilibrium sex ratio remains the same as under complete dispersal (the standard model of local mate competition). With dispersal before mating, there is a bias towards the sex with the higher dispersal rate due to lower competition between sibs of that sex.

Animals

Genome-wide SNP-based genomic diversity and population structure analysis in alpaca populations from Europe and Peru.

This study aimed to analyze the genetic diversity and population structure of alpacas in Germany, Switzerland, and Austria (German-speaking regions, GSR) and to compare with that of the country of origin of the species (Peru). A total of 179 animals from GSR and 151 from Peru were genotyped with a species-specific 76k SNP array. The observed and expected heterozygosity was 0.305 and 0.311 for GSR and 0.310 and 0.312 for Peru. The mean FROH values were 0.029 for GSR and 0.023 for Peru. In general, results show that breeders in both analyzed regions efficiently maintain genetic diversity. Principal component analysis identified the GSR and Peru populations as separate from each other, but the relative proximity of both clusters indicates the shared genetic heritage. FST and XPEHH methods identified genomic regions under selection for traits such as coat color and adaptation. Genome-wide association studies comparing black and brown with white or gray alpacas identified associated genome regions containing the ASIP and KIT genes, respectively. The association of a recently identified keratin locus on chromosome 16 with differences in fleece type in alpacas was confirmed, while the putative causality of a TRPV3 variant was rejected.

Animals

SPC: a SPectral Component approach leveraging Identity-by-Descent graphs to address recent population structure in genomic analysis.

Population structure is a well-known confounder in statistical genetics, particularly in genome-wide association studies (GWAS), where it can lead to inflated test statistics and spurious associations. Traditional methods, such as principal components (PCs), commonly used to adjust for population structure, are limited in capturing fine-scale, non-linear patterns that arise from recent demographic events - patterns that are crucial for understanding rare variant effects. To address this challenge, we propose a novel method called SPectral Components (SPCs), which leverages identity-by-descent (IBD) graphs to capture and transform local, non-linear fine-scale population structure into continuous representations that can be seamlessly integrated into genetic analysis pipelines. Using both simulated datasets and empirical data from the UK Biobank (N ≈ 420,000), we demonstrate that SPCs outperform PCs in adjusting for fine-scale population structure. In simulations, SPCs explained over 90% of the fine-scale population structure with fewer components, while PCs captured less than 5%. In the UK Biobank, SPCs reduced the inflation of p-values in the GWAS of an environmental-driven phenotype by 12% compared to PCs, while maintaining a similar performance to PCs in height, a highly heritable phenotype. Additionally, SPCs improved rare variant association analyses, reducing genomic inflation (e.g., from 7.6 to 1.2 in one analysis), and provided more accurate heritability estimates. Spatial autocorrelation analysis further confirmed the ability of SPCs to account for environmental effects, reducing Moran's I for both environmental and heritable phenotypes more effectively than PCs. Overall, our findings demonstrate that SPCs provide a robust, scalable adjustment for recent population structure, offering a powerful alternative or complement to PCs in large-scale biobank studies.

GWAS

Mitochondrial genome-derived microsatellites reveal genetic diversity and population structure in Callery pear populations.

Callery pear (Pyrus calleryana Decne.; PC) possesses many desirable characteristics valued in managed landscapes. This has driven the release of numerous cultivars, including both hybrids and selections derived from native populations. The extensive planting of PC cultivars in managed areas has contributed to the widespread occurrence of invasive individuals across a broad range of habitats in the eastern United States (US). Self-incompatibility, tolerance to various environmental conditions, pathogen and pest resistance, intraspecific hybridization among the cultivars, possible interspecific hybridization with other Pyrus species, and seed dispersal by various vertebrates have contributed to the spread and persistence of PC across diverse environments. Because effective and environmentally appropriate management options remain limited, improved understanding of PC genetics may help inform management strategies. Previous studies have characterized PC diversity using nuclear genomic short sequence repeats (gSSRs), however, neither a mitochondrial genome resource nor mitochondrial short sequence repeats (mtSSRs) have been developed for this purpose. Here, we assembled a mitochondrial genome of 485,892 bp and used five mtSSRs to characterize mitochondrial diversity and population structure among accessions from the species' native range in Asia (n = 72), southeastern US escapees (SNesc; n = 90), Tennessee escapees (TNesc; n = 90), and US-released commercial cultivars (UScult; n = 69 representing 14 unique cultivars). We found a high genetic diversity (He = 0.728) and evidence of genetic structure in PC. In distance-based and multivariate analyses, UScult occupied an intermediate position between the Asian populations and the US escapees. The observed mitochondrial diversity among samples assigned to PC cultivars is consistent with a complex genetic landscape and may reflect distinct maternal lineages, cultivar-labeling or record-keeping discrepancies, and/or technical variation. This study underscores the need for broader genomic investigations using authenticated cultivar reference material and high-resolution nuclear markers to resolve cultivar ancestry, validate true-to-name identity, and inform species management.

Genetic Variation

Population structure in Kanoya population, Japan.

The mean inbreeding coefficients found for Minami-cho (366 couples) and Shinsei-cho (511 couples) were 0.00307 and 0.00191, respectively. The mean inbreeding coefficient decreased and the mean marital distance increased as the year of marriage becomes more recent. The mean distances and their standard deviations between birthplaces of mates, father-offspring, mother-offspring, and sibs are 69.05 +/- 229.64, 73.09 +/- 246.66, 49.81 +/- 158.43, and 39.53 +/- 159.51 km, respectively, at Minami-cho. These values are 188.45 +/- 387.05, 187.79 +/- 562.59, 148.26 +/- 326.35, and 73.93 +/- 225.92 km, respectively, at Shinsei-cho. The dimensionality of migration is closest to one dimension.

Consanguinity

treestructure: an R package to detect population structure in phylogenetic trees.

MOTIVATION: How population structure can shape genetic diversity is a longstanding problem in population genetics. While the use of geographic locations, when available, can help answer some of these questions, it is still difficult to determine population structure when such metadata are not available or when the potential population structure is not easily observed. Here, we present an updated version of treestructure, an R package that implements a statistical test based on coalescent theory to detect unobserved population structure in a time-scaled phylogenetic tree. AVAILABILITY: treestructure is available at CRAN at https://cloud.r-project.org/web/packages/treestructure/ and at https://emvolz-phylodynamics.github.io/treestructure/.

Phylogeny

Recovering the precolonial population structure of Khoe-San descendant populations.

San populations from Botswana and Namibia retain exceptional linguistic, cultural, and genetic diversity, but few Khoisan-speaking groups remain south of the Kalahari Desert. However, historically, far southern Africa was home to many San and Khoekhoe groups. Popular opinion often implies that such populations do not contribute to the ancestry of contemporary South Africans. Here, we characterize the genetic ancestry of self-identified South African Coloured groups and reconstruct precolonial and colonial population structures from 620 newly sampled individuals. These groups retain the majority of Khoe-San genetic ancestry (>48%), suggesting the persistence of Khoe-San ancestry to the present day. By isolating the Khoe-San ancestry component, we show that it is intermediate between the ≠Khomani San and Nama and distinct from Kalahari Khoe-San populations. We also find that signatures of the Indian Ocean slave trade can be traced to Indonesian islands such as Sulawesi, Java, and Flores, while the South Asian ancestry is regionally nonspecific.

Humans

Genealogy of neutral genes and spreading of selected mutations in a geographically structured population.

In a geographically structured population, the interplay among gene migration, genetic drift and natural selection raises intriguing evolutionary problems, but the rigorous mathematical treatment is often very difficult. Therefore several approximate formulas were developed concerning the coalescence process of neutral genes and the fixation process of selected mutations in an island model, and their accuracy was examined by computer simulation. When migration is limited, the coalescence (or divergence) time for sampled neutral genes can be described by the convolution of exponential functions, as in a panmictic population, but it is determined mainly by migration rate and the number of demes from which the sample is taken. This time can be much longer than that in a panmictic population with the same number of breeding individuals. For a selected mutation, the spreading over the entire population was formulated as a birth and death process, in which the fixation probability within a deme plays a key role. With limited amounts of migration, even advantageous mutations take a large number of generations to spread. Furthermore, it is likely that these mutations which are temporarily fixed in some demes may be swamped out again by non-mutant immigrants from other demes unless selection is strong enough. These results are potentially useful for testing quantitatively various hypotheses that have been proposed for the origin of modern human populations.

Animals

Quantitative traits in relation to population structure: why and how are they used and what do they imply?

I describe the basic ingredients of a population structure analysis and the rationale for using polygenic quantitative traits in such analyses. The complexity of inheritance and the population dynamics of quantitative traits, however, imply that inferences regarding population structure based on such traits must be evaluated with appropriate cautions. Although many studies of quantitative traits in relation to population structure analysis underscore the importance of gene flow between subpopulations, I show that the role of selection in the evolution of a quantitative trait and its relationship to the inferred population structure cannot be overlooked. Finally, I review some recent advances in human quantitative genetic methodologies that can be used profitably in population structure analysis.

Anthropology, Physical

Kin groups and trait groups: population structure and epidemic disease selection.

A Monte Carlo simulation based on the population structure of a small-scale human population, the Semai Senoi of Malaysia, has been developed to study the combined effects of group, kin, and individual selection. The population structure resembles D.S. Wilson's structured deme model in that local breeding populations (Semai settlements) are subdivided into trait groups (hamlets) that may be kin-structured and are not themselves demes. Additionally, settlement breeding populations are connected by two-dimensional stepping-stone migration approaching 30% per generation. Group and kin-structured group selection occur among hamlets the survivors of which then disperse to breed within the settlement population. Genetic drift is modeled by the process of hamlet formation; individual selection as a deterministic process, and stepping-stone migration as either random or kin-structured migrant groups. The mechanism for group selection is epidemics of infectious disease that can wipe out small hamlets particularly if most adults become sick and social life collapses. Genetic resistance to a disease is an individual attribute; however, hamlet groups with several resistant adults are less likely to disintegrate and experience high social mortality. A specific human gene, hemoglobin E, which confers resistance to malaria, is studied as an example of the process. The results of the simulations show that high genetic variance among hamlet groups may be generated by moderate degrees of kin-structuring. This strong microdifferentiation provides the potential for group selection. The effect of group selection in this case is rapid increase in gene frequencies among the total set of populations. In fact, group selection in concert with individual selection produced a faster rate of gene frequency increase among a set of 25 populations than the rate within a single unstructured population subject to deterministic individual selection. Such rapid evolution with plausible rates of extinction, individual selection, and migration and a population structure realistic in its general form, has implications for specific human polymorphisms such as hemoglobin variants and for the more general problem of the tempo of evolution as well.

Communicable Diseases

The decay of genetic variability in geographically structured populations.

The geographical structure of a population distributed continuously and homogeneously along an infinite linear habitat is explored. The analysis is restricted to a single locus in the absence of selection, and every mutant is assumed to be new to the population. An explicit formula is derived for the probability that two homologous genes separated by a given distance at any time t are the same allele. The ultimate rate of approach to equilibrium is shown to be t(-3/2)e(-2ut), where u is the mutation rate. An approximation is given for the stationary probability of allelism in an infinite two-dimensional population, which, unlike previous expressions, is finite everywhere. For a finite habitat of arbitrary shape and any number of dimensions, it is proved that if the population density is very high, then asymptotically the transient part of the probability of allelism is spatially uniform and decays at the rate e(-[2u+1/(2N)]t), where N is the total population size. Thus, in this respect the population behaves as if it were panmictic. The dependence of the amount of local gene frequency differentiation on population density and habitat size and dimensionality is discussed.

Alleles

Examining population structure through the use of surname matrices: methodology for visualizing nonrandom mating.

The analysis of nonrandom mating using the frequency of marital isonymy indirectly measures the degree of population structure. However, population structure is the result of all matings in a population. Difficulties with large surname matrices have resulted in data being summarized into a single statistic or collapsed into brief tables, with considerable loss of information. By using sophisticated computer graphing procedures and displays, it is possible to directly analyze the mating structure of a community. If P is a vector of proportions for each male surname i (i = 1, 2, 3, ..., n), Q a similar vector of female surnames j(j = 1, 2, 3, ...,m), then the expected frequency matrix E of each possible mating is P x Q. The difference D between the observed frequency matrix O and the expected matrix is O-E. The D matrix is graphed with the x axis containing the male surnames, the y axis the female surnames, and the z axis the difference values dij. Negative values represent negative nonrandom mating and positive values positive nonrandom mating. From 5417 marriages (1840-1963) in the Midlands of Tasmania, those between spouses having 1 of 194 core names were extracted. We analyze these marriages utilizing the new technique and examine the surface of the graph and statistical analysis of its finer structure. Among the results was the demonstration of frequency-dependent selection of surnames. This finding has significant implications for microevolution of human populations, as surnames have existed for possibly 700 years.

Computer Simulation

Hierarchical selection theory and sex ratios. I. General solutions for structured populations.

Models of sex-ratio evolution in structured populations are derived with G.R. Price's covariance form for the hierarchical analysis of natural selection (1970, Nature 227, 520-521). Previous work on competition among related males for mates (local mate competition), competition among related females for a limiting resource (local resource competition), inbreeding, group selection, and asymmetry of genetic inheritance between males and females, are subsumed under a general formulation for sex-ratio biases in structured populations. I found that the evolutionarily stable strategy sex ratio (males:females) for diploids is 1 - rho m:1 - rho f, where rho m is the regression coefficient of relatedness of the controlling genotypes on males competing for mates, rho f is the regression of controlling genotypes on females that compete for a fixed, limiting resource, and there is no inbreeding. For inbreeding and no competition among females, the evolutionarily stable strategy is 1 - rho m:1 + rho mf, where rho mf is the regression of controlling genotypes on females' mates.

Animals

Partial correlation of distance matrices in studies of population structure.

Anthropological studies of human population structure commonly compare various monogenic and polygenic (metric) distance matrices to distance matrices obtained from measures of geographical dispersion, linguistic differences, and migration patterns in an attempt to infer something about the effects of evolutionary factors (drift and differential selection, in particular). It is, though, commonly recognized that geography, language, and migration patterns may be intercorrelated due to the common effects of historical and social processes. Previous attempts to deal with the problems of assessing relative effects among such sets of intercorrelated factors using partial correlations have resulted in coefficients that are either not well defined or have no known sampling distribution or both. Here, we outline a general approach to partialling distance matrices that results in well-defined coefficients and valid significance testing procedures. Application of the matrix partialling methods to a variety of distance matrices obtained for a sample of eight ethnolinguistic groups from the Harvard Solomon Islands Expedition (Friedlaender et al., 1986) reveals a close association between language dissimilarity and dermatoglyphics controlling for geography, thus reinforcing earlier suggestions that dermatoglyphics, properly used, reflect historical relationships of groups in this region better than do anthropometry, odontometrics, or small batteries of blood polymorphisms.

Anthropology

Life not lived due to disequilibrium in heterogeneous age-structured populations.

Three models of age-structured populations with demographically heterogeneous subpopulations are analyzed. In the first model, each subpopulation has its own age-specific vital rates which are fixed in time. In the second model, the vital rates of each subpopulation are uniformly inhibited by increasing total numbers of individuals. In the third, the vital rates of groups of subpopulations are inhibited by the total numbers of individuals in other groups of subpopulations with an intensity that depends on the interacting pair of groups. Three functions are defined to measure disequilibrium in the subpopulation frequencies, subpopulation age structures, and total population size. For the first model, we show that disequilibrium will shift the trajectory of the total numbers of individuals forward or backward in time by an asymptotic constant that is proportional to the sum of the disequilibrium measures. For the second model, we establish sufficient conditions for the existence of a globally stable equilibrium and we show that disequilibrium will result in a finite loss or gain in life which is proportional to the sum of the disequilibrium measures. For the last model, we show that the loss or gain in life for each group of subpopulations is a linear combination over all groups of the sums of the three disequilibrium measures. We illustrate these results with numerical examples and give possible biological interpretations of the models. We relate these new results to previous work on the cost of natural selection and measures of demographic disequilibrium.

Age Factors

Density-dependent migration and human population structure in historical Massachusetts.

Studies of population structure often focus on the effects of population size and migration rates on genetic variation. Few studies, however, have investigated the relationship between these two factors. The purpose of this paper is to determine the extent to which migration (and gene flow) is density-dependent (that is, affected by population size) for populations in historical Massachusetts. Data from 4,859 marriage records were analyzed from four populations in north-central Massachusetts during the time period 1741 to 1849. These data were placed into 29 samples defined in terms of population and time cohort. Within each cohort the overall exogamy rate was computed along with three estimates of gene flow based on marital migration: local migration (k), long-distance migration (m), and effective migration rate (me). Three samples show unusually low rates that reflect the history of settlement. Regression analyses were used with the remaining samples, and they show nonlinear density-dependent migration that is unrelated to temporal trends. Migration is highest in samples with small population sizes (less than 800) and large population sizes (greater than 1,600). Migration is lowest in medium-sized populations. Two processes are suggested to explain this curvilinear relationship of migration and population size. In small populations, the lack of suitable potential mates and/or availability of settled land leads to an increase in migration into the population. As population size increases, this migration decreases. After populations reach a certain size, migration increases again, most likely reflecting the economic pull of larger populations. These patterns could act to enhance, or counter, genetic drift, depending on the direction of density dependence.

Female