Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “population stratification”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Effect of population stratification on case-control association studies. I. Elevation in false positive rates and comparison to confounding risk ratios (a simulation study).

OBJECTIVES: This is the first of two articles discussing the effect of population stratification on the type I error rate (i.e., false positive rate). This paper focuses on the confounding risk ratio (CRR). It is accepted that population stratification (PS) can produce false positive results in case-control genetic association. However, which values of population parameters lead to an increase in type I error rate is unknown. Some believe PS does not represent a serious concern, whereas others believe that PS may contribute to contradictory findings in genetic association. We used computer simulations to estimate the effect of PS on type I error rate over a wide range of disease frequencies and marker allele frequencies, and we compared the observed type I error rate to the magnitude of the confounding risk ratio. METHODS: We simulated two populations and mixed them to produce a combined population, specifying 160 different combinations of input parameters (disease prevalences and marker allele frequencies in the two populations). From the combined populations, we selected 5000 case-control datasets, each with either 50, 100, or 300 cases and controls, and determined the type I error rate. In all simulations, the marker allele and disease were independent (i.e., no association). RESULTS: The type I error rate is not substantially affected by changes in the disease prevalence per se. We found that the CRR provides a relatively poor indicator of the magnitude of the increase in type I error rate. We also derived a simple mathematical quantity, Delta, that is highly correlated with the type I error rate. In the companion article (part II, in this issue), we extend this work to multiple subpopulations and unequal sampling proportions. CONCLUSION: Based on these results, realistic combinations of disease prevalences and marker allele frequencies can substantially increase the probability of finding false evidence of marker disease associations. Furthermore, the CRR does not indicate when this will occur.

Bias↗

Detecting association in a case-control study while correcting for population stratification.

Case-control studies are subject to the problem of population stratification, which can occur in ethnically mixed populations and can lead to significant associations being detected at loci that have nothing to do with disease. Here, we describe a way to measure and correct for stratification by genotyping a moderate number of unlinked genetic markers in the same set of cases and controls in which a candidate association was found. The average of association statistics across the markers directly measures stratification. By dividing the candidate association statistic by this average, a P-value can be obtained that corrects for stratification.

Case-Control Studies↗

Effect of population stratification on case-control association studies. II. False-positive rates and their limiting behavior as number of subpopulations increases.

There has been considerable debate in the literature concerning bias in case-control association mapping studies due to population stratification. In this paper, we perform a theoretical analysis of the effects of population stratification by measuring the inflation in the test's type I error (or false-positive rate). Using a model of stratified sampling, we derive an exact expression for the type I error as a function of population parameters and sample size. We give necessary and sufficient conditions for the bias to vanish when there is no statistical association between disease and marker genotype in each of the subpopulations making up the total population. We also investigate the variation of bias with increasing subpopulations and show, both theoretically and by using simulations, that the bias can sometimes be quite substantial even with a very large number of subpopulations. In a companion simulation-based paper (Heiman et al., Part I, this issue), we have focused on the CRR (confounding risk ratio) and its relationship to the type I error in the case of two subpopulations, and have also quantified the magnitude of the type I error that can occur with relatively low CRR values.

Bias↗

Continuous covariates in genetic association studies of case-parent triads: gene and gene-environment interaction effects, population stratification, and power analysis.

We propose a multinomial logistic regression method which permits estimation and likelihood ratio tests for allele effects, their interactions with continuous covariates, and assessment of the degree of population stratification in genetic association studies of case-parent triads. Our approach overcomes the constraint imposed by the categorical nature of explanatory variables in the log-linear model. We also demonstrate that the multinomial logistic method can yield efficient inference in the presence of missing parental genotype data via the use of the Expectation-Maximization (EM) algorithm. We performed simulations to compare the multinomial logistic model with the case-pseudosibling conditional logistic model approach, both of which permit the incorporation of continuous covariates. Simulation results indicate that the multinomial logistic model and the conditional logistic model lead to similar estimates in large samples. A simulation-based method of sample size estimation is also used to show that the two models are approximately equivalent in sample size requirements. When parental genotype data are missing, either completely at random or dependent on covariates, the use of the EM algorithm gives multinomial logistic model greater power. Since the multinomial logistic model offers the possibility of assessing the degree of population stratification in the sample and can also provide efficient inference in the presence of missing parental genotypes, the proposed model has an important application in epidemiological family-based association studies.

Journal Article↗

Examining population stratification via individual ancestry estimates versus self-reported race.

Population stratification has the potential to affect the results of genetic marker studies. Estimating individual ancestry provides a continuous measure to assess population structure in case-control studies of complex disease, instead of using self-reported racial groups. We estimate individual ancestry using the Federal Bureau of Investigation CODIS Core short tandem repeat set of 13 loci using two different analysis methods in a case-control study of early-onset lung cancer. Individual ancestry proportions were estimated for "European" and "West African" groups using published allele frequencies. The majority of Caucasian, non-Hispanics had >50% European ancestry, whereas the majority of African Americans had <20% European ancestry, regardless of ancestry estimation method, although significant overlap by self-reported race and ancestry also existed. When we further investigated the effect of ancestry and self-reported race on the frequency of a lung cancer risk genotype, we found that the frequency of the GSTM1 null genotype varies by individual European ancestry and case-control status within self-reported race (particularly for African Americans). Genetic risk models showed that adjusting for individual European ancestry provided a better fit to the data compared with the model with no group adjustment or adjustment for self-reported race. This study suggests that significant population substructure differences exist that self-reported race alone does not capture and that individual ancestry may be confounded with disease status and/or a candidate gene risk genotype.

Black People↗

[Use of unlinked genetic markers to detect population stratification: a case-control multicenter study].

OBJECTIVE: To evaluate whether population stratification exists in microsatellite polymorphic sites. METHODS: Eight microsatellite markers not related to stroke, D11S1361, D19S927, D7S483, D14S990, D15S993, D1S2622, D1S2876 and D3S3560 located in different chromosomes were selected based on the GenBank database (http://www.gdb.org/). PCR assay was used to detect the alleles of these microsatellite in 294 patients of cerebral apoplexy, aged 58, and 325 sex, age and geographically-match controls in 7 cities in China. The PCR products were subjected to electrophoresis on the ABI 377 DNA sequencer and analyzed with Genescana and Genotypera software. The frequencies of these markers in these stroke patients and controls were compared by c2 test. Total 294 patients with stroke and 325 age, sex, and geographically matched controls were randomly selected from 973 case-control population. Eight loci including D11S1361, D19S927, D7S483, D14S990, D15S993, D1S2622, D1S2876 and D3S3560 were detected by fluorescence-based genotyping approach. RESULTS: Except for the alleles of D11S1361, the size, frequency, heterozygosity, and polymorphism information in the alleles of the other 7 microsatellite sites were in the range of 0.63 - 0.81 and their distribution varies in different races and populations. between Chinese and Caucasians. There were no differences in the frequencies of main alleles between the case group and control group (all P > or = 0.05). CONCLUSION: The selected seven markers are highly polymorphic and variable between human populations, and the genomic control data indicate that there is no unequal genetic admixture and stratification in the case-control cohort.

Aged↗

An evaluation of power and type I error of single-nucleotide polymorphism transmission/disequilibrium-based statistical methods under different family structures, missing parental data, and population stratification.

Researchers conducting family-based association studies have a wide variety of transmission/disequilibrium (TD)-based methods to choose from, but few guidelines exist in the selection of a particular method to apply to available data. Using a simulation study design, we compared the power and type I error of eight popular TD-based methods under different family structures, frequencies of missing parental data, genetic models, and population stratifications. No method was uniformly most powerful under all conditions, but type I error was appropriate for nearly every test statistic under all conditions. Power varied widely across methods, with a 46.5% difference in power observed between the most powerful and the least powerful method when 50% of families consisted of an affected sib pair and one parent genotyped under an additive genetic model and a 35.2% difference when 50% of families consisted of a single affection-discordant sibling pair without parental genotypes available under an additive genetic model. Methods were generally robust to population stratification, although some slightly less so than others. The choice of a TD-based test statistic should be dependent on the predominant family structure ascertained, the frequency of missing parental genotypes, and the assumed genetic model.

Computer Simulation↗

Robustness of case-control studies of genetic factors to population stratification: magnitude of bias and type I error.

Case-control studies of genetic factors are prone to a special form of confounding called population stratification, whenever the existence of one or more subpopulations may lead to a false association, be it positive or negative. We quantify both the bias (in terms of confounding risk ratio) and the probability of false association (type I error) in the most unfavorable situation in which only one high-risk subpopulation is hidden within the studied population, considering different scenarios of population structuring and varying sample sizes. In accord with previous work, we find that the bias is likely to be small in most cases. In addition, we show that the same applies to the associated type I error whenever the subpopulation is small in proportion. For instance, when the hidden subpopulation makes up 5% of the entire population, with an allelic frequency of 0.25 (versus 0.10) and a disease rate that is double, then the estimated bias is 1.07 and the type I error associated with a sample of 500 cases and 500 controls is 8% (instead of 5%). We also show that the type I error is substantially greater for a rare allele (frequency of 0.1) than for a common allele (frequency of 0.5) and analyze the pattern of increase of vulnerability to stratification bias with sample size. Based on our findings, we may therefore conclude that with moderate sample sizes the type I error associated with population stratification remains very limited in most realistic scenarios.

Bias↗

Properties of structured association approaches to detecting population stratification.

OBJECTIVE: To examine the properties of the structured association approach for the detection and correction of population stratification. METHOD: A method is developed, within a latent class analysis framework, similar to the methods proposed by Satten et al. (2001) and Pritchard et al. (2000). A series of simulations illustrate the relative impact of number and type of loci, sample size and population structure. RESULTS: The ability to detect stratification and assign individuals to population strata is determined for a number of different scenarios. CONCLUSION: The results underline the importance of careful marker selection.

Algorithms↗

Hierarchical modeling of the relation between sequence variants and a quantitative trait: addressing multiple comparison and population stratification issues.

When analyzing the relation between genetic sequence information and disease traits, false-positive associations can arise due to multiple comparisons and population stratification. In an attempt to address these issues, we incorporate into a conventional analytic model higher-level--or "prior"--models that use additional information to improve estimates while allowing for differing population structures. We apply this hierarchical model to simulated data from the Genetic Analysis Workshop 12. We focus on the effects of common candidate gene sequence variants on quantitative risk factor 5 (Q5) levels. In particular, we compare the regression coefficients (and 95% confidence intervals) obtained from conventional (one-stage) analyses versus the corresponding results from the hierarchical analyses. When examining either the marry-ins or all subjects in the general and isolate populations, the conventional model detected numerous sites in candidate genes 1-5 and 7 that had statistically significant regression coefficients (alpha level = 0.05). In contrast, our hierarchical model primarily only detected associations for variants in candidate gene 2, which is the casual gene for Q5.

Chromosome Mapping↗

Centralizing the non-central chi-square: A new method to correct for population stratification in genetic case-control association studies.

We present a new method, the delta-centralization (DC) method, to correct for population stratification (PS) in case-control association studies. DC works well even when there is a lot of confounding due to PS. The latter causes overdispersion in the usual chi-square statistics which then have non-central chi-square distributions. Other methods approach the noncentrality indirectly, but we deal with it directly, by estimating the non-centrality parameter tau itself. Specifically: (1) We define a quantity delta, a function of the relevant subpopulation parameters. We show that, for relatively large samples, delta exactly predicts the elevation of the false positive rate due to PS, when there is no true association between marker genotype and disease. (This quantity delta is quite different from Wright's F(ST) and can be large even when F(ST) is small.) (2) We show how to estimate delta, using a panel of unlinked "neutral" loci. (3) We then show that delta2 corresponds to tau the noncentrality parameter of the chi-square distribution. Thus, we can centralize the chi-square using our estimate of 6; this is the DC method. (4) We demonstrate, via computer simulations, that DC works well with as few as 25-30 unlinked markers, where the markers are chosen to have allele frequencies reasonably close (within +/- .1) to those at the test locus. (5) We compare DC with genomic control and show that where as the latter becomes overconservative when there is considerable confounding due to PS (i.e. when delta is large), DC performs well for all values of delta.

Bayes Theorem↗

Joint modeling of genetic association and population stratification using latent class models.

We show how latent class log-linear models can be used to test for an association between a candidate gene and a disease phenotype in a stratified population when the stratification is unobserved. The stratification may arise because of several ethnic groups or immigration and may lead to spurious associations between several loci and the disease. The information about the stratification is drawn from additional markers that are chosen to be independent of the disease and unlinked to the candidate gene and to each other within each population stratum. We use the EM algorithm to simultaneously estimate all the model parameters, including proportions of individuals in the latent population strata. The latent class model is used to test the phenotype association of single nucleotide polymorphism markers in four candidate regions in population-based case-control data selected from simulated Genetic Analysis Workshop (GAW) 12 population isolate 30. The analysis clearly demonstrates how the number of false positive associations can be reduced when the model accounts for population stratification.

Case-Control Studies↗

Population stratification in Northern shrimp (Pandalus borealis) off Iceland evident from RADseq analysis.

The northern shrimp Pandalus borealis (ice. St&#xf3;ri kampalampi) is a North Atlantic crustacean of significant commercial interest which has been harvested consistently in Icelandic waters since 1936. In Icelandic waters, the length at which this protandrous species transitions from male to female differs between the inshore and offshore populations, suggesting a biologically meaningful stratification which may or may not be plastic. Using reduced representative genomes assembled from RADseq data, sampled from 96 individuals collected at two time points (2018 and 2021), we compare the level of genetic structure across a gradient extending out of Skj&#xe1;lfandi bay, north Iceland. These data are compared to samples from a far offshore site, some 65&#xa0;km out from the bay, as well as another inshore fjord in Arnarfj&#xf6;r&#xf0;ur, in northwestern Iceland. Since 1999, no harvesting of inshore populations of P. borealis in Skj&#xe1;lfandi has been allowed due to stock decline, but harvesting of offshore stocks has continued. Uncertainty surrounding the extent of structure between the in- and offshore aggregations has remained. Here we report distinct genetic structure defining the inshore and offshore populations of northern shrimp, but find significant admixture between the two. Most importantly, we see that genetically inshore populations of northern shrimp extend far outside the harvest boundaries of inshore shrimp, and offshore individuals may exhibit punctuated migration into the inshore areas.

Animals↗

Secular change in height in Austria: an effect of population stratification?

The records of height on some 700,000 18-year-old Austrian males, called for examination as to their fitness for conscription, were analysed. The data covered 14 conscription years (1980-93), representing males born in the years 1962-75. The sample covered over 90% of the total male population of those cohorts in Austria. The data were analysed by birth year to show the secular trend in 18-year-old stature and its rate, which overall amounted to 0.53 cm/decade. Analysis by urban-rural residence showed that both participated in the trend, but that the urban-rural differences were appreciably less than the differences between the types of school the young men had attended. The rate of increase over the 14 years was less within each of the eight subgroups (urban-rural, four school categories). It is argued that the secular trend in height that has occurred is largely attributable to the change in social stratification, as evidenced by the changed proportion of subjects who attended schools of different types.

Adolescent↗