Search PubMedSearch

SEARCH · Search PubMed

Results for “population stratification”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Population stratification in Northern shrimp (Pandalus borealis) off Iceland evident from RADseq analysis.

The northern shrimp Pandalus borealis (ice. Stóri kampalampi) is a North Atlantic crustacean of significant commercial interest which has been harvested consistently in Icelandic waters since 1936. In Icelandic waters, the length at which this protandrous species transitions from male to female differs between the inshore and offshore populations, suggesting a biologically meaningful stratification which may or may not be plastic. Using reduced representative genomes assembled from RADseq data, sampled from 96 individuals collected at two time points (2018 and 2021), we compare the level of genetic structure across a gradient extending out of Skjálfandi bay, north Iceland. These data are compared to samples from a far offshore site, some 65 km out from the bay, as well as another inshore fjord in Arnarfjörður, in northwestern Iceland. Since 1999, no harvesting of inshore populations of P. borealis in Skjálfandi has been allowed due to stock decline, but harvesting of offshore stocks has continued. Uncertainty surrounding the extent of structure between the in- and offshore aggregations has remained. Here we report distinct genetic structure defining the inshore and offshore populations of northern shrimp, but find significant admixture between the two. Most importantly, we see that genetically inshore populations of northern shrimp extend far outside the harvest boundaries of inshore shrimp, and offshore individuals may exhibit punctuated migration into the inshore areas.

Animals

Sparse multitask group Lasso for genome-wide association studies.

A critical hurdle in Genome-Wide Association Studies (GWAS) involves population stratification, wherein differences in allele frequencies among subpopulations within samples are influenced by distinct ancestry. This stratification implies that risk variants may be distinct across populations with different allele frequencies. This study introduces Sparse Multitask Group Lasso (SMuGLasso) to tackle this challenge. SMuGLasso is based on MuGLasso, which formulates this problem using a multitask group lasso framework in which tasks are subpopulations, and groups are population-specific Linkage-Disequilibrium (LD)-groups of strongly correlated Single Nucleotide Polymorphisms (SNPs). The novelty in SMuGLasso is the incorporation of an additional [Formula: see text]-norm regularization for the selection of population-specific genetic variants. As MuGLasso, SMuGLasso uses a stability selection procedure to improve robustness and gap-safe screening rules for computational efficiency. We evaluate MuGLasso and SMuGLasso on simulated data sets as well as on a case-control breast cancer data set and a quantitative GWAS in Arabidopsis thaliana. We show that SMuGLasso is well suited to addressing linkage disequilibrium and population stratification in GWAS data, and show the superiority of SMuGLasso over MuGLasso in identifying population-specific SNPs. On real data, we confirm the relevance of the identified loci through pathway and network analysis, and observe that the findings of SMuGLasso are more consistent with the literature than those of MuGLasso. All in all, SMuGLasso is a promising tool for analyzing GWAS data and furthering our understanding of population-specific biological mechanisms.

Genome-Wide Association Study

Estimating population structure using epigenome-wide methylation data.

Population stratification is one of the source of inflation in epigenome-wide association studies (EWAS) when not properly accounted for. To address this, we developed methylation population scores (MPSs) to predict genetic principal components (GPCs) using a feature selection approach. We used multi-ethnic DNA methylation data from Illumina EPIC arrays across five cohorts, including MESA (n&#xa0;=&#xa0;929), CARDIA (n&#xa0;=&#xa0;1123), JHS (n&#xa0;=&#xa0;1365), ARIC (n&#xa0;=&#xa0;2338), and HCHS/SOL (n&#xa0;=&#xa0;1475), randomly splitting participants into training (85%) and test (15%) sets. Within each cohort, associations between GPCs and CpG sites were estimated using linear regression adjusting for age, sex, smoking and alcohol use, race/ethnicity, body mass index, and cell type proportions, followed by meta-analysis and selection of CpGs with FDR <0.05. We then applied a two-stage weighted least squares Lasso regression to construct MPSs, adjusting for the aforementioned covariates. In the test dataset, MPSs showed strong correlation with GPCs, with R&#xb2; ranging from 0.27 (MPS7 vs. GPC7) to 0.98 (MPS1 vs. GPC1). Visualization demonstrated that MPSs recapitulated the pattern shown by GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups and outperformed methylation-based principal components constructed using alternative published methods. Additionally, MPSs showed comparable performance to GPCs in reducing inflation in EWAS. Overall, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations, and provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent.

Humans

Effects of population structure on DNA fingerprint analysis in forensic science.

DNA fingerprints are used in forensic sciences to identify individuals. However, current analyses could underestimate the probability of two individuals sharing the same profile because the effect of population structure is not incorporated. An alternative analysis is proposed to take into account population stratification. The analysis uses studies of inbreeding in human populations to obtain an empirical upper bound on the magnitude of the effect.

Alleles

Estimating population structure using epigenome-wide methylation data.

INTRODUCTION: In epigenome-wide association analysis (EWAS), unaddressed population stratification often leads to inflation. We aimed to compute methylation population scores (MPSs) that predict genetic principal components (GPCs) using a feature selection and regression approach. METHODS: We used multi-ethnic methylation data (Illumina 450K/EPIC array) from unrelated MESA (n=929), CARDIA (n=1123), JHS (n=1365), ARIC (n=2338), and HCHS/SOL (n=1475) individuals, randomly assigning 85% of participants from each cohort to a training dataset and the remaining 15% to a test dataset. First, we estimated the associations of GPCs with each available CpG methylation site using linear regression within each cohort, adjusting for age, sex, smoking status, race/ethnic background (as a proxy for background information associated with lifestyle and other environmental exposures that may impact methylation), alcohol use status, body mass index, and cell type proportions. We meta-analyzed the associations across cohorts and selected CpG sites with association FDR-adjusted q-value <0.05. We next aggregated individuallevel data across the cohort-specific training datasets, and applied two-stage weighted least squares Lasso regression, with the GPCs as the outcomes and the selected CpG sites as penalized predictors, adjusting for the aforementioned covariates. The developed MPSs are the weighted sum of selected CpG sites from the Lasso. To evaluate the developed MPSs, we constructed them in the test dataset, and compared them with GPCs, and with MPSs constructed based on a previously-published paper. Comparison was based on correlation analysis and data visualization. We demonstrate the use of the MPSs in EWAS. RESULTS: In the test dataset, the MPSs were highly correlated with GPCs, with correlation decreasing, though not monotonically, for later components. Specifically, MPS1 and GPC1 had R2= 0.99, while MPS7 and GPC7 had R2=0.27 (the lowest observed correlation). In data visualization, MPSs had similar patterns as GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups, while outperforming MPC constructed using alternative published methods. MPSs showed comparable performance to GPCs in reducing some of the inflation in EWAS. CONCLUSIONS: Methylation-based population scores provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent. Unlike previous methods based on unsupervised methylation PCA, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations. The weights for each GPCs derived in our study can be applied to generate MPSs in other studies.

Journal Article

The expected polygenic risk score (ePRS) framework: an equitable metric for quantifying polygenetic risk via modeling of ancestral makeup.

Polygenic risk scores (PRSs) depend on genetic ancestry due to differences in allele frequencies between ancestral populations. This leads to implementation challenges in diverse populations. We propose a framework to calibrate PRS based on ancestral makeup. We define a metric called "expected PRS" (ePRS), the expected value of a PRS based on one's global or local admixture patterns. We further define the "residual PRS" (rPRS), measuring the deviation of the PRS from the ePRS. Simulation studies confirm that it suffices to adjust for ePRS to obtain nearly unbiased estimates of the PRS-outcome association without further adjusting for PCs. Using the TOPMed dataset, the estimated effect size of the rPRS adjusting for the ePRS is similar to the estimated effect of the PRS adjusting for genetic PCs. Similarly, we applied the ePRS framework to six cardiovascular-related traits in the All of Us dataset, and the results are consistent with those from the TOPMed analysis. The ePRS framework can protect from population stratification in association analysis and provide an equitable strategy to quantify genetic risk across diverse populations.

Journal Article

National genomic projects in Asia and Africa: a review.

National genome projects (NGPs) are increasingly shaping precision medicine by improving representation of population-specific genetic diversity. This review compiles findings from NGPs across Asia and Africa, regions that remain underrepresented in global genomic databases despite their extensive demographic and genetic diversity. A total of 53 studies from 24 countries were identified to understand (1) the genomic approach utilized, (2) novel findings that have emerged, and (3) strategies for improving research in these regions. The NGPs implement population-based variome databases (20 NGPs), linear reference genome assemblies (8 NGPs), and graph-based pangenome assemblies (1 NGP). Novel variants ranged between 0.28% (China) and 19.6% (Iran), whereas rare variants accounted for up to 88.9% of the detected variants in the Chinese population. Each NGP documents its country's evolutionary and migration history, which impacts disease frequency and pharmacogenomic variants. Clinically, NGPs revealed strong population stratification in disease-associated and pharmacogenomic variants. For example, the&#xa0;GJB2&#xa0;rs72474224 hearing-loss variant ranged from 13% in Vietnam and 12% in Hong Kong to 0.0894% in Turkey, while the&#xa0;VKORC1&#xa0;rs9923231 pharmacogenomic variant reached 89.2% in Taiwan but was 20%-25% in European-related Russian subpopulations. These findings demonstrate that clinically relevant allele frequencies, pathogenicity assessments, and drug-response markers differ substantially across ancestries. This review highlights ongoing efforts and strategies to enhance the representativeness of genomic data through NGPs in Asia and Africa. We also suggest future directions for national projects, including integrating family-based studies, multi-omic data, and standardized pipelines to accelerate discovery and support the equitable implementation of precision medicine.

Humans

Histocompatibility (HL-A) factors in familial multiple sclerosis. Is multiple sclerosis susceptibility inherited via the HL-A chromosome?

In order to study a possible hereditary factor leading to multiple sclerosis (MS) susceptibility, histocompatibility (HL-A) types were studied in families where two or more first-degree relatives had MS. Neither the inheritance of a particular parental HL-A chromosome, nor the occurrence of any specific HL-A antigens, could be shown to be necessary or sufficient for the development of MS in family members. The distribution of HL-A chromosomes was essentially the same for affected and unaffected family members. An excess of 3,7 haplotype and W21 antigen was demonstrated, both in affected patients and in unaffected family members, in equal proportions. We conclude that the HL-A chromosome has no direct causal relationship to MS susceptibility, although it may be indirectly associated by population stratification, maternal factors, or some other mechanism.

Adolescent

Population genetics of four hypervariable loci.

Populations of white Caucasians, Afro-Caribbeans and Asians residing within the UK have been analysed at 4 different hypervariable loci. A computerised system was used to store and to analyse the data. Simulation experiments were carried out in order to determine whether there was any evidence for population stratification, which would lead to non-independence of allelic distributions.

Alleles

Mining of important genetic loci and evaluation of genetic effects for growth traits in Baicheng You Chicken.

The Baicheng You Chicken is a precious indigenous breed in Xinjiang, China, prized for its strong disease and stress resistance and superior meat quality. However, the lack of scientific breeding and conservation has led to poor production performance, particularly in growth traits. In this study, we collected phenotypic and whole-genome resequencing data from 1,535 18-week-old Baicheng You Chickens (180 males and 1,355 females). After stringent quality control (SNP call rate > 95%, minor allele frequency > 1%), we constructed the breed's first comprehensive SNP-based genome-wide variation map, which comprised 2,020,743 high-quality SNPs across the genome. The filtered SNPs had high mapping quality (99.73% mapped to the bGalGal1.mat.broiler.GRCg7b reference genome, Q30 = 93.26%) and a reasonable Ti/Tv ratio (2.596), guaranteeing the reliability of subsequent analyses. We estimated genetic effects (SNP-based heritability and phenotypic variance explained (PVE) by individual loci) via the restricted maximum likelihood (REML) method, and performed a genome-wide association study (GWAS) using a mixed linear model (MLM) - with sex as a fixed effect and principal components to correct for population stratification - to identify significant loci and their effect sizes (Beta). All eight growth traits showed moderate to high heritability: body weight (BW) had the highest heritability (0.86&#xb1;0.11), while chest width (CW, 0.41&#xb1;0.08) and body slanting length (BSL, 0.43&#xb1;0.09) were the lowest; keel length (KL), chest girth (CG), pelvic width (PW), chest depth (CD) and shank length (SL) had heritabilities of 0.50&#xb1;0.09, 0.46&#xb1;0.09, 0.54&#xb1;0.09, 0.67&#xb1;0.10 and 0.74&#xb1;0.10, respectively. GWAS identified 145 significant SNPs, with a maximum Beta value of 0.39 and PVE ranging from 1.25% to 6.25%. We annotated 22 candidate genes, with TAPT1, IGF2BP1, ADGRB3, LDB2, NCAPG and LCORL as key candidates. These quantifiable genetic markers and effect estimates provide direct targets for marker-assisted selection (MAS) and valuable resources for future genomic selection (GS) programs, offering a practical approach to improve the breed's slow growth while preserving its unique meat quality.

Baicheng You Chicken

Differential gene expression study in whole blood identifies candidate genes for psychosis in African American individuals.

Genome-wide association has identified regions of the genome that mediate risk for psychosis. It is possible that variants in these regions confer risk by altering gene expression. This work has predominantly been conducted in individuals of European descent and has focused narrowly on schizophrenia rather than psychosis as a syndrome. In the present study we investigated alterations in gene expression in African American individuals with a range of psychotic diagnoses to increase understanding of the etiology in an underserved population. We performed RNA-seq in whole bloody to survey the transcriptome in 126 patients with a psychosis-spectrum disorder and 217 healthy controls and applied differential gene expression analyses across the genome while controlling for age, sex, population stratification and batch. We found 18 differentially expressed genes (DEGs), some of the locations of the corresponding genes overlap with previously implicated regions for psychosis, but many of which were novel associations. Enrichment analysis of nominally significant genes (p&#xa0;<&#xa0;0.05) revealed overrepresentation of biological processes relating to platelet, immune and cellular function, and sensory perception. Weighted gene co-expression network analysis, applied to identify modules of co-expressed genes associated with psychosis, revealed 10 modules, one of which was significantly associated with psychosis. This module was significantly enriched for DEGs, and for platelet function. These results support the potential role of immune function in the etiology of psychosis, identify novel candidate gene expression phenotypes that correspond to both established and new genomic regions, in individuals of African American ancestry.

Humans

Non-HLA region genes in insulin dependent diabetes mellitus.

The focus of this chapter is on the contribution of genes outside the HLA region to insulin dependent diabetes mellitus (IDDM) susceptibility. We review laboratory evidence for such genes from published studies and also present unpublished data from our recent research. The existence of genes predisposing to IDDM in the region of the insulin (INS) gene now appears established. Association analysis has demonstrated an increased frequency of class 1 alleles of the 5' INS polymorphism in diabetics compared with controls, and a new method of analysis (AFBAC) has shown that this association is not an artefact of population stratification. Interestingly, the effect of INS region susceptibility on IDDM cannot be detected by linkage analysis, suggesting that if a genetic marker locus is close to a disease susceptibility locus, association analysis may be a more sensitive method than linkage analysis for detecting the susceptibility locus. There is no convincing evidence that genes in the T cell receptor beta chain (TCRB) or alpha chain (TCRA) regions influence predisposition to IDDM, either directly, or indirectly through interaction with HLA region genes. However, we present new evidence for interaction between TCRB and immunoglobulin heavy chain (Gm) region genes in IDDM: diabetics who are positive for the IgG2 allotype G2m(23) have significantly different frequencies of a TCRB restriction fragment length polymorphism (RFLP) than those who are negative for the allotype. Gm region genes also appear to have indirect effects on IDDM susceptibility through interaction with HLA and INS region genes: DR3/4 and non-DR3/4 diabetics have significantly different frequencies of G2m(23), and INS1/1 and non-INS1/1 diabetics also have significantly different frequencies of this allotype. To our knowledge, there are no other studies of Gm-TCRB or Gm-INS interaction in IDDM susceptibility. Evidence for Gm-HLA interaction in IDDM has been published by several other groups of investigators, however the specific phenotypic interaction effects reported have differed. Nevertheless, pooled data from three studies of Gm/HLA haplotype segregation in affected sib pairs shows significantly increased sharing of Gm haplotypes in affected pairs who share both HLA haplotypes. The biological mechanisms underlying the direct (HLA, INS) and indirect (Gm-TCRB, Gm-HLA, Gm-INS) effects of these genetic regions on IDDM susceptibility remain to be elucidated.

Diabetes Mellitus, Type 1

Polygenic risk of coronary artery disease for long-term survivors of breast cancer.

BACKGROUND: Cardiovascular disease is a leading cause of death for long-term breast cancer survivors. We evaluated whether a polygenic risk score for coronary artery disease (CAD-PRS) was associated with the risk of incident CAD for survivors of unilateral or contralateral breast cancer. METHODS: The study included 1307 women with breast cancer first diagnosed at younger than 55&#x2009;years of age who participated in the Women's Environmental&#xa0;Cancer and Radiation Epidemiology Follow-up Study. The CAD-PRS was based on a PRS developed and validated in a separate population. We modeled the association between incident CAD and the CAD-PRS, adjusting for age, CAD risk factors, first (and second) breast cancer treatment, study recruitment phase, and genetic population stratification. We also explored whether the risk of CAD depended on interactions between the CAD-PRS and cardiotoxic cancer treatment. RESULTS: There were 66 incident CAD diagnoses reported at a median of 16&#x2009;years after breast cancer diagnosis. Participants with CAD-PRS&#x2009;at or above the median had a 2.48-times increased risk of CAD (95% confidence interval [CI]&#x2009;=&#x2009;1.44 to 4.29) relative to participants with CAD-PRS&#x2009;below the&#x2009;median. Anthracycline-based chemotherapy was associated with increased CAD risk (hazard ratio [HR]&#x2009;=&#x2009;2.04, 95% CI&#x2009;=&#x2009;1.04 to 3.98), and the association was not modified by the CAD-PRS. The association between incident CAD and left-sided radiation therapy (RT) was increased for those with CAD-PRS&#x2009;at or above the median (HR&#x2009;=&#x2009;2.90, 95% CI&#x2009;=&#x2009;1.26 to 6.68) but not for those with CAD-PRS&#x2009;below the median (HR&#x2009;=&#x2009;0.96, 95% CI&#x2009;=&#x2009;0.32 to 2.88). There was evidence of super-additive interaction between the CAD-PRS and left-sided RT (relative excess risk due to interaction&#x2009;=&#x2009;2.06, 95% CI&#x2009;=&#x2009;0.05 to 4.06). CONCLUSION: A genome-wide CAD-PRS was associated with nonfatal CAD risk for long-term breast cancer survivors, providing potential utility for personalized cardiovascular care, particularly after RT.

Humans

Two-sample Mendelian randomization study of gut microbiota and inflammatory proteins: Predictive, preventive, and personalized treatment for migraine.

The human gut microbiota is increasingly recognized as a significant factor in the pathogenesis of migraine, potentially via inflammatory pathways. Identifying specific human gut microbiota components associated with migraines, along with the investigation of particular inflammatory proteins, is essential for advancing primary prediction, targeted prevention, and personalized treatment strategies for migraines. We conducted a two-sample Mendelian randomization study using publicly available summary statistics from genome-wide association studies. Data for 473 human gut microbiota taxa were obtained from the Finnish national health survey conducted by the National Institute for Health and Welfare study (FINRISK, n = 5959 European participants). Genome-wide association study data (https://www.ebi.ac.uk/gwas/) for 91 circulating inflammatory proteins were obtained from 14,824 participants across 11 cohorts using the Olink Target 96 Inflammation panel. Migraine outcome data were obtained from the FinnGen R12 release, with cases defined using ICD-10 code G43. All genome-wide association study analyses were adjusted for sex, age, genotyping batch, and 10 genetic principal components to control population stratification (genomic inflation factors: 1.00&#x2013;1.05). Inverse variance-weighted Mendelian randomization was the primary analysis method, with Mendelian randomization-Egger, weighted median, and mode-based methods as sensitivity analyses. Two-step Mendelian randomization mediation analysis quantified the proportion of the effects of human gut microbiota on migraine that are mediated through inflammatory proteins. Thirty-seven bacterial genera were found to be associated with migraine using the inverse variance-weighted method. Of these, 18 genera exhibited a negative association, while 19 genera demonstrated a positive association with migraine risk. Additionally, eight inflammatory proteins were found to increase the risk of migraine. Among human gut microbiota, four were observed to reduce inflammatory protein levels, whereas another four were associated with increased inflammatory protein levels. Additionally, five gut microbiota were identified to influence migraine through inflammatory proteins in both Mendelian randomization analyses. Specifically, Actinobacteria, Brachyspiraceae, CAG-269 sp001915995, and Paraglaciecola were found to affect migraine outcomes via inflammatory proteins, with mediation proportions of 12%, 19%, 15.5%, and 6.7%, respectively. Lawsonibacter sp002161175 was identified to influence migraine risk through Oncostatin-M and SLAM, with mediation proportions of 15.6% and 11.3%, respectively. Our study elucidated the role of specific human gut microbiota alterations in the pathogenesis of migraine and highlighted the mediating effects of inflammatory proteins. Targeting these particular human gut microbiota alterations offers a promising strategy for predictive, preventive, and personalized medicine in migraine management, resulting in substantial clinical advancements.

causality

Complex structural variation, phylogeny, and disease associations of the mucin pangenome.

Mucins are large glycoproteins that provide hydration and barrier function to epithelial tissues. Although genetically heterogeneous, all mucins harbor a large exon composed of variable number tandem repeats (VNTRs). Short-read sequencing has limited our understanding of mucin VNTR diversity and makes disease association studies challenging. We leverage 296 long-read phased genome assemblies to characterize 14 mucin family members, achieving &#x2265;97% accuracy across 572 haplotypes. Phylogenetic haplogroup analysis reveals extraordinary structural heterozygosity, with MUC4 harboring the greatest allelic diversity (n=240 distinct lengths) and MUC12 the greatest size range (&#x394; = 55,233 bp; 23,080 amino acids). Ten mucins show significant population stratification (pFDR < 0.05). At the MUC4/MUC20 locus, we characterize higher-order structural variation, including a recurrent inversion, copy number variation, and interlocus gene conversion. Optimized genotyping achieves &#x2265;95% haplogroup concordance across 10 loci. We apply this to 4,637 deeply phenotyped cystic fibrosis patients and identify a significant association between short MUC1 VNTRs and severe disease (p=0.0056), demonstrating the pangenome's utility for complex locus genotyping and disease discovery.

Journal Article

Seasonality of birth in schizophrenia. An insufficient stratification of control population?

The seasonality of birth in 142 schizophrenics from Barcelona (Spain) has been studied, showing a pattern similar to that described by other authors, with a maximum in winter. If sample is compared with the local general population there are significant differences, but these disappear when the comparison is made with a control population properly stratified by age, sex, place of birth, place of residence and social status. The same results are obtained using statistical methods suggested by other authors. The seasonality pattern usually obtained in schizophrenia could be due to a selective stratification of the ill population based on variables that might or might not be related to illness.

Adolescent