Search PubMedSearch

SEARCH · Search PubMed

Results for “Gene Imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

36 records · Page 2Linked to original sources

Genome-wide diversity of chromosomal inversions and their disease relationships.

Chromosomal inversions shape evolution and are implicated in human disease, yet their effects on genomic variation and health outcomes remain poorly understood. We analyze genome-wide human inversion polymorphisms, contrasting single-event and recurrent loci. Inversion recurrence is validated using structured-coalescent simulations. We show that single-event inversions evolve in near-complete isolation: inverted haplotypes show ~16-fold lower diversity and strong differentiation from direct haplotypes (median FST = 0.33). By contrast, recurrent inversions maintain gene flow, resulting in similar diversity across orientations and ~4-fold lower differentiation. We further find marked differences in coding sequence conservation between single-event and recurrent inversions. Using the NIH All of Us biobank, we impute inversions and identify four inversions with significant disease associations. Notably, the 17q21 inversion is associated with reduced risk of cognitive decline (OR=0.919) and breast cancer (OR=0.910) but with increased obesity risk (OR=1.097), consistent with pleiotropic selection. These findings establish inversions as major drivers of human genetic diversity and disease, with evolutionary outcomes critically dependent on recurrence.

Evolution

Genetic Determinants of Pulmonary Artery Size in over 50,000 Subjects with and without COPD.

RATIONALE: Pulmonary artery (PA) enlargement is a non-invasive imaging biomarker associated with pulmonary hypertension and mortality in COPD; however, its genetic determinants remain incompletely understood. OBJECTIVES: To characterize the genetic architecture of PA size across COPD-enriched and population-based cohorts. METHODS: We performed genome-wide association analyses of PA diameter using whole-genome sequencing in COPDGene (n=9,418) and ECLIPSE (n=1,859), and imputed-genotype data from the UK Biobank (n=37,073). We replicated lead variants in the Framingham Heart Study (FHS; n=3,289), incorporated all four studies into a joint meta-analysis, and identified independent signals through conditional analyses. Candidate effector genes were prioritized using coding variant annotation, colocalization, and integrative regulatory evidence. MEASUREMENTS AND MAIN RESULTS: We identified 44 independent genome-wide significant PA diameter signals within 39 loci, including 8 variants replicated in FHS, novel associations near FRMD4B, SLC20A2, BORCS7-ASMT, and KCNRG, and 5 signals in conditional analysis including multiple signals at ANO1. Genetic effects were concordant across imaging modalities and cohorts of differing COPD burden. Effector-gene prioritization nominated ABCC8, PDGFD, HMCN1, CCNE1, and TBX20, implicating pathways in vascular remodeling, developmental regulation, smooth muscle and endothelial function, ion-channel signaling, and extracellular matrix organization. Colocalization with pulse pressure GWAS demonstrated substantial shared causal variation between pulmonary and systemic vascular biology. CONCLUSIONS: In this largest genetic study of pulmonary vascular imaging to date, PA diameter exhibits a polygenic architecture consistent across imaging modalities and cohorts of differing COPD burden. The prioritized effector genes bridge rare-variant pulmonary hypertension biology with common-variant systemic vascular biology.

Pulmonary artery diameter

Reimagining research papers as interactive and reliable AI agents.

Here we introduce Paper2Agent, an automated framework that converts research papers into artificial intelligence (AI) agents. Paper2Agent transforms research output from passive artefacts into active systems that accelerate use and discovery. Conventional research papers require readers to understand and adapt the paper's code, data and methods to their work, creating barriers to dissemination and reuse. Paper2Agent addresses this challenge by converting a paper into an AI agent that functions as a virtual corresponding author, exposing its manuscript, supplementary materials, datasets, code and workflows as active, agent-native knowledge rather than static text. It analyses the paper and codebase using multiple agents to construct a model context protocol (MCP) server, then generates and runs tests to refine and increase robustness of the MCP. These paper MCPs can be connected to a chat agent (such as Claude Code) to carry out complex scientific queries through natural language while invoking tools and workflows from the paper. We demonstrate Paper2Agent's effectiveness through case studies. Paper2Agent created an agent that leveraged AlphaGenome1 to interpret genomic variants and agents based on Scanpy2 and TISSUE (transcript imputation with spatial single-cell uncertainty estimation)3 to conduct single-cell and spatial transcriptomics analyses. We validate that these agents reproduce the results of the original papers and carry out novel user queries. Paper2Agent created multiple agents that collaborate to prioritize a causal gene for psoriasis. By turning static papers into interactive AI agents, Paper2Agent introduces a paradigm for knowledge dissemination and a collaborative ecosystem of AI co-scientists.

Journal Article

Enhancing pan-cancer spatial transcriptomics at single-cell resolution with stPainter.

Subcellular spatial transcriptomics can resolve tissue architecture at cellular scale, but sparse gene panels and limited detection sensitivity constrain downstream analysis. Existing enhancement methods often require tissue-matched single-cell RNA sequencing (scRNA-seq) references and dataset-specific retraining. Here we show that stPainter, a conditional generative model pretrained on a pan-cancer scRNA-seq atlas, can enhance spatial transcriptomics data without matched references or retraining. Using a latent diffusion architecture guided by Stochastic Differential Equations (SDE), stPainter reconstructs expanded expression profiles from sparse measurements and produces latent representations for clustering and cell-state analysis. When we apply stPainter upon 6 spatial transcriptomics datasets of different cancer types, we demonstrate that our model empowers downstream biological analyses, including fine-grained subpopulation clustering and pathway enrichment. Comparison with spatially resolved proteomics (CODEX) provided independent support for regional agreement between imputed cellular compositions and protein-level tissue organization. These results establish stPainter as a scalable approach for analyzing tumor microenvironments without auxiliary sequencing data.

Spatial Transcriptomics

Multi-omics analysis identifies key genes and functional loci affecting teat number in American Large White and Landrace pigs and their application in optimizing genomic selection models.

BACKGROUND: Teat number is a crucial economic trait in pigs. It directly affects the ability of sows to lactate, which in turn influences the survival and health of piglets. The teat number of French Large White pigs is close to 16, while the teat number of American Large White and Landrace pigs is about 14. In order to improve the teat number of American Landrace and Large White pigs through molecular approaches and precise breeding techniques, we genotyped 2,131 American Landrace and 4,564 American Large White with teat number phenotype using a 50 K SNP chip. Then, the SNP-chip data was imputed to the level of whole-genome sequencing (iWGS). Based on iWGS data, we conducted GWAS to identify novel, significant SNPs associated with teat number and to incorporate them into genomic selection. RESULTS: In Landrace pigs, significant SNPs for TTN mapped to SSC2, SSC7, SSC8, and SSC14; the SSC8 and SSC14 effects are novel. LTN mapped to SSC7, RTN to SSC7 and SSC8. The lead SSC7 SNP explained 2.60% of TTN phenotypic variance. In Large White pigs, significant SNPs were detected on SSC7 and SSC10 for TTN; SSC7, SSC10, and SSC12 for LTN; and SSC7 and SSC10 for RTN. The most significant locus on SSC7 accounted for 2.99% of the phenotypic variance in TTN. Additionally, a multi-population meta-analysis detected significant novel SNPs for LTN on SSC1 and SSC8. By utilizing Bayesian fine mapping, the most precise QTL confidence interval on SSC7 for both TTN and RTN in Large White pigs was reduced to 40 kb. By integrating functional gene annotation with RNA-seq and ATAC-seq data from Erhualian and Bamaxiang pigs mammary placodes at embryonic day 26, we prioritized PTPN13, TRPV3, ZDHHC13, and BRD2 as novel candidate genes for teat number. We then incorporated the significant SNPs to GBLUP and benchmarked genomic-selection accuracy. In both breeds, fitting the top SNP as fixed maximized prediction for TTN and RTN, whereas treating all significant loci as an additional random effect optimized LTN. CONCLUSIONS: Our findings provide a theoretical basis for dissecting new key genes affecting teat number and for advancing molecular breeding of teat number in pigs.

Animals

Genetic architecture and analysis practices of circulating metabolites in the NHLBI Trans-Omics for Precision Medicine Program.

Circulating metabolite levels partly reflect the state of human health and diseases and can be impacted by genetic determinants. Hundreds of loci associated with circulating metabolites have been identified; however, most findings focus on predominantly European ancestry or single-study analyses. Leveraging the rich metabolomics resources generated by the National Heart, Lung, and Blood Institute (NHLBI) Trans-Omics for Precision Medicine (TOPMed) Program, we harmonized and accessibly cataloged 1,729 circulating metabolites among 25,058 ancestrally diverse samples. From our comparison of multiple methods, we provided a set of reasonable strategies for outlier and imputation handling to process metabolite data and show that inverse normalization by study and half-minimum imputation provide mostly similar results for pooled or meta-analysis. Following the practical analysis framework, we further performed a genome-wide association analysis on 1,135 selected metabolites using whole-genome sequencing data from 16,359 individuals passing the quality-control filters and discovered 1,775 independent loci associated with 667 metabolites. Among 160 unreported locus-metabolite pairs, we identified associations with loci locating within previously implicated metabolite-associated genes, as well as associations with loci locating in genes such as GAB3 and VSIG4 (located on the X chromosome) that may play a role in metabolic regulation. In the sex-stratified analysis, we revealed 85 independent locus-metabolite pairs with evidence of sexual dimorphism, which were located in well-known metabolic genes such as FADS2, D2HGDH, SUGP1, and UGT2B17, strongly supporting the importance of exploring sex difference in the human metabolome. Taken together, our study depicted the genetic contribution to circulating metabolite levels, providing additional insight into the understanding of human health.

Humans

Novel HLA class I and II insights into the pathogenesis of systemic sclerosis-associated interstitial lung disease.

OBJECTIVES: Systemic sclerosis-associated interstitial lung disease (SSc-ILD) is the leading cause of mortality in systemic sclerosis (SSc), yet its genetic architecture remains incompletely understood. Therefore, given the key role of the major histocompatibility complex (MHC) in SSc, we aimed to perform a comprehensive MHC-wide association study in the largest SSc-ILD cohort to date. METHODS: We analysed 2412 patients with SSc-ILD⁺, 3550 patients with SSc-ILD⁻, and 15,076 controls of European ancestry from 10 international cohorts. After quality control, the MHC region was imputed, and inverse variance weighted meta-analysis was performed. Subsequently, conditional stepwise analyses, adjustment for antitopoisomerase autoantibody (ATA) status, and functional annotation of significant single-nucleotide polymorphisms were performed. Finally, we constructed a composite score combining genetic, clinical, and demographic variables to predict SSc-ILD. RESULTS: After conditional analysis, we detected 12 significant associations within class I and class II human leukocyte antigen (HLA) genes. ATA adjustment reduced the significance of class II HLA variants, whereas class I HLA variants remained unaffected. Finally, the built composite score had an area under the curve of 0.754, significantly outperforming the models including any of the variables alone. CONCLUSIONS: In this study, we identify genetic mechanisms underlying SSc-ILD that support the potential implication of CD8+ T cells and ATAs in its pathogenesis. Moreover, we also demonstrate the enhanced efficacy of integrating genetic information into predictive models to detect patients at high risk of SSc-ILD. These findings provide new insights into disease pathogenesis and suggest potential biomarkers and therapeutic targets for improved patient management.

Humans

CYClones: a highly powered, fully genotyped, eight-parent yeast mapping population.

The budding yeast Saccharomyces cerevisiae is a remarkably adaptable organism that thrives in diverse environments. Global sequencing of natural isolates has revealed extensive genetic diversity within the species. Here, we describe the construction and characterization of CYClones (Collaborative Yeast Cross clones), a library of 11,392 segregants generated from a multiparent funnel cross of eight genetically diverse parental strains. To enable the genetic dissection of complex traits, we imputed whole-genome sequences for all segregants and show that CYClones captures a substantial fraction of the global genetic diversity of S. cerevisiae. Haplotype representation is well maintained, with each parental haplotype present at >5% frequency across >95% of the genome. Simulations demonstrate that CYClones has ≥95% power to detect variants with heritability as low as 0.36%, with mapping resolution often finer than the length of a single gene. In summary, CYClones is a powerful community resource for dissecting the genetic architecture of complex and quantitative traits, uncovering context-dependent mutational effects, and identifying causal variants underlying phenotypic diversity.

Saccharomyces cerevisiae

Adjustment for Genotype Imputation Uncertainty Corrects for Inflated Type I Error in Family-Based Association Testing.

Genotype imputation is a widely-used data augmentation approach that is applied to samples of related and/or unrelated individuals. Association testing may then be carried out on the complete data with commonly-used methods. This approach has typically not accounted for the mix of observed and imputed data, although recent work has noted the potential for introduction of confounding in case-control studies. In the Alzheimer's Disease Sequencing Project family sample we found severe inflation of the test statistics in logistic regression analysis following genotype imputation, even after standard covariate adjustments. Here we dissect sources of this inflation, which is driven by three factors: frequency-dependent bias in imputation-induced allele frequencies, differential measurement error, and differential genotyping rates in cases versus controls that introduces confounding. To address the problem, we propose a statistic, imputation deviance (), which can be easily computed from the observed and imputed genotype probabilities. We show that, as an additional fixed-effect covariate, controls the genome-wide inflation in analysis of this family-based sample, and we speculate that use of imputation deviance may also provide a practical approach to correct for genotype imputation effects in other settings, particularly when a data set is unbalanced and includes related individuals.

Humans

Genetic Analysis of Asymptomatic Antinuclear Antibody Production.

OBJECTIVE: Antinuclear antibodies (ANA) are detected in up to 14% of the population, and many individuals with ANA are asymptomatic. The literature on the genetic contribution to asymptomatic ANA positivity is limited. In this study, we aimed to perform a genome-wide association study of asymptomatic ANA positivity in multiple populations. METHODS: Asymptomatic individuals who were either ANA positive or ANA negative from the All of Us Research Program were included in this study, selecting those with an ANA test performed by immunofluorescence and no evidence of autoimmune disease. Imputation was performed, and a multipopulation meta-analysis including approximately 6 million single-nucleotide polymorphisms (SNPs) was conducted. Genome-wide SNP-based heritability was estimated using the Genome-wide Complex Trait Analysis&#xa0;software. A cumulative genetic risk score for lupus was constructed using previously reported genome-wide significant loci. RESULTS: A total of 1,955 asymptomatic ANA positive and 3,634 asymptomatic ANA negative individuals across three populations were included. The multipopulation meta-analysis revealed SNPs with a suggestive association (P <1 &#xd7; 10-5) across 8 different loci, but no genome-wide significant loci were identified. A gene variant upstream of HLA-DQB1, (rs17211748, P = 1.4 &#xd7; 10-6, odds ratio 0.82, 95% confidence interval 0.76-0.89), showed the most significant association. The heritability of asymptomatic ANA positivity was estimated to be 24.9%. Individuals who were asymptomatic and ANA positive did not exhibit increased cumulative genetic risk for lupus compared with individuals who were ANA negative. CONCLUSION: ANA production is not associated with significant genetic risk and is primarily determined by environmental factors.

Humans

Alterations in ether lipid metabolism in obesity revealed by systems genomics of multi-omics datasets.

Ratios between two metabolites are sensitive indicators of metabolic changes. Lipidomic profiling studies have revealed that plasma ether lipids, a class of glycero- and glycerophospho-lipids with reported health benefits, are negatively associated with obesity. Here, we utilized lipid ratios as surrogate markers of lipid metabolism to explore the processes underlying the inverse relationship between ether lipid metabolism and obesity. Plasma lipidomics data from two independent human cohorts (n&#x2009;=&#x2009;10,339 and n&#x2009;=&#x2009;4,492) were integrated to assess the associations between 82 lipid ratios and obesity-related markers in males and females. Results were externally validated using mouse transcriptomics data from the Hybrid Mouse Diversity Panel (n&#x2009;=&#x2009;152-227 across 74 strains). Genome-wide association studies using imputed genotypes from a population cohort (n&#x2009;=&#x2009;4,492) were performed to examine the genetic architecture of the ratios. Findings showed that waist circumference (WC), body mass index, and waist-hip ratio were inversely associated with total plasmalogens relative to total phospholipids in both sexes. Ratios comprising product-substrate pairs positioned either side of enzymes involved in plasmalogen synthesis and degradation showed positive and negative associations with WC, respectively. Branched-chain fatty acids negatively correlated with WC, while omega-6 polyunsaturated fatty acids exhibited differing associations depending on their position within the pathway. Mouse transcriptomics corroborated these results. Genomics data showed strong associations between ratios containing choline-plasmalogens and single-nucleotide polymorphisms in the transmembrane protein 229B (TMEM229B) gene region. This work demonstrates the utility of lipid ratios in understanding lipid metabolism. By applying the ratios to multi-omic datasets, we identified alterations in enzymatic activity and genetic variants likely affecting ether lipid synthesis in obesity that could not have been obtained from lipidomics data alone. Additionally, we characterized a potential role for TMEM229B, offering new perspectives on ether lipid metabolism and regulation.

Humans

Common genetic variants associated with urinary phthalate levels in children: A genome-wide study.

INTRODUCTION: Phthalates, or dieters of phthalic acid, are a ubiquitous type of plasticizer used in a variety of common consumer and industrial products. They act as endocrine disruptors and are associated with increased risk for several diseases. Once in the body, phthalates are metabolized through partially known mechanisms, involving phase I and phase II enzymes. OBJECTIVE: In this study we aimed to identify common single nucleotide polymorphisms (SNPs) and copy number variants (CNVs) associated with the metabolism of phthalate compounds in children through genome-wide association studies (GWAS). METHODS: The study used data from 1,044 children with European ancestry from the Human Early Life Exposome (HELIX) cohort. Ten phthalate metabolites were assessed in a two-void pooled urine collected at the mean age of 8&#xa0;years. Six ratios between secondary and primary phthalate metabolites were calculated. Genome-wide genotyping was done with the Infinium Global Screening Array (GSA) and imputation with the Haplotype Reference Consortium (HRC) panel. PennCNV was used to estimate copy number variants (CNVs) and CNVRanger to identify consensus regions. GWAS of SNPs and CNVs were conducted using PLINK and SNPassoc, respectively. Subsequently, functional annotation of suggestive SNPs (p-value&#xa0;<&#xa0;1E-05) was done with the FUMA web-tool. RESULTS: We identified four genome-wide significant (p-value&#xa0;<&#xa0;5E-08) loci at chromosome (chr) 3 (FECHP1 for oxo-MiNP_oh-MiNP ratio), chr6 (SLC17A1 for MECPP_MEHHP ratio), chr9 (RAPGEF1 for MBzP), and chr10 (CYP2C9 for MECPP_MEHHP ratio). Moreover, 115 additional loci were found at suggestive significance (p-value&#xa0;<&#xa0;1E-05). Two CNVs located at chr11 (MRGPRX1 for oh-MiNP and SLC35F2 for MEP) were also identified. Functional annotation pointed to genes involved in phase I and phase II detoxification, molecular transfer across membranes, and renal excretion. CONCLUSION: Through genome-wide screenings we identified known and novel loci implicated in phthalate metabolism in children. Genes annotated to these loci participate in detoxification, transmembrane transfer, and renal excretion.

Humans

Low-pass whole-genome sequencing reveals genomic diversity and ecotype-specific adaptation in indigenous Tigrayan chickens.

Indigenous chickens play a critical role in food security and climate resilience in smallholder systems, yet their genomic diversity and adaptive potential remain insufficiently characterised. This study employed low-pass whole-genome sequencing (LP-WGS; 0.2-1.99&#xd7;) to investigate genomic diversity, population structure, inbreeding and candidate environment-associated genomic variation in 33 chickens from highland, midland, and lowland agroecologies in the Tigray region of northern Ethiopia. After imputation and stringent filtering, 23.4 million high-confidence SNPs were retained, including&#x2009;~&#x2009;17% novel variants, indicating substantial uncharacterised genetic diversity in these populations. SNP density (13.8&#x2009;&#xb1;&#x2009;8.6 SNPs/kb) was comparable to values reported from high-coverage Ethiopian chicken datasets, demonstrating the suitability of LP-WGS for population genomics in resource-limited settings. Marked differences in genomic diversity were observed among ecotypes: midland chickens showed the highest nucleotide diversity (&#x3c0;&#x2009;=&#x2009;0.00267), followed by lowland (&#x3c0;&#x2009;=&#x2009;0.00233), whereas highland chickens showed the lowest diversity (&#x3c0;&#x2009;=&#x2009;0.00203) and elevated genomic inbreeding (FROH and FHOM &#x2248; 0.18). Population structure analyses revealed clear genetic separation among ecotypes. PCA (13.91% variation explained) distinguished lowland chickens along PC1 and separated highland from midland along PC2, while ADMIXTURE and FST patterns supported three major ancestral genomic backgrounds. Functional annotation of private missense variants uncovered distinct adaptive signatures reflecting the contrasting agroecological conditions. Highland chickens showed enrichment of candidate genes potentially involved in physiological processes relevant to high-altitude environments, including cold response, angiogenesis, cardiovascular regulation and metabolic homeostasis (eg., PARP1, ACOX2, ITGB3, EDNRB, SOX8, and SOX10). Midland chickens exhibited candidate signals of selection in genes with known roles in innate antiviral immunity, bacterial defence and inflammatory regulation (eg., BAK1, CLSTN1, CYSLTR1, CYSLTR2, CXCR7, GIPR, DSCAM, GDAP1, TLR3, TLR4, TLR7, IFIH1, ADORA1, EPHB1, and TMPRSS2). Lowland chickens displayed candidate variants associated with heat-stress response, DNA damage repair, oxidative balance and cardiovascular support under extreme temperatures (e.g., MLH1, BDKRB1, GPR19, FLT1, CCL18, TGM2, and RAMP3). Overall, the results indicate substantial genomic differentiation among ecotypes and suggest candidate environment-associated genetic divergence across Tigray's diverse agroecological zones. These populations may represent important reservoirs of adaptive genetic variation for climate-resilient poultry breeding, warranting further functional validation and conservation-oriented management.

Animals

Assessing data size requirements for training generalizable sequence-based TCR specificity models via pan-allelic MHC-I point-mutation ligandome evaluation.

Rapid identification of T cell receptors (TCRs) that specifically bind patient-unique neoepitopes is a critical challenge for personalized TCR-based therapies in oncology. Due to enormous diversity of both TCR and neoepitope repertoires, a machine learning predictor of TCR-pMHC specificity for personalized therapy must generalize to TCRs and epitopes not seen in the training data. We estimate the necessary size of such training data. We first confirm that published models fail to generalize beyond a single-residue dissimilarity to the epitope training set distribution. We then impute the point-mutation ligandome across the 34 most prevalent human MHC alleles and represent it as a graph based on our established dissimilarity cutoff. By finding the dominating set of this graph, we estimate that between one and 100 million epitopes are required to train a generalizable sequence-based TCR specificity prediction model-1000 times the size of current public data.

Humans

The Soifua Manuia reference panel with 2,570 Samoan haplotypes improves genotype imputation quality among Samoans.

Genotype imputation is fundamental to association studies, and yet even gold standard panels like TOPMed are limited in the populations for which they yield good imputation. Specifically, Pacific Islanders are poorly represented in extant panels. To address this, we used whole-genome sequencing from 1,285 Samoan individuals combined with 1000 Genomes Project (1KGP) individuals to construct an imputation reference panel that better represents Pacific Islander, specifically Samoan, genetic variation. Here we show that this panel yielded up to two times more well-imputed (r2&#x2009;&#x2265;&#x2009;0.80) variants than TOPMed-R3 and 1KGP and was enriched for moderate and high impact variants. There was improved imputation accuracy across the minor allele frequency (MAF) spectrum; accuracy (r2) was greater for population-specific variants (high fixation index, FST) and those from larger haplotypes (high LD score). However, the gain in accuracy over TOPMed-R3 was largest for small haplotypes, reflecting the Samoan panel's ability to capture variation not well tagged by other panels.

Haplotypes

Towards a Standard Threshold for Genome Wide Significance in Dogs.

Genome-wide association studies (GWAS) are a foundational step in tying phenotype to genotype, relying on statistical significance thresholds to distinguish true- from false-positive signals of association. Dog genomics has long relied on per-study Bonferroni thresholds of significance, basing these on SNP chip levels of markers (~100&#x2009;k to >&#x2009;14&#x2009;M variable sites). However, as the field progresses into whole genome imputation analyses and more powerful meta-analyses, there is a clear need to develop a standard significance threshold for common-variant GWAS. Using 1591 dogs from the broad-ancestry Dog10K dataset, we performed permutation analysis and developed GWAS thresholds for datasets using either 1% or 5% minor allele frequencies. The resultant p-values, 4.2&#x2009;&#xd7;&#x2009;10-7 and 5.0&#x2009;&#xd7;&#x2009;10-7 respectively, are similar to previous Bonferroni levels (p-value ~6&#x2009;&#xd7;&#x2009;10-7), but less restrictive than the standard human p-value, 5&#x2009;&#xd7;&#x2009;10-8, which is sometimes used in dog studies. Given the diverse haplotypes from the >&#x2009;320 breeds in the Dog10K input dataset, we suggest a p-value of 4&#x2009;&#xd7;&#x2009;10-7 as a standard significance threshold that could be applied to any dog GWAS.

Animals

Anthropometric and cardio-metabolic trait variation and genetic associations in sub-Saharan Africa.

The genetics of complex traits in Africa has been historically understudied, which can contribute to healthcare inequalities. Here, we present observations of 27 anthropometric, cardiovascular, and blood biomarker measurements across 2,124 individuals from sub-Saharan Africa for whom we also have dense genotype data. First, we identified trait values that differ significantly across populations and subsistence lifestyles (e.g., hemoglobin levels and height). We then identified traits with high degrees of sexual dimorphism (e.g., weight and grip strength). ADMIXTURE analyses revealed substantial population structure in our dataset, and many of the phenotypes studied here are correlated with genetic ancestry components, particularly skin color and body size traits. A variance partitioning approach further revealed traits in which much of the SNP heritability is due to polymorphisms that also contribute to differences between ancestry components. Following genomic imputation, we performed genome-wide association studies (GWASs) for all 27 traits and identified >100 independent autosomal SNPs with genome-wide significant associations for at least one trait (p < 5 &#xd7; 10-8). Many of these trait-associated variants are rare outside of Africa (minor-allele frequency [MAF] < 1%). We found that 100 kb windows surrounding the top GWAS hits from our African-ancestry cohort were enriched for trait associations in an identically sized European cohort and vice versa. We performed a more detailed analysis of height prediction from genetic data, finding that genome-wide admixture proportions predict height in Africans better than polygenic predictors based on large-scale European height GWASs.

Female

Genetics of Latin American Diversity Project: Insights into population genetics and association studies in admixed groups in the Americas.

Latin Americans are underrepresented in genetic studies, increasing disparities in personalized genomic medicine. Despite available genetic data from thousands of Latin Americans, accessing and navigating the bureaucratic hurdles for consent or access remains challenging. To address this, we introduce the Genetics of Latin American Diversity (GLAD) Project, compiling genome-wide information from 53,738 Latin Americans across 39 studies representing 46 geographical regions. Through GLAD, we identified heterogeneous ancestry composition and recent gene flow across the Americas. Additionally, we developed GLAD-match, a simulated annealing-based algorithm, to match the genetic background of external samples to our database, sharing summary statistics (i.e., allele and haplotype frequencies) without transferring individual-level genotypes. Finally, we demonstrate the potential of GLAD as a critical resource for evaluating statistical genetic software in the presence of admixture. By providing this resource, we promote genomic research in Latin Americans and contribute to the promises of personalized medicine to more people.

Humans