Search PubMedSearch

SEARCH · Search PubMed

Results for “Gene Imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

CoxFormer enables spatial omics inference with multimodal generative modeling.

Gene co-expression maps transcriptome-wide gene-gene relationships, yet high-quality estimates cover less than half the genome. Meanwhile, spatial omics either profiles restricted in situ panels or lacks cellular resolution. Extending co-expression transcriptome-wide could overcome these limitations by inferring unassayed gene expression at subcellular resolution. Here we show that CoxFormer integrates literature-derived gene knowledge with co-expression networks from bulk tissues and large-scale single-cell atlases to learn 512-dimensional representations for 32,016 human genes. These embeddings capture functional gene relationships and serve as a generative prior for spatial inference across platforms and modalities. Without requiring a matched single-cell RNA-sequencing reference, CoxFormer supports four applications beyond measured genes: histology-based expression imputation, gene activity prediction from chromatin accessibility, subcellular super-resolution inference, and pathological region detection. Together, CoxFormer extends gene embedding from gene- and cell-level tasks to whole-transcriptome spatial inference, providing a unified framework for biological analysis beyond the limited gene coverage of current spatial omics technologies.

Humans

A novel reusable transcriptome-wide association study workflow used to map key genes linked to important cattle traits.

Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.

Animals

LungGENIE: the lung gene-expression and network imputation engine.

BACKGROUND: Few cohorts have study populations large enough to conduct molecular analysis of ex vivo lung tissue for genomic analyses. Transcriptome imputation is a non-invasive alternative with many potential applications. We present a novel transcriptome-imputation method called the Lung Gene Expression and Network Imputation Engine (LungGENIE) that uses principal components from blood gene-expression levels in a linear regression model to predict lung tissue-specific gene-expression. METHODS: We use paired blood and lung RNA sequencing data from the Genotype-Tissue Expression (GTEx) project to train LungGENIE models. We replicate model performance in a unique dataset, where we generated RNA sequencing data from paired lung and blood samples available through the SUNY Upstate Biorepository (SUBR). We further demonstrate proof-of-concept application of LungGENIE models in an independent blood RNA sequencing data from the Genetic Epidemiology of COPD (COPDGene) study. RESULTS: We show that LungGENIE prediction accuracies have higher correlation to measured lung tissue expression compared to existing cis-expression quantitative trait loci-based methods (median Pearson's r = 0.25, IQR 0.19-0.32), with close to half of the reliably predicted transcripts being replicated in the testing dataset. Finally, we demonstrate significant correlation of differential expression results in chronic obstructive pulmonary disease (COPD) from imputed lung tissue gene-expression and differential expression results experimentally determined from lung tissue. CONCLUSION: Our results demonstrate that LungGENIE provides complementary results to existing expression quantitative trait loci-based methods and outperforms direct blood to lung results across internal cross-validation, external replication, and proof-of-concept in an independent dataset. Taken together, we establish LungGENIE as a tool with many potential applications in the study of lung diseases.

Humans

SIGEL: a context-aware genomic representation learning framework for spatial genomics analysis.

Spatial transcriptomics (ST) integrates spatial information into genomics, yet methods for generating spatially-informed gene representations are limited and computationally intensive. We present SIGEL, a cost-effective framework that derives gene manifolds from ST data by exploiting spatial genomic context. The resulting SIGEL-generated gene representations (SGRs) are context-aware, biologically meaningful, and robust across samples, making them highly effective for key downstream tasks, including imputing missing genes, detecting spatial expression patterns, identifying disease-related genes and interactions, and improving spatial clustering. Extensive experiments across diverse ST datasets validate SIGEL's effectiveness and highlight its potential in advancing spatial genomics research.

Genomics

Predicted brain-regional gene expression patterns in individuals living with Alzheimer's disease.

Studying brain gene expression in Alzheimer's Disease (AD) remains difficult as postmortem brain is difficult to access, cannot be used to guide donor treatment, may be confounded by environmental factors before and after death, and is difficult to link to early AD states or disease progression. To circumvent these limitations, several studies have tested blood transcriptome biomarkers for AD. However, gene-expression levels in the blood have limited correlation with those in the brain. To evaluate the potential of monitoring Alzheimer's progression with peripheral data, we used transcriptome-imputation to identify brain-region-specific AD-associated gene-expression differences in cohorts with blood-based transcriptome data. This approach provides a high-resolution image of AD-associated molecular differences in the brains of individuals actively living with disease. We analyzed eight AD studies (777 AD cases, 779 cognitively unimpaired controls), imputing transcriptomes in 10 brain regions via the Brain Gene Expression and Network Imputation Engine (BrainGENIE). Hundreds of differentially expressed genes (DEGs) associated with AD were identified in nine brain regions, with anterior cingulate cortex and amygdala showing the most differential expression. AD-associated genes were enriched in pathways such as proteostasis, mitochondrial dysfunction, and immune activation. We observed significant yet moderate concordance between imputed AD-associated changes and those directly measured in the dorsolateral prefrontal cortex and cerebellum. These transcriptomic changes can guide future in vitro studies focused on pathogenesis or be targets of novel therapeutic development. In conclusion, we demonstrated the scope and utility of brain expression imputation from the peripheral transcriptome, laying the groundwork for biomarker discovery and prospective AD studies.

Alzheimer Disease

Integrative cross-tissue transcriptome-wide association and metabolomic analysis reveals novel genetic risk loci for aortic aneurysm.

BACKGROUND: Aortic aneurysm (AA) is a life-threatening cardiovascular condition with a strong genetic component, however, its molecular mechanisms remain poorly understood. Although genome-wide association studies (GWAS) have identified numerous risk loci, most prior studies have investigated genetic and metabolic factors separately, leaving the causal pathways from genetic variants to disease largely unexplored. METHODS: We established an integrative framework combining cross-tissue transcriptome-wide association studies (TWAS) with metabolomic mediation analysis. First, we integrated GWAS data from FinnGen R12 with multi-tissue expression quantitative trait loci (eQTL) data from Genotype-Tissue Expression Project (GTEx) V8, then performed cross-tissue TWAS using the Unified Test for MOlecular SignaTures (UTMOST) and single-tissue validation with the Functional Summary-based Imputation (FUSION) to prioritize susceptibility genes. Second, we applied Mendelian randomization (MR), colocalization, and Fine-mapping Of CaUsal gene Sets (FOCUS) to assess causality and identify high-confidence genes. Third, we performed metabolite mediation analysis to uncover metabolic pathways linking genetic variants to disease risk. Finally, we validated key findings in mouse models of thoracic aortic aneurysm (TAA) and abdominal aortic aneurysm (AAA) using Quantitative Real-Time Reverse Transcription Polymerase Chain Reaction (RT-qPCR) and Western blotting. RESULTS: We identified multiple novel susceptibility genes for AA and its subtypes. Key genes included ADH family members (ADH1A, ADH1B, ADH4, ADH6) and ZNF827, which showed cross-subtype associations with strong colocalization evidence in vascular tissues. Metabolite mediation analysis revealed significant pathways involving N-acetylphenylalanine and methionine sulfoxide. Functional enrichment revealed distinct biological mechanisms: AA and AAA were primarily associated with metabolic pathways, whereas TAA-related genes were enriched in developmental and contractile processes. PheWAS indicated no significant off-target associations. Critically, experimental validation in mouse models confirmed significant upregulation of ZNF827 in TAA and ADH6 in AAA at both mRNA and protein levels, corroborating the genetic predictions. CONCLUSION: This integrated cross-omics analysis identifies novel genetic loci and, crucially, uncovers specific nutrient-related metabolic pathways that mediate genetic risk. These findings provide a mechanistic basis for future nutritional and metabolic intervention studies in AA and its subtypes.

MAGMA

Revealing the Shared Genetic Architecture of Metabolic Dysfunction-Associated Steatotic Liver Disease-Related Traits Through Genomic Structural Equation Modeling.

Although individual traits related to metabolic dysfunction-associated steatotic liver disease (MASLD) have been investigated through large-scale genome-wide association studies (GWASs), the shared genetic susceptibility across these traits remains unclear. We therefore conducted a multivariate GWAS of key MASLD-related traits to elucidate their common genetic architecture. We applied genomic structural equation modeling to model a latent genetic factor (MASLD-F) underlying genetically correlated MASLD-related traits, leveraging their GWAS-derived genetic correlations. We then performed functional annotations, including fine-mapping, transcriptome-wide association study, and cell- and tissue-type-specific enrichment analyses, and conducted Mendelian randomization analyses to identify modifiable risk factors. Our multivariate MASLD-F GWAS identified 50 independent variants across 48 genomic loci. Transcriptomic imputation identified several MASLD-F-associated genes, including ARNTL, NPC1, BTBD10, VDAC2, TSKU, SFMBT1, and ABHD17C. We observed significant enrichment of MASLD-F-related genetic signals predominantly in brain tissues, pancreatic islets, and the adrenal gland. Additionally, six modifiable risk factors and four modifiable protective factors for MASLD-F were identified. These findings reveal a complex shared genetic architecture underlying MASLD components, thereby expanding our understanding of disease pathogenesis and providing novel insights for precision medicine and public health interventions.

Humans

Deep generative neural network for accurate drug response imputation.

Drug response differs substantially in cancer patients due to inter- and intra-tumor heterogeneity. Particularly, transcriptome context, especially tumor microenvironment, has been shown playing a significant role in shaping the actual treatment outcome. In this study, we develop a deep variational autoencoder (VAE) model to compress thousands of genes into latent vectors in a low-dimensional space. We then demonstrate that these encoded vectors could accurately impute drug response, outperform standard signature-gene based approaches, and appropriately control the overfitting problem. We apply rigorous quality assessment and validation, including assessing the impact of cell line lineage, cross-validation, cross-panel evaluation, and application in independent clinical data sets, to warrant the accuracy of the imputed drug response in both cell lines and cancer samples. Specifically, the expression-regulated component (EReX) of the observed drug response achieves high correlation across panels. Using the well-trained models, we impute drug response of The Cancer Genome Atlas data and investigate the features and signatures associated with the imputed drug response, including cell line origins, somatic mutations and tumor mutation burdens, tumor microenvironment, and confounding factors. In summary, our deep learning method and the results are useful for the study of signatures and markers of drug response.

Antineoplastic Agents

New Genetic Loci Implicated in Cardiac Morphology and Function Using Three-Dimensional Population Phenotyping.

BACKGROUND: Cardiac remodeling occurs in the mature heart and is a cascade of adaptations in response to stress, which are primed in early life. A key question remains as to the processes that regulate the geometry and motion of the heart and how it adapts to stress. METHODS: We performed spatially resolved phenotyping using machine learning-based analysis of cardiac magnetic resonance imaging in 47 549 UK Biobank participants. We analyzed 16 left ventricular spatial phenotypes, including regional myocardial wall thickness and systolic strain in both circumferential and radial directions. In up to 40 058 participants, genetic associations across the allele frequency spectrum were assessed using genome-wide association studies with imputed genotype participants, and exome-wide association studies and gene-based burden tests using whole-exome sequencing data. We integrated transcriptomic data from the GTEx project and used pathway enrichment analyses to further interpret the biological relevance of identified loci. To investigate causal relationships, we conducted Mendelian randomization analyses to evaluate the effects of blood pressure on regional cardiac traits and the effects of these traits on cardiomyopathy risk. RESULTS: We found 42 loci associated with cardiac structure and contractility, many of which reveal patterns of spatial organization in the heart. Whole-exome sequencing revealed 3 additional variants not captured by the genome-wide association study, including a missense variant in CSRP3 (minor allele frequency 0.5%). The majority of newly discovered loci are found in cardiomyopathy-associated genes, suggesting that they regulate spatially distinct patterns of remodeling in the left ventricle in an adult population. Our causal analysis also found regional modulation of blood pressure on cardiac wall thickness and strain. CONCLUSIONS: These findings provide a comprehensive description of the pathways that orchestrate heart development and cardiac remodeling. These data highlight the role that cardiomyopathy-associated genes have on the regulation of spatial adaptations in those without known disease.

Humans

Rare variant analyses in 51,256 type 2 diabetes cases and 370,487 controls reveal the pathogenicity spectrum of monogenic diabetes genes.

Type 2 diabetes (T2D) genome-wide association studies (GWASs) often overlook rare variants as a result of previous imputation panels' limitations and scarce whole-genome sequencing (WGS) data. We used TOPMed imputation and WGS to conduct the largest T2D GWAS meta-analysis involving 51,256 cases of T2D and 370,487 controls, targeting variants with a minor allele frequency as low as 5 × 10-5. We identified 12 new variants, including a rare African/African American-enriched enhancer variant near the LEP gene (rs147287548), associated with fourfold increased T2D risk. We also identified a rare missense variant in HNF4A (p.Arg114Trp), associated with eightfold increased T2D risk, previously reported in maturity-onset diabetes of the young with reduced penetrance, but observed here in a T2D GWAS. We further leveraged these data to analyze 1,634 ClinVar variants in 22 genes related to monogenic diabetes, identifying two additional rare variants in HNF1A and GCK associated with fivefold and eightfold increased T2D risk, respectively, the effects of which were modified by the individual's polygenic risk score. For 21% of the variants with conflicting interpretations or uncertain significance in ClinVar, we provided support of being benign based on their lack of association with T2D. Our work provides a framework for using rare variant GWASs to identify large-effect variants and assess variant pathogenicity in monogenic diabetes genes.

Diabetes Mellitus, Type 2

MetaGLIMPSE: Meta-imputation of low-coverage sequencing data for modern and ancient genomes.

The advent of efficient and accurate imputation for low-coverage sequencing offers an unbiased alternative to SNP array imputation, increasing the accuracy of rare variant imputation across all populations. Since imputation accuracy generally increases with larger reference panels and closer ancestry match between target and reference samples, leveraging imputation from multiple reference panels improves imputation accuracy; however, individual reference panel genotypes are often privacy protected. Meta-imputation bypasses individual-level data by combining single-panel imputed genotypes through estimating panel- and marker-specific weights. We present a meta-imputation method, MetaGLIMPSE, that combines estimates from multiple reference panels for low-coverage sequencing imputation. Across all our scenarios, for both modern and ancient DNA samples, MetaGLIMPSE consistently outperforms the best single-panel imputation for coverages of 0.1×-8× and across all minor-allele frequencies, equaling the combined panel imputation for some parameters. Finally, MetaGLIMPSE is computationally efficient, meta-imputing 500 whole genomes in 16% of the time of GLIMPSE2.

Humans

Penalised regression improves imputation of cell-type specific expression using RNA-seq data from mixed cell populations compared to domain-specific methods.

Gene expression studies often use bulk RNA sequencing of mixed cell populations because single cell or sorted cell sequencing may be prohibitively expensive. However, mixed cell studies may miss expression patterns that are restricted to specific cell populations. Computational deconvolution can be used to estimate cell fractions from bulk expression data and infer average cell-type expression in a set of samples (e.g., cases or controls), but imputing sample-level cell-type expression is required for more detailed analyses, such as relating expression to quantitative traits, and is less commonly addressed. Here, we assessed the accuracy of imputing sample-level cell-type expression using a real dataset where mixed peripheral blood mononuclear cells (PBMC) and sorted (CD4, CD8, CD14, CD19) RNA sequencing data were generated from the same subjects (N=158), and pseudobulk datasets synthesised from eQTLgen single cell RNA-seq data. We compared three domain-specific methods, CIBERSORTx, bMIND and debCAM/swCAM, and two cross-domain machine learning methods, multiple response LASSO and ridge, that had not been used for this task before. We also assessed the methods according to their ability to recover differential gene expression (DGE) results. LASSO/ridge showed higher sensitivity but lower specificity for recovering DGE signals seen in observed data compared to deconvolution methods, although LASSO/ridge had higher area under curves than deconvolution methods. Machine learning methods have the potential to outperform domain-specific methods when suitable training data are available.

Humans

Genetic Architecture of Idiopathic Inflammatory Myopathies From Meta-Analyses.

OBJECTIVE: Idiopathic inflammatory myopathies (IIMs, myositis) are rare systemic autoimmune disorders that lead to muscle inflammation, weakness, and extramuscular manifestations, with a strong genetic component influencing disease development and progression. Previous genome-wide association studies identified loci associated with IIMs. In this study, we imputed data from two prior genome-wide myositis studies and analyzed the largest myositis data set to date to identify novel risk loci and susceptibility genes associated with IIMs and its clinical subtypes. METHODS: We performed association analyses on 14,903 individuals (3,206 patients and 11,697 controls) with genotypes and imputed data from the Trans-Omics for Precision Medicine reference panel. Fine-mapping and expression quantitative trait locus colocalization analyses in myositis-relevant tissues indicated potential causal variants. Functional annotation and network analyses using the random walk with restart (RWR) algorithm explored underlying genetic networks and drug repurposing opportunities. RESULTS: Our analyses identified novel risk loci and susceptibility genes, such as FCRLA, NFKB1, IRF4, DCAKD, and ATXN2 in overall IIMs; NEMP2 in polymyositis; ACBC11 in dermatomyositis; and PSD3 in myositis with anti-histidyl-transfer RNA synthetase autoantibodies (anti-Jo-1). We also characterized effects of HLA region variants and the role of C4. Colocalization analyses suggested putative causal variants in DCAKD in skin and muscle, HCP5 in lung, and IRF4 in Epstein-Barr virus (EBV)-transformed lymphocytes, lung, and whole blood. RWR further prioritized additional candidate genes, including APP, CD74, CIITA, NR1H4, and TXNIP, for future investigation. CONCLUSION: Our study uncovers novel genetic regions contributing to IIMs, advancing our understanding of myositis pathogenesis and offering new insights for future research.

Humans

Selphi, a tool for improving genotype imputation accuracy.

Genotype imputation is a powerful tool for inferring missing genotype data in large-scale genetic studies. Over the last two decades, multiple imputation algorithms have been developed, steadily improving in speed and overall accuracy. However, accurate imputation of rare and infrequent variants remains a challenge, largely because existing methods rely on local haplotype matching within genomic windows and do not fully exploit the extended patterns of haplotype sharing that span entire chromosomes. Here we present Selphi, a new genotype imputation algorithm that combines the Positional Burrows-Wheeler Transform (PBWT) with a multi-stage haplotype selection heuristic operating across entire chromosomes. When compared to state-of-the-art methods Beagle 5.4, IMPUTE5, and Minimac4, Selphi showed higher accuracy on the 1000 Genomes Project and TOPMed datasets, across all super-populations and allele frequencies. Similarly, Selphi achieved higher accuracy than Beagle 5.4 on the UK Biobank dataset, which translated into improved concordance with hc-WGS GWAS summary statistics at known trait-associated loci and more accurate polygenic risk scores (PRS). Selphi outputs standard VCF files with genotype dosages (DS), haplotype-specific allele probabilities (AP1, AP2), and a per-variant dosage R-squared quality score (DR2), enabling direct integration with downstream analytical pipelines including standard post-imputation quality filtering.

Genome-Wide Association Study

The cellular associates of late life changes in white matter microstructure.

The microstructural architecture of white matter supporting information flow across local circuits and large-scale networks changes throughout the lifespan. However, the genetic and cellular factors underlying age-related variations in white matter microstructure have yet to be established. Here, we examined the genetic associates of individual differences in diffusion-based measures of white matter in a population-based cohort (N=29,862) from the UK Biobank. Estimates of heritability from Genome-Wide Association Study (GWAS) data revealed that genetic factors are linked to population variability in 96.1% of 432 tract microstructural measures. The presence of shared genetic influences was observed to be greater within, relative to between, broad tract classes (commissural, association, projection, and complex cerebellar). Age associations with microstructural changes were estimated across diffusivity measures, with association class tracts showing the greatest vulnerability to age-related decline in older adults. Analyses of imputed cellular associates of age-related changes in white matter revealed a preferential relationship with cell gene markers of oligodendrocytes and other glial cell types, with sparse relationships observed for inhibitory and excitatory cells. These data indicate that white matter tract microstructure is shaped by genetic factors and suggest a role for glial cell-related transcripts in late-life changes in the structural wiring properties of the human brain.

Aging

A model of assortative mating with partial dominance.

A model of assortative mating incorporating partial dominance is proposed for a single locus with two alleles. It is derived by starting from an arbitrary genotypic distribution and finding symmetric and non-selective mating frequencies which duplicate this distribution. Numerical values are imputed to genotypes, the homozygotes having numerically equal values, opposite in sign, and the heterozygote having a value determined by the gene and heterozygote frequencies. The model is specified in a canonical form which reveals the correlation between mates based on genotypic values, and relates the correlation to the fixation index. It permits negative as well as positive values of the fixation index. It is shown that this general model includes several particular cases, in equilibrium phase, occurring in the literature.

Alleles

Genome-wide association studies for feed efficiency, production and feeding behavior traits in Canadian purebred Duroc pigs.

This study aimed to identify potential genetic variants and candidate genes associated with feed efficiency (FE), production, and feeding behavior traits in Canadian purebred Duroc pigs. Genome-wide association studies (GWAS) were conducted using 8,861 individuals and an imputed Affymetrix PigGen Canada 50K panel v2.0 using a linear mixed model (LMM) and a Bayesian B model. This analysis used an adjusted P-value threshold (ranging from 6.6 × 10-5 to 1.3 × 10-4) using a false-discovery rate to determine significance. The number of significant SNPs identified for each trait was as follows: average daily gain (ADG, 48), daily feed intake (DFI, 85), feed conversion ratio (FCR, 101), residual feed intake (RFI, 37), residual gain (RG, 64), residual intake and gain (RIG, 55), backfat thickness (BF, 100), loin depth (LD, 6), Kleiber's ratio (KR, 0), total time spent eating per day (TPD, 7), and number of visits to the feeder per day (NVD, 6). Several traits (BF, DFI, FCR, RFI, RG, and RIG) showed strong overlapping signals on chromosomes 7 and 10 with 24 shared significant SNPs, indicating potential shared genetic mechanisms. These traits also had 71 overlapping candidate genes, such as PACSIN1, PTCH1, ADIPOR1, and ITPR3, associated with glucose, lipid, and cholesterol metabolism. Well-known candidate genes in literature associated with growth and fatness such as MC4R and CDH20 were also identified to be associated with ADG, BF, FCR, and DFI in this study. Gene ontology enrichment analysis revealed that a set of the candidate genes were involved in the gonadotropin-releasing hormone (GnRH) and the platelet-derived growth factor (PDGF) signaling pathways. Overall, this study contributed to understanding the genetic architecture and provided a biological foundation for improving FE, production, and feeding behavior traits in Canadian Duroc pigs, facilitating the selection of more efficient pigs.

Sus scrofa

Distinct Effects of Complement C4A and C4B Copy Numbers in Systemic Sclerosis Serological and Clinical Subtypes.

OBJECTIVE: Complement component 4 (C4), encoded by C4A and C4B within the major histocompatibility complex (MHC) on chromosome 6, regulates the immune response and clears immune complexes. The variable copy number (CN) of C4 genes and retroviral human endogenous retrovirus K (HERV-K) element influence its function. Given the relationship of C4 CN with systemic sclerosis (SSc) risk, we assessed associations with SSc clinical and serologic subtypes. METHODS: We compared imputed C4 CNs across SSc subgroups (4,049 anticentromere positive [ACA+]; 2,200 anti-topoisomerase I [ATA+]; 577 anti-RNA polymerase [ARA+]; 1,078 triple-negative [TN] patients; 6,295 limited cutaneous SSc [lcSSc]; and 2,946 diffuse cutaneous SSc [dcSSc]) and 17,991 controls. We evaluated associations with SSc subtypes, identifying C4-independent HLA alleles. RESULTS: Lower C4 CN and higher HERV-K CN were associated with increased risk in all SSc subgroups. ATA+ patients showed the strongest association, particularly with C4A (odds ratio = 1.88), and differences in C4A CN association were more pronounced between autoantibody subgroups (ATA+ vs ACA+, P = 4 × 10-11) than between clinical subgroups (dcSSc vs lcSSc, P = 1 × 10-4). In ACA+ patients, only low C4B CN showed a significant association to SSc risk (P = 1.23 × 10-5). We also observed sex-biased associations: dcSSc, ATA+, and ARA+ male patients showed stronger effects for C4A and ACA+ and lcSSc female patients for C4B. Finally, our results suggest that the HLA alleles associated with SSc subgroups are independent of C4 CN. CONCLUSION: This study highlights distinct genetic contributions of C4A and C4B in SSc subtypes susceptibility. Our findings suggest that lower C4 CNs, particularly C4A, increase the risk of the severe dcSSc subtype, potentially through a mechanism involving immune complex clearance.

Humans