Search PubMedSearch

SEARCH · Search PubMed

Results for “United Kingdom Biobank”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Sugar rationing during the first 1000 days and early onset cancer: a natural experiment.

BACKGROUND: The "first 1000 days" of life is a critical window for metabolic programming, while the long-term oncological consequences of nutritional exposures during this period remain understudied. OBJECTIVES: We aimed to evaluate whether restricted sugar intake in utero and during early childhood reduces risk of early onset cancer diagnosis and mortality in adulthood, utilizing a natural experiment. METHODS: We analyzed 63,819 United Kingdom Biobank participants born between October 1951 and March 1956, spanning the end of United Kingdom sugar rationing (September 1953). Leveraging a quasi-experimental birth cohort design, we compared participants exposed to sugar rationing in utero and during infancy with those unexposed. Early onset cancer incidence (&#x2264;50 y) and mortality were ascertained via integrated national Cancer Registry and hospital inpatient records. Multivariable Cox proportional hazards models (including Gompertz distribution) were used to estimate hazard ratios (HRs), with exploratory site-specific analyses. RESULTS: Among 63,819 participants (56.3% female), 40,397 were exposed to rationing and 23,422 were unexposed. Early life sugar restriction significantly reduced early onset cancer risk (HR: 0.66; 95% confidence interval: 0.53, 0.81; P < 0.001). A dose-response relationship was observed, with peak protection in individuals exposed for &#x2264;24 mo postnatally. This protection was observed systemically across solid tumors, independent of specific cancer sites. Specificity was corroborated by null associations with negative controls (herpes zoster and cataract). No significant difference was found for cancer-specific mortality. CONCLUSIONS: Restricting sugar intake during the first 1000 days is associated with a reduced risk of early onset cancer, extending the disease-free lifespan. The divergence between reduced incidence and unchanged mortality suggests early life metabolic environments primarily influence tumor latency rather than biological aggressiveness. These findings highlight the potential long-term public health implications of early life dietary guidelines against the rising burden of early onset cancer.

Humans

Biological Mechanisms Underlying the Cardiovascular Effects of Branched-Chain Amino Acids: A Proteome-Wide Mendelian Randomization Study.

BACKGROUND: Ischemic heart disease (IHD) is the leading cause of morbidity and mortality. Branched-chain amino acids (BCAAs) are associated with higher IHD risk, but the underlying biological pathways remain unclear. OBJECTIVES: This study aims to explore these pathways using 2-step proteome-wide Mendelian randomization. METHODS: We examined the associations between genetic proxies for BCAAs and 2922 proteins in the United Kingdom Biobank Pharma Proteomics Project, supplemented by a meta-analysis with data from deCODE to identify proteins associated with BCAAs. Next, we tested their effects on IHD risk using Coronary Artery Disease Genome-wide Replication and Meta-analysis plus Coronary Artery Disease Genetics Consortium (122,733 cases and 424,528 controls) and replicated in FinnGen (31,640 cases and 187,152 controls). We conducted sensitivity analyses using genetic instruments from deCODE. Proteins associated with IHD risk and, in a consistent direction, with genetically predicted BCAAs were considered potential mediators. RESULTS: Genetic proxies for BCAAs were associated with 40 proteins. Among these, 6 proteins showed consistent evidence of mediation, including complement C1s subcomponent, coagulation factor II, granulin, proprotein convertase subtilisin/kexin type 9, sex hormone-binding globulin, and V-set and transmembrane domain-containing protein 2-like. These proteins are involved in inflammation, coagulation, lipid metabolism, and cellular stress response. All associations were robust across different analytical methods and replicated in independent datasets. Mediation analysis showed that these proteins accounted for 6.5% to 32.1% of the association between BCAAs and IHD risk. CONCLUSIONS: This study identified 6 proteins that potentially link BCAAs to IHD, implicating pathways related to inflammation, coagulation, lipid metabolism, and cellular stress responses. To our knowledge, these findings provide novel mechanistic insights into the BCAA-IHD relationship and highlight potential protein targets for future prevention and intervention strategies.

Amino Acids, Branched-Chain

Height variation independent of known genetic variants and health in later life: a cohort study.

BACKGROUND: Adult-attained height is associated with later-life health, but it reflects both genetic and nongenetic influences. The health implications of height variation not explained by known common height-associated genetic variants remain unclear. OBJECTIVES: This study aimed to examine associations of residual height (height variation independent of known genetic variants) with multiple disease incidence and all-cause mortality in later life. METHODS: In this cohort study of 407,366 adults of European ancestry (aged 40-70 y) in the United Kingdom Biobank (2006-2010), sex- and age-specific genetically predicted height was estimated from 9863 height-associated variants, adjusted for 30 principal components of ancestry. Residual height was calculated as the difference between observed and genetically predicted height. Plasma proteomics (2054 proteins; Olink Explore) were profiled. Deaths and 49 incident diseases were ascertained through national registries. Multivariable Cox models estimated associations of residual height and related proteins with disease incidence and mortality. RESULTS: Higher residual height [mean (standard deviation, SD), 0.0 (4.8)] was associated with more favorable self-reported preadulthood exposures (e.g., later birth years, no maternal smoking around birth, being breastfed as an infant, no adoption experience, and lower childhood adversity scores) and lower hazard ratios (HRs) of 32 out of 49 diseases (median follow-up = &#x223c;12.5 y). Using participants with residual height within &#xb1;0.5 SDs from the mean as reference, those with residual height < -2 SDs had higher adjusted HRs of mortality [1.61; 95% confidence interval (CI): 1.50, 1.72], multimorbidity (1.28; 95% CI: 1.12, 1.46), cardiovascular disease (1.45; 95% CI: 1.32, 1.60), psychiatric/neurological disease (1.38; 95% CI: 1.28, 1.48), and other disease categories (e.g., diabetes, digestive, and musculoskeletal diseases). In contrast, higher genetically predicted height was associated with a higher incidence of 19 diseases, including subtypes of cancer, non-atherosclerotic cardiovascular diseases, and musculoskeletal diseases, as well as higher all-cause mortality. We identified 806 plasma proteins related to inflammation, immune response, and autophagy via tumor necrosis factor, Nuclear factor-kappa B, phosphoinositide-3 kinase/protein kinase B, and Janus kinase/signal transducer and activator of transcription signaling pathways, which were associated with residual height and multiple diseases and mortality. CONCLUSIONS: Higher residual height is associated with lower disease incidence and mortality, with associations that are distinct from those for genetically predicted height.

Humans

Early life sugar rationing and ageing related diseases, biological ageing and mortality.

Early-life nutrition may influence lifelong ageing, yet human evidence is scarce. Using Britain's postwar sugar rationing as a natural experiment, we examine its long-term effects in 64,809 United Kingdom Biobank participants. Exposure to sugar rationing during the first 1,000 days of life is associated with a 9% lower incidence of hallmark-related disease, with a hazard ratio of 0.91 and a 95% confidence interval of 0.88-0.94, and a 19% lower risk of all-cause mortality, with a hazard ratio of 0.81 and a 95% confidence interval of 0.69-0.93. Mediation analysis indicates that the survival association is statistically mediated, by approximately 60%, through differences in incident hallmark-related disease. Rationed individuals show 1.0-1.2-year younger biological ages across multiple clocks and lower organ ages, particularly in the lung, heart, and liver. Proteomic profiling identifies 47 altered proteins, with enrichment of adenosine monophosphate-activated protein kinase and longevity pathways and suppression of mechanistic target of rapamycin signaling. These findings are consistent with international recommendations to limit free or added sugars from the World Health Organization, United States Dietary Guidelines, and American Heart Association, and may inform policy discussions related to sugar taxation and infant food and marketing policies under the United Nations 2030 Agenda.

Humans

Causal association between non-steroidal anti-inflammatory drugs use and the risk of benign prostatic hyperplasia: a univariable and multivariable Mendelian randomization study.

BACKGROUND: The results of earlier observational research on the relationships between the usage of non-steroidal anti-inflammatory medicines (NSAIDs) and the risk of benign prostatic hyperplasia (BPH) have been inconsistent. METHODS: To assess these associations, we performed both univariable and multivariable Mendelian randomization (MR) studies. Instrumental variables (IVs) associated with exposures at the significance level (p&#x2009;<&#x2009;5&#x2009;&#xd7;&#x2009;10-6) were selected from a comprehensive meta-analysis conducted by the United Kingdom Biobank (UKB). Summary data for BPH were obtained from the FinnGen consortium, which comprised 30,066 cases and 119,297 controls. Sensitivity analyses were performed to evaluate heterogeneity and pleiotropy. RESULTS: We found evidence by univariable MR (UVMR) that genetically predicted NSAIDs use increased the risk of BPH (odds ratio [OR] per unit increase in log odds NSAIDs use: 1.164, 95% confidence interval [CI]: 1.041-1.302, p&#x2009;=&#x2009;0.008). After controlling for inflammation in multivariable MR (MVMR), the link persisted (OR: 1.165, 95% CI: 1.049-1.293, p&#x2009;=&#x2009;0.004). There were no indications of potential heterogeneity and pleiotropy in UVMR and MVMR analyses. CONCLUSION: The results of the MR estimates suggest that genetically predicted NSAIDs use may elevate the risk of BPH. This outcome prompts the imperative for deeper exploration into potential underlying mechanisms.

Humans

Genome-Wide Association Study of Accessory Atrioventricular Pathways.

IMPORTANCE: Understanding of the genetics of accessory atrioventricular pathways (APs) and affiliated arrhythmias is limited. OBJECTIVE: To investigate the genetics of APs and affiliated arrhythmias. DESIGN, SETTING, AND PARTICIPANTS: This was a genome-wide association study (GWAS) of APs, defined by International Classification of Diseases (ICD) codes and/or confirmed by electrophysiology (EP) study. Genome-wide significant AP variants were tested for association with AP-affiliated arrhythmias: paroxysmal supraventricular tachycardia (PSVT), atrial fibrillation (AF), ventricular tachycardia, and cardiac arrest. AP variants were also tested in data on other heart diseases and measures of cardiac physiology. Individuals with APs and control individuals from Iceland (deCODE Genetics), Denmark (Copenhagen Hospital Biobank, Danish Blood Donor Study, and SupraGen/the Danish General Suburban Population Study [GESUS]), the US (Intermountain Healthcare), and the United Kingdom (UK Biobank) were included. Time of phenotype data collection ranged from January 1983 to December 2022. Data were analyzed from August 2022 to January 2024. EXPOSURES: Sequence variants. MAIN OUTCOMES AND MEASURES: Genome-wide significant association of sequence variants with APs. RESULTS: The GWAS included 2310 individuals with APs (median [IQR] age, 43 [28-57] years; 1252 [54.2%] male and 1058 [45.8%] female) and 1&#x202f;206&#x202f;977 control individuals (median [IQR] year of birth, 1955 [1945-1970]; 632&#x202f;888 [52.4%] female and 574&#x202f;089 [47.6%] male). Of the individuals with APs, 909 had been confirmed in EP study. Three common missense variants were associated with APs, in the genes CCDC141 (p.Arg935Trp: adjusted odds ratio [aOR], 1.37; 95% CI, 1.24-1.52, and p.Ala141Val: aOR, 1.55; 95% CI 1.34-1.80) and SCN10A (p.Ala1073Val: OR, 1.22; 95% CI, 1.15-1.30). The 3 variants associated with PSVT and the SCN10A variant associated with AF, supporting an effect on AP-affiliated arrhythmias. All 3 AP risk alleles were associated with higher heart rate and shorter PR interval, and have reported associations with chronotropic response. CONCLUSIONS AND RELEVANCE: Associations were found between sequence variants and APs that were also associated with risk of PSVT, and thus likely atrioventricular reentrant tachycardia, but had allele-specific associations with AF and conduction disorders. Genetic variation in the modulation of heart rate, chronotropic response, and atrial or atrioventricular node conduction velocity may play a role in the risk of AP-affiliated arrhythmias. Further research into CCDC141 could provide insights for antiarrhythmic therapeutic targeting in the presence of an AP.

Humans

A genome-wide association study of stroke risk in Asian statin users: evidence from KoGES and UK Biobank.

BACKGROUND: Despite proven efficacy of statins in stroke prevention, genetic factors may influence individual stroke risk among statin users. With increasing precision medicine approaches and growing evidence of population-specific genetic variations, identifying genetic markers that predict stroke risk in statin-treated Asian populations has become critically important for personalized cardiovascular prevention strategies. METHODS: We conducted a genome-wide association study of 1,678 participants using lipid-lowering agents in the Korean Genome and Epidemiology Study (KoGES) cohort. Significant findings were replicated in 2,170 Asian participants on statins from the UK Biobank using an additive genetic model adjusted for relevant covariates. RESULTS: In the discovery analysis, 83 single nucleotide polymorphisms were suggestively associated with stroke (p&#x2009;<1.0&#x2009;&#xd7;&#x2009;10-5). Among these, 21 SNPs in the CDH13 gene were associated with increased stroke risk. The lead SNP, rs7201829, was significantly replicated in the UK Biobank (odds ratio: 2.29, p&#x2009;=&#x2009;2.39&#x2009;&#xd7;&#x2009;10-5). CONCLUSIONS: This study identified CDH13 as a significant genetic marker associated with stroke risk among Asian statin users. These findings provide the first genome-wide evidence for genetic determinants of stroke susceptibility during statin therapy, supporting the development of personalized prevention strategies in Asian populations.

Aged

A genome-wide association study identified 10 novel genomic loci associated with intrinsic capacity.

BACKGROUND: Intrinsic capacity (IC) is a multidimensional concept within the World Health Organization framework for healthy aging. It refers to the composite of an individual's physical and mental capacities that enable them to maintain well-being, functional ability, and engagement in valued activities throughout life. While substantial evidence supports the biological basis of IC and its subdomains, the extent to which genetic factors influence IC remains largely unexplored, with no studies currently available. METHODS: Using datasets from the UK Biobank (UKB; N&#x2009;=&#x2009;44 631) and the Canadian Longitudinal Study on Aging (CLSA; N&#x2009;=&#x2009;13 085), we implemented the restricted maximum likelihood method to estimate SNP-based heritability (h2snp), followed by a Genome-Wide Association Study (GWAS) to identify genetic variants associated with IC, and post-GWAS analyses to pinpoint biological implications. RESULTS: The h2snp for IC was estimated at 25.2% in UKB and 19.5% in CLSA. Our GWAS identified 38 independent SNPs for IC across 10 genomic loci and 4289 candidate SNPs, mapped to 197 genes. Post-GWAS analysis revealed the role of these genes in cellular processes such as cell proliferation, immune function, metabolism, and neurodegeneration, with high expression in muscle, heart, brain, adipose, and nerve tissues. Of the 52 traits tested, 23 showed significant genetic correlations with IC, and a higher genetic loading for IC was associated with higher IC scores. CONCLUSIONS: Overall, this study provides comprehensive evidence on the genetic architecture of IC, identifying novel genetic variants and biological pathways, advancing our current knowledge and laying the foundation for ongoing and future research on healthy aging.

Adult

Proteomic Profiling Captures Residual Cardiovascular Risk Beyond the PREVENT Model in Individuals With Cardiovascular-Kidney-Metabolic Syndrome Stages 2-3.

BACKGROUND: Cardiovascular-kidney-metabolic (CKM) syndrome reflects complex pathobiological interactions among metabolic disorders, kidney injury, and cardiovascular disease (CVD). Stages 2 and 3 represent critical phases of disease progression characterised by high pathological heterogeneity. This study aimed to develop a CVD protein risk score (PRS) for this population and evaluate its incremental predictive value over the PREVENT model. METHODS: This study included 24&#x2009;017 participants with CKM Stages 2-3 from the UK Biobank. Using 2923 plasma proteins measured via the Olink platform, a PRS was developed in a training set (n&#x2009;=&#x2009;19&#x2009;218) using the LASSO method. In the validation set (n&#x2009;=&#x2009;4799), the incremental predictive performance of this score over the PREVENT model was assessed using Harrell's C-statistic, net reclassification improvement (NRI) and integrated discrimination improvement (IDI). RESULTS: A risk score comprising 63 proteins was constructed, primarily reflecting inflammation, kidney injury and matrix remodelling. Key proteins included growth differentiation factor 15 (GDF15), hepatitis A virus cellular receptor 1 (HAVCR1), matrix metallopeptidase 12 (MMP12) and NT-proBNP. In the validation set, after adjusting for PREVENT risk factors, individuals in the high PRS group had a 2.56-fold higher risk of CVD compared to those in the low score group (HR: 2.56, 95% CI: 1.96-3.37). Integrating the score into the PREVENT model improved the C-statistic by 0.034 (0.672-0.706) and achieved a 10-year NRI of 15.8% (95% CI: 9.5%-20.9%) and an IDI of 2.2% (95% CI: 1.3%-3.3%). CONCLUSION: Combining the PREVENT model with the PRS developed in this study enhances the prediction of future CVD events in the CKM Stages 2-3 population. This approach facilitates the capture of residual risk and supports precision risk stratification and management for this high-risk group.

Humans

Machine Learning-Driven Prediction of Coronary Artery Disease Risk Based on UK Biobank Plasma Proteomics.

BACKGROUND: Coronary artery disease (CAD) is a leading global cause of mortality, yet the predictive accuracy of conventional risk models is limited. Here, we integrate conventional risk factors, polygenic risk scores, and large-scale proteomics to develop a unified model for enhanced CAD risk prediction. METHODS: Using data from UK Biobank, participants with plasma proteomics and genetic risk data were included after excluding prevalent CAD. Participants from England were split into training (n=32&#x2009;330) and internal validation (n=13&#x2009;857) sets, and Scotland/Wales participants formed an external validation set (n=5775). Incident CAD was ascertained from linked health records. A 202-protein proteomic risk score was derived by least absolute shrinkage and selection operator Cox regression, and CatBoost models were trained using conventional risk factors alone and with incremental addition of polygenic risk scores and protein proteomic risk scores; Shapley Additive Explanations-guided forward selection identified a compact protein panel. RESULTS: Across cohorts, the median age was 58&#x2009;years and &#x223c;45% were men. Protein proteomic risk score was dose-dependently associated with CAD risk. Compared with conventional risk factors alone, integrating polygenic risk scores and protein proteomic risk scores improved discrimination, with the area under the curve increasing from 0.750 (95% CI, 0.732-0.767) to 0.789 (95% CI, 0.772-0.805) in internal validation and from 0.717 (95% CI, 0.683-0.750) to 0.762 (95% CI, 0.732-0.791) in external validation. A 9-protein panel (GDF15 [growth differentiation factor 15], MMP12 [matrix metalloproteinase 12], NPPB [natriuretic peptide B], PGF [placental growth factor], REN [renin], ADGRG2 [adhesion G-protein coupled receptor], ACE2 [angiotensin-converting enzyme 2], CDCP1 [CUB domain-containing protein 1], CXCL17 [C-X-C motif chemokine ligand 17)]) captured most proteomic predictive information. CONCLUSIONS: Our findings demonstrate that integrating conventional risk factors, polygenic risk scores, and proteomic data improves CAD risk prediction. This study highlights the utility of proteomics in precision cardiovascular medicine and simplified risk stratification tools.

Humans

Prostate Cancer, Genetic Susceptibility, and Risk of Chronic Non-Urological Complications.

BACKGROUND: Chronic nonurological complications are common among prostate cancer (PCa) survivors; however, their spectrum, magnitude, and genetic contribution remain poorly characterized. METHODS: We evaluated 15 commonly reported nonurological complications and tested their associations with exposure to PCa and disease-specific polygenic risk scores (PRS) in the UK Biobank (UKB; N&#x2009;=&#x2009;219,133). Analyses were conducted using cause-specific Cox proportional-hazards models within a full-cohort framework with time-updated PCa status, delayed entry at study recruitment, and age as the underlying time scale. RESULTS: After recruitment, incident PCa was diagnosed in 13,780 men, of which 1,656 (12.02%) had metastatic PCa (mPCa). Risks for seven complications were higher among men with PCa, adjusting for genetic background (all p&#x2009;<&#x2009;0.05), including osteoporosis, venous thromboembolism, depression, and four primary cancers (bladder, kidney, colorectal, and pancreatic). Risks were consistently stronger among men exposed to mPCa than non-mPCa; for example, the hazard ratio (HR) (95% confidence interval) for osteoporosis was 1.88 (1.63-2.18) for any PCa, 4.67 (3.11-7.01) for mPCa, and 1.75 (1.50-2.04) for non-mPCa. Elevated risks for three additional complications (coronary artery disease, type 2 diabetes and chronic obstructive pulmonary disease) were observed only among men with mPCa. The risks for these ten complications were further increased among men with higher disease-specific PRS; for example, the HR for osteoperosis in mPCa patients in the top PRS quartile was 13.57 (8.24-22.34) compared with men without PCa (p&#x2009;<&#x2009;0.001). CONCLUSION: PCa diagnosis and inherited genetic susceptibility jointly contribute to increased risks of multiple chronic non-urological complications among survivors.

Humans

Proteomic signatures for sudden cardiac death and related intermediate phenotypes.

BACKGROUND: Novel markers for sudden cardiac death (SCD) are needed. OBJECTIVE: This study aimed to explore whether a protein risk score derived from a large-scale proteomics dataset improves risk prediction of SCD in the general population. METHODS: A total of 52,705 individuals with 1459 unique plasma protein measurements were included from the UK Biobank Pharma Proteomics Project. A protein risk score was developed using lasso-penalized Cox regression on 40,722 participants enrolled at the English centers and validated on 11,983 participants enrolled at the remaining centers. RESULTS: The protein risk score formula developed from the derivation set comprised 64 unique plasma proteins including latent-transforming growth factor beta-binding protein 2, protein tyrosine phosphatase receptor sigma, and spondin-1. In the test set, a per standard deviation increase in protein risk score was associated with a hazard ratio of 2.60 (95% confidence interval [CI] 2.12-3.18) for SCD. Adding a protein risk score to SCD clinical risk factors resulted in a concordance index increase of 0.063 (95% CI 0.037-0.105) for SCD. For ventricular arrhythmia-mediated SCDs, an increase in concordance index when a protein risk score was added to SCD clinical risk factors was 0.070 (95% CI 0.010-0.188). A protein risk score added to SCD clinical risk factors resulted in a risk reclassification of 16.9% (95% CI 9.0-24.7) at a 10-year risk threshold of 5%. A protein risk score was significantly associated with intermediate phenotypes of SCD including corrected QT prolongation, an increase in left ventricular mean myocardial thickness, and a decrease in left ventricular global longitudinal strain. CONCLUSION: A protein risk score derived from a single plasma sample significantly improved risk prediction of SCD and related intermediate phenotypes.

Humans

Proteomic Signatures Related to Physical Activity Are Associated with Risks of Future Disease.

PURPOSE: Physical activity (PA) can lower the risk of developing chronic diseases. However, few studies have examined the proteomic signatures linked to PA, and the role of these signatures in the connection between PA levels and future disease risk remains unclear. This study aimed to investigate whether proteomic signatures indicative of PA are associated with the risk of developing common chronic diseases and to explore their role as statistical links in the relationship between PA levels and disease development. METHODS: We used data from a subcohort of UK Biobank participants. PA intensity data were collected from accelerometers worn by each participant. Plasma proteomics results were obtained through Olink analysis. The risks of developing each primary chronic disease were evaluated for types of PA and their associated proteomic signatures, adjusting for age, sex, ethnicity, socioeconomic status, lifestyle factors, and key measurement time-lag covariates. RESULTS: Based on the UK Biobank, we identified significant differences among the proteomic signatures of accelerometer-measured light PA, moderate-to-vigorous PA, and total PA. The main enriched pathways of these proteomic signatures included cell adhesion, cell migration, and immune response. Higher levels of accelerometer-measured PA and their associated proteomic signatures correlated with a lower risk of developing cardiometabolic disorders, cancers, psychological or neurological disorders, and respiratory diseases. CONCLUSIONS: Our findings show that PA and PA-related proteomic signatures are statistically associated with lower risks of chronic diseases. Further analyses identified proteins that were correlated with both PA and disease risk. These results need to be confirmed through longitudinal studies involving diverse populations.

Humans

A genome-wide investigation of depression among individuals with and without irritability.

Individuals presenting with both depression and irritability may constitute a different group of individuals with respect to those presenting without irritability, but their biological differences remain unknown. We aimed to identify genetic variants associated with depression among individuals with and without irritability, highlight biological pathways, and test for genetic associations with other traits. We conducted a genome-wide association study (GWAS) using data from the UK Biobank (N&#x2009;=&#x2009;487,409). We identified a group of individuals presenting with depression and reporting never having experienced irritability (depression without irritability, n&#x2009;=&#x2009;35,857, 11.8%), and another with depression and reporting having experienced irritability (depression with irritability, n&#x2009;=&#x2009;23,613, 8.1%) and compared them to controls with no depression or irritability (n&#x2009;=&#x2009;268,012). The GWAS of depression without irritability identified 2 SNPs which reached genome-wide significance (P&#x2009;<&#x2009;5&#xd7;10-8; rs72795440 and rs1233494). The GWAS of depression with irritability (NGWAS&#x2009;=&#x2009;292,485) identified 3 SNPs reaching genome-wide significance (rs2815748, rs102275, and rs7227069). When comparing SNPs between depression phenotypes, 15 SNPs had significantly different effect sizes. Patterns of genetic correlation with 44 complex traits were overall similar between the 2 depression phenotypes, with the highest genetic overlap observed with anxiety for depression without irritability (rg&#x2009;=&#x2009;.77) and neuroticism for depression with irritability (rg&#x2009;=&#x2009;.76). This study shed light into common and distinct biological factors characterizing depression among individuals with and without irritability and contribute to better understanding the genetic architecture of depression to potentially inform treatment and personalized medicine.

Humans

Enhanced Prediction of Peripheral Artery Disease Using Plasma Proteomics Among Individuals Without Diabetes.

BACKGROUND: Although peripheral artery disease (PAD) is an important diabetes complication, a substantial proportion of cases occur among individuals without diabetes. This study aimed to assess the predictive value of plasma proteomics in the long-term risk of PAD among individuals initially free of diabetes. METHODS: Included were 46&#x2009;508 participants (6046 with prediabetes) without diabetes or major cardiovascular disease at recruitment of the UK Biobank. Using multivariable Cox regression models, a total of 2923 unique plasma proteins were assessed for the associations with incident PAD. Significant proteins were subsequently processed by a trained light gradient boosting machine classifier to determine important proteins. Using receiver operating characteristic analyses, the performance of these important proteins in predicting incident PAD were evaluated, in the whole sample and by glycemic status (normoglycemia and prediabetes). RESULTS: During a median follow-up of 12.7&#x2009;years, 461 participants developed PAD. There were 107 proteins associated with incident PAD, with 103 positive associations. The LGBM approach identified 9 proteins (eg, WFDC2 [WAP 4-disulfide core domain protein 2], MMP12 [macrophage metalloelastase], and GDF15 [growth differentiation factor 15]) as the top-ranked proteins based on their importance ordering. Whereas glycated hemoglobin showed very modest predictive accuracy, a panel incorporating these top proteins showed good performance in the prediction of PAD risk (area under the curve 0.820), and it significantly enhanced the prediction beyond traditional risk factors (raising area under the curve from 0.803 to 0.837, DeLong test P=5.21&#xd7;10-3). These observations were consistent for participants with normoglycemia or prediabetes. CONCLUSIONS: Plasma protein biomarkers enhance the prediction of long-term risk for PAD among individuals without diabetes, regardless of glycemic status.

Humans

Proteomics-Based Soluble Urokinase Plasminogen Activator Receptor Levels Are Associated With Adverse Cardiovascular Outcomes in the General Population: Insights From the UK Biobank.

BACKGROUND: Elevated soluble urokinase plasminogen activator receptor (suPAR) levels are associated with inflammation, immune activation, and major adverse cardiovascular events in coronary artery disease. Encoded by the PLAUR gene, suPAR levels are influenced by the rs4760 genetic variant. Whether proteomics-based suPAR levels predict adverse outcomes in the general population remains unknown. METHODS: Proteomics-based suPAR levels were measured using the Olink Immunoassay in 33&#x2009;963 UK Biobank participants without known coronary artery disease. Fine-Gray and Cox proportional hazards models assessed associations between suPAR and major adverse cardiovascular events (primary outcome: cardiovascular mortality, nonfatal myocardial infarction, or stroke), cardiovascular mortality, and all-cause mortality (secondary outcomes), after adjustment for demographic and clinical risk factors, hs-CRP (high-sensitivity C-reactive protein), and the rs4760 variant. Incremental discrimination was evaluated using C-statistics. RESULTS: Participants were aged 56.4 (SD, 8.2) years; 45% were men, and 93.4% were White. Over a median follow-up of 14&#x2009;years (476&#x2009;177 person-years), 10.7% experienced major adverse cardiovascular events, 2.6% experienced cardiovascular mortality, and 9.4% experienced all-cause mortality. Each 1-SD increment in proteomics-based suPAR was associated with significantly higher risk of major adverse cardiovascular events (hazard ratio [HR], 3.2 [95% CI, 2.9-4.5]), cardiovascular mortality (HR, 5.9 [95% CI, 5.1-6.9]), and all-cause mortality (HR, 5.0 [95% CI, 4.6-5.5]), independent of clinical risk factors and hs-CRP. Additional adjustment for rs4760 did not attenuate these associations. Proteomics-based suPAR significantly improved discrimination beyond clinical risk factors (C-statistic: 0.719 versus 0.732; P<0.001). CONCLUSIONS: Proteomics-based suPAR independently predicts adverse cardiovascular outcomes in the general population, beyond conventional risk factors, hs-CRP, and genetic predisposition to elevated suPAR levels.

Humans

Large-Scale Plasma Proteomics Enhances Prediction of Liver-Related Events Among Individuals With Prediabetes and Type 2 Diabetes: A Prospective Cohort Study in the UK Biobank.

OBJECTIVE: To develop a protein risk score (ProRS) for predicting liver-related events (LREs) in patients with diabetes and compare its predictive performance with the Fibrosis-4 Index (FIB-4) and an established polygenic risk score. RESEARCH DESIGN AND METHODS: This prospective cohort study included 13&#x2009;516 individuals with prediabetes and type 2 diabetes (T2D) from the UK Biobank. Cox proportional hazards models and LASSO regression were applied to identify proteins associated with incident LREs and construct the ProRS. Predictive performance was assessed using Harrell's C-index, time-dependent area under the receiver operating characteristic curve, net reclassification improvement and integrated discrimination improvement. RESULTS: Over a median follow-up of 13.5&#x2009;years, 171 (1.3%) incident LREs occurred. We identified 877 proteins associated with LRE risk, primarily enriched in inflammatory signalling, extracellular matrix remodelling and complement/coagulation cascades. In the training set, we developed a 24-protein ProRS (C-index, 0.842; 95% CI 0.797-0.884) that stratified individuals into low-, medium- and high-risk groups, with 10-year cumulative incidences of LREs of 0.2%, 1.2% and 14.2%, respectively. Compared with the low-risk group, the hazard ratio for LREs was 57.1 (95% CI 31.9-102) in the high-risk group. In the internal validation set, the ProRS model (C-index, 0.876; 95% CI 0.827-0.920) accurately predicted both short- and long-term LREs and outperformed FIB-4 index (C-index, 0.733; 95% CI 0.657-0.807) and polygenic risk score (C-index, 0.636; 95% CI 0.564-0.706). CONCLUSIONS: The protein risk score demonstrated superior performance compared with the FIB-4 index and the polygenic risk score in predicting incident LREs among individuals with prediabetes and T2D. The score allows stratification of individuals according to liver-related risk, though external validation in multi-ethnic cohorts is warranted.

Humans

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models