Search PubMedSearch

Biomedical subjects

Tian Ge

Publications and source records attributed to Tian Ge.

10 recordsLinked to original sources

All of Us diversity and scale yield context-dependent improvements in polygenic prediction.

Polygenic risk scores (PRSs) trained on multiancestry data can improve prediction in under-represented groups, but large linked genetic and health datasets capturing broad human diversity remain limited. Using 245,388 whole-genome sequences from the All of Us research program (AoU) together with UK Biobank data, we developed multiancestry PRSs for 32 traits and diseases. We evaluated how ancestry, methodology and genetic architecture influenced PRS performance across ancestrally diverse AoU participants. Increased diversity in the AoU improved PRS accuracy for several traits, especially in under-represented populations. However, maximizing sample size by meta-analyzing AoU and UK Biobank was not universally optimal: for less polygenic traits, AoU-only training performed best in African ancestry participants, consistent with ancestry-enriched effects. Individual PRS accuracy declined linearly with increasing ancestry divergence from the discovery GWAS, but this decay was attenuated using multiancestry training data. These findings underscore the value of more representative biobanks for equitable PRS performance.

Journal Article

Contribution of copy number variants to schizophrenia in East Asian populations.

Studies on schizophrenia-associated rare copy number variants (CNVs) have predominantly focused on people of European (EUR) ancestry. Here we present a rare CNV study of schizophrenia in East Asian (EAS) populations, comprising 20,903 cases and 23,258 controls. We observed a significantly elevated genome-wide rare CNV burden in EAS cases compared with controls. Cross-population comparisons showed largely consistent rare CNV effects on schizophrenia risk. In the EAS sample, we identified nine genome-wide-significant schizophrenia-associated rare CNV loci. Meta-analysis with EUR data yielded 14 significant loci, including 8 that reached genome-wide significance for the first time. Genes within these 14 loci were significantly less tolerant to loss-of-function variants than genes in other CNV loci. The new rare CNVs associated with schizophrenia in EAS populations showed higher carrier frequencies in EAS than in EUR populations (0.38% versus 0.0017%). Overall, this study underscores the importance of increasing population diversity to fully capture the genetic underpinnings of schizophrenia.

Journal Article

Distinguishing different psychiatric disorders using DDx-PRS.

Despite great progress on case-control polygenic prediction, an unmet need remains for a method that genetically distinguishes clinically related disorders (e.g., schizophrenia (SCZ) versus bipolar disorder (BIP) versus major depressive disorder (MDD) versus controls). We introduce differential diagnosis-polygenic risk score (DDx-PRS), which jointly estimates the posterior probabilities of each diagnostic category (e.g., SCZ = 50%, BIP = 25%, MDD = 15%, control = 10%) by modeling variance-covariance structure across disorders, leveraging case-control polygenic risk scores and prior clinical probabilities for each diagnostic category. We applied DDx-PRS to Psychiatric Genomics Consortium SCZ, BIP, MDD and control data, including summary-level training data from three case-control genome-wide association studies (n = 41,917-173,140 cases; total n = 1,048,683) and held-out test data from different cohorts with equal numbers for each diagnostic category (total n = 11,460). DDx-PRS was well calibrated and well powered (consistent with simulations) and produced comparable results to methods that require tuning data. True diagnosis probabilities in the top deciles of predicted diagnosis probabilities were considerably larger than prior baseline probabilities, implying appreciable potential for clinical utility in certain settings.

Humans

Characterizing the impact of plasma protein levels on human brain structure and disorders leveraging integrative multi-omics analysis.

With recent advances in high-throughput proteomic technologies, population-scale plasma proteomics datasets, often linked to extensive genetic and phenotypic information, have become increasingly accessible. Yet the relationships between circulating protein levels, brain imaging phenotypes, and risk for neurological and psychiatric disorders remain largely unexplored. Proteome-wide association studies offer a promising approach for elucidating biological mechanisms that connect genetic variation to complex brain-related traits and diseases. In this study, we integrated protein quantitative trait loci (pQTLs) from the two largest plasma proteomic resources (the UK Biobank Pharma Proteomics Project [UKB-PPP] and Ferkingstad et al. [deCODE]) with genome-wide association studies of brain imaging-derived phenotypes in UK Biobank using Mendelian randomization and colocalization analyses. We identified 120 cis and 20 trans associations between plasma proteins and imaging phenotypes and validated these findings using brain tissue-derived proteomic and transcriptomic datasets. Multivariable Mendelian randomization revealed eleven plasma proteins (coding genes APOE, ARL3, MICB, NSF, RHOC, RSPO3, ENPP2, BTN2A1, EIF2AK3, MRVI1, and OPLAH) with significant direct effects on the risk of Alzheimer's disease, Parkinson's disease, multiple sclerosis, bipolar disorder, and schizophrenia. Single-cell expression and pathway enrichment analyses further revealed cell-type-specific effects and distinct biological processes underlying these protein-disease associations. Together, these findings demonstrate robust links between plasma protein variation and brain structure, delineate protein-disease pathways, and highlight the cellular and molecular mechanisms that contribute to neurobiological diversity and pathology.

Journal Article

Rare Coding Variants Reveal Distinct Genetic Architectures Across Multidimensional Sleep Phenotypes.

Sleep and circadian traits have been widely studied using common variants, but the contribution of rare coding variation remains unclear. We analyzed rare coding variants in 397,065 whole-exome sequenced UK Biobank participants across 36 sleep phenotypes from self-report, diagnoses, sleep medication use and accelerometry, and meta-analyzed results with 171,536 whole-genome sequenced All of Us participants of diverse ancestries, with replication in the Mass General Brigham Biobank (N = 31,275). We identified 260 genes associated with sleep phenotypes, including novel associations with sleep medication use in 29 genes and 24 out of 29 have not previously been reported with any sleep phenotypes. We observed modest but significant rare variant heritability and strong genetic correlations between sleep medication use, insomnia and fatigue. Temporal gene expression trajectory analyses indicate that genes associated with self-reported sleep traits show constant high prenatal expression, whereas genes linked to sleep medication phenotypes exhibit peak expression in the late prenatal period. These findings highlight distinct biological mechanisms captured by different measurement sources of sleep phenotypes and reveal rare-variant-informed targets for therapeutic discovery.

exome sequencing

Association of Genetic Liability to Psychiatric Disorders with Peripheral Metabolic Dysregulation.

IMPORTANCE: Individuals with psychiatric disorders face elevated cardiometabolic risk which is linked to increased mortality. The extent to which this reflects shared pathogenesis or the downstream effects of illness and treatment remains poorly understood. OBJECTIVE: To characterize the direct pleiotropic effects of psychiatric genetic liability on circulating metabolites and aggregate cardiometabolic risk, independent of psychiatric diagnosis and psychotropic medication use. DESIGN SETTING AND PARTICIPANTS: Cross-sectional analysis of Mass General Brigham Biobank participants with metabolomic profiling, genomic data, and linked electronic health records. EXPOSURES: Genetic liability to nine psychiatric disorders quantified using polygenic risk scores (PRS): attention deficit/hyperactivity disorder (ADHD), anorexia nervosa (ANO), anxiety disorder (ANX), autism spectrum disorder (ASD), bipolar disorder (BD), major depressive disorder (MDD), PTSD, schizophrenia (SCZ), and substance use disorder (SUD). MAIN OUTCOMES AND MEASURES: 249 circulating metabolites and four metabolomic risk scores (MRS) for type 2 diabetes, myocardial infarction, ischemic stroke, and vascular dementia. PRS-metabolite associations were estimated using nested models adjusting for lifetime psychiatric diagnosis and psychotropic medication use. RESULTS: Across 25,290 participants, we identified 604 significant PRS-metabolite associations (Bonferroni p< 1.36 x 10-4), of which 89% persisted after adjustment for lifetime diagnosis and medication use, suggesting that the direct genetic effects on metabolism are largely independent of illness or treatment. PRS for MDD, PTSD, and ADHD showed the most extensive dysregulation, with a transdiagnostic pattern of elevated lipids and systemic inflammation, specifically triglycerides (&#x3b2; = 0.04 to 0.05, all p< 4.4 x10-13) and glycoprotein acetyls (&#x3b2; = 0.05, all p< 2.2 x10-16). Notably, PRS for SCZ and BD showed minimal metabolite dysregulation despite having the strongest association with their target diagnoses. PRS for MDD, PTSD, ADHD, and SUD were associated with increased MRS across cardiometabolic conditions (&#x3b2; = 0.03 to 0.08, all p< 2.1 x10-4). Sensitivity analyses controlling for BMI or excluding participants without any psychiatric history (N: 21,305 and 11,150, respectively) showed a similar pattern. CONCLUSIONS AND RELEVANCE: Psychiatric genetic liability is associated with systemic metabolic dysregulation independent of illness onset or treatment, supporting a partially pleiotropic basis for psychiatric-cardiometabolic comorbidity.

Journal Article

Characterizing the Uncertainty, Misclassification and Inconsistency of Polygenic Prediction.

Polygenic risk scores (PRSs) hold promise for precision medicine, yet their clinical translation is hindered by substantial uncertainty in individual risk estimates and often limited agreement in risk stratification across multiple PRSs for the same disease. We develop a unified inferential framework to calibrate PRS point estimates and uncertainties for both quantitative traits and binary phenotypes, and to characterize how PRS accuracy, uncertainty, pairwise correlation jointly determine misclassification and classification inconsistency. We show, both theoretically and empirically, that individual- and population-level misclassification and inconsistency rates are highly predictable in independent datasets. We further evaluate PRS integration and uncertainty-aware probabilistic thresholding strategies that reduce misclassification and improve concordance in risk stratification. Together, these results demonstrate that instability in PRS-based classification is a predictable statistical consequence of uncertainty and establish a principled foundation for incorporating uncertainty into PRS-based risk interpretation, communication, and clinical decision-making.

Journal Article

Optimizing Control Definitions in Opioid Use Disorder Genetic Research Using Electronic Health Records.

Amidst the opioid crisis, understanding the genetic basis of opioid use disorder (OUD) is crucial for identifying biological mechanisms and intervention points. However, genome-wide association studies (GWASs) have been hampered by inadequate sample sizes and often the use of control populations not assessed for prior opioid exposure. Because opioid exposure is a prerequisite for the development of OUD, consideration of exposure history in controls is important. Electronic health record data (EHR) paired with genomic information allow a broader sampling of patients with OUD and exposed controls. We leveraged data across two healthcare systems to evaluate the impact of using controls not screened for opioid exposure ('generic') versus minimally opioid-exposed control ('exposed'). First, at the phenotypic level, we conducted phenome-wide association studies (PheWAS) to compare the medical comorbidity profiles of OUD cases when using generic versus exposed controls. While PheWAS results for OUD-related comorbidities were more pronounced when using the generic group, 83% of the disease associations were overlapping and of similar effect sizes. Second, at the genetic level, we conducted GWAS (cases vs. generic; cases vs. exposed) and assessed differences in genetic correlations and degrees of phenotypic misclassification. Genetic results were concordant across control groups based on heritability (generic: 0.16&#x2009;&#xb1;&#x2009;0.07 vs. 0.10&#x2009;&#xb1;&#x2009;0.07), associations with the coding OPRM1 variant rs1799971 (pgeneric&#x2009;=&#x2009;8.83E-03 vs. pexposed&#x2009;=&#x2009;1.83E-02) and genetic correlations with prior OUD GWAS (rg-generic&#x2009;=&#x2009;0.83&#x2009;&#xb1;&#x2009;0.26 vs. rg-exposed&#x2009;=&#x2009;0.78&#x2009;&#xb1;&#x2009;0.27). Although GWASs were limited by sample size (Ngeneric&#x2009;=&#x2009;6269, Nexposed&#x2009;=&#x2009;6365), compared to an independent OUD GWAS (N&#x2009;=&#x2009;425&#x2009;944), the dilution value for the two GWAS was not different from 1, suggesting no major impact of phenotypic misclassification. This study represents the first effort to enhance OUD genetic research through optimization of control definitions using EHR data. Generic controls ascertained within the US health systems, where exposure to prescription opioids is high, offer a practical alternative for genetic studies of OUD.

Humans

Distinguishing different psychiatric disorders using DDx-PRS.

Despite great progress on methods for case-control polygenic prediction (e.g. schizophrenia vs. control), there remains an unmet need for a method that genetically distinguishes clinically related disorders (e.g. schizophrenia (SCZ) vs. bipolar disorder (BIP) vs. depression (MDD) vs. control); such a method could have important clinical value, especially at disorder onset when differential diagnosis can be challenging. Here, we introduce a method, Differential Diagnosis-Polygenic Risk Score (DDx-PRS), that jointly estimates posterior probabilities of each possible diagnostic category (e.g. SCZ=50%, BIP=25%, MDD=15%, control=10%) by modeling variance/covariance structure across disorders, leveraging case-control polygenic risk scores (PRS) for each disorder (computed using existing methods) and prior clinical probabilities for each diagnostic category. DDx-PRS uses only summary-level training data and does not use tuning data, facilitating implementation in clinical settings. In simulations, DDx-PRS was well-calibrated (whereas a simpler approach that analyzes each disorder marginally was poorly calibrated), and effective in distinguishing each diagnostic category vs. the rest. We then applied DDx-PRS to Psychiatric Genomics Consortium SCZ/BIP/MDD/control data, including summary-level training data from 3 case-control GWAS ( N =41,917-173,140 cases; total N =1,048,683) and held-out test data from different cohorts with equal numbers of each diagnostic category (total N =11,460). DDx-PRS was well-calibrated and well-powered relative to these training sample sizes, attaining AUCs of 0.66 for SCZ vs. rest, 0.64 for BIP vs. rest, 0.59 for MDD vs. rest, and 0.68 for control vs. rest. DDx-PRS produced comparable results to methods that leverage tuning data, confirming that DDx-PRS is an effective method. True diagnosis probabilities in top deciles of predicted diagnosis probabilities were considerably larger than prior baseline probabilities, particularly in projections to larger training sample sizes, implying considerable potential for clinical utility under certain circumstances. In conclusion, DDx-PRS is an effective method for distinguishing clinically related disorders.

Journal Article

1q21.1 distal copy number variants are associated with cerebral and cognitive alterations in humans.

Low-frequency 1q21.1 distal deletion and duplication copy number variant (CNV) carriers are predisposed to multiple neurodevelopmental disorders, including schizophrenia, autism and intellectual disability. Human carriers display a high prevalence of micro- and macrocephaly in deletion and duplication carriers, respectively. The underlying brain structural diversity remains largely unknown. We systematically called CNVs in 38 cohorts from the large-scale ENIGMA-CNV collaboration and the UK Biobank and identified 28 1q21.1 distal deletion and 22 duplication carriers and 37,088 non-carriers (48% male) derived from 15 distinct magnetic resonance imaging scanner sites. With standardized methods, we compared subcortical and cortical brain measures (all) and cognitive performance (UK Biobank only) between carrier groups also testing for mediation of brain structure on cognition. We identified positive dosage effects of copy number on intracranial volume (ICV) and total cortical surface area, with the largest effects in frontal and cingulate cortices, and negative dosage effects on caudate and hippocampal volumes. The carriers displayed distinct cognitive deficit profiles in cognitive tasks from the UK Biobank with intermediate decreases in duplication carriers and somewhat larger in deletion carriers-the latter potentially mediated by ICV or cortical surface area. These results shed light on pathobiological mechanisms of neurodevelopmental disorders, by demonstrating gene dose effect on specific brain structures and effect on cognitive function.

Brain