Search PubMedSearch

SEARCH · Search PubMed

Results for “Ancestry-informative markers (AIMs)”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,083 recordsLinked to original sources

Quo vadis, BGA? A collaborative EDNAP exercise on the challenges and progress in forensic biogeographical ancestry inference.

There is a broad consensus that forensic tests for the prediction of externally visible characteristics (EVC) and analysis of biogeographic ancestry (BGA) of an individual are technically reliable. However, interpretation of the results and population-specific genotype distribution patterns remains challenging. EVC and BGA analyses provide valuable information for population genetics studies and as investigative leads for criminal cases, as well as for historical and contemporary identification tests. However, inaccurate or incorrect predictions, for example, from subjective bias in the interpretations made, have the potential to misdirect police investigations. The legal situation regarding EVC and BGA testing varies by country: ranging from countries where it is explicitly prohibited, to those without specific regulations on biogeographic ancestry prediction, and others that have already enacted laws governing its use. The reluctance to utilize these analyses is not only due to legal restrictions and data protection concerns, but also to initial limited sets of sufficiently comprehensive forensic DNA assays. Forensic BGA marker panels typically contain up to ∼300 SNPs. This relatively small number of genetic markers, along with limited reference population data, complicates the interpretation of results from donors of unknown origin. This paper presents the results of a collaborative EDNAP study, which, for the first time, evaluated the approach to reporting EVC and BGA data between international laboratories. For the study, DNA from nine individuals with self-reported ancestry was collected and analysed using various forensic panels differing in the number and composition of ancestry-informative markers genotyped, comprising: the Precision ID mtDNA Whole Genome Panel, the VISAGE Basic Tool and the VISAGE Enhanced Tool for Appearance and Ancestry Prediction, and the Ion AmpliSeq™ PhenoTrivium Panel. To ensure full data protection, all SNP genotypes and uniparental marker haplotypes obtained were not shared with third parties. Instead, the genetic data were analysed using a range of commonly used population analysis software packages. These analysis outcomes were then distributed to twelve European forensic laboratories (both academic and law enforcement institutions), who were asked to prepare reports based on their interpretation of the phenotypes and ancestry they inferred from the analysis data. A questionnaire sent alongside the genetic information, aimed to evaluate which difficulties were encountered by the participants in processing the BGA analysis data they were given.

Humans

Genetic Ancestry and Carrier Variant Frequency Enrichment in a Colombian Andean Population: Insights From the Eje Cafetero.

Colombia is one of the most genetically diverse populations in Latin America, and its demographic process has promoted the persistence and local enrichment of deleterious alleles, increasing the frequency of autosomal recessive disorders, particularly in semi-isolated Andean populations such as the Eje Cafetero. However, exome-based reference data from this region remain scarce, limiting ancestry-aware variant interpretation and carrier screening strategies. We aimed to characterize the ancestry proportions of this population using exome data, and to estimate the carrier frequency and distribution of pathogenic and likely pathogenic (P/LP) variants in clinically relevant recessive genes. We conducted a cross-sectional study with whole-exome sequencing (WES) in 316 unrelated individuals from the Colombian Eje Cafetero. P/LP variants were evaluated in 454 genes associated with autosomal recessive disorders. The global ancestry proportions were estimated using a validated panel of 250 exome-compatible ancestry-informative markers. Carrier frequencies were compared against Non-Finnish Europeans (NFE) and Admixed Americans (AMX) from gnomAD v4. The cohort showed predominant European ancestry (mean 51%), followed by Native American (36%) and African (13%) components. We identified 151 carriers of 89 distinct pathogenic variants across autosomal recessive genes. The most frequent variants were SERPINA1 c.863A>T (5.5%), CFTR c.1210-11T>G (3.5%), and PYGM c.1094C>T (1.5%). Also, recurrent variants were significantly enriched compared with both NFE and AMX populations, supporting regional founder effects. This study represents one of the most comprehensive exome-based genetic characterizations of the Colombian Eje Cafetero, revealing ancestry-specific enrichment of clinically relevant autosomal recessive variants driven by founder effects.

Female

Reduced FOXP3 expression and its association with genetic ancestry in neuromyelitis Optica spectrum disorder: a Colombian cohort.

BACKGROUND: Autoimmune disorders are characterized by impaired immune tolerance, largely mediated by CD4⁺ regulatory T cells (Tregs), whose function depends on the transcription factor FOXP3. OBJECTIVES: To compare FOXP3 expression levels between patients with neuromyelitis optica spectrum disorder (NMOSD) and healthy controls from Bogotá, Colombia, and to explore their association with genetic ancestry. METHODS: FOXP3 expression was quantified from peripheral blood mononuclear cells using RNA-based analysis. Genomic ancestry proportions were estimated using ancestry-informative markers. RESULTS: FOXP3 mRNA expression was significantly reduced in NMOSD patients compared with controls (∼2.4-fold decrease in median expression; p = 0.027). Lower FOXP3 mRNA expression was associated with optic neuritis (p = 0.0059), but not with historical or current AQP4-IgG seropositivity. In a sensitivity analysis restricted to patients with documented historical AQP4-IgG seropositivity, the direction of reduced FOXP3 mRNA expression was preserved but did not reach statistical significance. In pre-specified exploratory ancestry-related analyses, African (p = 0.035) and Amerindian ancestry (p = 0.012) were associated with FOXP3 mRNA expression levels. CONCLUSIONS: Reduced FOXP3 mRNA expression in Peripheral Blood Mononuclear Cells (PBMCs) suggests an altered immune regulatory profile in NMOSD, potentially involving mechanisms beyond antibody-mediated immunity. The association between genetic ancestry and FOXP3 mRNA expression suggests that population-specific genetic background may influence immune regulatory pathways. These findings should be interpreted as exploratory and require validation in larger, clinically homogeneous cohorts.

Adolescent

The Childhood Cancer and Leukemia International Consortium (CLIC): Expanding global collaboration in pediatric cancer etiology research.

Childhood cancers are rare, but incidence has risen modestly in countries with robust registration, partly reflecting improved diagnosis. In high-income countries, cancer is the leading cause of disease-related death in children. Marked inequities in incidence, survival, and research capacity underscore the need for large-scale collaboration to identify environmental, genetic, and contextual determinants of risk. The Childhood Cancer and Leukemia International Consortium (CLIC) was established in 2007 to study the etiology of childhood leukemia and later expanded in 2019 to include other childhood cancers, principally solid tumors. CLIC pools harmonized, individual-level data from case-control and cohort studies, obtained through interviews, record linkage (insurance claims, registries), or geographic information systems, and integrates germline genomic data where available. Membership has grown from 13 studies in 9 countries to 57 studies in 21 countries; recruitment spans the early 1960s to the present and encompasses approximately 150,000 cases across all tumor types and 300,000 controls with clinical, demographic, and exposure data, centralized via harmonized data dictionaries at the Data Coordination Center, established in 2014 at the International Agency for Research on Cancer, and supported by a secure analysis platform. Pooled analyses across diverse populations have implicated parental age, prenatal vitamin or folic acid use, mode of delivery, fetal growth, selected congenital anomalies, occupational or household exposures (e.g., pesticides), paternal smoking, and markers of early-life immune modulation (e.g., breastfeeding, daycare attendance) in leukemia risk, informing carcinogen evaluation and prevention. The integration of genetic ancestry and germline susceptibility data is clarifying ancestry-related differences in leukemia biology and outcomes, while confirming risk loci with population-specific effects. CLIC is now adding polygenic risk scores and exposomic data to refine etiologic subtyping and identify modifiable pathways, while broadening representation from underserved regions through partnership-building and capacity-strengthening.

Humans

Assessment of Genetic Correlations Between Tobacco or Alcohol Use and Neurodegenerative Diseases Using East Asian Genetic Ancestry Genome-Wide Association Study Results.

Alzheimer's disease (AD) and Parkinson's disease (PD) are the most prevalent late-onset neurodegenerative diseases worldwide. Both are influenced in part by genetic factors and are currently incurable. Tobacco and alcohol, the two most common substances used among the general adult population, are potential AD/PD risk factors and are also heritable. Although important progress has been made, most existing research on the genetics of AD and PD has been carried out in individuals of European genetic ancestry. Investigations in a broad range of groups are crucial to understand disease mechanisms. Given the current availability of ancestry-specific tobacco and alcohol use as well as AD and PD genome-wide association study summary statistics, we performed global and local genetic correlation analyses using East Asian datasets. Genes within the correlated genetic regions were subsequently used to identify potentially enriched biological pathways between substance use and neurodegenerative diseases. We identified a global genetic correlation between smoking cessation and PD, which we confirmed in complementary European genetic ancestry data. Gene set enrichment analyses highlighted potentially shared genetic mechanisms between breast cancer and AD, which warrants further exploration. This work aims to promote further analyses across genetic ancestry groups.

Female

Statistical test to compare the linkage model and the admixture model based on central limit results.

In the Admixture Model, the probability that an individual carries a certain allele at a specific marker depends on the allele frequencies in K ancestral populations and the proportion of the individual's genome originating from these populations. The markers are assumed to be independent. The Linkage Model is a Hidden Markov Model that extends the Admixture Model by incorporating linkage between neighboring loci. We prove consistency and asymptotic normality of maximum likelihood estimators for the ancestry of individuals in the Linkage Model, complementing earlier results by (Pfaff et al., 2004; Pfaffelhuber and Rohde, 2022; Heinzel, 2025) for the Admixture Model. These results are used to prove that a statistical test that allows for model selection between the Admixture Model and the Linkage Model is an asymptotic level-α-test. Finally, we demonstrate the practical relevance of our results by applying the test to real-world data from The 1000 Genomes Project Consortium (2015).

Genetic Linkage

The impact of sex, age, and genetic ancestry on DNA methylation across tissues.

Understanding the consequences of individual DNA methylation variation is crucial for advancing our knowledge of human biology and disease, yet the collective impact of individual traits on DNA methylation and their downstream effects on gene expression across human tissues remains poorly understood. Here, we quantify the contributions of sex, age, genetic ancestry, and BMI on autosomal DNA methylation variation across nine human tissues and 424 individuals from the Genotype-Tissue Expression project. We show that genetic ancestry and age have a greater impact on DNA methylation compared with sex, with aging effects being more widespread but less pronounced. On average, <10% of the gene expression variation in sex, age, and ancestry is mediated by DNA methylation differences, with ancestry showing the largest proportion of mediation. We further show that ancestry-associated DNA methylation differences accumulate at CpG sites with extreme methylation states and are largely under genetic control. The female autosomal genome exhibits consistent hypermethylation across tissues at Polycomb-repressed regions. Ultimately, we show that age-related Polycomb target hypermethylation is observed across multiple tissues but not in the gonads. Our multi-individual, multitissue approach defines the key drivers of human DNA methylation variation in healthy conditions, establishing a baseline for the interpretation of DNA methylation changes in disease contexts.

Humans

Upscaling Genotyping by Amplicon Sequencing With GBAS-GUI.

Genotyping by amplicon sequencing (GBAS) is a relatively low-cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large-scale genetic monitoring projects. However, most existing analytical pipelines are either marker-specific, insufficiently scalable, or lacking efficient data management systems for the long-term integration of genotypic information, limiting the full potential of GBAS. Here, we address this gap by introducing GBAS-GUI (https://github.com/sonnenbe-dot/GBAS-GUI), a pipeline capable of generating GBAS-based genotypic data for a wide variety of loci at scale. GBAS-GUI integrates a graphical user interface with multiple checkpoints to improve accessibility and robustness. It implements multiprocessing architecture and a relational database that links genotypic data with associated sample metadata to enhance scalability and data management. The pipeline further enables marker screening through automated calculation of polymorphism information content (PIC) and implements a strategy to recover homologous genotypic information from paralogous loci with non-overlapping amplicon length ranges. Using multiple empirical datasets, we demonstrate substantial improvements in processing speed, database management and handling artefacts related to co-amplification of unspecific regions and duplicates of the same genomic region. We further show that incorporating the full sequence information captured by an amplicon increases marker information content beyond what is achievable with length-based genotyping alone and expands the analytical versatility of GBAS. Overall, GBAS-GUI provides a robust, scalable and versatile framework that unlocks the potential of GBAS for large-scale population genetic and phylogeographic studies.

Genotyping Techniques

Germline variants and impact on lung cancer outcomes following chemotherapy: A systematic review.

BACKGROUND: Lung cancer is the primary cause of cancer deaths in the UK and globally, and the main subtypes are non-small cell lung cancer (NSCLC) and small cell lung cancer (SCLC). Many treatment options are available, with platinum-based chemotherapy being a key component for many patients. However, variation in survival outcomes exists among individuals of European ancestry, which makes it important to identify germline genetic variants that help guide decision-making and optimise patient treatment and outcomes. METHOD: A systematic literature search was conducted in PubMed and Web of Science for lung cancer studies investigating the impact of germline genetic variants on systemic anti-cancer therapy (SACT) outcomes in populations of European ancestry. The review was conducted according to the Preferred Reporting Items of Systematic Review and Meta-Analysis (PRISMA) and Synthesis without Meta-Analysis (SWiM) guidelines. RESULTS: A total of 20 studies were included in the review out of 4469 on NSCLC and SCLC, encompassing 3639 patients. The most thoroughly investigated area was NSCLC treated with platinum-based chemotherapy. Genetic variants associated with overall survival and/or progression-free survival included XPD Lys751Gln, XPD Asp312Asn, ERCC1 C118T, and XRCC1 Arg399Gln. For non-platinum-treated NSCLC and SCLC, there was insufficient evidence to conduct a meaningful investigation. CONCLUSION: The XPD Lys751Gln, XPD Asp312Asn, ERCC1 C118T, and XRCC1 Arg399Gln variants showed potential associations with survival outcomes among patients of European ancestry with NSCLC after platinum-based chemotherapy. To support clinical implementation, large real-world pharmacogenomics studies stratified by ancestry are needed to overcome statistical power and heterogeneity limitations.

Humans

Genome-wide association study of estimated glomerular filtration rate using repeated measurements in the Taiwan Biobank.

BACKGROUND: Chronic kidney disease (CKD) is a major global public health issue, with genetic factors playing a significant role in kidney function. Although genome-wide association studies (GWAS) have identified numerous loci associated with estimated glomerular filtration rate (eGFR), most studies relied on a single time-point measurement, which limits the capacity to account for within-individual measurement variability. METHODS: We performed a repeated-measurement GWAS in the prospective Taiwan Biobank (Taiwanese ancestry; n = 25,004) using two repeated creatinine-based eGFR measurements. Repeated eGFR values were analyzed using a linear mixed-effects model with a subject-specific random intercept and time-varying covariates, providing a more precise estimate of eGFR level. Identified loci underwent functional annotation (expression quantitative trait locus, deleteriousness prediction, and epigenetic markers) and were compared with results from a single-measurement GWAS. RESULTS: Six loci associated with eGFR were identified, including four previously reported regions (1q22, 4q21.1, 11p14.1, and 17q21.2) and two additional loci (6p21.32 and 15q24.2). Functional annotation implicated several candidate genes-such as MUC1/EFNA1, SHROOM3, HLA-DQB1, MPPED2, NRG4, and PGAP3/FBXL20-in the regulation of kidney function. CONCLUSION: Incorporating repeated eGFR measurements into GWAS may improve phenotypic precision for identifying genetic associations with kidney function. This study identified eGFR-associated loci and biologically plausible candidate genes in a Taiwanese population, which require further replication and functional validation.

Chronic kidney disease

Pharmacogenomics of antipsychotic-induced weight gain: A systematic review.

BACKGROUND: Antipsychotic-induced weight gain (AIWG) is a major clinical concern, affecting approximately 30% of patients. Clinical predictors explain only part of AIWG risk. Genetic and molecular variations are hypothesized to contribute to susceptibility. The purpose of this review is to summarize recent results to identify replicated and novel findings. STUDY DESIGN: Applying PRISMA guidelines, we searched MEDLINE, Embase, and PsycINFO (May 2018-May 2026) for studies on genetic and molecular associations with AIWG, extending our prior review. Reviews, editorials, and conference abstracts were excluded. We extracted study characteristics (design, diagnosis, antipsychotic exposure, sample size, ancestry, genetic variants, and AIWG outcomes) (e.g., &#x2265;7% weight gain, BMI change). RESULTS: Fifty-three studies met inclusion criteria. In candidate gene studies, the most consistently replicated genes associated with AIWG were observed for DRD2, HTR2C, and MC4R. Multiple novel associations were identified by genome-wide association studies (GWAS) (e.g., MAP2K1, ZDBF2, PEPD), polygenic risk scores (PRS) (e.g., body mass index PRS), gene expression (e.g., CYP3A4, EP300), and epigenetic analyses (e.g., cg12034943 at CRTC1). CONCLUSIONS: Polymorphisms in candidate genes related to neurotransmission and appetite regulation continue to be investigated for associations with AIWG, while novel findings have emerged from GWAS, gene expression, and epigenetic studies. Evidence remains inconsistent due to limited replication, methodological variability, sparse ancestry data, and geographical underrepresentation. No single genetic variant is ready for clinical use, and multi-omic and multi-ancestry models are needed to improve prediction and clinical utility.

Humans

Can't see the forest for the trees: The influence of marker type on inferred phylogenetic relationships in a cosmopolitan bat genus.

Fine-resolution information on species relationships and biological diversity is critically needed to guide conservation efforts amidst rapid environmental changes. Systematics, which forms the foundation of this knowledge, has been revolutionized by phylogenomics, utilizing genome-scale datasets. However, the use of diverse marker types, non-comparable taxon sampling, and outgroup selection can lead to conflicting phylogenetic hypotheses. These inconsistencies complicate study comparisons and hinder our ability to assess marker-specific impacts on phylogenetic resolution. The phylogenetic reconstruction of the bat genus Myotis, encompassing over 140 species and characterized by a rapid radiation in the last 20 million years, has been particularly influenced by these challenges. Achieving phylogenetic resolution in Myotis is particularly complex due to subtle interspecific differences in both morphological and molecular traits. Mitochondrial and nuclear markers often produce discordant trees, influenced by hybridization, introgression, and methodological variations. In this study, we employed a consistent taxonomic sample set of 44 Myotis taxa to evaluate the impact of five different genetic marker types on phylogenetic reconstruction. We observed significant discordance between topologies derived from conserved nuclear and mitochondrial markers and found that transposable elements were inadequate for resolving relationships across the entire genus. Our results also clarify the placement of previously problematic taxa within the genus. These findings emphasize the importance of aligning genetic marker choice with specific phylogenetic questions and highlight the influence of taxonomic and methodological variation on phylogenomic outcomes. This work provides a framework for improving phylogenetic inference in rapidly radiating groups and enhances our understanding of evolutionary history in Myotis.

Animals

Test-retest reliability of spatiotemporal, kinematic, and kinetic measures in marker-based 3D gait analysis: A systematic review.

BACKGROUND: Marker-based 3D gait analysis (3DGA) is widely used to quantify impairments and evaluate treatment effects. For longitudinal clinical interpretation, clinicians and researchers need reference values for inter-session measurement error. For this purpose, this systematic review synthesized Standard Error of Measurement (SEM) values for spatiotemporal, kinematic, and kinetic (moments) outcomes obtained from marker-based 3DGA studies. METHODS: PubMed and Scopus were searched (final search: 11 December 2025). Studies reporting inter-session test-retest SEM and/or MDC for steady-state overground or treadmill walking using marker-based motion capture were included. Two authors screened records and appraised methodological/reporting quality using a custom tool informed by COSMIN, GRRAS, and biomechanics-specific items. Due to heterogeneity, results were synthesized descriptively using study-level median SEM values, stratified by joint, plane, population (healthy, pathological, single subgroups), and walking condition. Minimal Detectable Change (MDC) values were computed for all available data. RESULTS: Thirty-four studies (762 participants, 44.2% females) were included, with substantially more evidence for overground than treadmill walking. Overground spatiotemporal outcomes showed low errors (walking speed SEM of 0.06 m/s; timing typically &#x2264;0.03 s; spatial parameters generally &#x2264;0.03 m). For joint kinematics during overground walking, median SEMs were 2.4&#xb0; (sagittal), 1.9&#xb0; (frontal), and 3.3&#xb0; (transverse). The corresponding joint-kinetic SEMs were approximately 0.06, 0.04, and 0.03 Nm/kg, respectively. Treadmill data followed similar patterns. SIGNIFICANCE: Marker-based 3DGA allows for accurate assessment of spatiotemporal, kinematic, and kinetic gait features. We provided detailed SEM/MDC lookup tables to support clinical decision-making. Results further offer a benchmark for validating emerging gait assessment technologies (e.g., markerless systems) against realistic limits of marker-based 3DGA.

Humans

First insights into the role of evolutionary history in shaping venom composition of Vipera ammodytes.

Understanding intraspecific venom variation requires distinguishing the contributions of neutral population history from natural selection. This study aims to determine whether venom variation in Vipera ammodytes species complex is structured across eight phylogenetic groups. Despite a complex evolutionary history, venom composition did not differ among phylogenetic groups within the analytical framework used, suggesting that shared ancestry alone does not explain venom variation. Whether local adaptation to environmental conditions explains the observed variation remains an open question for future studies.

Animals

Exploratory proteomic and metabolomic profiling of pleural effusions identifies histone H4 and alanine as promising complementary markers for pleural tuberculosis.

The diagnosis of pleural tuberculosis (Pl-TB) remains challenging. Histopathological analysis and pathogen detection in pleural biopsies are informative but limited. We investigated differentially expressed proteins and metabolites in pleural effusions from patients with Pl-TB, malignancies, and other pathologies. A proteomic analysis of pooled pleural effusions identified 45 proteins exclusively detected or upregulated in Pl-TB samples, many linked to infectious processes. Conversely, 18 proteins were uniquely found or upregulated in malignant pleural effusions, mainly associated with detoxification and hemostasis. To validate these findings, we employed targeted proteomics in individual samples. Eight proteins were validated: S100-A9, histone H4, insulin-like growth factor-binding protein 2, fibrinogen beta chain, ficolin-3, immunoglobulin heavy constant alpha 1, sulfhydryl oxidase 1, and histidine-rich glycoprotein. Additionally, NMR-based metabolomics identified 13 metabolites with differential abundance between Pl-TB and non-TB samples. Notably, N-acetyl-glycoprotein and the branched-chain amino acids, alanine and lysine differed between groups. Proteomic and metabolomic analyses revealed distinct molecular profiles between Pl-TB and non-TB patients, despite intra-group variability. To address this, we applied classification models. Histone H4 and alanine consistently emerged as discriminative features. Overall, this study provides novel insights into the molecular landscape of Pl-TB. The combined quantification of proteins and metabolites may improve differential diagnosis, although should be further validated in larger, independent cohorts before clinical application.

Humans

Height variation independent of known genetic variants and health in later life: a cohort study.

BACKGROUND: Adult-attained height is associated with later-life health, but it reflects both genetic and nongenetic influences. The health implications of height variation not explained by known common height-associated genetic variants remain unclear. OBJECTIVES: This study aimed to examine associations of residual height (height variation independent of known genetic variants) with multiple disease incidence and all-cause mortality in later life. METHODS: In this cohort study of 407,366 adults of European ancestry (aged 40-70 y) in the United Kingdom Biobank (2006-2010), sex- and age-specific genetically predicted height was estimated from 9863 height-associated variants, adjusted for 30 principal components of ancestry. Residual height was calculated as the difference between observed and genetically predicted height. Plasma proteomics (2054 proteins; Olink Explore) were profiled. Deaths and 49 incident diseases were ascertained through national registries. Multivariable Cox models estimated associations of residual height and related proteins with disease incidence and mortality. RESULTS: Higher residual height [mean (standard deviation, SD), 0.0 (4.8)] was associated with more favorable self-reported preadulthood exposures (e.g., later birth years, no maternal smoking around birth, being breastfed as an infant, no adoption experience, and lower childhood adversity scores) and lower hazard ratios (HRs) of 32 out of 49 diseases (median follow-up = &#x223c;12.5 y). Using participants with residual height within &#xb1;0.5 SDs from the mean as reference, those with residual height < -2 SDs had higher adjusted HRs of mortality [1.61; 95% confidence interval (CI): 1.50, 1.72], multimorbidity (1.28; 95% CI: 1.12, 1.46), cardiovascular disease (1.45; 95% CI: 1.32, 1.60), psychiatric/neurological disease (1.38; 95% CI: 1.28, 1.48), and other disease categories (e.g., diabetes, digestive, and musculoskeletal diseases). In contrast, higher genetically predicted height was associated with a higher incidence of 19 diseases, including subtypes of cancer, non-atherosclerotic cardiovascular diseases, and musculoskeletal diseases, as well as higher all-cause mortality. We identified 806 plasma proteins related to inflammation, immune response, and autophagy via tumor necrosis factor, Nuclear factor-kappa B, phosphoinositide-3 kinase/protein kinase B, and Janus kinase/signal transducer and activator of transcription signaling pathways, which were associated with residual height and multiple diseases and mortality. CONCLUSIONS: Higher residual height is associated with lower disease incidence and mortality, with associations that are distinct from those for genetically predicted height.

Humans

Menopause in the All of Us Research Program: a descriptive summary of electronic health record and survey response across sociodemographic characteristics.

OBJECTIVES: Menopause is a significant physiological transition with implications for health outcomes (eg, cardiometabolic disease), yet gaps remain in understanding this transition, including how menopause timing and type influence health outcomes. Large-scale cohort studies in midlife (age=40-60) females, including the All of Us Research Program (AoURP), provide opportunities to study menopause across diverse populations and data modalities. We characterized menopause-related data in AoURP, focusing on age distributions and concordance between electronic health record (EHR) diagnosis codes and survey responses. METHODS: We analyzed menopause-related surveys, EHR diagnostic codes, and genomic data among ~396,000 AoURP female participants. We summarized menopause-related variables across data sources, evaluated overlap between survey, EHR, and genomic data sets, and described age distributions overall and across sociodemographic characteristics. RESULTS: Among ~396,000 females, survey responses captured ~193,000 menopause observations, nearly seven times more than EHR diagnoses (~28,000), suggesting under-ascertainment in EHR data. Nearly all females (~99%) with an EHR menopause diagnosis reported menopause in the survey. Approximately 22,000 participants had overlapping menopause-related EHR, survey, and genomic data. Survey age patterns matched expectations, with participants predominantly <40 years reporting premenopausal status and those >60 years reporting postmenopausal status. A small subset with age >70 years (N&#x2248;1,700; 4%) reported no menopause, suggesting response or recall bias. EHR menopause codes were concentrated after age 45 years, with a notable spike at age 65. Modest differences in survey-based menopause age distributions were observed across sociodemographic characteristics (eg, race and ancestry). CONCLUSIONS: These findings inform sampling strategies, power calculations, phenotype definition, and study design for menopause research using AoURP data.

Age