Search PubMedSearch

SEARCH · Search PubMed

Results for “Biological Specimen Banks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models

Rapid derivation of cloning-competent cells from peripheral blood advances conservation biobanking.

Establishing viable cell lines from endangered species is essential for conservation, yet traditional fibroblast derivation from skin biopsies faces challenges including contamination risk and extended culture timelines. Here, we demonstrate that endothelial progenitor cells (EPCs) and pericytes isolated from peripheral blood represent robust alternatives to fibroblasts for biobanking. Compared to canid fibroblasts, canid blood-derived cells exhibit 2- to 3-fold faster doubling rates (15 to 20&#xa0;h vs. ~35&#xa0;h for fibroblasts) and reduced time to banked cell lines (1.5 to 2&#xa0;wks vs. 3 to 4&#xa0;wks for fibroblasts). Proteomic profiling of 32 canonical markers confirmed EPCs and pericytes represent distinct populations with lineage-specific molecular signatures. Optical genome mapping demonstrated equivalent genomic stability across cell types with no detectable structural variants or aneuploidies. Finally, interspecific somatic cell nuclear transfer (iSCNT) experiments confirmed both EPCs and pericytes generate viable canid embryos with efficiency meeting or exceeding fibroblasts. As a proof of concept for conservation cloning, iSCNT embryos made with gray wolf blood-derived cells had a 15% implantation rate following embryo transfer and resulted in six viable fetuses. These findings support integrating blood-derived cell banking into conservation programs, which enables opportunistic genetic preservation during standard management activities and expands options for genetic rescue through assisted reproductive technologies.

Animals

Integration of Genetic Information to Improve Brain Age Gap Estimation Models in the UK Biobank.

Neurodegeneration occurs when the body's central nervous system becomes impaired as a person ages, which can happen at an accelerated pace. Neurodegeneration impairs quality of life, affecting essential functions, including memory and the ability to self-care. Genetics play an important role in neurodegeneration and longevity. Brain age gap estimation (BrainAGE) is a biomarker that quantifies the difference between a machine learning model-predicted biological age of the brain and the true chronological age for healthy subjects; however, a large portion of the variance remains unaccounted for in these models, attributed to individual differences. This study focuses on predicting the BrainAGE more accurately, aided by genetic information associated with neurodegeneration. To achieve this, a BrainAGE model was developed based on MRI measures, and then the associated genes were determined with a Genome-Wide Association Study. Subsequently, genetic information was incorporated into the models. The incorporation of genetic information yielded improvements in the model performances by 7% to 12%, showing that the incorporation of genetic information can notably reduce unexplained variance. This work helps to define new ways of determining persons susceptible to neurological aging decline and reveals genes for targeted precision medicine therapies.

Humans

Dynamic Fusion of Genomics and Functional Network Connectivity in UK Biobank Reveals Schizophrenia-Related SNP Manifolds.

Many mental disorders show strong genetic influence. In parallel, dynamic functional network connectivity (dFNC) has shown high sensitivity to brain changes related to mental disorders. However, previous studies linking dFNC to genetics largely follow a paradigm to identify associations between one set of genetic factors and multiple sets of connectivity features from different dFNC states, ignoring the potential variability in genetic correlates across states. We propose a novel joint ICA (jICA)-based "dynamic fusion" framework to identify dynamically tuned genetic manifolds. A sliding window approach was utilized to estimate four dFNC states and compute subject-level state-average dFNC (sa-dFNC) features. The sa-dFNC features of each state were combined with schizophrenia risk single nucleotide polymorphisms (SNPs) within a jICA fusion framework, resulting in four parallel fusions in 32,861 individuals of the UK Biobank cohort. The extracted four sets of joint SNP-dFNC components were further validated for clinical relevance in a combined schizophrenia cohort of 820 individuals (348 patients). The similarity of SNP-dFNC components across four parallel fusions was evaluated as a measure of state variability. We observed a mixture of "state-invariant" and "state-variant" components for SNP and dFNC modalities. Particularly, the schizophrenia-related state-variant SNP components, or manifolds, complemented each other by capturing different SNPs involved in the same biological functions, revealing a partition of genomic risk particularly elicited by the dynamics of brain function. By augmenting the SNP factors to state-variant manifolds, this dynamic fusion framework promises additional insights into the underlying genetic risk of disease-related alterations in dynamic brain function.

Humans

The phenotype-genotype reference map: Improving biobank data science through replication.

Population-scale biobanks linked to electronic health record data provide vast opportunities to extend our knowledge of human genetics and discover new phenotype-genotype associations. Given their dense phenotype data, biobanks can also facilitate replication studies on a phenome-wide scale. Here, we introduce the phenotype-genotype reference map (PGRM), a set of 5,879 genetic associations from 523 GWAS publications that can be used for high-throughput replication experiments. PGRM phenotypes are standardized as phecodes, ensuring interoperability between biobanks. We applied the PGRM to five ancestry-specific cohorts from four independent biobanks and found evidence of robust replications across a wide array of phenotypes. We show how the PGRM can be used to detect data corruption and to empirically assess parameters for phenome-wide studies. Finally, we use the PGRM to explore factors associated with replicability of GWAS results.

Humans

Extracting and calibrating evidence of variant pathogenicity from population biobank data.

Genomic medicine requires a robust evidence base of variant phenotypic impacts, which remains incomplete even in extensively studied genes with monogenic disease associations. Here, we evaluated the broad potential of using population cohort data to identify evidence that can be used in variant assessment. Across 41 genes related to 18 clinically actionable monogenic phenotypes, we calculated variant-level odds ratios of disease enrichment using data from 469,803 UK Biobank participants. We found significant differences in odds ratio values between ClinVar-labeled pathogenic and benign variants in 11 phenotypes, spanning both common and rare disorders. To facilitate clinical translation, we calibrated the strength of evidence provided by variant-level odds ratios to align with American College of Medical Genetics and Genomics and the Association for Molecular Pathology (ACMG/AMP) interpretation guidelines (PS4 criterion) and found that odds ratios may reach "moderate," "strong," or "very strong" evidence, varying by phenotype and gene. Overall, we found that 2.6% (N = 12,350) of participants harbor a rare variant of uncertain significance (VUS) with at least moderate evidence of pathogenicity-an indication of potentially unrecognized disease risk. Finally, by incorporating computational and functional data alongside population-based odds ratios, we identified variants that met the criteria for clinical reclassification. Notably, using this approach, we identified that 12.4% of rare VUSs in LDLR seen in participants meet diagnostic criteria to be classified as likely pathogenic, demonstrating its potential to scale the reclassification of VUSs.

Humans

Evolving knowledge of red flag clinical features associated with TTR p.(Val142Ile) in a diverse electronic health-record-linked biobank.

PURPOSE: Previous studies have established red flags that raise clinical suspicion for the hereditary form of transthyretin amyloidosis (ATTRv). However, these have not been specifically evaluated for the most common associated variant, TTR p.(Val142Ile). METHODS: Using an ancestrally diverse electronic health-record-linked biobank with exome sequence data from 27,630 unrelated adults, we evaluated 9 ATTRv-related clinical features among TTR p.(Val142Ile)-positive and -negative individuals. RESULTS: Among 337 variant-positive individuals (median age 63, 60% female), 10 (3.0%) were diagnosed with amyloidosis. TTR p.(Val142Ile) was associated with increased odds of cardiomyopathy/heart failure (CM/HF), atrial fibrillation, polyneuropathy, carpal tunnel syndrome, and proteinuria, but only in individuals &#x2265;60 years. These features were evident 1.7 to 7.7 years earlier in variant-positive vs -negative individuals (hazard ratio [HR] 1.37, P&#xa0;= 3.99&#xa0;&#xd7; 10-2; HR 1.78, P&#xa0;= 2.52&#xa0;&#xd7; 10-3; HR 1.78, P&#xa0;= 1.70&#xa0;&#xd7; 10-3; HR 1.81, P&#xa0;= 5.14&#xa0;&#xd7; 10-3; HR 1.60, P&#xa0;= 1.94&#xa0;&#xd7; 10-2, respectively). By age 50, the cumulative incidence of CM/HF was 3.5-fold higher, and by age 60, the incidences of CM/HF, polyneuropathy, and proteinuria were 2-fold higher in variant-positive individuals. CONCLUSION: This study clarifies red flags that are associated with TTR p.(Val142Ile) in an age-dependent manner. With modifying therapies being available, early diagnosis of ATTRv in variant-positive individuals through the recognition of key clinical features is paramount.

Humans

Prevalence and determinants of profound vitamin D deficiency (25-hydroxyvitamin D <10 nmol/L) in the UK Biobank and potential implications for disease association studies.

BACKGROUND: 25-hydroxyvitamin D (25OHD) is the principal biomarker of vitamin D status. Values below the assay detection limit (<10 nmol/L) are often reported as missing. Thus the most severely deficient participants are excluded from research which can lead to inaccurate findings such as underestimated prevalence of deficiency, overlooked risk factors, and biased evaluation of disease associations. METHODS: In total 369,626 individuals from the UK Biobank cohort were included in this study. Data on 25OHD concentration and relevant demographic and lifestyle factors such as age, supplement intake, diet, and time spent outdoors were used in the analyses. Ambient UVB radiation was approximated for each participant. 25OHD was evaluated as a categorical outcome and we reintroduced participants with 25OHD values <&#x202f;10 nmol/L (conventionally reported as missing values) back to the dataset. Adjusted regression models were used to investigate the determinants of profound (25OHD <10 nmol/L) and severe (10-25 nmol/L) vitamin D deficiency and to assess disease associations (with 25-50 nmol/L as the reference category). RESULTS: 1,784 (0.48&#x202f;%) individuals were profoundly deficient and a further 47,226 (12.78 %) individuals were severely vitamin D deficient. The proportions of profoundly and severely deficient were highest among Asians, 9&#x202f;% and 47&#x202f;%, respectively. Ambient UVB radiation was the second strongest predictor: comparing the lowest vs. highest quartile, the risk of profound deficiency was 17-fold increased and that of severe deficiency 7.5-fold increased. Use of vitamin D supplements substantially reduced risk of profound (4.4-fold) and severe (2.5-fold) deficiency, as did fish intake (5- and 1.9-fold, respectively). Profound deficiency was more strongly associated with chronic illness, diabetes, and emphysema compared to severe deficiency. CONCLUSION: The prevalence of profound and severe vitamin D deficiency among Asian and Black ethnicities in the UK is high and requires targeted action. Solar radiation is potent in protecting against profound and severe vitamin D deficiency. Studies evaluating the relationship between vitamin D status and other health outcomes may be biased if profoundly deficient participants are excluded.

Humans

Global biological sample collections from tuberculosis studies: a scoping review.

Progress in tuberculosis vaccine development is hindered by the incomplete understanding of protective immunity and other disease mechanisms. An interconnected network of sample biorepositories from tuberculosis studies could help to address these gaps. To assess the feasibility of such a resource, we conducted a scoping review of tuberculosis observational studies and vaccine clinical trials. The included studies collected at least one biological sample from tuberculosis cases, contacts, or controls and had more than 100 participants. We contacted the corresponding authors of these studies to determine the sample availability and interest in interconnected biorepositories. For the period 2014-24, we identified 104 observational studies and 18 vaccine trials that collected biological samples from 35&#x2009;075 tuberculosis cases, 39&#x2009;450 contacts or controls, and 45&#x2009;628 trial participants across 43 countries. The commonly collected samples were blood, human genomic DNA, RNA, and sputum. Interest among the contacted investigators was high. Interconnected sample biorepositories could facilitate large-scale investigations and accelerate progress towards tuberculosis vaccine development.

Humans

Quantifying and improving rheumatoid arthritis algorithm performance in biobank settings.

OBJECTIVE: To quantify and improve the performance of standard rheumatoid arthritis (RA) algorithms in a biobank setting. METHODS: This retrospective cohort study within the Mayo Clinic (MC) Biobank and MC Tapestry Study identified RA cases by presence of at least two RA codes OR positive anti-cyclic citrullinated peptide antibodies (CCP) plus disease-modifying anti-rheumatic drug (DMARD) prescription as of 7/18/2022. Rheumatology physicians manually verified all RA cases using RA criteria and/or rheumatology physician diagnosis plus DMARD use. All other biobank participants served as non-RA controls. We defined seropositivity as rheumatoid factor and/or anti-CCP positivity. We assessed rules-based and Electronic Medical Records and Genomics (eMERGE) RA algorithms using positive predictive value (PPV). Finally, we developed a novel RA algorithm using a LASSO-based machine learning approach with five-fold cross validation. RESULTS: We identified 1,316 confirmed RA cases (968 MC Biobank, 348 Tapestry, 70 % seropositive) and 82,123 non-RA controls (mean age 65, 61 % female). The PPV of 3 RA codes was 43 %, codes plus DMARD was 54 %, and codes plus DMARD plus seropositivity was 85 %. The PPV of eMERGE was 77 %. Available in the MC Biobank, self-reported RA (PPV 10 %) only minimally improved algorithm performance (PPV from 83 % to 85 %), whereas family history of RA (PPV 3 %) worsened performance. At 90 % PPV, the novel RA algorithm incorporating key variables such as anti-CCP and DMARD use increased sensitivity by 4-11 % compared to eMERGE. CONCLUSION: Rules-based and eMERGE RA algorithms had worse performance in biobank than administrative settings. Our novel RA algorithm outperformed these standard algorithms.

Humans

Using Large Genomic Biobanks to Generate Insights into Genetic Kidney Disease.

Chronic kidney disease (CKD) affects approximately 9% of the global population, leading to increased risks of end-stage kidney disease (ESKD), cardiovascular disease (CVD), and mortality. Patients with CKD are a huge burden on health care resources globally. CKD is a complex condition influenced by a combination of genetic, environmental, and traditional risk factors. Family studies have suggested heritability rates for CKD ranging from 30% to 75%, and large genomic biobank studies have proven essential in identifying genes with substantial effects on CKD risk and in capturing cumulative genetic risk through polygenic risk scores. These biobanks are crucial for discovering new genes associated with kidney health and disease, and their growing size enhances the power to detect novel genetic associations. Integrating multi-omics technologies such as transcriptomics, metabolomics, and proteomics further enriches our understanding of CKD, while advanced computational tools continue to expand our insights into genetic data. Polygenic risk scores, derived from hundreds of genetic variants with small effect sizes, can help identify individuals at high risk of CKD. Genomic biobanks offer valuable opportunities for early identification and personalized treatment of monogenic kidney disorders, such as autosomal dominant polycystic kidney disease and Alport syndrome. These biobanks help fill knowledge gaps, particularly in individuals with milder or asymptomatic presentations who are often underrepresented in traditional studies. Expanding genomic biobank efforts globally, especially in diverse populations, is vital to enhancing our understanding of the genetic underpinnings of kidney disease. This review highlights the significant contributions of genomic biobanks to advancing our comprehension of the genetics of CKD.

Humans

A pancreatic cancer organoid biobank links multi-omics signatures to therapeutic response and clinical evaluation of statin combination therapy.

Chemotherapy remains the primary treatment for pancreatic ductal adenocarcinoma (PDAC), but most patients ultimately develop resistance. Here, we established 260 pancreatic cancer organoid lines, followed by extensive multi-omics profiling and therapeutic sensitivity assessments. Integrated analyses uncovered 6 novel coding and 35 noncoding driver candidates. We discovered 2,794 multi-omics features associated with drug sensitivity and 322 features linked to radiation sensitivity. Pharmacogenomic analyses revealed that chemoresistant organoids exhibited enrichment in protein glycosylation and cholesterol metabolism pathways. Notably, statins effectively targeted chemoresistant PDAC organoids. Statin treatment attenuated protein glycosylation, cholesterol levels, and the epithelial-to-mesenchymal transition (EMT) signature in PDAC organoids. We conducted a single-center, single-arm, phase 2 clinical trial (NCT06241352) combining atorvastatin with chemotherapy in patients with advanced pancreatic cancer. Among 37 patients, 26 (70.3%) demonstrated a response, with tumor markers decreasing by more than 20%, suggesting durable responses and potential clinical benefits in this challenging patient population.

Humans

Ribosomal DNA copy number variation associates with hematological profiles and renal function in the UK Biobank.

The phenotypic impact of genetic variation of repetitive features in the human genome is currently understudied. One such feature is the multi-copy 47S ribosomal DNA (rDNA) that codes for rRNA components of the ribosome. Here, we present an analysis of rDNA copy number (CN) variation in the UK Biobank (UKB). From the first release of UKB whole-genome sequencing (WGS) data, a discovery analysis in White British individuals reveals that rDNA CN associates with altered counts of specific blood cell subtypes, such as neutrophils, and with the estimated glomerular filtration rate, a marker of kidney function. Similar trends are observed in other ancestries. A range of analyses argue against reverse causality or common confounder effects, and all core results replicate in the second UKB WGS release. Our work demonstrates that rDNA CN is a genetic influence on trait variance in humans.

Humans

Exome-wide evidence of compound heterozygous effects across common phenotypes in the UK Biobank.

The phenotypic impact of compound heterozygous (CH) variation has not been investigated at the population scale. We phased rare variants (MAF &#x223c;0.001%) in the UK Biobank (UKBB) exome-sequencing data to characterize recessive effects in 175,587 individuals across 311 common diseases. A total of 6.5% of individuals carry putatively damaging CH variants, 90% of which are only identifiable upon phasing rare variants (MAF&#xa0;<&#xa0;0.38%). We identify six recessive gene-trait associations (p&#xa0;<&#xa0;1.68&#xa0;&#xd7;&#xa0;10-7) after accounting for relatedness, polygenicity, nearby common variants, and rare variant burden. Of these, just one is discovered when considering homozygosity alone. Using longitudinal health records, we additionally identify and replicate a novel association between bi-allelic variation in ATP2C2 and an earlier age at onset of chronic obstructive pulmonary disease (COPD) (p&#xa0;<&#xa0;3.58&#xa0;&#xd7;&#xa0;10-8). Genetic phase contributes to disease risk for gene-trait pairs: ATP2C2-COPD (p&#xa0;= 0.000238), FLG-asthma (p&#xa0;= 0.00205), and USH2A-visual impairment (p&#xa0;= 0.0084). We demonstrate the power of phasing large-scale genetic cohorts to discover phenome-wide consequences of compound heterozygosity.

Humans

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans

Joint effects of childhood adversity and genetic risk for psychosis on psychopathology in the UK Biobank.

BACKGROUND: The individual effects of genetic factors and adverse childhood experiences (ACEs) on risk of psychosis, including schizophrenia (SCZ) and bipolar disorder (BIP), have been widely acknowledged, but their interaction effects on individual psychopathological symptoms remain unclear. METHODS: Based on data from 163,704 individuals in the UK Biobank, we investigated the joint effects of polygenic risk scores (PRSs) of SCZ and BIP and ACEs on psychopathology. ACEs status and 55 psychopathological symptoms from seven domains were measured retrospectively using an online mental health questionnaire in 2016. Recent genome-wide association studies for SCZ and BIP were combined with genotype data to generate PRSs. Logistic regression analyses were then conducted to explore univariate and joint main effects of PRSs and ACEs on psychopathological symptoms, as well as their additive and multiplicative interaction effects. RESULTS: The interaction mechanisms for PRSs and ACEs varied across symptom domains: additive interactions were observed on the depression (RERIBIP-ACEs&#xa0;=&#xa0;0.20-0.25), anxiety (RERISCZ-ACEs&#xa0;= 0.20; RERIBIP-ACEs&#xa0;=&#xa0;0.22-0.26), help-seeking (RERISCZ-ACEs&#xa0;=&#xa0;0.24; RERIBIP-ACEs&#xa0;=&#xa0;0.23), and cognition domains (RERISCZ-ACEs&#xa0;=&#xa0;-0.23 to -0.17), whereas multiplicative interactions were only detected on the psychotic (betaSCZ-ACEs&#xa0;=&#xa0;-0.543; betaBIP-ACEs&#xa0;=&#xa0;-0.181), mania (betaBIP-ACEs&#xa0;= -0.195), self-harm or suicide (betaSCZ-ACEs&#xa0;=&#xa0;-0.118), and cognitive domains (betaSCZ-ACEs&#xa0;=&#xa0;-0.204 to -0.157). CONCLUSIONS: The interplay mechanisms for genetic liability to SCZ and BIP and ACEs vary across symptom domains. This study reveals heterogeneity in gene-ACEs interaction mechanisms underlying psychosis and may provide personalized guidance for psychological care after ACEs.

Humans

Clinical implications of bone marrow adiposity identified by phenome-wide association and Mendelian randomization in the UK Biobank.

Bone marrow adiposity changes in diverse diseases, but the full scope of these, and whether they are directly influenced by marrow adiposity, remains unknown. To address this, we previously measured the bone marrow fat fraction of the femoral head, total hip, femoral diaphysis, and spine of over 48,000 UK Biobank participants. Here, we first use these data for PheWAS to identify diseases associated with marrow adiposity at each site. This reveals associations with 47 incident diseases across 12 disease categories, including osteoporosis, fracture, type 2 diabetes, cardiovascular diseases, cancers, and other conditions that burden public health worldwide. Intriguingly, type 2 diabetes associates positively with spine bone marrow adiposity but negatively with marrow adiposity at femoral sites. We then establish PRSs based on bone-marrow-fat-fraction-associated SNPs and use PRS-PheWAS and Mendelian randomization to explore causal associations between marrow adiposity and disease. PRS-PheWAS reveals that genetic predisposition to increased marrow adiposity is positively associated with osteoporosis and fractures. Mendelian randomization further suggests that increased marrow adiposity at the diaphysis and total hip is causally associated with osteoporosis. Our findings substantially advance understanding of how marrow adiposity impacts human health and highlight its potential as a biomarker and/or therapeutic target for diverse human diseases.

Humans

The importance of family-based sampling for biobanks.

Biobanks aim to improve our understanding of health and disease by collecting and analysing diverse biological and phenotypic information in large samples. So far, biobanks have largely pursued a population-based sampling strategy, where the individual is the unit of sampling, and familial relatedness occurs sporadically and by chance. This strategy has been remarkably efficient and successful, leading to thousands of scientific discoveries across multiple research domains, and plans for the next wave of biobanks are underway. In this Perspective, we discuss the strengths and limitations of a complementary sampling strategy for future biobanks based on oversampling of close genetic relatives. Such family-based samples facilitate research that clarifies causal relationships between putative risk factors and outcomes, particularly in estimates of genetic effects, because they enable analyses that reduce or eliminate confounding due to familial and demographic factors. Family-based biobank samples would also shed new light on fundamental questions across multiple fields that are often difficult to explore in population-based samples. Despite the potential for higher costs and greater analytical complexity, the many advantages of family-based samples should often outweigh their potential challenges.

Humans