Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Benchmark”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

A unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification.

The rapid growth of genomic sequencing demands fast, accurate, and scalable analysis methods. In viral genomic classification, expanding labeled reference collections can make supervised models costly to update and dependent on fixed label sets, motivating retrieval-based genomic classification as a simpler, more flexible alternative. We present a unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification across three viral classification tasks: hepatitis C virus (HCV) genotyping, COVID-19 discrimination, and human papillomavirus (HPV) genotyping. We compare standard sequence encodings (one-hot, k-mers, FCGR) with dense embeddings (dna2vec, DNABERT). For each representation, we evaluate supervised classifiers (Random Forest, Decision Tree, XGBoost) and retrieval-based classification, where sequence vectors are indexed with FAISS and labels are assigned via similarity-weighted k-NN. Furthermore, we benchmark multiple FAISS index types (Flat, IVF, HNSW, IVFPQ, OPQ) to characterize accuracy-speed-memory trade-offs at scale. The results show that XGBoost and retrieval using Flat or IVF indexes achieve strong classification performance under different computational profiles. Compressed indexes such as IVFPQ and OPQ substantially reduce memory usage, although their accuracy loss depends on the dataset and representation. Overall, supervised XGBoost provides a favorable accuracy-size trade-off, while retrieval-based classification remains competitive and allows labeled reference sequences to be incorporated without retraining a global classifier. This benchmark provides practical guidance for selecting sequence representations, classifiers, and vector-search indexes under different accuracy, memory, and update requirements.

Genome, Viral↗

MIPS bacterial genomes functional annotation benchmark dataset.

MOTIVATION: Any development of new methods for automatic functional annotation of proteins according to their sequences requires high-quality data (as benchmark) as well as tedious preparatory work to generate sequence parameters required as input data for the machine learning methods. Different program settings and incompatible protocols make a comparison of the analyzed methods difficult. RESULTS: The MIPS Bacterial Functional Annotation Benchmark dataset (MIPS-BFAB) is a new, high-quality resource comprising four bacterial genomes manually annotated according to the MIPS functional catalogue (FunCat). These resources include precalculated sequence parameters, such as sequence similarity scores, InterPro domain composition and other parameters that could be used to develop and benchmark methods for functional annotation of bacterial protein sequences. These data are provided in XML format and can be used by scientists who are not necessarily experts in genome annotation. AVAILABILITY: BFAB is available at http://mips.gsf.de/proj/bfab

Bacterial Proteins↗

LAMBDA: a prophage detection benchmark for genomic language models.

Transformer-based genomic sequence models represent an emerging frontier in computational biology. Yet, their embeddings have not yet shown the same level of predictive power as natural and protein language models, highlighting a gap between current implementations and theoretical promise. Existing benchmarks for DNA language models primarily focus on classifying regulatory elements in eukaryotic genomes, leaving open the fundamental question of whether these models learn sequence-level features across whole genomes. We introduce LAMBDA, a benchmark designed to rigorously evaluate genome language model embeddings through phage-bacteria sequence discrimination across four categories of increasing complexity: probing tasks, fine-tuning assessments, diagnostic tests, and genome-wide prophage detection. Our comprehensive analysis of current genomic language models provides insight into the importance of training data selection relative to model size, the need for domain-specific training, and the capabilities and limitations of genomic language models for detecting prophage sequences. This benchmark represents a challenging genomic annotation task in the bacterial domain and addresses a key computational problem with direct relevance to microbiology and medicine.

Prophages↗

Benchmarks of support in internal medicine residency training programs.

PURPOSE: To identify benchmarks of financial and staff support in internal medicine residency training programs and their correlation with indicators of quality. METHOD: A survey instrument to determine characteristics of support of residency training programs was mailed to each member program of the Association of Program Directors of Internal Medicine. Results were correlated with the three-year running average of the pass rates on the American Board of Internal Medicine certifying examination using bivariate and multivariate analyses. RESULTS: Of 394 surveys, 287 (73%) were completed: 74% of respondents were program directors and 20% were both chair and program director. The mean duration as program director was 7.5 years (median = 5), but it was significantly lower for women than for men (4.9 versus 8.1; p =.001). Respondents spent 62% of their time in educational and administrative duties, 30% in clinical activities, 5% in research, and 2% in other activities. Most chief residents were PGY4s, with 72% receiving compensation additional to base salary. On average, there was one associate program director for every 33 residents, one chief resident for every 27 residents, and one staff person for every 21 residents. Most programs provided trainees with incremental educational stipends, meals while oncall, travel and meeting expenses, and parking. Support from pharmaceutical companies was used for meals, books, and meeting expenses. Almost all programs provided meals for applicants, with 15% providing travel allowances and 37% providing lodging. The programs' board pass rates significantly correlated with the numbers of faculty fulltime equivalents (FTEs), the numbers of resident FTEs per office staff FTEs, and the numbers of categorical and preliminary applications received and ranked by the programs in 1998 and 1999. Regression analyses demonstrated three independent predictors of the programs' board pass rates: number of faculty (a positive predictor), percentage of clinical work performed by the program director (a negative predictor), and financial support from pharmaceutical companies (also a negative predictor). CONCLUSIONS: These results identify benchmarks of financial and staff support provided to internal medicine residency programs. Some of these benchmarks are correlated with board pass rate, an accepted indicator of quality in residency training. Program directors and chairs can use this information to identify areas that may benefit from enhanced financial and administrative support.

Benchmarking↗

Benchmarking Veterans Affairs Medical Centers in the delivery of preventive health services: comparison of methods.

OBJECTIVE: To identify consistent provision of clinical preventive services, we sought to benchmark all acute care Veterans Affairs Medical Centers (VAMCs) against each other nationally on the basis of multiple evidence-based, performance measures to identify facilities performing consistently higher and lower than expected. METHODS: The 1998 Veterans Health Survey assessed the self-reported delivery of evidence-based clinical preventive services in a stratified national sample of 450 ambulatory care patients seen at each VAMC. Proportions appropriately receiving each service within the recommended time interval were calculated for 138 VAMCs. Percentile ranks for each outcome were assigned. Two approaches were used for benchmarking performance. First, a scaled score for each facility was calculated across the set of 12 measures. Second, facilities were ranked based on the sum of the percentile ranks over a range of specific high cutoffs (eg, 70-80%) and above a range of lower cutoffs (eg, 40-50%). Ranking was validated by comparing with deciles of ranks on chart audit (External Peer Review Program, EPRP) data using Kendall's tau-b and chi2 quality-of-fit test. Differences between consistently high adherence (CHA) and low adherence (CLA) facilities were compared using the Wilcoxon rank sum test on 14 VHS and 11 EPRP outcomes. RESULTS: Data from 39,939 patients (67% response rate) were examined. In combination, cutoffs of greater than 50th percentile and greater than 75th percentile rank yielded 12 of 14 VHS and 6 of 11 EPRP measures different between CHA and CLA facilities. The scaled-score approach resulted in 20 CHA and 14 CLA facilities. The sum of outcomes ranked above 50th percentile and over 75th percentile for CHA facilities (n = 17) was 15 or more. The sum of outcomes ranked above the same cutoffs for CLA facilities (n = 16) was 3 or less. EPRP and 1998 VHS data demonstrated that the survey measures and benchmarking approaches were both reliable and valid. Both approaches resulted in multiple differences between CHA and CLA facilities; differences were greater using the percentile rank approach. CONCLUSIONS: The VA has successfully encouraged adoption of evidence-based clinical preventive services throughout its health care system. However, facilities show wide variation in their levels of delivery and can be distinguished on the basis of their consistently high or low levels of adherence. Examining service delivery across multiple performance indicators allows identification of opportunities to improve clinical practice guideline implementation and the delivery of preventive services. This approach identifies model institutions where focused investigation of factors associated with consistent performance may be particularly fruitful.

Adult↗

Benchmarking best practices in Web-based nursing courses.

This article describes the framework and process to determine best practices in online learning communities for Web-based nursing courses. The benchmarks for best practices were determined based on evidence-based research in higher education. These quality indicators were then used to develop and pilot test a benchmarking survey across three state schools of nursing. The results of the pilot test, as well as the applications and implications for benchmarking best practices, are discussed.

Adult↗

Benchmarking applications: linking state strategic planning, quality improvement, and consumer reporting.

This article demonstrates the value of using benchmark patient satisfaction data for Medicaid program quality improvement. The authors compare surveys of Maryland Medicaid and federal employees in Maryland, utilizing the latter as an external benchmark. Unadjusted and adjusted analyses found a significantly lower percentage of Medicaid than federal respondents rated telephone access excellent, very good, or good, whereas more Medicaid respondents rated advice on prevention and choice of primary care doctor highly. Patient satisfaction external benchmark data provide managed care organizations (MCOs) and state policy makers with goals to improve quality and standards to measure care objectively in vulnerable populations.

Benchmarking↗

Toward benchmarks for stroke rehabilitation in Ontario, Canada.

OBJECTIVE: Canadian benchmarking data do not exist for stroke rehabilitation services. This study used the FIM-function-related group (FIM-FRG) classification system to group patients and to describe the outcomes within each group. The intent was to begin to develop benchmarks for persons recovering from stroke in Canadian rehabilitation facilities. DESIGN: 561 patients were stratified into the nine categories of the FIM-FRG system. Length of stay (LOS), total FIM gain, total FIM at discharge, and discharge location were described for each category. RESULTS: Mean waiting time to rehabilitation admission was 29.7 days. Mean LOS was 49.2 days. Mean admission and discharge total FIM ratings were 78.1 and 103.1, respectively. FIM gain ranged from 8 to 37. Seventeen percent of patients were discharged to nursing homes, with rates ranging from a low of 0% (FRG 8 and 9) to a high of 60% (FRG 2). CONCLUSIONS: For the nine FIM-FRG groups, LOS was considerably longer in the Canadian facility than in the United States, and total FIM score at discharge was higher in Canada. This is likely related to differences in the healthcare systems of the two countries and confirms the need to develop benchmarks based on Canadian data.

Activities of Daily Living↗

Dental nomograms for benchmarking based on the study of health in Pomerania data set.

AIM: Benchmarking is a means of setting goals or targets. On an oral health level, it denotes retaining more teeth and/or improving the quality of life. The goal of this pilot investigation was to assess whether the data generated by a population-based study (SHIP 0) can be used as a benchmark data set to characterize different practice profiles. MATERIAL AND METHODS: The data collected in the population-based study SHIP (n=4310) in eastern Germany were used to generate nomograms of tooth loss, attachment loss, and probing depth. The nomograms included twelve 5-year age strata (20-79 years) presented as quartiles, and additional percentiles of the dental parameters for each age group. Cross-sectional data from a conventional dental office (n=186) and from a periodontology unit (n=130, Greifswald) in the study region as well as longitudinal data set of a another periodontology unit (n=135, Kiel) were utilized in order to verify whether the given practice profile was accurately reflected by the nomogram. RESULTS: In terms of tooth loss, the data from the conventional dental office agree with the median from the nomogram. For attachment loss and probing depth, some age groups yielded slight but not uniform deviations from the median. Cross-sectional data from the periodontology unit Greifswald showed attachment loss higher than the median in younger but not in older age groups. The probing depth was uniformly less than the median and tended toward the 25th percentile with increasing age. The longitudinal data of the Unit of Periodontology in Kiel showed a pronounced trend towards higher percentiles of residual teeth, meaning that the patients retained more teeth. CONCLUSION: The profile of the Pomeranian dental office does not deviate noticeably from the population-based nomograms. The higher attachment loss of the Unit of Periodontology in Greifswald in younger age strata clearly reflects their selection because of periodontal disease; the combination of higher attachment loss and decreased probing depth may reflect the success of the treatment. The tendency of attachment loss towards the median with increasing age may indicate that the Unit of Periodontology in Greifswald does not fulfill its function as a special care unit in the older subjects. The longitudinal data set of the Unit of Periodontology in Kiel impressively reflects the potential of population-based data sets as a means for benchmarking. Thus, nomograms can help to determine the practice profile, potentially yielding benefits for the dentist, health insurance company, or--as in the case of the special care unit--public health research.

Adult↗

Performance benchmarks for diagnostic mammography.

PURPOSE: To evaluate a range of performance parameters pertinent to the comprehensive auditing of diagnostic mammography examinations, and to derive performance benchmarks therefrom, by pooling data collected from large numbers of patients and radiologists that are likely to be representative of mammography practice in the United States. MATERIALS AND METHODS: Institutional review board approval was met, informed consent was not required, and this study was Health Insurance Portability and Accountability Act compliant. Six mammography registries contributed data to the Breast Cancer Surveillance Consortium (BCSC), providing patient demographic and clinical information, mammogram interpretation data, and biopsy results from defined population-based catchment areas. The study involved 151 mammography facilities and 646 interpreting radiologists. The study population included women 18 years of age or older who underwent at least one diagnostic mammography examination between 1996 and 2001. Collected data were used to derive mean performance parameter values, including abnormal interpretation rate, positive predictive value (for abnormal interpretation, biopsy recommended, and biopsy performed), cancer diagnosis rate, invasive cancer size, and the percentages of minimal cancers, axillary node-negative invasive cancers, and stage 0 and I cancers. Additional benchmarks were derived for these performance parameters, including 10th, 25th, 50th (median), 75th, and 90th percentile values. RESULTS: The study involved 332,926 diagnostic mammography examinations. Mean performance parameter values were abnormal interpretation rate, 8.0%; positive predictive value for abnormal interpretation, 31.4%; positive predictive value for biopsy recommended, 31.5%; positive predictive value for biopsy performed, 39.5%; cancer diagnosis rate, 25.3 per 1000 examinations; invasive cancer size, 20.2 mm; percentage of minimal cancers, 42.0%; percentage of axillary node-negative invasive cancers, 73.6%; and percentage of stage 0 and I cancers, 62.4%. CONCLUSION: The presented BCSC outcomes data and performance benchmarks may be used by mammography facilities and individual radiologists to evaluate their own performance for diagnostic mammography as determined by means of periodic comprehensive audits.

Adult↗

[Consensus on a process of benchmarking in primary care in Barcelona].

OBJECTIVE: To define the strategy, the conceptual framework, the methodology and the indicators that are needed to promote and consolidate the culture of external reference (benchmarking) as a strategy for change in Primary Care teams (PCT). DESIGN: Cross-sectional, descriptive study. SETTING: Primary care services of the Barcelona City Health Region. METHOD: Two stages were distinguished. At the first stage, an adviser group was set up. This was divided into 4 focus groups in which the main lines, the conceptual framework, the sizes, the indicators and the methodology for comparing PCTs were agreed. The second stage, that of prioritization, was conducted by means of a questionnaire to opinion-formers. For each of the indicators proposed, they appraised the degree of agreement, the suitability and relevance of indicators, the capacity of PC to modify results and the practicality of the information for composing the indicators. RESULTS: The involvement of professionals, their approach to improvement, and the transparency and dissemination of the evaluation were identified as strategic elements of benchmarking dynamics. In line with the basic principles of PC and the health system, 6 dimensions for evaluation were set: accessibility, effectiveness, capacity to resolve problems, longitudinality, cost-efficiency, and results. 43 of the 57 indicators prioritized gained the consensus of over 90% of the consultants. CONCLUSIONS: Evaluation as a useful tool for managing PC quality has to generate improvements or changes in PCTs. The involvement of professionals in the design and development of evaluation may help both its acceptance and the implementation of the changes arising from it. The indicators used and the effect of benchmarking policy on the results of PC service delivery require evaluation.

Benchmarking↗

OXBench: a benchmark for evaluation of protein multiple sequence alignment accuracy.

BACKGROUND: The alignment of two or more protein sequences provides a powerful guide in the prediction of the protein structure and in identifying key functional residues, however, the utility of any prediction is completely dependent on the accuracy of the alignment. In this paper we describe a suite of reference alignments derived from the comparison of protein three-dimensional structures together with evaluation measures and software that allow automatically generated alignments to be benchmarked. We test the OXBench benchmark suite on alignments generated by the AMPS multiple alignment method, then apply the suite to compare eight different multiple alignment algorithms. The benchmark shows the current state-of-the art for alignment accuracy and provides a baseline against which new alignment algorithms may be judged. RESULTS: The simple hierarchical multiple alignment algorithm, AMPS, performed as well as or better than more modern methods such as CLUSTALW once the PAM250 pair-score matrix was replaced by a BLOSUM series matrix. AMPS gave an accuracy in Structurally Conserved Regions (SCRs) of 89.9% over a set of 672 alignments. The T-COFFEE method on a data set of families with <8 sequences gave 91.4% accuracy, significantly better than CLUSTALW (88.9%) and all other methods considered here. The complete suite is available from http://www.compbio.dundee.ac.uk. CONCLUSIONS: The OXBench suite of reference alignments, evaluation software and results database provide a convenient method to assess progress in sequence alignment techniques. Evaluation measures that were dependent on comparison to a reference alignment were found to give good discrimination between methods. The STAMP Sc Score which is independent of a reference alignment also gave good discrimination. Application of OXBench in this paper shows that with the exception of T-COFFEE, the majority of the improvement in alignment accuracy seen since 1985 stems from improved pair-score matrices rather than algorithmic refinements. The maximum theoretical alignment accuracy obtained by pooling results over all methods was 94.5% with 52.5% accuracy for alignments in the 0-10 percentage identity range. This suggests that further improvements in accuracy will be possible in the future.

Amino Acid Sequence↗

Benchmarking tools for the alignment of functional noncoding DNA.

BACKGROUND: Numerous tools have been developed to align genomic sequences. However, their relative performance in specific applications remains poorly characterized. Alignments of protein-coding sequences typically have been benchmarked against "correct" alignments inferred from structural data. For noncoding sequences, where such independent validation is lacking, simulation provides an effective means to generate "correct" alignments with which to benchmark alignment tools. RESULTS: Using rates of noncoding sequence evolution estimated from the genus Drosophila, we simulated alignments over a range of divergence times under varying models incorporating point substitution, insertion/deletion events, and short blocks of constrained sequences such as those found in cis-regulatory regions. We then compared "correct" alignments generated by a modified version of the ROSE simulation platform to alignments of the simulated derived sequences produced by eight pairwise alignment tools (Avid, BlastZ, Chaos, ClustalW, DiAlign, Lagan, Needle, and WABA) to determine the off-the-shelf performance of each tool. As expected, the ability to align noncoding sequences accurately decreases with increasing divergence for all tools, and declines faster in the presence of insertion/deletion evolution. Global alignment tools (Avid, ClustalW, Lagan, and Needle) typically have higher sensitivity over entire noncoding sequences as well as in constrained sequences. Local tools (BlastZ, Chaos, and WABA) have lower overall sensitivity as a consequence of incomplete coverage, but have high specificity to detect constrained sequences as well as high sensitivity within the subset of sequences they align. Tools such as DiAlign, which generate both local and global outputs, produce alignments of constrained sequences with both high sensitivity and specificity for divergence distances in the range of 1.25-3.0 substitutions per site. CONCLUSION: For species with genomic properties similar to Drosophila, we conclude that a single pair of optimally diverged species analyzed with a high performance alignment tool can yield accurate and specific alignments of functionally constrained noncoding sequences. Further algorithm development, optimization of alignment parameters, and benchmarking studies will be necessary to extract the maximal biological information from alignments of functional noncoding DNA.

Animals↗

Benchmark concentrations for methylmercury obtained from the Seychelles Child Development Study.

Methylmercury is a neurotoxin at high exposures, and the developing fetus is particularly susceptible. Because exposure to methylmercury is primarily through fish, concern has been expressed that the consumption of fish by pregnant women could adversely affect their fetuses. The reference dose for methylmercury established by the U.S. Environmental Protection Agency was based on a benchmark analysis of data from a poisoning episode in Iraq in which mothers consumed seed grain treated with methylmercury during pregnancy. However, exposures in this study were short term and at much higher levels than those that result from fish consumption. In contrast, the Agency for Toxic Substances and Disease Registry (ATSDR) based its proposed minimal risk level on a no-observed-adverse-effect level (NOAEL) derived from neurologic testing of children in the Seychelles Islands, where fish is an important dietary staple. Because no adverse effects from mercury were seen in the Seychelles study, the ATSDR considered the mean exposure in the study to be a NOAEL. However, a mean exposure may not be a good indicator of a no-effect exposure level. To provide an alternative basis for deriving an appropriate human exposure level from the Seychelles study, we conducted a benchmark analysis on these data. Our analysis included responses from batteries of neurologic tests applied to children at 6, 19, 29, and 66 months of age. We also analyzed developmental milestones (age first walked and first talked). We explored a number of dose-response models, sets of covariates to include in the models, and definitions of background response. Our analysis also involved modeling responses expressed as both continuous and quantal data. The most reliable analyses were considered to be represented by 144 calculated lower statistical bounds on the benchmark dose (BMDLs; the lower statistical bound on maternal mercury hair level corresponding to an increase of 0.1 in the probability of an adverse response) derived from the modeling of continuous responses. The average value of the BMDL in these 144 analyses was 25 ppm mercury in maternal hair, with a range of 19 to 30 ppm.

Animals↗

Benchmark dose for cadmium-induced renal effects in humans.

OBJECTIVES: Our goal in this study was to explore the use of a hybrid approach to calculate benchmark doses (BMDs) and their 95% lower confidence bounds (BMDLs) for renal effects of cadmium in a population with low environmental exposure. METHODS: Morning urine and blood samples were collected from 820 Swedish women 53-64 years of age. We measured urinary cadmium (U-Cd) and tubular effect markers [N-acetyl-beta-d-glucosaminidase (NAG) and human complex-forming protein (protein HC) ] in 790 women and estimated glomerular filtration rate (GFR; based on serum cystatin C) in 700 women. Age, body mass index, use of nonsteroidal anti-inflammatory drugs, and blood lead levels were used as covariates for estimated GFR. BMDs/BMDLs corresponding to an additional risk (benchmark response) of 5 or 10% were calculated (the background risk at zero exposure was set to 5%) . The results were compared with the estimated critical concentrations obtained by applying logistic models used in previous studies on the present data. RESULTS: For both NAG and protein HC, the BMDs (BMDLs) of U-Cd were 0.5-1.1 (0.4-0.8) microg/L (adjusted for specific gravity of 1.015 g/mL) and 0.6-1.1 (0.5-0.8) microg/g creatinine. For estimated GFR, the BMDs (BMDLs) were 0.8-1.3 (0.5-0.9) microg/L adjusted for specific gravity and 1.1-1.8 (0.7-1.2) microg/g creatinine. CONCLUSION: The obtained benchmark doses of U-Cd were lower than the critical concentrations previously reported. The critical dose level for glomerular effects was only slightly higher than that for tubular effects. We suggest that the hybrid approach is more appropriate for estimation of the critical U-Cd concentration, because the choice of cutoff values in logistic models largely influenced the obtained critical U-Cd.

Benchmarking↗

Clinical practice benchmarking: implications for tissue viability.

Healthcare provision has been described as a 'lottery', reflecting a deficiency in equity of care (Ellis, 2000). Recent Department of Health (DoH, 1998, 1999) attempt to overcome this disparity and assure quality care by quality assessment and continuous quality improvement (Ellis and Morris, 1997). These documents advocate clinical benchmarking as a means to supporting this practice. This article provides an overview of the benchmarking process, with particular focus being applied to the pressure ulcer element. It reflects on the coalescence that appears to exist between implementing clinical benchmarking and the characteristics of specialist practice. It also analyses the effects of establishing the process on an existing tissue viability service.

Benchmarking↗

Surgical wound benchmark tool and best practice guidelines.

Surgical wounds that heal by primary intention are expected to heal without complication (Watret and White, 2001). Patients are frequently subjected to a variety of treatment regimens, often based on individual practitioners' preferences. This article discusses how one acute hospital trust developed a multidisciplinary approach to devise best practice guidelines. This was achieved through consensus and expertise of a working party, and a clinical practice benchmark tool for patients with surgical wounds to standardize and ensure the implementation of evidence-based practice. Clinical practice benchmarking is "a process through which best practice is identified and continuous improvement pursued though comparison and sharing" (Department of Health (DoH), 1999). The work has led to the development of a standardized assessment and documentation tool, which the working party hopes will be used trust-wide. In addition, ward staff are encouraged to undertake the benchmark process as a method of identifying areas where the use of this tool would ensure that standards of care could be improved.

Acute Disease↗

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models↗