Search PubMedSearch

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Identification of variables needed to risk adjust outcomes of coronary interventions: evidence-based guidelines for efficient data collection.

OBJECTIVES: Our objectives were to identify and define a minimum set of variables for interventional cardiology that carried the most statistical weight for predicting adverse outcomes. Though "gaming" cannot be completely avoided, variables were to be as objective as possible and reproducible and had to be predictive of outcome in current databases. BACKGROUND: Outcomes of percutaneous coronary interventions depend on patient risk characteristics and disease severity and acuity. Comparing results of interventions has been difficult because definitions of similar variables differ in databases, and variables are not uniformly tracked. Identifying the best predictor variables and standardizing their definitions are a first step in developing a universal stratification instrument. METHODS: A list of empirically derived variables was first tested in eight cardiac databases (158,273 cases). Three end points (in-hospital death, in-hospital coronary artery bypass graft surgery, Q wave myocardial infarction) were chosen for analysis. Univariate and multivariate regression models were used to quantify the predictive value of the variable in each database. The variables were then defined by consensus by a panel of experts. RESULTS: In all databases patient demographics were similar, but disease severity varied greatly. The most powerful predictors of adverse outcome were measures of hemodynamic instability, disease severity, demographics and comorbid conditions in both univariate and multivariate analyses. CONCLUSIONS: Our analysis identified 29 variables that have the strongest statistical association with adverse outcomes after coronary interventions. These variables were also objectively defined. Incorporation of these variables into every cardiac dataset will provide uniform standards for data collected. Comparisons of outcomes among physicians, institutions and databases will therefore be more meaningful.

Adult

Calculation of benchmark doses from teratology data.

The benchmark dose approach has several potential advantages over the no observed adverse effect level (NOAEL) as a basis for risk assessment of toxic chemicals, based upon animal toxicity data. The practical use of the benchmark dose has been evaluated by applying dose-response models to an extensive historical database of teratology bioassays. Doses corresponding to 1 and 5% increases in incidence of lesions are calculated and compared to NOAELs. The statistical accuracy of these estimates was determined by calculating confidence intervals. The lower confidence limit on the 5% benchmark dose (LED05) is found to be comparable to the NOAEL for most datasets, and slightly higher on average. Benchmark doses at the 1% level could not be estimated accurately (i.e., they had wide confidence intervals) for a significant fraction of the datasets. LED01 values were lower on average than the NOAEL. Based on these results, it is concluded that benchmark doses for a 5% increases in incidence can be calculated for most datasets, and could be used as a satisfactory basis for risk assessment, e.g., to set reference doses or acceptable daily intakes. An exception occurs when the benchmark dose exceeds the highest dose of the study. This is only likely to occur when the chemical causes a small, but significant, increase in a finding that is uncommon in untreated animals.

Abnormalities, Drug-Induced

The Evolutionary Significance of Leaf Nodulation: Evidence from Ardisia and Its Relatives (Primulaceae: Myrsinoideae).

Interactions between plants and microorganisms have long been a central topic in biological research. Bacterial symbiosis on leaf surfaces represents a distinctive and mutually beneficial system within the phyllosphere microbiome. Leaf nodules are the visible manifestation of the symbiosis and confer ecological advantages to host plants by enhancing host resistance against pathogens and herbivores. It has been hypothesized that these advantages promote higher diversification rates in host lineages, but this remains uncertain. Ardisia subg. Crispardisia and its close relatives (Amblyanthopsis and Amblyanthus) within Primulaceae are typical plant groups with leaf nodule symbiosis, making them an ideal system for testing this hypothesis. In this study, we conducted extensive sampling of "Ardisioids" (Ardisia and its allies) and reconstructed their phylogenetic relationships and evolutionary history using plastid genomes and nuclear datasets (i.e., nuclear ribosomal DNA (nrDNA) and genome-wide single nucleotide polymorphisms (SNPs)). We clarified the phylogenetic positions of several "Ardisioids" genera (e.g., Sadiria, Tapeinosperma, Amblyanthus, and Amblyanthopsis) and multiple subgenera within Ardisia. We further detected a rapid radiation during the middle Miocene in Ardisia and its allies. Notably, we found that the leaf-nodulated clade appears to have originated during this period, approximately 11-8 Ma. BAMM (Bayesian Analysis of Macroevolutionary Mixtures) analyses revealed elevated diversification rates in leaf-nodulated lineages, while HiSSE (Hidden State Speciation and Extinction) analyses indicated that leaf nodule symbiosis might have increased speciation rates without significantly affecting extinction rates. These results provide strong evidence that leaf nodule symbiosis, together with other abiotic and biotic factors, represents a key evolutionary innovation that has promoted diversification in Ardisia and its close relatives.

diversification rate

An approach to the use of the DO IT Study Group guidelines for supporting the optimal implementation of information systems in diabetes care.

The usefulness of the 1992 DO IT Study Group guidelines for diabetes data information systems was assessed using two established diabetes databases designed for different purposes. The recommendations detailed in the guidelines, written in four separate but overlapping modules, were applied individually to each database in turn. Percentage compliance with the recommendation to collect the DIABCARE dataset was high, after discounting specialist areas. While on the whole the information systems complied with the guidelines within the purposes for which they were designed, areas highlighted as demanding further action in at least one of the two systems included password protection, data validation checks, screen design, and communication with those whose records were held on the systems. Application of the guidelines is already stimulating attention to some of these areas. Some of the guidelines proved rather vague in construction to be applied in any formal sense, while others (for example in relation to international accreditation of datasets) were not applicable to individual systems. The results suggest the importance of a structured approach to the design, development, and ongoing assessment of information systems in diabetes, but require the present guidelines to develop a more formal structure to be fully effective. The widespread adoption and further testing and refinement of these guidelines (both within and outside Europe) should promote the ultimate goal of improved diabetes care.

Communication

The NCIC-Manitoba Breast Tumor Bank: a resource for applied cancer research.

The NCIC-Manitoba Breast Tumor Bank is one of several tumour banks that have been established through the Molecular Epidemiology Program of the National Cancer Institute of Canada (NCIC). The NCIC-Manitoba Breast Tumor Bank is an example of one model developed to facilitate research designed to translate the findings of basic science into information useful in the clinical arena. The tumour bank's mandate is to provide a national resource that consists of a preassembled dataset of matched samples of paraffin-embedded and frozen tumour tissue with corresponding pathological and clinical data. In the first 3 years the tumour bank has accrued data and samples from over 1800 cases of breast cancer and has provided support for 20 research projects across Canada and the United States.

Academies and Institutes

Artificial intelligence for dental caries detection: An umbrella review.

Artificial intelligence (AI) has been proposed as a tool to improve dental caries detection across imaging modalities; however, its clinical value remains uncertain. This umbrella review aimed to synthesize and critically appraise systematic reviews evaluating AI for caries detection and diagnosis. An umbrella review was conducted following PRIOR guidance (PROSPERO CRD420261340728). Searches were performed in MEDLINE, Embase, Scopus, Web of Science, and Google Scholar up to 15 March 2026. Methodological quality was assessed using AMSTAR 2, and overlap of primary studies was quantified using the corrected covered area (CCA). Seventeen systematic reviews were included, of which five reported diagnostic test accuracy meta-analyses using bivariate or HSROC models. Across these meta-analyses, pooled sensitivity ranged from 0.76 to 0.94 and specificity from 0.85 to 0.91. Most systems were based on deep learning models applied to bitewing radiographs and intraoral photographs. However, substantial heterogeneity was observed in imaging modalities, lesion thresholds, analytical tasks, and evaluation metrics. In addition, a high degree of overlap across reviews and recurrent methodological limitations, including reliance on retrospective datasets, limited external validation, and inconsistent reporting, substantially weaken the reliability of the evidence. Although AI models demonstrate high diagnostic performance under experimental conditions, current evidence does not support their use as stand-alone diagnostic tools. Their clinical applicability remains limited, and implementation should be restricted to decision-support contexts until robust prospective validation demonstrates meaningful impact on clinical decision-making and patient outcomes.

Dental Caries

The hospital information system as a source for the planning and feed-back of specialized health care.

1. INTRODUCTION. In university hospitals, choices are made to which extend specialized health care will be supported. It is characteristic, for this type of care, that it takes place in a process of the continual advance of medical technology and the growing awareness by consumers and payors. Specialized healthcare contributes to the hospital qualifiers having a political and strategic impact. The hospital board needs information for planning and budgeting these new tasks. Much of the information will be based on data stored in the Hospital Information System (HIS). Due to load limitations, instant retrieval is not preferred. A separate executive information system, uploaded with HIS data, features statistics, on a corporate level, with the power to drill-down to detailed levels. However, the ability to supply information on new types of healthcare is limited since most of these topics require a flexible system for new dedicated cross-sections, like medical treatment from several specialisms and functional levels. 2. DATA RETRIEVAL AND DISTRIBUTION. During the information analysis, details were gathered on the necessary working procedures and the administrative organization, including the data registration in the HIS. In the next phase, all relevant data was organized in a relational datamodel. For each topic of care, dedicated views were developed at both low and high aggregation levels. It revealed that a matching change of the administrative organization was required, with an emphasis on financial registration aspects. For the selection of relevant data, a bottom-up approach was applied, which was based on the registrations starting from the patient administrative subsystem, through several transactional systems, ending at the general ledger in the HIS. Data on all levels was gathered, resulting in medical details presented in quantities, up to financial figures expressed in amounts of money. This procedure distinguishes from the predefined top-down techniques generally used for management and executive information systems. Data was regularly collected from the HIS, then converted and reorganized into relational datasets using XBase protocols. After having performed central quality controls and privacy protection measures, the datasets were distributed electronically to local PCs. Standard low-cost software packages enable analyses by user-friendly selection and presentation facilities. 3. EVALUATION. The method developed for data retrieval is flexible and easy to implement. If all basic data is registered in the HIS, the procedure can be applied for all strategic hospital functions that require planning and controlling during a certain time. Critical success factors and pitfalls will be presented in the poster. Using one consistent dataset, the information required about production and budget is presented at several functional levels and is quantified in units familiar to that level. The motivation for fast and accurate registration in the HIS was improved from the moment the medical and administrative staff recognized their own data in the feed-back on specialized health care.

Decision Making, Organizational

On the use of a hospital information system in evaluating clinical care: a case report.

In this paper we describe, as an example, how we obtained the information needed to evaluate a newly introduced protocol for ordering X-rays for ankle trauma patients. Extensive use was made of available data and facilities of the hospital information system (HIS). Procedures for collecting the required additional data, which were not recorded in the HIS but were needed to evaluate the protocol, were embedded in the current medical and administrative routine of the emergency room. These additional data were also stored in the HIS. Periodically all data were downloaded to a personal computer to analyse the impact of using the protocol on quality of care and costs. In total 1241 patients entered the study, and for 1149 patients a complete dataset was obtained. The sensitivity and specificity of the protocol at the threshold value which was used during the initial study period was 0.77 and 0.80. The reduction in the number of ankle X-rays due to the protocol was significant when compared with a strategy of ordering an X-ray for every ankle trauma patient visiting the emergency room.

Ankle Injuries

Optimum numerical integration methods for estimation of area-under-the-curve (AUC) and area-under-the-moment-curve (AUMC).

Eleven numerical methods for estimation of AUC (including 4 new methods) and 22 methods for AUMC (including 8 new methods) were tested on large simulated noisy datasets representing bolus, oral and infusion concentration-time profiles. Some methods were unacceptable because their mean error was large; these included a commonly recommended form of the linear trapezoidal rule for AUMC. Others, notably Lagrange and cubic spline methods, were unacceptable because the variance of their estimates was large. These methods should be abandoned. A simple and easily programmed new method, parabolas-through-the-origin then log-trapezoidal rule, performed especially well.

Numerical Analysis, Computer-Assisted

Experiments to determine whether recursive partitioning (CART) or an artificial neural network overcomes theoretical limitations of Cox proportional hazards regression.

New computationally intensive tools for medical survival analyses include recursive patitioning (also called CART) and artificial neural networks. A challenge that remains is to better understand the behavior of these techniques in effort to know when they will be effective tools. Theoretically they may overcome limitations of the traditional multivariable survival technique, the Cox proportional hazards regression model. Experiments were designed to test whether the new tools would, in practice, overcome these limitations. Two datasets in which theory suggests CART and the neural network should outperform the Cox model were selected. The first was a published leukemia dataset manipulated to have a strong interaction that CART should detect. The second was a published cirrhosis dataset with pronounced nonlinear effects that a neural network should fit. Repeated sampling of 50 training and testing subsets was applied to each technique. The concordance index C was calculated as a measure of predictive accuracy by each technique on the testing dataset. In the interaction dataset, CART outperformed Cox (P < 0.05) with a C improvement of 0.1 (95% CI, 0.08 to 0.12). In the nonlinear dataset, the neural network outperformed the Cox model (P < 0.05), but by a very slight amount (0.015). As predicted by theory, CART and the neural network were able to overcome limitations of the Cox model. Experiments like these are important to increase our understanding of when one of these new techniques will outperform the standard Cox model. Further research is necessary to predict which technique will do best a priori and to assess the magnitude of superiority.

Databases, Factual

Some conceptual and statistical issues in analysis of longitudinal psychiatric data. Application to the NIMH treatment of Depression Collaborative Research Program dataset.

Longitudinal studies have a prominent role in psychiatric research; however, statistical methods for analyzing these data are rarely commensurate with the effort involved in their acquisition. Frequently the majority of data are discarded and a simple end-point analysis is performed. In other cases, so called repeated-measures analysis of variance procedures are used with little regard to their restrictive and often unrealistic assumptions and the effect of missing data on the statistical properties of their estimates. We explored the unique features of longitudinal psychiatric data from both statistical and conceptual perspectives. We used a family of statistical models termed random regression models that provide a more realistic approach to analysis of longitudinal psychiatric data. Random regression models provide solutions to commonly observed problems of missing data, serial correlation, time-varying covariates, and irregular measurement occasions, and they accommodate systematic person-specific deviations from the average time trend. Properties of these models were compared with traditional approaches at a conceptual level. The approach was then illustrated in a new analysis of the National Institute of Mental Health Treatment of Depression Collaborative Research Program dataset, which investigated two forms of psychotherapy, pharmacotherapy with clinical management, and a placebo with clinical management control. Results indicated that both person-specific effects and serial correlation play major roles in the longitudinal psychiatric response process. Ignoring either of these effects produces misleading estimates of uncertainty that form the basis of statistical tests of hypotheses.

Analysis of Variance

Epigenetic Profiling for Early Detection and Treatment Response Monitoring in Non-Small Cell Lung Cancer: Protocol for a Prospective Translational Biomarker Study.

BACKGROUND: Non-small cell lung cancer (NSCLC) is the leading cause of cancer-related mortality worldwide and continues to have poor survival outcomes, with most patients diagnosed at advanced stages of disease. In New Zealand, NSCLC contributes substantially to cancer inequities, with M&#x101;ori communities experiencing disproportionately high incidence and mortality rates. Although low-dose computed tomography screening can improve early detection, major limitations remain, including false-positive findings, overdiagnosis, high infrastructure costs, and limited accessibility for rural and underserved populations. Liquid biopsy approaches using circulating tumor DNA (ctDNA), particularly DNA methylation profiling, have emerged as promising, minimally invasive strategies for improving cancer detection, treatment monitoring, and precision oncology. OBJECTIVE: This study aims to establish integrated genomic and epigenomic predictive and prognostic biomarkers using ctDNA, tumor tissue, and transcriptomic profiling to improve early detection, risk stratification, treatment selection and response prediction, and longitudinal monitoring, with particular emphasis on identifying molecular mechanisms associated with treatment resistance and disease progression. METHODS: This prospective observational translational biomarker study is being conducted through the University of Otago and associated respiratory and oncology services in New Zealand. The study will recruit participants with NSCLC (including squamous and nonsquamous subtypes), individuals referred to fast-track lung nodule assessment clinics, and nonmalignant respiratory controls. Serial peripheral blood sampling will be performed in selected participants at predefined clinical follow-up time points to evaluate treatment response and disease progression. The availability of formalin-fixed paraffin-embedded archival tissues will be recorded, but will not be mandatory for enrollment. Genome-scale DNA methylation profiling will be performed using cell-free reduced representation bisulfite sequencing (cfRRBS), while targeted genomic profiling and transcriptomic analyses will be conducted using targeted sequencing panels and RNA sequencing. Integrative bioinformatic analyses will be used to identify molecular biomarkers associated with early-stage disease, advanced disease, treatment response, and therapeutic resistance. RESULTS: Ethics approval for the study has been obtained from the New Zealand Health and Disability Ethics Committee (2022 EXP 12566). This study commenced in 2022, and recruitment and biospecimen collection are ongoing. The study aims to recruit approximately 450 participants, including patients with NSCLC, individuals referred through respiratory diagnostic pathways, and nonmalignant controls. As of July 31, 2026, 205 participants have been recruited, with recruitment continuing until the target sample size is reached. Molecular and data analyses are ongoing, with additional publications expected as the cohort matures. CONCLUSIONS: This study will generate one of the first integrated genomic, epigenomic, and transcriptomic liquid biopsy datasets for NSCLC in New Zealand. The findings are expected to support the development of sensitive, accessible, and equitable blood-based biomarkers for NSCLC detection and treatment monitoring while also contributing to improved precision oncology approaches and reducing NSCLC inequities among M&#x101;ori populations.

Humans

Evaluating virtual endoscopy for clinical use.

Virtual endoscopy is a term used to describe computer simulated endoscopy procedures derived from high resolution images of patient anatomy. By simulating the endoscopic examination, the patient is spared the discomfort and possible complications of an actual examination. The physician also has more flexibility in a virtual endoscopic examination of 3D patient data in comparison to a real endoscopic examination. Virtual endoscopy removes the physical and physiologic constraints of real endoscopy and can create views that are not possible in an actual endoscopic examination. This may enhance the performance of actual endoscopic examinations. Virtual endoscopy may also be used to perform "numerical biopsies"; anatomic measurements such as size, distance, shape, and density. Virtual endoscopy allows the physician to comprehensively explore the patient anatomy using an intuitive and interactive interface. There are currently two technical approaches to performing virtual endoscopy: perspective volume rendering and surface rendering of polygonal models. Perspective volume rendering uses traditional volumetric rendering algorithms to create visualizations directly from the volumetric dataset. Polygonal models require a preprocessing step to convert the segmented volume information into a polygonal surface that may be displayed at real time frame rates. Both paradigms have inherent strengths and weaknesses. We illustrate and compare the methods on actual patient data, including simulated endoscopic examinations of the airways, colon and esophagus. Preliminary results in virtual endoscopy show promise and will continue to be an area of active research leading to useful clinical applications.

Algorithms

Development of predictive models of laboratory animal growth using artificial neural networks.

Traditional regression analysis of body weight growth curves encounters problems when the data are extremely variable. While transformations are often employed to meet the criteria of the analysis, some transformations are inadequate for normalizing the data. Regression analysis also requires presuppositions regarding the model to be fit and the techniques to be used in the analysis. An alternative approach using artificial neural networks is presented which may be suitable for developing predictive models of growth. Neural networks are simulators of the processes that occur in the biological brain during the learning process. They are trained on the data, developing the necessary algorithms within their internal architecture, and produce a predictive model based on the learned facts. A dataset of Sprague-Dawley rat (Rattus norvegicus) weights is analyzed by both traditional regression analysis and neural network training. Predictions of body weight are made from both models. While both methods produce models that adequately predict the body weights, the neural network model is superior in that it combines accuracy and precision, being less influenced by longitudinal variability in the data. Thus, the neural network provides another tool for researchers to analyze growth curve data.

Algorithms

Automated seed localization from CT datasets of the prostate.

With the increasing utilization of permanent brachytherapy implants for treating carcinoma of the prostate, the importance of accurate post-treatment dose calculation also increases for assessing patient outcome and planning future treatments. An automatic method for seed localization of permanent brachytherapy implants, using CT datasets of the prostate, has been developed and tested on a phantom using an actual patient planned seed distribution. This method was also compared to results with the three-film technique for three patient datasets. The automatic method is as accurate or more accurate than the three film technique for 1 mm, 3 mm, and 5 mm contiguous CT slices, and eliminates the inter- and intra-observer variability of the manual methods. The automated method improves the localization of brachytherapy seeds while reducing the time required for the user to input information, and is demonstrated to be less operator dependent, less time consuming, and potentially more accurate than the three-film technique.

Biophysical Phenomena

The influence of radiotherapy treatment time on the control of laryngeal cancer: a direct analysis of data from two British Institute of Radiology trials to calculate the lag period and the time factor.

This study analyses node-negative laryngeal tumour control data from two clinical trials conducted by the British Institute of Radiology in order to determine the time factors and the presence or absence of a lag period before the time factor takes effect. A direct maximum likelihood approach is used to fit a double-logarithmic model including a repopulation term which commences after an initial lag period, Tk. The analysis yields a time factor of 0.8 Gy per day (95% confidence interval 0.5-1.1 Gy per day) as the extra dose required to counteract the reduction in tumour control probability (TCP) with extension of the treatment time. The latter reduction amounted to between 5 and 12% TCP per week, depending on the stage and time period. With this dataset, where few patients were treated for short times, no statistically significant lag phase can be demonstrated. However, the best estimate of Tk is 21 days (95% confidence interval 0-27 days), which is consistent with estimates from other studies on other datasets. If a lag phase exists, this study would indicate that the duration is less than 27 days. Other studies have used retrospective data and are subject to a number of potential biases. The present study, using data from multicentre prospective randomized clinical trials, is free from some of these sources of bias. The fact that very similar estimates of the radiobiological parameters are obtained lends credence to these other studies and suggests that the potential biases may be small in practice.

Clinical Trials as Topic

Preliminary assessment of three-dimensional magnetic resonance imaging for various colonic disorders.

BACKGROUND: Improvements in magnetic resonance imaging (MRI) technology have enabled the acquisition of three-dimensional MRI datasets in a single breath hold. We adopted this technique to make a three dimensional intraluminal and extraluminal assessment of the colon in three patients with various colonic disorders. METHODS: One patient was studied after having a double-contrast barium enema. Two patients had MRI scans after colonoscopy, which showed three colonic tumours in one and multiple polyps in the ascending colon of the other. The process of rectal filling with 1.5-2.0 L water mixed with 15-20 mL 0.5 mol/L gadolinium-diethylenetriaminepentaacetic acid (Gd-DTPA) was monitored with MR fluoroscopic sequence. Three-dimensional datasets of the contrast-filled colon were taken with patients in prone (before and after intravenous administration of 0.1 mmol/kg bodyweight Gd-DTPA) and supine positions. 64 sections with a voxel-resolution of 2.0 x 2.0 x 1.25 mm3-were taken during a 28 s breath hold. Three-dimensional maximum intensity projection, multiplanar reconstruction, and virtual colonoscopic images of the colon were created from these. FINDINGS: Analysis of the coronal source images in conjunction with multiplanar reconstructions revealed all relevant abnormalities, including diverticula, carcinomas, and polyps. Three dimensional maximum-intensity projections gave a morphological overview of the whole colon. Targeted projections, made up of a limited number of coronal source images, showed diverticula and smaller polyps more clearly. After patients were given intravenous contrast all colonic mass lesions were enhanced. Datasets obtained in prone patients gave the best intraluminal views of the colon. Virtual magnetic resonance colonoscopy showed colonic haustra as well as the ileocaecal valve, but did not show clearly the diverticula. All intraluminal mass lesions, on the other hand, were easy to see. INTERPRETATION: The potential of three-dimensional colonic MRI to provide accurate, minimally invasive, cost-effective polyp screening, as well as comprehensive colonic tumour staging, warrants further investigation.

Aged

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models