Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Benchmark”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

An open benchmark and language models for AI in aging biology.

Over the past two decades, human aging has been characterized across DNA methylation, transcriptomic, proteomic, and clinical modalities, yet no benchmark evaluates whether AI systems can interpret these heterogeneous data types in the context of aging biology. We introduce LongevityBench, an open suite of 17 tasks spanning five biodata domains, and use it to assess 18 frontier AI systems from six developer teams. Despite recent advances in AI, no single model dominates all tasks, with omics-based age prediction being the hardest task regardless of scale. To test whether these gaps can be closed without frontier-scale resources, we fine-tuned a family of five multitask Longevity-LLMs on domain-specific aging data. The compact (0.6B-9B parameters) Longevity-LLMs matched or exceeded far larger frontier systems on LongevityBench, showing that general-purpose language models can be adapted to structured-omics tasks. We publicly release the benchmark, models, and Longevity Claw, an agentic research interface for aging researchers.

Aging↗

A new tool for benchmarking cardiovascular fluoroscopes.

This article reports the status of a new cardiovascular fluoroscopy benchmarking phantom. A joint working group of the Society for Cardiac Angiography and Interventions (SCA&I) and the National Electrical Manufacturers Association (NEMA) developed the phantom. The device was adopted as NEMA standard XR 21-2000, "Characteristics of and Test Procedures for a Phantom to Benchmark Cardiac Fluoroscopic and Photographic Performance," in August 2000. The test ensemble includes imaging field geometry, spatial resolution, low-contrast iodine detectability, working thickness range, visibility of moving targets, and phantom entrance dose. The phantom tests systems under conditions simulating normal clinical use for fluoroscopically guided invasive and interventional procedures. Test procedures rely on trained human observers.

Benchmarking↗

Benchmarking and quality in residential and nursing homes: lessons from the US.

BACKGROUND: Performance measurement and benchmarking are common concerns in the delivery of long term care. It is common to measure the performance of providers and to publicly report these data. This paper examines selected technical challenges facing those who design, implement and disseminate health care quality performance measures. METHOD: Review of the application of measures of performance in the US nursing home sector. RESULTS: Using examples drawn from the skilled nursing home arena, problems ranging from data reliability and validity, the multi-dimensional nature of quality measures and selection bias as well as differential measurement abilities are discussed. CONCLUSIONS: Benchmarking of performance is an inherently complex issue. However, to ensure that such comparisons are both fair and valid requires measures to be more technically sophisticated and sensitive to real changes attributable to changes in care.

Benchmarking↗

Estimation of benchmark dose for renal dysfunction in a cadmium non-polluted area in Japan.

Previously, the association between urinary cadmium (Cd) concentration and indicators of renal dysfunction, including beta(2)-microglobulin (beta(2)-MG), total protein and N-acetyl-beta-D-glucosaminidase (NAG) were investigated in 1270 inhabitants > or = 50 years of age (547 men, 723 women) in a Cd non-polluted area in Japan and showed that a dose-response relationship existed between renal effects and Cd exposure in the general environment without any known Cd pollution. However, the threshold levels of urinary Cd could not be estimated at that time. In the present study, the threshold levels of urinary Cd were estimated as the benchmark dose low (BMDL) using the benchmark dose (BMD) approach. Urinary Cd excretion was divided into 6-7 categories, and an abnormality rate was calculated for each. Cut-off values for urinary substances were defined as corresponding to the 84% upper limit values, which were calculated from 2034 persons who had been living in the non-polluted areas and did not smoke. Then the BMD and BMDL were calculated using a log-logistic model. The values of BMD and BMDL for all urinary substances could be calculated. The BMDL for the 84% cut-off value of beta(2)-MG, setting an abnormal value at 5%, was 2.0 microg g(-1) creatinine (cr) in men and 1.6 microg g(-1) cr in women. In conclusion, the present study demonstrated that the threshold level of urinary Cd could be estimated in people living in the general environment without any known Cd-pollution in Japan, and the value was inferred to be almost the same as that in Belgium and Sweden.

Acetylglucosaminidase↗

Evaluation of the benchmark dose method for dichotomous data: model dependence and model selection.

The benchmark dose (BMD) method was evaluated using the USEPA BMD software. Dose-response data on cleft palate and hydronephrosis for a number of related polyhalogenated aromatic compounds were obtained from the literature. According to chi(2) test statistics, each dichotomous USEPA model failed to adequately describe only 1 of 12 cleft palate data sets. For hydronephrosis, the models were discriminated to a higher extent according to global goodness-of-fit. NOAELs for cleft palate corresponded to BMDLs (the approximate lower confidence limit on the BMD) for extra risks in the range of 5% or below. Model dependence of the BMDL estimate was more pronounced at lower levels of benchmark response (BMR). A BMR of 5% (extra risk) is recommended for cleft palate since model differences at this level were limited for all data. In addition, at BMRs of 5-10% the BMDL for all models was little affected by the specified confidence limit size (in the 90-99% range). For BMDL determination a conservative model selection approach was applied. At the suggested level of BMR (5%) this procedure resulted in use of the same model (multistage model) for the cleft palate endpoint in general. Akaike's information criterion (AIC) was considered for comparison between models. Determination of appropriateness of use of such methods in dose-response applications requires further analysis.

Abnormalities, Drug-Induced↗

The National Outcomes Management Project: a benchmarking collaborative.

Traditional evaluation of health care quality usually involves the measurement of the structure, process, and outcome of care. Most quality improvement programs involve a cycle that includes a setting of goals, a measurement of either process or outcomes, and a real-time or retrospective feedback of the results of data measurement. Benchmarking, a well-known efficient business technology, can lead to practice innovations necessary to survive in an environment that has a need for decreasing cost and increasing quality. The purpose of this article is to present a novel use of benchmarking in managed ambulatory behavioral health care and its application in a model collaborative outcome management project at more than 16 sites and nine states in the United States.

Ambulatory Care Facilities↗

Establishing benchmarks for creation of a pro-forma economic model to evaluate filmless PACS operation.

The purpose of this study was to establish data points (benchmarks) to incorporate into a pro-forma cost analysis model, comparing film-based and filmless modes of operation. Prospective data were collected over a 6-year period at the Baltimore VA Medical Center (BVAMC) immediately before and after implementation of a hospital-wide PACS. These data were in turn compared with local and national VA centers during comparable time periods, to establish reference data between manual film-based (without PACS) and filmless operations (using PACS). Benchmarks utilized for the study fell into 2 broad categories: operational costs and revenues generated. Factors contributing to operational costs include space requirements, equipment, supplies, personnel, and maintenance. Factors contributing to revenues generated included examination volume, modality mix, and reimbursement rates. Collectively, these data points were incorporated into a pro-forma model that allows prospective PACS customers to compare total cost of ownership for film-based and filmless operations dependent on the unique variables of the respective institution.

Benchmarking↗

Useful benchmarks to evaluate outcomes after esophagectomy and pancreaticoduodenectomy.

BACKGROUND: Multiple publications have suggested that outcomes after complex operations are better at high-volume centers. However, of all the potential "outcomes" to measure, only mortality has been studied extensively. The broadest difference in mortality between low- and high-volume centers has been measured after esophagectomy (EG) and pancreaticoduodenectomy (PD). If a low-volume center recorded high mortality, then a broader set of outcomes beyond mortality would be useful for self-assessment. METHODS: Two single-surgeon prospective databases for outcomes of EG and PD were reviewed in a multispecialty clinic within a tertiary-referral, resident-training hospital. Between January 1996 and December 2002, 174 consecutive patients underwent EG performed by 1 surgeon (25 cases/y), and 232 consecutive patients underwent PD performed by another surgeon (34 cases/y). We measured hospital and 30-day mortality rate, mean operation time (OR time), mean estimated intraoperative blood loss (EBL), mean length of stay (LOS), and the anastomotic leak rate. These outcomes were compared with those of recently published cases for EG and PD. RESULTS: Mortality for both operations was zero. After EG, OR time was 394 minutes (literature = 336), EBL was 204 mL (literature = 964), transfusion rate was 3.5% (literature = 34%), LOS was 11.1 days (literature = 16.6), leak was 2.9% (literature = 9.1%), and reoperation was 1.7% (literature = not stated). After PD, OR time was 450 minutes (literature = 431), EBL was 382 mL (literature = 1,183), transfusion rate was 7.3% (literature = not stated), LOS was 11.2 days (literature = 17.8), leak was 6.5% (literature = 9.9%), and reoperation was 0.4% (literature = 3.8%). CONCLUSIONS: These 2 single-surgeon series provide benchmarks to help better define acceptable outcomes after EG and PD. This assessment demonstrated lower mortality and LOS in a high-volume surgical practice. These outcomes are not associated with OR time but with lower EBL, less need for transfusion, and lower need for reoperation. Anastomotic leaks occurred in both series; however, this was not associated with mortality because of early recognition and the use of nonsurgical minimally invasive techniques. If mortality is high at a low-volume center, then the additional benchmarks of this study, in addition to mortality and LOS, could be used to lower mortality through self-assessment by identifying specific outcomes that need improvement.

Adolescent↗

Double-blind evaluation and benchmarking of survival models in a multi-centre study.

Accurate modelling of time-to-event data is of particular importance for both exploratory and predictive analysis in cancer, and can have a direct impact on clinical care. This study presents a detailed double-blind evaluation of the accuracy in out-of-sample prediction of mortality from two generic non-linear models, using artificial neural networks benchmarked against a partial logistic spline, log-normal and COX regression models. A data set containing 2880 samples was shared over the Internet using a purpose-built secure environment called GEOCONDA (www.geoconda.com). The evaluation was carried out in three parts. The first was a comparison between the predicted survival estimates for each of the four survival groups defined by the TNM staging system, against the empirical estimates derived by the Kaplan-Meier method. The second approach focused on the accurate prediction of survival over time, quantified with the time dependent C index (C(td)). Finally, calibration plots were obtained over the range of follow-up and tested using a generalization of the Hosmer-Lemeshow test. All models showed satisfactory performance, with values of C(td) of about 0.7. None of the models showed a systematic tendency towards over/under estimation of the observed survival at tau=3 and 5 years. At tau=10 years, all models underestimated the observed survival, except for COX regression which returned an overestimate. The study presents a robust and unbiased benchmarking methodology using a bespoke web facility. It was concluded that powerful, recent flexible modelling algorithms show a comparative predictive performance to that of more established methods from the medical and biological literature, for the reference data set.

Benchmarking↗

Estimation of benchmark dose as the threshold levels of urinary cadmium, based on excretion of total protein, beta2-microglobulin, and N-acetyl-beta-D-glucosaminidase in cadmium nonpolluted regions in Japan.

Previously, we investigated the association between urinary cadmium (Cd) concentration and indicators of renal dysfunction, including total protein, beta2-microglobulin (beta2-MG), and N-acetyl-beta-D-glucosaminidase (NAG). In 2778 inhabitants 50 years of age (1114 men, 1664 women) in three different Cd nonpolluted areas in Japan, we showed that a dose-response relationship existed between renal effects and Cd exposure in the general environment without any known Cd pollution. However, we could not estimate the threshold levels of urinary Cd at that time. In the present study, we estimated the threshold levels of urinary Cd as the benchmark dose low (BMDL) using the benchmark dose (BMD) approach. Urinary Cd excretion was divided into 10 categories, and an abnormality rate was calculated for each. Cut-off values for urinary substances were defined as corresponding to the 84% and 95% upper limit values of the target population who have not smoked. Then we calculated the BMD and BMDL using a log-logistic model. The values of BMD and BMDL for all urinary substances could be calculated. The BMDL for the 84% cut-off value of beta2-MG, setting an abnormal value at 5%, was 2.4 microg/g creatinine (cr) in men and 3.3 microg/g cr in women. In conclusion, the present study demonstrated that the threshold level of urinary Cd could be estimated in people living in the general environment without any known Cd-pollution in Japan, and the value was inferred to be almost the same as that in Belgium, Sweden, and China.

Acetylglucosaminidase↗

Essence: A benchmarking-validated transformer framework for early diagnosis of Parkinson's disease using cerebrospinal fluid protein biomarkers.

Parkinson's disease (PD) is a progressive neurodegenerative disorder characterized by motor and non-motor symptoms. The lack of objective molecular biomarkers limits early diagnosis and personalized treatment. Here, we propose Essence, a benchmarking-validated framework integrating cerebrospinal fluid (CSF) proteomics with traditional and deep learning models to identify robust protein signatures for PD. Using data from two independent cohorts, 1266 high-confidence proteins are quantified, among which 178 exhibit differential abundance between PD and healthy controls (HC). Through systematic benchmarking of ten machine learning algorithms and four neural architectures, the Transformer model consistently outperforms alternatives across multiple feature selection strategies, achieving an area under the receiver operating characteristic curve (AUC) of 1.0000 with only 35 features. Functional analyses of the top-ranked 35 proteins reveal enrichment in neuroinflammatory, synaptic, and oxidative stress-related pathways. Importantly, spatial transcriptomic profiling based on the Allen Brain Atlas shows region-specific expression of these biomarkers in PD-relevant brain structures, including the striatum, subthalamic nucleus, hippocampus, and white matter tracts. This anatomical alignment supports the functional relevance of the identified markers and highlights their potential utility in early-stage diagnosis and mechanistic understanding of PD.

Benchmarking↗

A benchmark dose analysis for sodium monofluoroacetate (1080) using dichotomous toxicity data.

The use of a benchmark dose (BMD) as an alternative to a no-observed-adverse-effect-level (NOAEL) approach was investigated as a means to improve current risk assessment values of sodium monofluoroacetate (1080). The feasibility of implementing the two approaches was investigated for three critical toxicological end points, namely cardiomyopathy, testicular toxicity and teratogenic effects identified from the few available critical studies. The BMD provides better representation of the dose-response relationship, offering an advantage over the current NOAEL approach. The calculated BMDs and lower-bound confidence limits (BMDLs) for the three end points were estimated using the Weibull, probit and quantal linear models for each end point. All models passed the chi2 test statistics (p > or = 0.1) for all three toxicity endpoints tested. A benchmark response (BMR) of 10% (extra risk) was chosen and the Akaike's information criterion (AIC) was used in selecting the appropriate model. The BMDL estimates derived were found to be generally slightly higher but comparable to the NOAEL for those same endpoints. The BMD(10) and BMDL(10) for cardiomyopathy and testicular effects were 0.21 mgkg(-1) bw and 0.10 mgkg(-1) bw, respectively. These values are proposed for use in the eventual determination of the tolerable daily intake (TDI) for 1080.

Abnormalities, Drug-Induced↗

Regional data set of infection rates for long-term care facilities: description of a valuable benchmarking tool.

BACKGROUND: Surveillance for nosocomial infections has been clearly established as a key element of all infection control programs. Surveillance programs in long-term care facilities (LTCFs) have been described, but published infection rates vary widely depending on the type of facility studied, nature of resident population, definitions used for LTCF-acquired infections, and type of data analysis. The aim of this initial study was to create a standardized regional data set of infection rates that could provide an external benchmark for interfacility comparison. METHODS: The study included 6 LTCFs in close geographic proximity with similar patient populations. Surveillance in each facility was conducted by a licensed nurse supervised by an infectious diseases physician. Standard definitions for infections and uniform reporting forms were used. Data were pooled in an aggregate cumulative fashion, and data analysis was patterned after the National Nosocomial Infection Surveillance System. RESULTS: The data set consisted of 328,065 resident-days of care during 30 months, with a total of 1252 infections for a pooled mean rate of 3.82 infections per 1000 resident-days of care. Infections for specific categories were 496 urinary tract infections (rate 1.51), 376 respiratory tract infections (rate 1.15), 88 gastroenteritis infections (rate 0.27), 283 skin and soft tissue infections (rate 0.86), 2 bloodstream infections (rate 0.06), and 3 unexplained febrile illnesses (rate 0. 09). Data analysis for comparison included interfacility means +/-2 standard deviations and percentiles of distribution. CONCLUSIONS: A regional data set of infection rates for LTCFs allowed for meaningful interfacility comparison of overall and specific endemic rates and is a valuable benchmarking tool for participating facilities.

Aged↗

Internet-based monitoring and benchmarking in ambulatory surgery centers.

BACKGROUND: Each year the number of surgical procedures performed on an outpatient basis increases, yet relatively little is known about assessing and improving quality of care in ambulatory surgery. Conventional methods for evaluating outcomes, which are based on assessment of inpatient services, are inadequate in the rapidly changing, geographically dispersed field of ambulatory surgery. Internet-based systems for improving outcomes and establishing benchmarks may be feasible and timely. METHODS: Eleven freestanding ambulatory surgery centers (ASCs) reported process and outcome data for 3,966 outpatient surgical procedures to an outcomes monitoring system (OMS), during a demonstration period from April 1997 to April 1999. ASCs downloaded software and protocol manuals from the OMS Web site. Centers securely submitted clinical information on perioperative process and outcome measures and postoperative patient telephone interviews. Feedback to centers ranged from current and historical rates of surgical and postsurgical complications to patient satisfaction and the adequacy of postsurgical pain relief. RESULTS: ASCs were able to successfully implement the data collection protocols and transmit data to the OMS. Data security efforts were successful in preventing the transmission of patient identifiers. Feedback reports to ASCs were used to institute changes in ASC staffing, patient care, and patient education, as well as for accreditation and marketing. The demonstration also pointed out shortcomings in the OMS, such as the need to simplify hardware and software installation as well as data collection and transfer methods, which have been addressed in subsequent OMS versions. DISCUSSION: Internet-based benchmarking for geographically dispersed outpatient health care facilities, such as ASCs, is feasible and likely to play a major role in this effort.

Benchmarking↗

Using benchmarking research to locate agency best practices for African American clients.

Using a collective case study design with benchmarking features, research reported here sought to locate differences in agency practices between public mental health agencies in which African American clients were doing comparatively better on specific proxy outcomes related to community tenure, and agencies with less success on those same variables. A panel of experts from the Ohio Department of Mental Health matched four agencies on per capita spending, percentage of African American clients, and urban-intensive setting. The panel also differentiated agencies on the basis of racial group comparisons for a number of proxy variables related to successful community tenure. Two agencies had a record of success with this client group (benchmark agencies); and two were less successful based on the selected criteria (comparison agencies). Findings indicated that when service elements explicitly related to culture were similar across study sites, the characteristics that did appear to make a difference were aspects of organizational culture. Implications for administration practice and further research are discussed.

Black or African American↗

A frontier analysis approach for benchmarking hospital performance in the treatment of acute myocardial infarction.

This paper uses a non-parametric frontier model and adaptations of the concepts of cross-efficiency and peer-appraisal to develop a formal methodology for benchmarking provider performance in the treatment of Acute Myocardial Infarction (AMI). Parameters used in the benchmarking process are the rates of proper recognition of indications of six standard treatment processes for AMI; the decision making units (DMUs) to be compared are the Medicare eligible hospitals of a particular state; the analysis produces an ordinal ranking of individual hospital performance scores. The cross-efficiency/peer-appraisal calculation process is constructed to accommodate DMUs that experience no patients in some of the treatment categories. While continuing to rate highly the performances of DMUs which are efficient in the Pareto-optimal sense, our model produces individual DMU performance scores that correlate significantly with good overall performance, as determined by a comparison of the sums of the individual DMU recognition rates for the six standard treatment processes. The methodology is applied to data collected from 107 state Medicare hospitals.

Benchmarking↗

Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.

Spatial omics technologies have revolutionized the study of tissue architecture and cellular heterogeneity by integrating molecular profiles with spatial localization. In spatially resolved transcriptomics, delineating higher-order anatomical structures is critical for understanding how cellular organization affects function. However, the reliability of current benchmarks of spatially aware clustering (SAC) methods is undermined by their narrow focus on Visium and brain tissue datasets and the incorrect interpretation of manual annotation as ground truth. Here we present SACCELERATOR, a community-driven, extensible framework that standardizes data formatting, method integration and metric evaluation, enabling rapid inclusion of new methods and datasets. Our analysis revealed substantial limitations in the generalizability and reproducibility of SAC methods and shows that anatomical labels commonly used as ground truths are often biased, error prone and unsuitable for benchmarking. Rather than ranking methods, we propose a consensus-guided workflow where descriptive spatial metrics highlight high-entropy regions of method disagreement, enabling targeted feedback for tissue experts. Applied to brain and cancer datasets, this approach uncovered biologically meaningful patterns overlooked by individual SAC methods and manual annotations, highlighting the need for iterative, expert-in-the-loop evaluation.

Benchmarking↗