Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Benchmark”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Mortality within 30 days of chemotherapy: a clinical governance benchmarking issue for oncology patients.

No national benchmark figures exist for early mortality due to chemotherapy unlike for surgical interventions. Deaths within 30 days of chemotherapy during a 6-month period were identified from the Royal Marsden Hospital electronic patient records. Treatment intention--curative or palliative, cause of death and number of previous treatments--were documented. Between April 2005 and September 2005, 1976 patients received chemotherapy with 161 deaths within 30 days of chemotherapy (8.1%). Of these, 124 deaths (77.0%) were due to disease progression. Of the other 37 deaths, 12 (7.5%) were related to chemotherapy, six each for solid tumours and haematological malignancies, of which seven (4.3%) were due to neutropenic sepsis. For the remaining 25 deaths (15.5%) there was insufficient information. There were more deaths after third and subsequent lines of therapy than with first and secondlines of therapy. Only 12 of the 161 deaths occurred in patients who were receiving potentially curative chemotherapy to give a mortality rate in breast and gastrointestinal malignancy of 0.5 and 1.5%, respectively. It is possible to audit mortality within 30 days of chemotherapy and this should become a benchmark for standard practice nationally. Most deaths were due to disease progression in the palliative setting. We practice this form of audit each quarter and feed back to the treating teams so that deaths are discussed and practice monitored.

Adolescent↗

Trials and tribulations of vascular surgical benchmarking.

BACKGROUND: Benchmarking is a new tool to assess the efficiency of different hospitals. Classification of operations using healthcare resource groups (HRGs) is related to parameters including number of cases, length of stay (LOS) and age profile. METHODS: A National Comparative Database was used to compare three hospitals. Analysis was confined to the major HRGs involved with vascular/venous surgery. RESULTS: For high-volume low-complexity varicose vein surgery, all three hospitals had similar numbers of patients and LOS. In contrast, the LOS for routine vascular operations in hospital A was double that in hospital B (16.3 versus 7.4 days). Hospital A had three times as many patients classified as 'other - peripheral vascular disease' as hospital C and six times as many as hospital B (329, 49 and 111 for hospitals A, B and C respectively). LOS following major amputation in hospitals A and C was nearly double that in hospital B (32.4, 18.3 and 33.6 days for hospitals A, B and C respectively). CONCLUSION: There were a number of significant variations between the three hospitals during the 9-month interval. Explanations included the methods of coding, local facilities including availability of rehabilitation beds and difference in the patients' age profiles. Benchmarking in its present format reveals a number of variations which may not necessarily reflect real differences in clinical performance.

Benchmarking↗

Staff confidence in dealing with aggressive patients: a benchmarking exercise.

Interacting with potentially aggressive patients is a common occurrence for nurses working in psychiatric intensive care units. Although the literature highlights the need to educate staff in the prevention and management of aggression, often little, or no, training is provided by employers. This article describes a benchmarking exercise conducted in psychiatric intensive care units at two Western Australian hospitals to assess staff confidence in coping with patient aggression. Results demonstrated that staff in the hospital where regular training was undertaken were significantly more confident in dealing with aggression. Following the completion of a safe physical restraint module at the other hospital staff reported a significant increase in their level of confidence that either matched or bettered the results of their benchmark colleagues.

Adaptation, Psychological↗

Benchmarking in health care: a review of the literature.

This paper provides a review of the 10 significant publications related to benchmarking in health care. The discussion which follows is presented according to four headings: what the study did, how the study was conducted, what was learnt from the experience, and what the implications were for health care generally. The findings of this review are reassuring in that all studies provided valuable information, in terms of clinical practice and the health care service or the benchmarking process. They highlight the importance of the maintenance of quality health care, the reduction of health care costs and the need for improved efficiency and effectiveness in providing health care.

Australia↗

Using benchmarking to improve organizational communication.

Best practice refers to those practices that lead to superior performance in a company or enterprise relative to industry or international leaders. Benchmarking of those activities that are critical to organizational performance is an important part of the identification and implementation of best-practice approaches. This article looks at communication as one aspect in the development of best practice in the management of safety, environment, and quality. A number of barriers to effective communication are identified, and benchmarks for the evaluation of organizational communication are suggested.

Benchmarking↗

APDB: a novel measure for benchmarking sequence alignment methods without reference alignments.

MOTIVATION: We describe APDB, a novel measure for evaluating the quality of a protein sequence alignment, given two or more PDB structures. This evaluation does not require a reference alignment or a structure superposition. APDB is designed to efficiently and objectively benchmark multiple sequence alignment methods. RESULTS: Using existing collections of reference multiple sequence alignments and existing alignment methods, we show that APDB gives results that are consistent with those obtained using conventional evaluations. We also show that APDB is suitable for evaluating sequence alignments that are structurally equivalent. We conclude that APDB provides an alternative to more conventional methods used for benchmarking sequence alignment packages.

Algorithms↗

Systematic benchmarking of microarray data classification: assessing the role of non-linearity and dimensionality reduction.

MOTIVATION: Microarrays are capable of determining the expression levels of thousands of genes simultaneously. In combination with classification methods, this technology can be useful to support clinical management decisions for individual patients, e.g. in oncology. The aim of this paper is to systematically benchmark the role of non-linear versus linear techniques and dimensionality reduction methods. RESULTS: A systematic benchmarking study is performed by comparing linear versions of standard classification and dimensionality reduction techniques with their non-linear versions based on non-linear kernel functions with a radial basis function (RBF) kernel. A total of 9 binary cancer classification problems, derived from 7 publicly available microarray datasets, and 20 randomizations of each problem are examined. CONCLUSIONS: Three main conclusions can be formulated based on the performances on independent test sets. (1) When performing classification with least squares support vector machines (LS-SVMs) (without dimensionality reduction), RBF kernels can be used without risking too much overfitting. The results obtained with well-tuned RBF kernels are never worse and sometimes even statistically significantly better compared to results obtained with a linear kernel in terms of test set receiver operating characteristic and test set accuracy performances. (2) Even for classification with linear classifiers like LS-SVM with linear kernel, using regularization is very important. (3) When performing kernel principal component analysis (kernel PCA) before classification, using an RBF kernel for kernel PCA tends to result in overfitting, especially when using supervised feature selection. It has been observed that an optimal selection of a large number of features is often an indication for overfitting. Kernel PCA with linear kernel gives better results.

Algorithms↗

M@CBETH: a microarray classification benchmarking tool.

Microarray classification can be useful to support clinical management decisions for individual patients in, for example, oncology. However, comparing classifiers and selecting the best for each microarray dataset can be a tedious and non-straightforward task. The M@CBETH (a MicroArray Classification BEnchmarking Tool on a Host server) web service offers the microarray community a simple tool for making optimal two-class predictions. M@CBETH aims at finding the best prediction among different classification methods by using randomizations of the benchmarking dataset. The M@CBETH web service intends to introduce an optimal use of clinical microarray data classification.

Algorithms↗

'Emerge': Benchmarking of clinical performance and patients' experiences with emergency care in Switzerland.

OBJECTIVE: To assess the effects of uniform indicator measurement and group benchmarking followed by hospital-specific activities on clinical performance measures and patients' experiences with emergency care in Switzerland. DESIGN: Data were collected in a pre-post design in two measurement cycles, before and after implementation of improvement activities. Trained hospital staff recorded patient characteristics and clinical performance data. Patients completed a questionnaire after discharge/transfer from the emergency unit. SETTING: Emergency departments of 12 community hospitals in Switzerland, participating in the 'Emerge' project. SUBJECTS: Eligible patients were entered into the study (18 544 in total: 9174 and 9370 in the first and second cycles, respectively), and 2916 and 3370 patients returned the questionnaire in the first and second measurement cycles, respectively (response rates 32% and 36%, respectively). MAIN OUTCOME MEASURES: Clinical performance measures (concordance of prospective and retrospective assessment of urgency of care needs, and time intervals between sequences of events) and patients' reports about care provision in emergency departments (EDs), measured by a 22-item, self-administered questionnaire. RESULTS: Concordance of prospective and retrospective assignments to one of three urgency categories improved significantly by 1%, and both under- and over-prioritization, were reduced. The median duration between ED admission and documentation of post-ED disposition fell from 137 minutes in 2001 to 130 minutes in 2002 (P < 0.001). Significant improvements in the reports provided by patients were achieved in 10 items, and were mainly demonstrated in structures of care provision and perceived humanity. CONCLUSION: Undertaken in a real-world setting, small but significant improvements in performance measures and patients' perceptions of emergency care could be achieved. Hospitals accomplished these improvements mainly by averting strong outliers, and were most successful in preventing series of negative events. Uniform outcomes measurement, group benchmarking, and data-driven hospital-specific strategies for change are suggested as valuable tools for continuous improvement. Several hospitals have already implemented the developed measures in their internal quality systems and subsequent measurements are projected.

Adolescent↗

Quality indicators for international benchmarking of mental health care.

OBJECTIVE: To identify quality measures for international benchmarking of mental health care that assess important processes and outcomes of care, are scientifically sound, and are feasible to construct from preexisting data. DESIGN: An international expert panel employed a consensus development process to select important, sound, and feasible measures based on a framework that balances these priorities with the additional goal of assessing the breadth of mental health care across key dimensions. PARTICIPANTS: Six countries and one international organization nominated seven panelists consisting of mental health administrators, clinicians, and services researchers with expertise in quality of care, epidemiology, public health, and public policy. Measures. Measures with a final median score of at least 7.0 for both importance and soundness, and data availability rated as 'possible' or better in at least half of participating countries, were included in the final set. Measures with median scores </=3.0 or data availability rated as 'unlikely' were excluded. Measures with intermediate scores were subject to further discussion by the panel, leading to their adoption or rejection on a case-by-case basis. RESULTS: From an initial set of 134 candidate measures, the panel identified 12 measures that achieved moderate to high scores on desired attributes. CONCLUSIONS: Although limited, the proposed measure set provides a starting point for international benchmarking of mental health care. It addresses known quality problems and achieves some breadth across diverse dimensions of mental health care.

Benchmarking↗

Benchmarking radiation protection programs.

Results are presented of a modest benchmarking project with eight institutions as part of a program review. Four of the institutions had programs with both research and clinical components. Metrics were derived from the data acquired. They included Registered Users per total Technical Staff (mean 284, SD 42%); Principal Investigators per Senior IP staff (mean 68.7, SD 49.6%); and Principal Investigators per HP staff (mean 44.7, SD 34.4%). Reasons for the differences among institutions were identified. The benchmarking exercise was an effective tool in providing guidance in setting staffing levels.

Benchmarking↗

Benchmarking medical group practices using claims data: methodological and practical problems.

As claims data for physicians and groups of physicians has improved in quality and quantity, health information vendors have begun marketing information about medical groups' productivity, utilization, and quality. Based on interviews with product developers and our understanding of the evolution of their products, several methodological and practical issues remain. For now and the immediate future, health information vendors will continue to face the limitations of physicians' claims data. Vendors and purchasers should be aware of common data shortcomings such as inadequate monthly enrollment figures, possible physician upcoding to circumvent utilization management restrictions, and incorrect coding when a test is used to rule out a disease. In the longer term, several avenues seem likely to make medical groups' data better and richer because of computer-based medical records and efficiencies possible from the Internet. The field of benchmarking products for group practices is still an immature market. However, several trends suggest such products are highly desirable. Provider organizations which bear medical risk need benchmarking data to help improve their efficiency. There are many important nonprovider organizations that need good information on group practices' utilization patterns and outcomes to help them plan new products and negotiate with physicians.

Benchmarking↗

Clinical benchmarking: implications for perinatal nursing.

Health care is a dynamic environment where expectations of quality must be balanced with appropriateness of treatment and cost of care. Managers often have inadequate information on which to base decisions, policy, and practice. Clinical benchmarking is a tool and a process of continuously comparing the practices and performances of one's operations against those of the best in the industry or the focused area of service and then using that information to enhance and improve performance and productivity. The article discusses the advantages and disadvantages of benchmarking as well as the factors influencing the need for such tools in health care and in perinatal nursing.

Benchmarking↗

Health and productivity management: establishing key performance measures, benchmarks, and best practices.

Major areas considered under the rubric of health and productivity management (HPM) in American business include absenteeism, employee turnover, and the use of medical, disability, and workers' compensation programs. Until recently, few normative data existed for most HPM areas. To meet the need for normative information in HPM, a series of Consortium Benchmarking Studies were conducted. In the most recent application of the study, 1998 HPM costs, incidence, duration, and other program data were collected from 43 employers on almost one million workers. The median HPM costs for these organizations were $9992 per employee, which were distributed among group health (47%), turnover (37%), unscheduled absence (8%), nonoccupational disability (5%), and workers' compensation programs (3%). Achieving "best-practice" levels of performance (operationally defined as the 25th percentile for program expenditures in each HPM area) would realize savings of $2562 per employee (a 26% reduction). The results indicate substantial opportunities for improvement through effective coordination and management of HPM programs. Examples of best-practice activities collated from on-site visits to "benchmark" organizations are also reviewed.

Absenteeism↗

Using benchmarking data to determine vascular access device selection.

Benchmarking data has validated that patients with planned vascular access device (VAD) placement have fewer device placements, less difficulty with device insertions, fewer venipunctures, earlier assessment for placement of central VADs, and shorter hospital stays. This article will discuss VAD program planning, early assessment for VAD selection, and benchmarking of program data used to achieve positive infusion-related outcomes.

Benchmarking↗

National nosocomial infection surveillance system: from benchmark to bedside in trauma patients.

INTRODUCTION: Ventilator-associated pneumonia (VAP) is an important cause of morbidity and mortality in the injured patient. Identification of those with VAP is important both in immediate clinical decision making as well as for the epidemiologic evaluation of the disease and benchmarking of rates across institutions with variable practice patterns. Despite this, controversy exists over the optimal method of VAP diagnosis. Many centers currently use invasive culture methods such as bronchoalveolar lavage (BAL) for diagnosis. Another diagnostic method, and the most common epidemiologic tool used to track VAP, is the definition employed by the National Nosocomial Infections Surveillance (NNIS) system. This relies on a combination of clinical and culture data. Our goal was to evaluate the accuracy of the NNIS definition as compared with BAL diagnosis in trauma patients. METHODS: Records of all ventilated patients admitted to the trauma intensive care unit at a Level I center who were evaluated for the presence of pneumonia over a 2.5-year period were reviewed. VAP diagnosis was established if > or =10 cfu/mL were cultured on BAL. VAP rates and time of onset were compared with the hospital infection control database, which defines VAP by NNIS criteria. Assuming BAL to be correct, sensitivity, specificity, and positive and negative predictive values were calculated for NNIS VAP. RESULTS: From September 1, 2001, through December 31, 2003, 292 patients underwent BAL for suspected pneumonia. The pneumonia rate in this group was 34 per 1,000 ventilator days. The NNIS definition showed excellent overall agreement, with a rate of 36 per 1,000 ventilator days. The use of the NNIS definition for bedside decision making, however, is less accurate. Sensitivity and positive predictive value were reasonably good (84% and 83%, respectively), whereas specificity and negative predictive value suffer (69% and 69%, respectively). Most importantly, the use of NNIS would have led to no treatment in 16% of patients diagnosed with VAP by BAL. CONCLUSIONS: Compared with strict bacteriologic criteria for VAP, the NNIS definition has good overall agreement and seems to have utility as an epidemiologic benchmarking tool in trauma patients. However, the NNIS definition has less utility as a bedside decision-making tool in this population, leading to under-treatment in a significant number of patients.

Adult↗

Benchmarking large language models for genomic knowledge with GeneTuring.

Large language models (LLMs) show promise in biomedical research, but their effectiveness for genomic inquiry remains unclear. We developed GeneTuring, a benchmark consisting of 16 genomics tasks with 1,600 curated questions, and manually evaluated 48,000 answers from ten LLM configurations, including GPT-4o (via API, ChatGPT with web access, and a custom GPT setup), GPT-3.5, Claude 3.5, Gemini Advanced, GeneGPT (both slim and full), BioGPT, and BioMedLM. A custom GPT-4o configuration integrated with NCBI APIs, developed in this study as SeqSnap, achieved the best overall performance. GPT-4o with web access and GeneGPT demonstrated complementary strengths. Our findings highlight both the promise and current limitations of LLMs in genomics, and emphasize the value of combining LLMs with domain-specific tools for robust genomic intelligence. GeneTuring offers a key resource for benchmarking and improving LLMs in biomedical research.

Benchmark↗

Benchmarking patient improvement in physical therapy with data envelopment analysis.

PURPOSE: The purpose of this article is to present a case study that documents how management science techniques (in particular data envelopment analysis) can be applied to performance improvement initiatives in an inpatient physical therapy setting. DESIGN/METHODOLOGY/APPROACH: The data used in this study consist of patients referred for inpatient physical therapy following total knee replacement surgery (at a medium-sized medical facility in the Midwestern USA) during the fiscal year 2002. Data envelopment analysis (DEA) was applied to determine the efficiency of treatment, as well as to identify benchmarks for potential patient improvement. Statistical trends in the benchmarking and efficiency results were subsequently analyzed using non-parametric and parametric methods. FINDINGS: Our analysis indicated that the rehabilitation process was largely effective in terms of providing consistent, quality care, as more than half of the patients in our study achieved the maximum amount of rehabilitation possible given available inputs. Among patients that did not achieve maximum results, most could obtain increases in the degree of flexion gain and reductions in the degree of knee extension. RESEARCH LIMITATIONS/IMPLICATIONS: The study is retrospective in nature, and is not based on clinical trial or experimental data. Additionally, DEA results are inherently sensitive to sampling: adding or subtracting individuals from the sample may change the baseline against which efficiency and rehabilitation potential are measured. As such, therapists using this approach must ensure that the sample is representative of the general population, and must not contain significant measurement error. Third, individuals who choose total knee arthroplasty will incur a transient disability. However, this population does not generally fit the World Health Organization International Classification of Functioning, Disability and Health definition of disability if the surgical procedure is successful. Since the study focuses on the outcomes of physical therapy, range of motion measurements and circumferential measurements were chosen as opposed to the more global measures of functional independence such as mobility, transfers and stair climbing. Applying this technique to data on patients with different disabilities (or the same disability with other outcome variables, such as Functional Independence Measure scores) may give dissimilar results. PRACTICAL IMPLICATIONS: This case study provides an example of how one can apply quantitative management science tools in a manner that is both tractable and intuitive to the practising therapist, who may not have an extensive background in quantitative performance improvement or statistics. ORIGINALITY/VALUE: DEA has not been applied to rehabilitation, especially in the case where managers have limited data available.

Arthroplasty, Replacement, Knee↗