Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Benchmark”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

A closer look at the Medicare part D low-income benchmark premium: how low can it go?

This issue brief explains how the Medicare Part D low income benchmark premium is calculated, what factors influence the level of the low-income benchmark premium in any given year, and the implications of the benchmark amount for Medicare drug plans and beneficiaries as it changes from year to year. The paper provides a simplified, two-year example of how the low-income benchmark premium is calculated in order to illustrate the key factors that influence it.

Benchmarking↗

Laboratory benchmarking: the College of American Pathologists' experience.

Benchmarking is an important part of performance evaluation in the clinical laboratory. When used effectively, benchmarking can lead to significant changes and performance improvement. This article reviews the experience of the College of American Pathologists (CAP) with benchmarking clinical laboratory expenses and presents data that summarizes recent laboratory trends identified through CAP's benchmark data.

Benchmarking↗

[Benchmarks for surgical gynecology: results of the German Society of Gynecology and Obstetrics Quality Assurance Study].

Profiling of performance and quality in gynecological surgery is discussed. Unfortunately, most report cards miss valid clinical benchmarks. Within the German study on quality assurance in gynecological surgery we explored whether indicators of quality were suitable as clinical benchmarks. Using a factor analytic approach, we reduced the number of indicators and obtained in a set of 13 indicators of clinical quality. On the basis of the study data on post operative infections we show that these indicators are suitable as clinical benchmarks: the clinical benchmarks are able to make health care quality transparent and demonstrate opportunities for improvement of the processes of gynecological care.

Benchmarking↗

Protein-Protein Docking Benchmark 2.0: an update.

We present a new version of the Protein-Protein Docking Benchmark, reconstructed from the bottom up to include more complexes, particularly focusing on more unbound-unbound test cases. SCOP (Structural Classification of Proteins) was used to assess redundancy between the complexes in this version. The new benchmark consists of 72 unbound-unbound cases, with 52 rigid-body cases, 13 medium-difficulty cases, and 7 high-difficulty cases with substantial conformational change. In addition, we retained 12 antibody-antigen test cases with the antibody structure in the bound form. The new benchmark provides a platform for evaluating the progress of docking methods on a wide variety of targets. The new version of the benchmark is available to the public at http://zlab.bu.edu/benchmark2.

Algorithms↗

Benchmark concentrations for methyl mercury obtained from the 9-year follow-up of the Seychelles Child Development Study.

Methyl mercury (MeHg) is highly toxic to the developing nervous system. Human exposure is mainly from fish consumption since small amounts are present in all fish. Findings of developmental neurotoxicity following high-level prenatal exposure to MeHg raised the question of whether children whose mothers consumed fish contaminated with background levels during pregnancy are at an increased risk of impaired neurological function. Benchmark doses determined from studies in New Zealand, and the Faroese and Seychelles Islands indicate that a level of 4-25 parts per million (ppm) measured in maternal hair may carry a risk to the infant. However, there are numerous sources of uncertainty that could affect the derivation of benchmark doses, and it is crucial to continue to investigate the most appropriate derivation of safe consumption levels. Earlier, we published the findings from benchmark analyses applied to the data collected on the Seychelles main cohort at the 66-month follow-up period. Here, we expand on the main cohort analyses by determining the benchmark doses (BMD) of MeHg level in maternal hair based on 643 Seychellois children for whom 26 different neurobehavioral endpoints were measured at 9 years of age. Dose-response models applied to these continuous endpoints incorporated a variety of covariates and included the k-power model, the Weibull model, and the logistic model. The average 95% lower confidence limit of the BMD (BMDL) across all 26 endpoints varied from 20.1 ppm (range=17.2-22.5) for the logistic model to 20.4 ppm (range=17.9-23.0) for the k-power model. These estimates are somewhat lower than those obtained after 66 months of follow-up. The Seychelles Child Development Study continues to provide a firm scientific basis for the derivation of safe levels of MeHg consumption.

Animals↗

A statistical evaluation of toxicity study designs for the estimation of the benchmark dose in continuous endpoints.

The benchmark approach is gaining attention as an alternative to the No-Observed-Adverse-Effect-Level (NOAEL) approach. However, current guidelines for the design of toxicity tests are based on assessing a NOAEL. It has been suggested that the current study design may not be optimal for assessing a Benchmark Dose (BMD). To further investigate this we performed three simulation studies in which a large number of designs were compared, focusing on continuous endpoints. Four fictitious endpoints were considered, their underlying dose-response curves having a linear, sublinear, supralinear, or sigmoidal shape. In each simulation run the BMD was derived from a model fitted to the generated data, where the selection of the model was based on that particular data set (according to a formal likelihood ratio test procedure). Thus, the model used for deriving the BMD in a single generated data set may not be the same as the one used for generating the data. In this way, model uncertainty is taken into account as well. The results show that the performance of a design is, first of all, determined by the total number of animals used. Distributing them over more dose groups does not result in a poorer performance of the study, despite the smaller number of animals per dose group. Dose placement is another crucial factor, and to minimize the risk of inadequate dose placement, the use of multiple dose studies is favorable. As a concomitant advantage, the use of multiple doses mitigates the disturbing effect of potential systematic errors in single dose groups. However, for endpoints with large residual variation (CV > or = 18%) there is a substantial probability of not detecting the overall dose-response, and this probability increases in designs with increasing number of dose groups. In such situations, six dose groups may be used as a compromise. Designs with high dose levels (i.e., associated with relatively high effects) are helpful in estimating doses with smaller effects (such as the benchmark dose), and it appears bad practice to omit higher dose groups to improve the fit at lower doses. The typical 28-day study design of four dose groups with five animals (per sex) may not be adequate to assess endpoints with large residual variation (CV > or = 18%), both in assessing a benchmark dose and in assessing a NOAEL.

Animals↗

Benchmarking of the CAP-88 and GENII computer codes using 1990 and 1991 monitored atmospheric releases from the Idaho National Engineering Laboratory.

The CAP-88 environmental radiological assessment computer code was benchmark tested to establish confidence in its results. The results from CAP-88 were compared to the results from the GENII computer code, which has undergone rigorous testing. The codes were benchmarked using 1990 and 1991 monitored atmospheric releases from Idaho National Engineering Laboratory facilities and the results (the effective dose equivalent to the maximally exposed offsite individual) were quantitatively compared using a metric based on the uncertainty in the Gaussian plume model and terrestrial transport models. The results of the benchmark tests were within the 95% acceptance region specified in the test protocol. CAP-88 was found to overpredict effective dose equivalent relative to GENII for elevated releases, largely because CAP-88 calculates a larger atmospheric dispersion factor (chi/Q) than does GENII using the same meteorological data. However, CAP-88 consistently underpredicted effective dose equivalent relative to GENII for ground-level releases. This was because CAP-88 accounts for the processes of plume depletion by dry and wet deposition while GENII does not account for these processes. The effect of depletion was tested and found to be most important for a ground-level release of a highly depositing species such as radioiodine which implies that acceptable benchmark results would be difficult to obtain for a highly dopositing species.

Computer Simulation↗

Multiplicity-adjusted inferences in risk assessment: benchmark analysis with quantal response data.

A primary objective in quantitative risk or safety assessment is characterization of the severity and likelihood of an adverse effect caused by a chemical toxin or pharmaceutical agent. In many cases data are not available at low doses or low exposures to the agent, and inferences at those doses must be based on the high-dose data. A modern method for making low-dose inferences is known as benchmark analysis, where attention centers on the dose at which a fixed benchmark level of risk is achieved. Both upper confidence limits on the risk and lower confidence limits on the "benchmark dose" are of interest. In practice, a number of possible benchmark risks may be under study; if so, corrections must be applied to adjust the limits for multiplicity. In this short note, we discuss approaches for doing so with quantal response data.

Biometry↗

Role of the standard deviation in the estimation of benchmark doses with continuous data.

For continuous data, risk is defined here as the proportion of animals with values above a large percentile, e.g., the 99th percentile or below the 1st percentile, for the distribution of values among control animals. It is known that reducing the standard deviation of measurements through improved experimental techniques will result in less stringent (higher) doses for the lower confidence limit on the benchmark dose that is estimated to produce a specified risk of animals with abnormal levels for a biological effect. Thus, a somewhat larger (less stringent) lower confidence limit is obtained that may be used as a point of departure for low-dose risk assessment. It is shown in this article that it is important for the benchmark dose to be based primarily on the standard deviation among animals, s(a), apart from the standard deviation of measurement errors, s(m), within animals. If the benchmark dose is incorrectly based on the overall standard deviation among average values for animals, which includes measurement error variation, the benchmark dose will be overestimated and the risk will be underestimated. The bias increases as s(m) increases relative to s(a). The bias is relatively small if s(m) is less than one-third of s(a), a condition achieved in most experimental designs.

Animals↗

Characterizing dose-response: I: Critical assessment of the benchmark dose concept.

We present a critical assessment of the benchmark dose (BMD) method introduced by Crump as an alternative method for setting a characteristic dose level for toxicant risk assessment. The no-observed-adverse-effect-level (NOAEL) method has been criticized because it does not use all of the data and because the characteristic dose level obtained depends on the dose levels and the statistical precision (sample sizes) of the study design. Defining the BMD in terms of a confidence bound on a point estimate results in a characteristic dose that also varies with the statistical precision and still depends on the study dose levels. Indiscriminate choice of benchmark response level may result in a BMD that reflects little about the dose-response behavior available from using all of the data. Another concern is that the definition of the BMD for the quantal response case is different for the continuous response case. Specifically, defining the BMD for continuous data using a ratio of increased effect divided by the background response results in an arbitrary dependence on the natural background for the endpoint being studied, making comparison among endpoints less meaningful and standards more arbitrary. We define a modified benchmark dose as a point estimate using the ratio of increased effect divided by the full adverse response range which enables consistent placement of the benchmark response level and provides a BMD with a more consistent relationship to the dose-response curve shape.

Animals↗

Modification and benchmarking of MCNP for low-energy tungsten spectra.

The MCNP Monte Carlo radiation transport code was modified for diagnostic medical physics applications. In particular, the modified code was thoroughly benchmarked for the production of polychromatic tungsten x-ray spectra in the 30-150 kV range. Validating the modified code for coupled electron-photon transport with benchmark spectra was supplemented with independent electron-only and photon-only transport benchmarks. Major revisions to the code included the proper treatment of characteristic K x-ray production and scoring, new impact ionization cross sections, and new bremsstrahlung cross sections. Minor revisions included updated photon cross sections, electron-electron bremsstrahlung production, and K x-ray yield. The modified MCNP code is benchmarked to electron backscatter factors, x-ray spectra production, and primary and scatter photon transport.

Algorithms↗

An enhanced RNA alignment benchmark for sequence alignment programs.

BACKGROUND: The performance of alignment programs is traditionally tested on sets of protein sequences, of which a reference alignment is known. Conclusions drawn from such protein benchmarks do not necessarily hold for the RNA alignment problem, as was demonstrated in the first RNA alignment benchmark published so far. For example, the twilight zone - the similarity range where alignment quality drops drastically - starts at 60 % for RNAs in comparison to 20 % for proteins. In this study we enhance the previous benchmark. RESULTS: The RNA sequence sets in the benchmark database are taken from an increased number of RNA families to avoid unintended impact by using only a few families. The size of sets varies from 2 to 15 sequences to assess the influence of the number of sequences on program performance. Alignment quality is scored by two measures: one takes into account only nucleotide matches, the other measures structural conservation. The performance order of parameters--like nucleotide substitution matrices and gap-costs--as well as of programs is rated by rank tests. CONCLUSION: Most sequence alignment programs perform equally well on RNA sequence sets with high sequence identity, that is with an average pairwise sequence identity (APSI) above 75 %. Parameters for gap-open and gap-extension have a large influence on alignment quality lower than APSI < or = 75 %; optimal parameter combinations are shown for several programs. The use of different 4 x 4 substitution matrices improved program performance only in some cases. The performance of iterative programs drastically increases with increasing sequence numbers and/or decreasing sequence identity, which makes them clearly superior to programs using a purely non-iterative, progressive approach. The best sequence alignment programs produce alignments of high quality down to APSI > 55 %; at lower APSI the use of sequence+structure alignment programs is recommended.

Journal Article↗

Using a benchmarking system to improve patient care and assist in technology assessment.

Clinical benchmarking is a tool of CQI that can be used to improve outcomes in areas of strategic importance. While it is a simple tool, benchmarking requires a long-term commitment from the entire organization involved in its use to be successful. Benchmarking is a means of setting goals or targets. As a tool used for continuous quality management, benchmarking is an ongoing activity of comparing an organization's service, product, or process with similar ones outside the organization that are known to be the best. In attempting to emulate or surpass "best practice," an organization must set challenging but attainable goals and reach them with a plan of realistic and efficient actions.

Critical Pathways↗

Analysis and comparison of benchmarks for multiple sequence alignment.

The most popular way of comparing the performance of multiple sequence alignment programs is to use empirical testing on sets of test sequences. Several such test sets now exist, each with potential strengths and weaknesses. We apply several different alignment packages to 6 benchmark datasets, and compare their relative performances. HOMSTRAD, a collection of alignments of homologous proteins, is regularly used as a benchmark for sequence alignment though it is not designed as such, and lacks annotation of reliable regions within the alignment. We introduce this annotation into HOMSTRAD using protein structural superposition. Results on each database show that method performance is dependent on the input sequences. Alignment benchmarks are regularly used in combination to measure performance across a spectrum of alignment problems. Through combining benchmarks, it is possible to detect whether a program has been over-optimised for a single dataset, or alignment problem type.

Cluster Analysis↗

The critical phase inspection process: a benchmarking study in search of industries' best practices.

Benchmarking is the orderly process of measuring one's own products, services, and practices against those of companies recognized as leaders. Eli Lilly and Company's Quality Assurance Department formed the Critical Phase Inspection Team to benchmark the processes for selecting and conducting critical phase inspections and reporting inspection findings. The team developed a telephone survey that was conducted with 33 other quality assurance units across the country. Analysis of the phone survey responses resulted in the identification of 5 quality assurance units that we felt could provide valuable information to us on these activities. Site visits to these companies were arranged and information was shared. We present here the analysis and results of our benchmarking endeavor. Through the information sharing involved in the benchmarking process, namely, the telephone surveys and the site visits, fresh ideas emerged and new acquaintances were made. Comparisons and adaptations of our methods with others in the quality assurance business will lead us to breakthrough improvements that will allow us to improve our current processes.

Data Collection↗

Benchmarking: improving outcomes for the congestive heart failure population.

The benchmarking process has been used extensively to evaluate and improve performance in business and industry. There is currently increasing interest in utilizing this process in the health care field to maximize efficiency and improve patient outcomes. The article defines benchmarking in health care and lists characteristics of the benchmarking process. The process of conducting a clinical benchmarking project aimed at improving outcomes for patients with congestive heart failure is described and determined to be an effective means of reducing costs and improving both patient outcomes and quality of care.

Cardiology Service, Hospital↗

Result of a national audit of bariatric surgery performed at academic centers: a 2004 University HealthSystem Consortium Benchmarking Project.

HYPOTHESIS: Bariatric surgery performed at US academic centers is safe and associated with low mortality. DESIGN: Multi-institutional consecutive cohort study. SETTING: Academic medical centers. PATIENTS AND INTERVENTIONS: We audited the medical records from 40 consecutive bariatric surgery cases performed between October 1, 2003, and March 31, 2004, at each of the 29 institutions participating in the University HealthSystem Consortium Bariatric Surgery Benchmarking Project. All medical records that met inclusion criteria (patient age, >17 and <65 years; and body mass index [calculated as weight in kilograms divided by the square of height in meters], 35-70) and exclusion criteria (previous bariatric surgery) were reviewed and data were collected on a standardized form. MAIN OUTCOME MEASURES: Demographic data, operative time, blood loss, transfusion requirement, complications, readmission, reoperation, and in-hospital and 30-day mortality. RESULTS: Data from 1144 bariatric surgery cases were reviewed from 29 University HealthSystem Consortium institutions. The specific bariatric procedures included gastric bypass (91.7%), gastroplasty or gastric banding (8.2%), and biliopancreatic diversion (0.1%). For gastric bypass procedures (n = 1049), the mean patient age was 43 years and mean body mass index was 49; 76% of procedures were performed laparoscopically, with a conversion rate of 2.2%; the overall complication rate was 16%, with an anastomotic leakage rate of 1.6%; the 30-day readmission rate was 6.6%; and the 30-day mortality rate was 0.4%. For restrictive procedures (n = 94), the mean patient age was 45 years and mean body mass index was 45; 92% of procedures were performed laparoscopically with no conversion; the overall complication rate was 3.2%; the 30-day readmission rate was 4.3%; and the 30-day mortality rate was 0%. CONCLUSIONS: Within the context of the 2004 University HealthSystem Consortium Bariatric Surgery Benchmarking Project, the risk for death within 30 days after bariatric surgery at academic centers is less than 1%. In addition, the practice of bariatric surgery at these centers has shifted from open surgery to predominately laparoscopic surgery. These quality-controlled outcome data can be used as a benchmark for the practice of bariatric surgery at most US hospitals.

Adult↗

BAliBASE 3.0: latest developments of the multiple sequence alignment benchmark.

Multiple sequence alignment is one of the cornerstones of modern molecular biology. It is used to identify conserved motifs, to determine protein domains, in 2D/3D structure prediction by homology and in evolutionary studies. Recently, high-throughput technologies such as genome sequencing and structural proteomics have lead to an explosion in the amount of sequence and structure information available. In response, several new multiple alignment methods have been developed that improve both the efficiency and the quality of protein alignments. Consequently, the benchmarks used to evaluate and compare these methods must also evolve. We present here the latest release of the most widely used multiple alignment benchmark, BAliBASE, which provides high quality, manually refined, reference alignments based on 3D structural superpositions. Version 3.0 of BAliBASE includes new, more challenging test cases, representing the real problems encountered when aligning large sets of complex sequences. Using a novel, semiautomatic update protocol, the number of protein families in the benchmark has been increased and representative test cases are now available that cover most of the protein fold space. The total number of proteins in BAliBASE has also been significantly increased from 1444 to 6255 sequences. In addition, full-length sequences are now provided for all test cases, which represent difficult cases for both global and local alignment programs. Finally, the BAliBASE Web site (http://www-bio3d-igbmc.u-strasbg.fr/balibase) has been completely redesigned to provide a more user-friendly, interactive interface for the visualization of the BAliBASE reference alignments and the associated annotations.

Amino Acid Sequence↗