Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Benchmark”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Benchmark data from the literature for evaluation of new glucose sensing technologies.

New glucose sensors based on various technologies are being developed to provide information for improved therapy in diabetes. There is a need to establish rational performance standards for these sensors. Frequently sampled, direct blood glucose recordings representative of blood glucose excursions in diabetes are the "gold standard." An extensive literature search revealed a limited number of diabetic and nondiabetic blood glucose recordings suitable for this purpose. Certain blood glucose recordings reflect the diversity of glycemic dynamics and provide sufficient challenge for evaluation of sensor systems. These recordings were converted into an accessible electronic format. An example is given of the use of these benchmark data to estimate aliasing error, or the error due to insufficient sampling frequency, based on a hypothetical sensor system having some properties of conventional "fingerstick" systems. Discrete sampling systems accumulate substantial aliasing error as the sampling period increases.

Biosensing Techniques↗

Measuring food insecurity and hunger in the United States: development of a national benchmark measure and prevalence estimates.

Since 1992, the U.S. Department of Agriculture Food and Nutrition Service (FNS) has led a collaborative effort to develop a comprehensive benchmark measure of the severity and prevalence of food insecurity and hunger in the United States. Based on prior research and wide consultation, a survey instrument specifically relevant to U.S. conditions was designed and tested. Through its Current Population Survey (CPS), the U.S. Bureau of the Census has fielded this instrument each year since 1995. A measurement scale was derived from the data through fitting, testing and validating a Rasch scale. The unidimensional Rasch model corresponds to the form of the phenomenon being measured, i.e., the severity of food insufficiency due to inadequate resources as directly experienced and reported in U.S. households. A categorical measure reflecting designated ranges of severity on the scale was constructed for consistent comparison of prevalence estimates over time and across population groups. The technical basis and initial results of the new measure were reported in September 1997. For the 12 months ending April 1995, an estimated 11.9% of U.S. households (35 million persons) were food insecure. Among these, 4.1% of households (with 6.9 million adults and 4.3 million children) showed a recurring pattern of hunger due to inadequate resources for one or more of their adult and/or child members sometime during the period. The new measure has been incorporated into other federal surveys and is being used by researchers throughout the U.S. and Canada.

Adult↗

Static benchmarking of membrane helix predictions.

Prediction of trans-membrane helices continues to be a difficult task with a few prediction methods clearly taking the lead; none of these is clearly best on all accounts. Recently, we have carefully set up protocols for benchmarking the most relevant aspects of prediction accuracy and have applied it to >30 prediction methods. Here, we present the extension of that analysis to the level of an automatic web server evaluating new methods (http://cubic.bioc.columbia.edu/services/tmh_benchmark/). The most important achievements of the tool are: (i) any new method is compared to the battery of well-established tools; (ii) the battery of measures explored allows spotting strengths in methods that may not be 'best' overall. In particular, we report per-residue and per-segment scores for accuracy and the error-rates for confusing membrane helices with globular proteins or signal peptides. An additional feature is that developers can directly investigate any hydrophobicity scale for its potential in predicting membrane helices.

Hydrophobic and Hydrophilic Interactions↗

A benchmark of multiple sequence alignment programs upon structural RNAs.

To date, few attempts have been made to benchmark the alignment algorithms upon nucleic acid sequences. Frequently, sophisticated PAM or BLOSUM like models are used to align proteins, yet equivalents are not considered for nucleic acids; instead, rather ad hoc models are generally favoured. Here, we systematically test the performance of existing alignment algorithms on structural RNAs. This work was aimed at achieving the following goals: (i) to determine conditions where it is appropriate to apply common sequence alignment methods to the structural RNA alignment problem. This indicates where and when researchers should consider augmenting the alignment process with auxiliary information, such as secondary structure and (ii) to determine which sequence alignment algorithms perform well under the broadest range of conditions. We find that sequence alignment alone, using the current algorithms, is generally inappropriate <50-60% sequence identity. Second, we note that the probabilistic method ProAlign and the aging Clustal algorithms generally outperform other sequence-based algorithms, under the broadest range of applications.

Algorithms↗

A Protein Classification Benchmark collection for machine learning.

Protein classification by machine learning algorithms is now widely used in structural and functional annotation of proteins. The Protein Classification Benchmark collection (http://hydra.icgeb.trieste.it/benchmark) was created in order to provide standard datasets on which the performance of machine learning methods can be compared. It is primarily meant for method developers and users interested in comparing methods under standardized conditions. The collection contains datasets of sequences and structures, and each set is subdivided into positive/negative, training/test sets in several ways. There is a total of 6405 classification tasks, 3297 on protein sequences, 3095 on protein structures and 10 on protein coding regions in DNA. Typical tasks include the classification of structural domains in the SCOP and CATH databases based on their sequences or structures, as well as various functional and taxonomic classification problems. In the case of hierarchical classification schemes, the classification tasks can be defined at various levels of the hierarchy (such as classes, folds, superfamilies, etc.). For each dataset there are distance matrices available that contain all vs. all comparison of the data, based on various sequence or structure comparison methods, as well as a set of classification performance measures computed with various classifier algorithms.

Algorithms↗

Smoke yields of tobacco-specific nitrosamines in relation to FTC tar level and cigarette manufacturer: analysis of the Massachusetts Benchmark Study.

OBJECTIVES: This research assessed the relationship between the deliveries of carcinogenic tobacco-specific nitrosamines (TSNAs) and the Federal Trade Commission (FTC) "tar" ratings of US commercial cigarettes. METHODS: Analysis of covariance (ANCOVA) was used to assess the explanatory power of FTC tar, the particular manufacturer, and other cigarette characteristics to predict the yields of four TSNAs (N'-nitrosonornicotine [NNN], 4-(N-methyl-N-nitrosamino)-1-(3-pyridyl)-1-butanone [NNK], N'-nitrosoanatabine [NAT], and N'-nitrosoanabasine [NAB]) in 26 US commercial brands tested in the 1999 Massachusetts Benchmark Study. RESULTS: When FTC tar alone was used to predict TSNA yield, the squared correlation coefficient (R(2)) was only 38% for NNN, 76% for NNK, 46% for NAT, and 49% for NAB. Inclusion of manufacturer-specific variables significantly (p < 0.001) increased the estimated R(2) for three of the four species of nitrosamine to: 78% for NNN, 88% for NNK, and 81% for NAT. Inclusion of other cigarette characteristics (filter type, paper permeability, tobacco weight, tip dilution) did not reduce the significance of the manufacturer-specific effects. Federal Trade Commission nicotine and carbon monoxide (CO) yields were no better at predicting TSNA levels. CONCLUSIONS: FTC ratings for tar, nicotine, and carbon monoxide do not tell the entire story about the comparative yields of toxic agents in marketed cigarette brands. The significant manufacturer-specific effects suggest that proprietary blending and processing of tobacco matter as well. Public, brand-by-brand disclosure of the yields of TSNA and possibly other smoke constituents appears to be warranted.

Analysis of Variance↗

Dose-response modeling and benchmark calculations from spontaneous behavior data on mice neonatally exposed to 2,2',4,4',5-pentabromodiphenyl ether.

In this paper the benchmark dose (BMD) method was introduced for spontaneous behavior data observed in 2-, 5-, and 8-month-old male and female C57Bl mice exposed orally on postnatal day 10 to different doses of 2,2',4,4',5-pentabromodiphenyl ether (PBDE 99). Spontaneous behavior (locomotion, rearing, and total activity) was in the present work quantified in terms of a fractional response defined as the cumulative response after 20 min divided by the cumulative response produced over the whole 1-h test period. The fractional response contains information about the time-response profile (which differs between the treatment groups) and has appropriate statistical characteristics. In the analysis, male and female mice could be characterized by a common dose-response model (i.e., they responded equally to the exposure to PBDE 99). As a primary approach, the BMD was defined as the dose producing a 5 or 10% change in the mean fractional response. According to the Hill model, considering a 10% change the lower bound of the BMD for rearing, locomotion, and total activity was 1.2, 0.85, and 0.31 mg PBDE 99/kg body weight, respectively. A probability-based procedure for BMD modeling was also considered. Using this methodology, the BMD was defined as corresponding to an excess risk of 5 or 10% of falling below cutoff points representing adverse levels of fractional response.

Algorithms↗

A comparison of ratio distributions based on the NOAEL and the benchmark approach for subchronic-to-chronic extrapolation.

One approach to derive a data-based assessment factor (AF) for subchronic-to-chronic extrapolation is to determine ratios between the NOAEL(subchronic) and NOAEL(chronic) for the same compounds. Instead of using ratios of NOAELs, the distribution can also be estimated by ratios of subchronic and chronic Benchmark Doses (or Critical Effect Doses, CEDs, for continuous data). In this study 314 dose-response datasets on body weights and liver weights of mice and rats were selected providing dose-response information after both subchronic and chronic exposure. NOAEL ratios could be derived in only 68 of these datasets, while CED ratios could be derived in 189 datasets. When only the (53) datasets suitable for both approaches were evaluated the variation of the CED ratio distribution (GSD [geometric standard deviation]: 2.9) was smaller than the one of the NOAEL ratio distribution (GSD: 3.3). After correcting for the estimation error of the individual CED ratios the GSD of the CED distribution decreased to 2.3. The geometric means (GMs) of the NOAEL and CED distributions were similar (1.2 and 1.6, respectively). Comparing the NOAEL distribution based on all 68 datasets suitable for deriving NOAEL ratios with the CED distribution based on the 189 ratios suitable for deriving CED ratios resulted in similar GMs (1.5 and 1.7, respectively), but the GSDs differed considerably (5.3 and 2.3 respectively). It is concluded that usage of the CED approach results in less wide distributions. Furthermore, a larger fraction of available datasets is useful to inform the ratio distribution. This results in more accurate, and less conservative distributions of AFs in general compared to the distributions based on NOAEL ratios that have been proposed so far.

Algorithms↗

Objective and subjective impairment from often-used sedative/analgesic combinations in ambulatory surgery, using alcohol as a benchmark.

Impairment caused by different sedative/analgesic combinations commonly used in ambulatory settings was compared to that of alcohol at blood alcohol concentrations (BACs) higher than or equal to 0.10%. Impairment was measured via subjective (mood) and objective (psychomotor performance) assays. Twelve healthy human volunteers (10 males and 2 females; age range 21-34 yr) participated in this prospective, double-blind, randomized, cross-over study. Each subject was exposed to five drug conditions across 5 wk. Each of the following drug conditions were adjusted for body weight (per 70 kg):fentanyl 50 micrograms and propofol 35 mg (FP), fentanyl 50 micrograms and midazolam 2 mg (FM), fentanyl 50 micrograms, midazolam 2 mg, and propofol 35 mg (FMP), alcohol 56 g (orally administered), and placebo (PLC). With the exception of alcohol, the other drugs were administered via the intravenous route. Tests for psychomotor performance, subjective effects, and short-term memory were done at baseline, and at different intervals until 240 min postinjection. Psychomotor impairment caused by alcohol at 15 min postingestion (at a BAC of 0.11% +/- 0.03% [mean +/- SE]) was used as a benchmark with which impairment caused by other sedative/analgesic combinations was compared. All the study drug combinations produced impairment (i.e., impairment greater than that seen with PLC), similar to that observed with alcohol at a BAC of 0.11%. We have demonstrated that some sedative/analgesic drug combinations used in anesthesia for ambulatory procedures produce impairment similar to or greater than that observed with a large dose of alcohol.

Adult↗

Benchmark test of transport calculations of gold and nickel activation with implications for neutron kerma at Hiroshima.

A benchmark test of the Monte Carlo neutron and photon transport code system (MCNP) was performed using a 252Cf fission neutron source to validate the use of the code for the energy spectrum analyses of Hiroshima atomic bomb neutrons. Nuclear data libraries used in the Monte Carlo neutron and photon transport code calculation were ENDF/B-III, ENDF/B-IV, LASL-SUB, and ENDL-73. The neutron moderators used were granite (the main component of which is SiO2, with a small fraction of hydrogen), Newlight [polyethylene with 3.7% boron (natural)], ammonium chloride (NH4Cl), and water (H2O). Each moderator was 65 cm thick. The neutron detectors were gold and nickel foils, which were used to detect thermal and epithermal neutrons (4.9 eV) and fast neutrons (> 0.5 MeV), respectively. Measured activity data from neutron-irradiated gold and nickel foils in these moderators decreased to about 1/1,000th or 1/10,000th, which correspond to about 1,500 m ground distance from the hypocenter in Hiroshima. For both gold and nickel detectors, the measured activities and the calculated values agreed within 10%. The slopes of the depth-yield relations in each moderator, except granite, were similar for neutrons detected by the gold and nickel foils. From the results of these studies, the Monte Carlo neutron and photon transport code was verified to be accurate enough for use with the elements hydrogen, carbon, nitrogen, oxygen, silicon, chlorine, and cadmium, and for the incident 252Cf fission spectrum neutrons.

Californium↗

Benchmark test of neutron transport calculations: indium, nickel, gold, europium, and cobalt activation with and without energy moderated fission neutrons by iron simulating the Hiroshima atomic bomb casing.

A benchmark test of the Monte Carlo neutron and photon transport code system (MCNP) was performed using a bare- and energy-moderated 252Cf fission neutron source which was obtained by transmission through 10-cm-thick iron. An iron plate was used to simulate the effect of the Hiroshima atomic bomb casing. This test includes the activation of indium and nickel for fast neutrons and gold, europium, and cobalt for thermal and epithermal neutrons, which were inserted in the moderators. The latter two activations are also to validate 152Eu and 60Co activity data obtained from the atomic bomb-exposed specimens collected at Hiroshima and Nagasaki, Japan. The neutron moderators used were Lucite and Nylon 6 and the total thickness of each moderator was 60 cm or 65 cm. Measured activity data (reaction yield) of the neutron-irradiated detectors in these moderators decreased to about 1/1,000th or 1/10,000th, which corresponds to about 1,500 m ground distance from the hypocenter in Hiroshima. For all of the indium, nickel, and gold activity data, the measured and calculated values agreed within 25%, and the corresponding values for europium and cobalt were within 40%. From this study, the MCNP code was found to be accurate enough for the bare- and energy-moderated 252Cf neutron activation calculations of these elements using moderators containing hydrogen, carbon, nitrogen, and oxygen.

Cobalt↗

Risk-adjusted quality outcome measures: indexes for benchmarking rates of mortality, complications, and readmissions.

This article describes a risk-adjusted approach for profiling hospitals and physicians on quality outcomes using readily available administrative data. By comparing risk-adjusted rates of mortality, complications, and readmissions to rates for peers, national norms, and benchmarks, this approach enables purchasers and providers to compare the performance of providers in terms of both favorable and adverse outcomes.

Centers for Medicare and Medicaid Services, U.S.↗

Protected areas as biodiversity benchmarks for human impact: agriculture and the Serengeti avifauna.

Protected areas as biodiversity benchmarks allow a separation of the direct effects of human impact on biodiversity loss from those of other environmental changes. We illustrate the use of ecological baselines with a case from the Serengeti ecosystem, Tanzania. We document a substantial but previously unnoted loss of bird diversity in agriculture detected by reference to the immediately adjacent native vegetation in Serengeti. The abundance of species found in agriculture was only 28% of that for the same species in native savannah. Insectivorous species feeding in the grass layer or in trees were the most reduced. Some 50% of both insectivorous and granivorous species were not recorded in agriculture, with ground-feeding and tree species most affected. Grass-layer insect abundance and diversity was much reduced in agriculture, consistent with the loss of insectivorous birds. These results indicate that many species of birds will become confined to protected areas over time. We need to determine whether existing protected areas are sufficiently large to maintain viable populations of insectivorous birds likely to become confined to them. This study highlights the essential nature of baseline areas for assessing causes of change in human-dominated systems and for developing innovative strategies to restore biodiversity.

Agriculture↗

Synthetic community Hi-C benchmarking provides a baseline for virus-host inferences.

Microbiomes influence diverse ecosystems, and viruses increasingly appear to impose key constraints. While viromics has expanded genomic catalogs, host identification for these viruses remains challenging due to the limitations in scaling cultivation-based approaches and the uncertain reliability and relative low resolution of in silico predictions - particularly for understudied viral taxa. Towards this, Hi-C proximity ligation uses sequenced, cross-linked virus and host genomic fragments to infer virus-host linkages and has now been applied in at least ten studies. However, its accuracy remains unknown. Here we assess Hi-C performance in recovering virus-host interactions using synthetic communities (SynComs) composed of four marine bacterial strains and nine phages with known interactions and then apply optimized bioinformatic protocols to natural soil samples. In SynComs, standard Hi-C sample preparations and analyses showed poor normalized contact score performance (26% specificity, 100% sensitivity, incorrect matches up to class level) that could be dramatically improved by Z-score filtering (Z &#x2265; 0.5, 99% specificity), though at reduced sensitivity (62% down from 100%). Detection limits were established as reproducibility was poor below minimal phage abundances of 105 PFU/mL. Applying optimized bioinformatic protocols to natural soil samples, we compared virus-host linkages inferred from proximity-ligated Hi-C sequencing with predictions generated by in silico homology-based and machine learning-based bioinformatic approaches. Prior to Z-score thresholding, agreement was relatively high at the phylum to family levels (72%), but not at the genus (43%) or species (15%) levels. Z-score thresholding reduced sensitivity (only 34% of predictions were retained), with only modest improvements in congruence with bioinformatic methods (48% or 18% at genus or species levels, respectively). Regardless, this led to 79 genus-level-congruent virus-host linkages and 293 new ones revealed by Hi-C alone - i.e., providing many new virus-host interactions to explore in already well-studied climate-critical soils. Overall, these findings provide empirical benchmarks and methodological guidelines to improve the accuracy and reliability of Hi-C for virus-host linkage studies in complex microbial communities.

Genomics↗

Evaluating Language Models for Biomedical Fact-Checking: A Benchmark Dataset for Cancer Variant Interpretation Verification.

Accurate interpretation of genomic variants is critical for precision oncology but remains slow and dependent on specialized expertise. Public knowledgebases such as the Clinical Interpretation of Variants in Cancer (CIViC) help by curating literature-backed variant interpretations in a structured form, yet verification and review have become major bottlenecks. To address this, we developed CIViC-Fact, a benchmark dataset and pipeline for testing automated systems that verify the accuracy of cancer variant claims. CIViC-Fact links structured claims to sentence-level supporting or refuting evidence from full-text articles, and includes expert annotations and explanations. We evaluated multiple language models. Proprietary models performed well without training, but a smaller open-source model, fine-tuned on CIViC-Fact, achieved the highest accuracy (89%). Applying our fact-checking pipeline to real CIViC entries showed that reviewing less than 20% of content, focusing on flagged entries, would be sufficient to catch over half of all errors. This AI-assisted triage greatly accelerates the review process without replacing or reducing expert insight, ensuring that existing careful oversight remains in place while curators can work more efficiently. CIViC-Fact provides a realistic, high-consequence framework for biomedical fact-checking and a path toward more rigorous and efficient knowledgebase curation.

Journal Article↗

A benchmark for methods in reverse engineering and model discrimination: problem formulation and solutions.

A benchmark problem is described for the reconstruction and analysis of biochemical networks given sampled experimental data. The growth of the organisms is described in a bioreactor in which one substrate is fed into the reactor with a given feed rate and feed concentration. Measurements for some intracellular components are provided representing a small biochemical network. Problems of reverse engineering, parameter estimation, and identifiability are addressed. The contribution mainly focuses on the problem of model discrimination. If two or more model variants describe the available experimental data, a new experiment must be designed to discriminate between the hypothetical models. For the problem presented, the feed rate and feed concentration of a bioreactor system are available as control inputs. To verify calculated input profiles an interactive Web site (http://www.sysbio.de/projects/benchmark/) is provided. Several solutions based on linear and nonlinear models are discussed.

Algorithms↗

Benchmarking quantum computers: the five-qubit error correcting code.

The smallest quantum code that can correct all one-qubit errors is based on five qubits. We experimentally implemented the encoding, decoding, and error-correction quantum networks using nuclear magnetic resonance on a five spin subsystem of labeled crotonic acid. The ability to correct each error was verified by tomography of the process. The use of error correction for benchmarking quantum networks is discussed, and we infer that the fidelity achieved in our experiment is sufficient for preserving entanglement.

Journal Article↗

Benchmark nonperturbative calculations for the electron-impact ionization of Li(2s) and Li(2p).

Three independent nonperturbative calculations are reported for the electron-impact ionization of both the ground and first excited states of the neutral lithium atom. The time-dependent close-coupling, the R matrix with pseudostates, and the converged close-coupling methods yield total integral cross sections that are in very good agreement with each other, while perturbative distorted-wave calculations yield cross sections that are substantially higher. These nonperturbative calculations provide a benchmark for the continued development of electron-atom experimental methods designed to measure both ground and excited state ionization.

Journal Article↗