Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reference Standards”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

A reliability study for evaluating information extraction from radiology reports.

GOAL: To assess the reliability of a reference standard for an information extraction task. SETTING: Twenty-four physician raters from two sites and two specialties judged whether clinical conditions were present based on reading chest radiograph reports. METHODS: Variance components, generalizability (reliability) coefficients, and the number of expert raters needed to generate a reliable reference standard were estimated. RESULTS: Per-rater reliability averaged across conditions was 0.80 (95% CI, 0.79-0.81). Reliability for the nine individual conditions varied from 0.67 to 0.97, with central line presence and pneumothorax the most reliable, and pleural effusion (excluding CHF) and pneumonia the least reliable. One to two raters were needed to achieve a reliability of 0.70, and six raters, on average, were required to achieve a reliability of 0.95. This was far more reliable than a previously published per-rater reliability of 0.19 for a more complex task. Differences between sites were attributable to changes to the condition definitions. CONCLUSION: In these evaluations, physician raters were able to judge very reliably the presence of clinical conditions based on text reports. Once the reliability of a specific rater is confirmed, it would be possible for that rater to create a reference standard reliable enough to assess aggregate measures on a system. Six raters would be needed to create a reference standard sufficient to assess a system on a case-by-case basis. These results should help evaluators design future information extraction studies for natural language processors and other knowledge-based systems.

Evaluation Studies as Topic↗

Location of adenomas missed by optical colonoscopy.

BACKGROUND: Previous estimates of the adenoma miss rate with optical colonoscopy (OC) are hindered by the use of OC as its own reference standard. OBJECTIVE: To evaluate the frequency and characteristics of colorectal neoplasms that are missed prospectively on OC by using virtual colonoscopy (VC) as a separate reference standard. DESIGN: Prospective, multicenter screening trial. SETTING: 3 medical centers. PARTICIPANTS: 1233 asymptomatic adults who underwent same-day VC and OC. MEASUREMENTS: Colorectal neoplasms (adenomatous polyps) missed at OC before VC results were unblinded. RESULTS: Fourteen (93.3%) of 15 nonrectal neoplasms were located on a fold; 10 (71.4%) of these were located on the backside of a fold. Five (83.3%) of 6 rectal lesions were located within 10 cm of the anal verge. LIMITATIONS: Estimation of the OC miss rate depended on polyp detection on both VC and second-look OC and therefore underestimates the true OC miss rate, particularly for smaller polyps. CONCLUSIONS: Most clinically significant adenomas missed prospectively on OC are located behind a fold or near the anal verge. The 12% OC miss rate for large adenomas (>or=10 mm) when state-of-the-art 3-dimensional VC is used as a separate reference standard is increased from the previous 0% to 6% estimates derived by using OC as its own reference standard.

Adenoma↗

Clinical estimation of trunk position among mechanically ventilated patients.

OBJECTIVES: Trunk position at 45 degrees from the horizontal is associated with a decreased risk of gastroesophageal aspiration. The objectives of this study were to determine the accuracy of trunk flexion estimates compared to a reference standard measurement, and to determine agreement about trunk flexion among ICU clinicians. DESIGN: Prospective observational study. SETTING: Two university-affiliated medical-surgical ICUs. PATIENTS AND PARTICIPANTS: Thirty-three mechanically ventilated ICU patients, seven residents, two fellows, three intensivists, and twenty-eight bedside nurses. INTERVENTIONS: Prospectively, concurrently, and independently during rounds, one bedside nurse, one resident, one fellow, and one intensivist clinically estimated the trunk flexion of mechanically ventilated patients. To record the reference standard, a trained investigator measured trunk position in the vertical plane using a goniometer. MEASUREMENTS AND RESULTS: We made 438 clinical assessments on 33 patients aged 57.2+/-19.4 (SD) years with an APACHE II score of 27.3+/-9.4. Mean trunk flexion estimates were: nurses 24.3+/-12.3 degrees from the horizontal, residents 20.2+/-13.7, fellows 20.3+/-10.8, and intensivists 21.1+/-13.1 compared to the reference standard measurement 16.2+/-9.0 degrees. The accuracy of trunk flexion estimates was fair to moderate [intraclass correlation for reference standard versus nurses (ICC 0.42), residents (ICC 0.52), fellows (ICC 0.36), and intensivists (ICC 0.55)]. The agreement among different groups of clinicians was moderate. CONCLUSIONS: In mechanically ventilated patients, we found that clinical estimates of trunk position were moderately good, agreement amongst caregivers was moderately good, but that all clinicians tended to overestimate the angle of semirecumbency.

APACHE↗

Determination of serum ferritin by a one-step immunoenzymoassay, and comparison of four liver-ferritin standards.

We describe a one step "sandwich"-type immunoenzymoassay for ferritin in human serum. The solid-phase consists of glutaraldehyde-treated polypropylene tubes coated with rabbit antibody to human ferritin. Liver ferritin is the standard. Peroxidase-conjugated antiserum to ferritin and a sensitive chromogen, o-phenylenediamine, are used. The assay requires 90 min. The standard curve is linear up to 400 micrograms of ferritin per liter of serum. Within- and between-run CVs are less than 6% for low, high, and medium concentrations and are about 13.0% at the decision level for iron deficiency. Results by a two-step "sandwich" procedure (New England Immunology Associates kit) correlated well, r = 0.98. We assessed four liver ferritin standards from different manufacturers with the described method. The mean absorbance for the 40 micrograms/L ferritin standard was 1.5 for that from Diagnostics Biochem and National Institute for Biological Standards and Controls, 1.0 for that from Dako, and 0.4 for that from Sigma. Consequently, to standardize results, all liver ferritin standards should be calibrated vs the National Institute for Biological Standards and Controls reference standard.

Animals↗

Ventilator-associated pneumonia in intubated children: comparison of different diagnostic methods.

OBJECTIVES: To compare different methods for diagnosis of ventilator-associated pneumonia in intubated children. DESIGN: Prospective epidemiologic study. SETTING: Pediatric intensive care unit of a tertiary care university hospital. PATIENTS: All consecutive pediatric intensive care unit patients <18 yrs of age with suspected ventilator-associated pneumonia. INTERVENTIONS: For all patients, the following diagnostic methods were compared: 1) clinical data using Centers for Disease Control criteria; 2) blind protected bronchoalveolar lavage, evaluating quantitative cultures, bacterial index of >5, Gram stain, and presence of intracellular bacteria; and 3) nonquantitative cultures of endotracheal secretions. The reference standard used was clinical judgment of three independent experts (Delphi method) who retrospectively established by consensus the presence of ventilator-associated pneumonia based on clinical, microbiological, and radiologic data. Concordance between each diagnostic method and the reference standard was evaluated by concordance percentage and kappa score. Validity was evaluated using sensitivity, specificity, positive predictive value, negative predictive value, and global value. RESULTS: A total of 30 patients were included in the study. According to the reference standard, ventilator-associated pneumonia occurred in 10 of 30 patients (33%). Bacterial index of >5 in bronchoalveolar secretions showed the best concordance compared with the reference standard (concordance, 83%; kappa, 0.61). Bacterial index of >5 also showed the best validity (sensitivity, 78%; specificity, 86%; positive predictive value, 70%; negative predictive value, 90%; global value, 90%). Intracellular bacteria and Gram stain from bronchoalveolar secretions were very specific (95% and 81%, respectively) but not sensitive (30% and 50%, respectively). Clinical criteria and endotracheal cultures were very sensitive (100% and 90%, respectively) but poorly specific (15% and 40%, respectively). CONCLUSION: Our data show that the most reliable diagnostic method for ventilator-associated pneumonia is a bacterial index of >5, using blind protected bronchoalveolar lavage. Further studies should evaluate the validity of all these methods according to the gold standard (autopsy).

Bronchoalveolar Lavage Fluid↗

Can Medicare billing claims data be used to assess mammography utilization among women ages 65 and older?

BACKGROUND: Medicare data may be a useful source for determining the utilization of mammography among elderly women, but the accuracy of these data has not been established. OBJECTIVE: We determined whether Medicare physician billing claims are an accurate reflection of mammography utilization among women ages 65 and older and whether they can be used to assess the use of screening as compared with diagnostic mammography. DATA SOURCES: Mammography use was assessed using Medicare billing claims and radiology reports from 2 mammography registries; the San Francisco Mammography Registry and the New Mexico Mammography Registry. METHODS: Completeness of the Medicare data was assessed by comparing mammography use based on Medicare, with radiology reports from the mammography registries, which served as the referent standard. Capture rates for Medicare claims for individual mammograms were examined, and women were characterized as having undergone at least 1 mammogram within each 2-year period based on the Medicare data, and these rates were compared with the referent standard. To determine whether Medicare data can distinguish between screening and diagnostic mammography, we performed a classification analysis using the mammography registries screening/diagnostic designation as the referent standard (dependent variable) and Medicare claim information as the independent/predictor variable. On the basis of the mammogram level classification analysis, women were categorized as having been frequently screened (at least 2 screening mammograms spaced by 12 to 36 months), screened (at least 1 screening mammogram), or not screened. SUBJECTS: Women ages 65 and older, diagnosed with breast cancer between 1992-1999, who had at least 1 mammogram between 1992-1999 were examined. RESULTS: A total of 3340 mammograms were obtained in 1371 women between 1992 and 1999. Overall, 83% of mammograms obtained by these women had a corresponding billing claim in Medicare. This increased from 65% in 1992 to 90% in 1999. Of women who underwent at least 1 mammogram during each 2-year period per the referent standard, 94% of women were accurately classified by Medicare claims as also having undergone mammography during the same 2-year period. In multivariable analysis, a mammogram was more likely to be associated with a billing claim over time, for women 80 years or older, and for white and Asian as compared with Hispanic women. Neither socioeconomic status nor screening/diagnostic designation affected the likelihood that a mammogram would be associated with a billing claim. The Medicare data accurately categorized a given mammogram as screening or diagnostic for 87.5% of mammograms. Lastly, there was moderate to substantial agreement in the categorization of women as frequently screened, screened or not screened between the 2 data sets (weighted kappa 0.74, 95% confidence interval 0.70-0.78). CONCLUSION: Medicare administrative claims are reliable for assessment of mammography utilization and have become more accurate over time. Medicare claims data also provide a mechanism for designating mammography as screening or diagnostic, which subsequently may allow accurate description of a woman's screening history.

Age Distribution↗

Gene-specific dye bias in microarray reference designs.

The most widely used microarray experiment design includes the use of a reference standard. Comparisons of gene expression between samples are facilitated because each sample is directly measured against the reference standard, using two fluorescent dyes. Numerous reports indicate that some genes incorporate the two commonly used dyes with different efficiencies, contributing to inaccurate data. However, it is widely assumed that these effects will not corrupt results if the reference standard is labeled with the same dye on each microarray. We demonstrate that this assumption is not reliable and that dye orientation can significantly influence measured changes in gene expression.

Animals↗

Gastric-juice ammonia assay for diagnosis of Helicobacter pylori infection and the relationship of ammonia concentration to gastritis severity.

OBJECTIVES: To determine the test characteristics of gastric-juice ammonia concentration as measured by an ion-selective electrode and a rapid ammonia detection device for the diagnosis of Helicobacter pylori infection and to assess the relationship between gastric-juice ammonia concentration and the severity of gastritis. METHODS: Patients undergoing upper endoscopy had collection of gastric juice that was tested for ammonia using an ion-selective electrode and a rapid ammonia assay device that uses a pH-indicating membrane. A receiver operating characteristic curve was calculated for ammonia concentration. Severity of gastritis was graded using the Sydney classification (1) and correlated to gastric-juice ammonia concentration. Patients also underwent H. pylori testing by IgG serology, rapid urease testing, and histological special stain. Ammonia testing results were compared with a reference standard of two of three positive tests and with a second reference standard of a positive serology. RESULTS: 73 patients underwent endoscopy and collection of gastric juice. The receiver operating characteristic curve indicated an optimal cutoff value of 5 mM, yielding a sensitivity of 67%, specificity of 93%, positive predictive value of 67%, and negative predictive value of 93% (compared with the combined reference standard). The rapid NH3-testing device yielded a sensitivity of 83%, specificity 63%, positive predictive value 31%, and negative predictive value 95%. The severity of neutrophilic (p = 0.001) and mononuclear cell (p = 0.003) infiltration were significantly correlated with gastric-juice ammonia concentration. CONCLUSIONS: Measurement of gastric-juice ammonia concentration by ion-selective electrode or rapid detection device is a relatively insensitive and nonspecific means of H. pylori diagnosis. Gastritis severity increases with gastric-juice ammonia concentration.

Ammonia↗

The effects of standardization and reference values on patient classification for spine and femur dual-energy X-ray absorptiometry.

The effect of two methods for standardizing dual-energy X-ray absorptiometry (DXA) measurements on patient classification by the T-score has been determined for a group of over 2000 patients. The methods proposed by the International DXA Standardization Committee and the European Community's COMAC-BME group were used in conjunction with young reference data from the major DXA manufacturers, the COMAC-BME group and the third US National Health and Nutrition Examination Survey (NHANES III). The two standardization techniques produced dissimilar classifications as measured by the kappa statistic (kappa = 0.34-0.90), especially for the femoral neck, with up to 24.3% of patients reclassified from osteopenic to normal and 18.6% reclassified from osteoporotic to osteopenic when the standardization method was changed. Considering the effects of both reference data and standardization techniques together, there was a wide variation of patient classification, with the number of patients classified as osteoporotic varying from 9.6% to 21.1% for the postero-anterior spine L2-4 region and from 2.3% to 27.6% for the femoral neck. The agreement between different classifications ranged widely, from very poor to excellent (kappa = 0.02-0.98). The creation of standardized reference data must be an important priority in order to harmonize patient management using standardized BMD measurements. The choice of standardization technique, however, must be addressed in light of the results presented here.

Absorptiometry, Photon↗

A study on the reference electrode standardization technique for a realistic head model.

One oldest technical problem in EEG practice is the effect of an active reference on EEG recording, and it is especially important for identifying the temporal information of EEG recordings. To solve this problem, a reference electrode standardization technique (REST) has been proposed for a concentric three-sphere head model. REST, based on an equivalent distributed source model, reconstructs the potential with a reference at infinity from the potential with a scalp point reference or with the average reference. In this paper, investigated was the REST for a realistic head model. The results of simulation studies show that the potential reconstruction for the realistic head model is more sensitive to noise than that for the concentric three-sphere head model, so a regularized inverse by truncated singular value decomposition was introduced. The results confirm that REST is still an efficient method even for a realistic head model especially for the most important superficial cortex region.

Algorithms↗

The present status of serodiagnosis and seroepidemiology of schistosomiasis.

The literature of the past 4-5 yr on serodiagnosis and seroepidemiology of schistosomiasis is reviewed. A variety of assays with different antigens are being used for serodiagnosis. Several purified antigens appear to be sensitive and specific, but have little if any capability of indicating duration of infection, parasite burden, or effect of chemotherapy. The results of long-term posttherapy field studies indicate that serology has a role in monitoring control programs. Standardized serologic assays and the need for International Standard Reference Sera are emphasized. A standardized enzyme-linked immunosorbent assay based on the Falcon Assay Screening Test system (FAST-ELISA), and involving a standard reference serum pool, is suitable for both serodiagnosis and field studies. Measurement of circulating antigens as a parameter of active infection is considered to have increased potential, compared with antibody measurement, in management of clinical disease and in control programs. Recombinant DNA technology may be useful for producing standard antigens for use in assays measuring antibody or circulating antigen. Time-resolved immunofluorescence involving europium-labeled conjugates may provide the increased assay sensitivity needed for measurement of circulating antigen.

Animals↗

[Effect of the contour line on cup surface using the Heidelberg Retina Tomograph].

BACKGROUND: Significance of topometric follow-up examinations of the optic nerve head in glaucomatous eyes depends on the reproducibility of the calculated parameters. Since the definition of the standard reference plane in software version 1.11 of the Heidelberg Retina Tomograph has been changed, intrapapillary parameters depend directly on the position of the contourline in the sector between -10 degrees to -4 degrees, and therefore on the observer variability to determine the disc border. We evaluated intra- and interobserver variability and present a simple approach to increase reproducibility. METHOD: The disc border of 4 glaucomatous eyes, 3 ocular hypertensive eyes and 3 eyes of healthy subjects were traced by two observers, 5 times using the free draw mode and 5 times by the addition of contourline circles. RESULTS: We found a median variability of the mean disc radius in sector -10 degrees to -4 degrees of 51 microns, which defines the position of the standard reference plane, resulting in a median variability of the position of the standard reference plane of 33 microns which caused a variability of 81 microns2 of the cup area. Addition of contourline circles smoothing the final contourline along the border of the optic disc resulted in a decrease of the coefficient of variation of the standard reference plane of 3.76% (6.76% vs. 3.0%), of the cup area of 2.34% (3.87% vs. 1.53%) and of the rim volume of 3.41% (9.75% vs. 6.34%). CONCLUSION: The calculation of the cup area using software version 1.11 of the Heidelberg Retina Tomograph depends on observer variability. The addition of contourline circles to define the final contourline along the disc border increases reproducibility. However, in follow-up of topometric examinations of the optic nerve head the software supported transfer mode should be used. Comparing topometric data of an individual optic disc in follow-up suppose the same definition of the contourline. Therefore, topometric data evaluated using software version 1.10 or earlier needs to be recalculated.

Adolescent↗

Data collection for sexually transmitted disease diagnoses: a comparison of self-report, medical record reviews, and state health department reports.

PURPOSE: To compare three methods of data collection on case ascertainment of past chlamydia or gonorrhea diagnoses. METHODS: Data collection for 361 adolescent females between 1998 and 2000 included: 1) face-to-face interviews; 2) computerized and paper medical record reviews; and 3) chlamydia and gonorrhea reports to the state health department. Statistical methods include latent class and composite reference standard analyses. RESULTS: The estimated prevalence of past diagnoses did not differ significantly by data collection method for chlamydia (20.5%, 23.0%, and 19.7% by self-report, medical record reviews, and state health department reports, respectively) or gonorrhea (4.7%, 6.9%, and 5.5%, respectively) during the 2-year study period. The estimated latent class and composite reference standard prevalences for chlamydia were 23.5% and 26.9%, respectively (p=.04 and p < .01 for differences from self-report alone, respectively). For gonorrhea, the estimated latent class and composite reference standard prevalences were 7.8% and 6.9%, respectively (p < .01 for both differences from self-report alone). Kappa scores for self-report compared with the latent class and composite reference standard prevalences ranged from .67 to .80, and the magnitude of under-reporting ranged from 21% to 47%. CONCLUSIONS: The similar case ascertainment from the three sources separately and high reliability of self-report, coupled with its feasibility and low cost, suggest that self-report is a viable data collection method for STD diagnoses. However, using multiple sources may be preferable when time and resources permit given that under-reporting by self-report is likely to occur (particularly for gonorrhea) and that greater case ascertainment can be achieved.

Adolescent↗

Chondroitin product selection for the glucosamine/chondroitin arthritis intervention trial.

OBJECTIVE: To select a high-quality chondroitin dosage form and/or an appropriate source of sodium chondroitin for the National Institutes of Health's Glucosamine/Chondroitin Arthritis Intervention Trial (GAIT). DESIGN: Controlled experimental trials. SETTING: Laboratory. PATIENTS OR PARTICIPANTS: Not applicable. INTERVENTIONS: Commercially available chondroitin products were reviewed, and purified sodium chondroitin from two suppliers was evaluated through tests (infrared and near-infrared identification, moisture content, pH, optical rotation, color and clarity of aqueous solutions prepared from the powders, protein contamination, total residue following ignition and nitrogen content, determination of sodium chondroitin molecular weight, disaccharide analysis, and measurement of chondroitin, sodium, and total glycosaminoglycan content) and an onsite supplier audit. MAIN OUTCOME MEASURES: Purity, potency, and quality of sodium chondroitin powders. RESULTS: No commercially available chondroitin product was deemed appropriate for use in GAIT. Samples of sodium chondroitin powder from two suppliers exhibited similar disaccharide and glycosaminoglycan content. Each contained approximately 2% hyaluronic acid and 8%-9% unsulfated disaccharide. Potency was inconsistent across groups, which might have resulted from different analytical methods and choice of reference standard. Mean potency obtained by five separate methods ranged from 82.2% to 95.5% for one supplier, 92.5% to 110.1% for another, and 95.1% to 112.5% for a commercially obtained reference standard. Critical issues raised by the results include choice of reference standard, selection of assay method, and the consistent appearance of an unidentifiable contaminant present in all three lots from one supplier. CONCLUSION: This blinded study determined methods to identify acceptable agents and provided results, which, in addition to regulatory compliance supplier audits, formed the basis for chondroitin product selection in GAIT.

Chondroitin↗

The accuracy of electrocardiographic Q waves for the detection of prior myocardial infarction as assessed by a novel standard of reference.

BACKGROUND: The electrocardiogram (ECG) is valuable for the identification of prior myocardial infarction (MI) in individuals participating in epidemiologic studies or undergoing screening examinations. Although the Minnesota Code, a set of criteria for the interpretation of ECGs in such situations, is commonly used to identify MI in these settings, its accuracy is incompletely understood. HYPOTHESIS: We sought to test the accuracy of the Minnesota Code Q and QS criteria for MI against a new standard of reference, the presence of a perfusion defect on a resting myocardial scintigraphic image. METHODS: The resting myocardial scintigrams of all patients studied in our nuclear cardiology laboratory during 7 consecutive months were screened for the presence of perfusion defects. For each patient with such a defect, two individuals examined on the same day, who had no perfusion defect, were selected as controls. Electrocardiograms recorded within 30 days of the scintigraphy were read blindly by two of the authors using the Minnesota Code criteria for Q or QS waves indicative of MI. RESULTS: For 214 patients selected on the basis of their scintigraphic findings, a satisfactory ECG recorded within a month of the scintigraphy was also available. The overall sensitivity of the Q or QS criteria was 0.58 and the specificity was 0.75. As might be expected when only the most stringent criteria were applied, sensitivity was least and the specificity best. CONCLUSIONS: As in previous studies, in which necropsy material served as the standard of reference, sensitivity of the Q and QS criteria contained in the Minnesota Code is relatively modest and specificity is reasonable but not outstanding.

Electrocardiography↗

Analytical method for the determination of O,O-diethyl phosphate and O,O-diethyl thiophosphate in faecal samples.

A residue analytical method was developed for the determination of the dialkylphosphate metabolites of parathion in faecal samples obtained from rabbits. The faecal pieces were homogenised in water and highly water-soluble O,O-diethyl phosphate (DEP) and O,O-diethyl thiophosphate (DETP) were subsequently alkylated to pentafluorobenzyl esters by a phase transfer reaction. Derivatisation yields depend on the reaction time. The recovery rates were determined over the complete procedure using authentic reference standards in matrix solution. The reference standards allow to observe an effect of the sample matrix on the area of signals while GC-FPD is used. The recoveries over the concentration range from 0.05 to 5 microg/g were 47-62% for O,O-diethyl phosphate and 92-106% for O,O-diethyl thiophosphate potassium salt with FPD.

Alkylation↗

How to use diagnostic test articles in the intensive care unit: diagnosing weanability using f/Vt.

Medical diagnosis involves generating a set of hypotheses and obtaining information that modifies these hypotheses. Sources of this information include the history, physical examination, and laboratory investigations, all of which function as diagnostic tests. Studies of diagnostic tests are useful when a) the population under study is representative of those to whom we would like to apply the results; b) when an independent, blind comparison is made of the test results with a reference standard; and c) when the reference standard is performed on all patients, rather than restricted to those patients with particular test results. Clinicians can use the data from such high quality studies in the form of sensitivity and specificity, as well as likelihood ratios, which indicate the direction and magnitude of the change in probability of a target condition from pretest to posttest. Study results will be more easily applicable to practice when the performance and interpretation of the test is similar in study and clinical settings. We conduct diagnostic tests primarily to improve the process of patient care and patient outcome, and test ordering behavior ideally reflects these goals.

Critical Care↗