What is "appropriate" utilization of diagnostic tests? Technical performance, diagnostic performance, and clinical utility.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
PURPOSE: To determine the diagnostic performance of technetium bone scanning in the setting of possible osteomyelitis in the foot of a patient who has diabetes or other vasculopathy. DESIGN: Meta-analysis. Data identification and study selection: To be eligible for inclusion, a report must have used intravenous technetium-99m methylene diphosphonate or a similar agent in humans over the age of 16 years, must have addressed possible osteomyelitis of the lower extremity with ulcer or soft-tissue inflammation in the setting of diabetes, neuropathy, or vasculopathy, and must have allowed the generation of a two-by-two table. A structured search of the MEDLARS database found 296 possibly eligible reports; ten met all the inclusion criteria. DATA EXTRACTION AND SYNTHESIS: The reported sensitivity and specificity of each report were converted to their logistic transforms and a straight line was fitted by weighted least-squares regression. The line was then back-transformed to yield a summary receiver operating characteristic curve. The false-positive rate of the bone scan is at best in the range of 10 to 20%. This occurs at sensitivities between 70 and 80%. The studies with increased sensitivity also reported sizable increases in the false-positive rate ranging from 20 to over 90%. Even small increases in sensitivity have necessitated large sacrifices in specificity. Seven of the ten studies reported specificities under 70%. CONCLUSIONS: Published data defining the effectiveness of technetium bone scanning for the diagnosis of osteomyelitis in the impaired foot indicate relatively poor performance. In many clinical situations, the specificity of the bone scan will not be high enough to confirm the diagnosis of osteomyelitis.
Faculty evaluations of residents' diagnostic performance in radiology subspecialty rotations were examined in two studies in order to ascertain whether the relative capabilities of residents change during training and to predict residents' diagnostic performance using a three-dimensional form perception test. In the first study, numeric ratings by faculty members were averaged to provide interpretive/diagnostic scores for each of 16 residents in each of 5 consecutive half-years. Sixty-seven percent of the relative differences among residents' diagnostic proficiencies persisted during training. The magnitude of these unchanging differences between individuals in diagnostic performance strongly favors resident selection based upon diagnostic potential. An additional 22% of the relative differences correlated with an abrupt change in the rank order of residents in diagnostic performance at the beginning of the second year of residency. This rearrangement of ranks may have resulted from an abrupt change in the tasks and expectations assigned to residents. In the second study, correlations between scores on a recently described Form Test and monthly faculty ratings of diagnostic performance were computed. An average diagnostic performance score for each resident was generated for all rotations combined, each of eleven subspecialties, and all rotations completed within a given half-year. Scores on the Form Test correlated well with these combined diagnostic scores, further substantiating results reported previously. Form Test scores were highly correlated with subspecialty diagnostic performance scores in subspecialties using cross-sectional imaging methods, such as neuroradiology. Although Form Test scores were poorly correlated with half-year diagnostic performance during the first year, they were highly correlated with performance beginning in the second year.
UNLABELLED: Effective tuberculosis (TB) management relies on prompt diagnosis of Mycobacterium tuberculosis complex (MTBC) and associated drug resistance. The Sanity 2.0 assay is a high-resolution melting assay designed for direct respiratory sample testing, enabling simultaneous detection of MTBC and resistance to rifampicin (RIF), isoniazid (INH), and fluoroquinolones (FQ) in a single step. This study evaluated its diagnostic performance in two registered multicenter trials among bacteriologically confirmed TB patients. Diagnostic performance was evaluated for MTBC detection, as well as for the identification of resistance to RIF, INH, and FQ, using phenotypic drug susceptibility testing, whole-genome sequencing, and a composite reference standard. Agreement analyses were conducted between the Sanity 2.0 assay and Xpert MTB/RIF and Xpert MTB/XDR. Among 611 patients, the Sanity 2.0 assay detected MTBC in 563 patients, exhibiting a sensitivity of 92.1% (95% CI: 89.7-94.0). For detecting resistance to RIF, INH, and FQ, sensitivities exceeded 90%, with specificities of 95.8% (95% CI: 88.5-98.6), 100.0% (95% CI: 96.4-100.0), and 97.8% (95% CI: 93.8-99.3) against the composite reference standard, respectively. The agreement with Xpert MTB/RIF for RIF detection was 98.6% (95% CI: 96.9-99.3). For INH and FQ resistance, the agreement with Xpert MTB/XDR was 92.0% (95% CI: 88.5-94.5) and 94.3% (95% CI: 91.2-96.3), respectively. The Sanity 2.0 assay is a rapid and user-friendly platform capable of detecting both MTBC and key drug resistance. It demonstrated good diagnostic performance and could potentially be an effective alternative to guide individualized anti-TB treatment, especially in resource-limited settings. IMPORTANCE: Rapid and accurate detection of both Mycobacterium tuberculosis complex (MTBC) and key drug resistance is critical to improving tuberculosis treatment outcomes and reducing transmission. However, current molecular diagnostic workflows often require sequential testing, which can delay the initiation of effective and individualized therapy. We evaluated the Sanity 2.0 assay, an integrated high-resolution melting test that simultaneously detects MTBC and resistance to rifampicin, isoniazid, and fluoroquinolone resistance directly from respiratory samples in about 2-3 hours. The assay demonstrated excellent performance, with MTBC detection sensitivity of 92.1% and drug resistance sensitivities exceeding 90% and specificities over 95% against a composite reference standard, as well as strong concordance with World Health Organization-endorsed molecular assays. Implementation of the Sanity 2.0 assay could streamline TB diagnostic workflows; enable rapid, single-step resistance profiling; and facilitate timely, individualized treatment-particularly in resource-limited settings where rapid and comprehensive resistance testing remains a critical unmet need.
In a retrospective approach the medical records and radiographs of 618 patients with midfacial trauma were reviewed by radiologists who were classified into four groups according to the length of their training and experience. Initially reported diagnoses were compared with discharge diagnoses, and the impact of training and experience on diagnostic performance was assessed. The same parameters were also investigated in a prospective study. Neither of the investigations revealed any statistically significant improvement in diagnostic performance with increased training and experience. After the basic radiologic education, innate perceptive and cognitive abilities seem to have a larger influence on radiologic diagnostic performance than training and experience.
BACKGROUND: Cancer antigen 125 (CA125) is widely recognized as a useful biomarker for the surveillance of patients with ovarian and other cancers. Prior genome-wide association studies have identified variants that influence CA125 levels. We evaluated the utility of stratifying CA125 levels by such variants and evaluated diagnostic performance in control subjects and patients with pancreatic ductal adenocarcinoma (PDAC). METHODS: We measured CA125 levels in 807 control subjects and 450 patients with PDAC and genotyped 10 variants involving four genes (GAL3ST2, MSLN, D2HGDH, and MUC16). We compared CA125 levels in controls by variant and generated variant-defined CA125 cutoffs and then classified cases and controls into functional groups based on their variant profile. We used this variant classification to evaluate the diagnostic performance of CA125 in patients with PDAC. RESULTS: Six variants associated with CA125 levels were used to group controls into one of four groups. Mean CA125 levels in the highest variant group were approximately fourfold higher than in the lowest group. African Americans were more likely to have a variant group associated with low CA125 levels. After setting diagnostic cutoffs by variant group, the diagnostic sensitivity of CA125 for PDAC was 20.2% at 98% specificity (areas under the ROC curve, 0.702), not significantly different from a uniform CA125 diagnostic cutoff (areas under the ROC curve, 0.700). CONCLUSIONS: Gene variants can be used to generate personalized CA125 reference ranges. This approach did not significantly improve CA125's diagnostic performance for pancreatic cancer, but it merits evaluation in other diagnostic settings, such as detecting ovarian cancer. IMPACT: Gene variants can be used to personalize CA125 levels.
To determine the diagnostic performance of magnetic resonance (MR) imaging in the evaluation of suspected rotator cuff tears, eight asymptomatic volunteers and 32 patients with rotator cuff tendonopathy who underwent surgery were examined with MR imaging. Twenty-four of these patients also underwent contrast arthrography. The ability of MR imaging to depict the size of cuff tears and the quality of torn tendon edges was also evaluated. The MR imaging and arthrographic studies were reviewed without knowledge of surgical results or of the other studies. A scoring system was developed and a score assigned to each patient's MR study. The sensitivity of MR imaging for all tears (partial and full thickness) was 0.91, and the specificity was 0.88; whereas the sensitivity and specificity of arthrography were each 0.71. The scoring system improved the sensitivity to 1.0 and the specificity to 0.92. Linear regression analysis showed excellent correlation between preoperative assessment of the size of rotator cuff tears and measurement at surgery (r = .95).
This study investigated the diagnostic performance of the MMPI validity and clinical scales, and especially of Scale 7 (Pt), for the DSM-IIII--R obsessive-compulsive personality disorder by comparing the MMPI variables for 24 obsessive-compulsive with those for 58 nonobsessive-compulsive inpatients. Both groups were diagnosed by semistructured interview (SCID-II). The obsessive-compulsive group obtained for the mean MMPI profile a 2-(6-1) (D-Pa-Hs) code, with a tendency for a lowered Scale 4 (Pd) score, compared to the nonobsessive-compulsive group. Neither the ROC analysis of the individual MMPI scales, including Scale 7 (Pt), nor the analyses of frequency of two-point codes and elevated (T greater than 69) scales showed any clear indications of good diagnostic performance for the DSM-III--R obsessive-compulsive personality disorder.
BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2×2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I²=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.
ROC analysis permits an evaluation of the diagnostic performance of imaging systems by incorporating the individual decision threshold of the diagnostician that determines the relation between sensitivity and specificity of a tested system. It therefore facilitates the comparison between different imaging systems and observers. A thorough knowledge and sophistication of this methodology is of special importance in the comparison of digital and conventional projection radiography because the images in question may have a very similar character. The method is explained and special points of interest (standard of "truth", size of sample, lesion conspicuity and localisation, viewing time and determination of significance) are emphasised.
To determine the effect of clinical information on the radiological diagnostic performance in middle-face injury, the medical records and radiographs of 618 patients with middle-face injury were reviewed. The information value of clinical data given in each x-ray requisition was evaluated. The radiological diagnoses were compared with the final clinical diagnoses both retrospectively and prospectively (with and without clinical data). Knowledge of clinical data with statistical significance changed the reader's decision threshold towards improved sensitivity but poorer specificity. Clinical data did not improve the radiologists' performance. Clinical data may be helpful in diagnostic tasks rather by increasing the sensitivity than the specificity.
This paper presents a method for performing diagnostic x-ray shielding calculations which is somewhat different than that given in NCRP Report No. 49. The method computes exposure at the location to be shielded from multiple sources of radiation in the room and accounts for differences in transmission characteristics of leakage and primary/scatter radiation. Also discussed in the paper is a method for determining shielding credit for existing common structural materials. Methods and results derived in the paper are then discussed and compared with NCRP Report No. 49.
Literature data on the diagnostic performance of phlebography, myelography, and CT scan applied to patients with suspected lumbar disk herniation (LDH) are analyzed to extract maximal information about their relative discriminatory power. Seventeen papers meeting the selection criteria contain 13 reports on myelography, 6 on phlebography, and 5 on CT. Sensitivity and specificity are considered simultaneously in logistic ROC space. The reports of each procedure are effectively summarized by a linear regression in logistic ROC space. Taking into account the individual confidence regions of sensitivity and specificity obtained from each report, the slope of the regression line is estimated by Generalized Least Squares (ML). This approach also allows to test the assumption of a common odds ratio (i.e., of a unit slope). The simply to determine common odds ratio as well as the perpendicular distance between the origin and the regression line (as a good approximation to the area under the ROC curve) are used as a measure for the discriminatory power of the procedures. For CT, homogeneity of sensitivity turns out to be much more likely than a common odds ratio. Based on the available, retrospective data, phlebography appears to have the highest performance in visualizing an LDH, followed by myelography and CT.
The diagnostic performance of the Hearing Handicap Inventory for the Elderly--Screening Version (HHIE-S) was evaluated against five definitions of hearing loss in 178 elderly subjects screened in primary care. Hearing loss was assessed by pure-tone audiometry. Using a score of greater than 8 as a cut point, the HHIE-S had sensitivities ranging from 53 to 72% and specificities ranging from 70 to 84% with the different definitions. The HHIE-S receiver-operating characteristics and likelihood ratios were similar regardless of hearing loss definition used. The HHIE-S is a valid, robust test for identifying hearing-impaired elderly, irrespective of the audiometric definition used to finally diagnose hearing difficulties.
We present a general spreadsheet model for evaluating diagnostic performance of clinical tests. Our model depicts test results as an r X c matrix, with r possible test results and c possible clinical states. Analysis of this matrix is based on the Ri/Cj ratio, calculated as a number of subjects having a specified result Ri within a given clinical state Cj, divided by total subjects within this clinical state. From this model, we can identify three special cases: (1) a 2 X c matrix, with two possible test results of T+ or T-, over c possible clinical states; (2) an r X 2 matrix, with r possible test results, over two possible clinical states of D+ or D-; and (3) a 2 X 2 matrix, with two possible test results over two possible clinical states. Application of the Ri/Cj ratio to the r X c matrix provides a useful approach to graphic analysis of multiple test results over multiple clinical states. The Ri/Cj ratio also provides a general approach to Bayesian analysis, in which likelihood ratio, relative operating characteristic analysis, sensitivity, and specificity represent special cases or special applications.
A digital system for chest radiography based on a large image intensifier was compared with a conventional film-screen system. The diagnostic performance was evaluated with special reference to the digital monitor images with a modified version of receiver operating characteristic (ROC) analysis--free response ROC (FROC) analysis--on a chest equivalent phantom. Measurements of spatial resolution and energy imparted were also performed. The detectability of low-contrast objects as well as spatial resolution was better for the full-size film-screen radiographs than for both the digital monitor images and the 100 mm photofluorograms. The image-intensifier system has a potential for considerable dose savings in relation to the conventional technique provided that fluoroscopy is excluded in the positioning of the patients.
Immunoassay of creatine kinase-MB provides numerical information, which makes it possible to estimate quantitatively the diagnostic performance following myocardial infarction. Two graphical methods for such an evaluation are presented. The relationship between technical sensitivity and specificity was analyzed using a continuous function, the receiver-operator characteristic curve. Using this function, a diagnostic threshold ("upper normal limit") was chosen, which balances both technical specificity and sensitivity on the first day following infarction. On the second and third days this threshold caused a progressive loss of technical sensitivity. The relationship between effectiveness and the prevalence of myocardial infarction in the tested population was evaluated with a nomogram correlating these two quantities. With the chosen diagnostic threshold, effectiveness is independent of prevalence on the first day, but the loss in technical sensitivity on subsequent days causes effectiveness to decay when the prevalence is high.
The survey programs in surgical pathology and cytopathology of the College of American Pathologists are described, along with a summary of the data produced on diagnostic performance in cytopathology. The levels of agreement between individual diagnoses and consensus diagnoses were highest for slides in the "negative" and "positive" categories and somewhat less for slides in the "suspicious" categories. Submission of slides with 70% or less agreement and slides on which a consensus had not been reached to a new group of observers yielded similar levels of agreement, indicating that there is a fairly small percentage of cervical smears that are diagnostic problems. Agreement levels, especially in relation to false negatives and false positives, represent a commendable achievement, especially considering that cervical cytology is a screening procedure and that "suspicious" smears result in further investigations. It is also important to recognize that additional clinical data and studies are often taken into consideration before definitive treatment is begun or additional follow-up smears are taken.