Search PubMed⌕ Search

Biomedical subjects

David Gur

Publications and source records attributed to David Gur.

At least 37 records · Page 2Linked to original sources

A conditional nonparametric test for comparing two areas under the ROC curves from a paired design.

RATIONALE AND OBJECTIVES: To develop a conditional nonparametric procedure for comparing two correlated areas under receiver operating characteristic (ROC) curves (AUC). MATERIALS AND METHODS: A nonparametric conditional test to compare areas under two ROC curves was developed using the distribution of the elements of the nonparametric AUC estimators in a permutation space. The conditioning is made on the observed discordances between the relative orderings of ratings of the normal and abnormal cases for the two modalities taken over all possible pairs. The type I error of the procedure was verified using computer simulations. The power of the test was compared with an existing unconditional procedure on simulated datasets from binormal distributions as well as from a mixture of binormal distributions of ratings. RESULTS: The proposed test is conservative for low sample sizes, large AUC, and high correlation between modalities. It possesses a reasonable type I error for sample sizes as low as 20 actually positive and 20 actually negative cases. In plausible situations in which the sample in observer performance studies can not be monotonically transformed into a binormal distribution, this approach may have modest power advantages over the conventional nonparametric test. CONCLUSION: The conditional nonparametric test presented here is an alternative approach to existing unconditional procedures and may offer advantages in certain types of observer performance studies.

Area Under Curve↗

"Memory effect" in observer performance studies of mammograms.

RATIONALE AND OBJECTIVE: To evaluate breast radiologists' recognition of mammograms showing cancers that they correctly detected or "missed" during clinical interpretations. MATERIALS AND METHODS: Two similar experiments were conducted. In the first, 33 bilateral screening mammograms were reviewed by four breast imagers. These included five cancers that each radiologist had detected, two cancers that each radiologist had "missed," and five mammograms recalled by other radiologists that were not cancer. Radiologists were asked if they had interpreted the mammogram in clinic and if the mammogram was suspicious for cancer. In the second experiment, four different breast imagers reviewed 48 mammograms that included five cancers that each radiologist had detected, two cancers that each radiologist had "missed," and five mammograms that were recalled by each radiologist but were not cancer. Using chi-square analysis, the performance of the radiologists on screening mammograms they had read in clinic was compared with their performance on mammograms read in clinic by other radiologists. RESULTS: Seven of eight radiologists did not remember interpreting any of the mammograms in clinic. One radiologist correctly remembered interpreting one mammogram in clinic, but interpreted it incorrectly. Average performance showed no significant difference (P = .60) between mammograms they had interpreted in clinic and those interpreted by others. CONCLUSION: Radiologists do not remember most mammograms showing cancer that they have interpreted, either correctly or incorrectly, after they are mixed with mammograms showing cancer that were interpreted by other radiologists. Screening mammograms can be used in observer performance studies in which the interpreting radiologist participates as an observer.

Breast Neoplasms↗

Incorporating utility-weights when comparing two diagnostic systems: a preliminary assessment.

RATIONALE AND OBJECTIVES: We sought to develop a new index that incorporates utility-weights when assessing the overall performance of a diagnostic system and to provide a statistical test for comparing two indices in a paired study design. MATERIALS AND METHODS: The area under the receiver operating characteristic (ROC) curve (AUC) was used as the basis for constructing a new index. The index we propose represents a weighted average of class-specific AUCs each of which relates to a class of pairs of actually negative (normal) and actually positive (abnormal) cases with a specific predetermined utility (or clinical importance). For each pair of normal-abnormal cases, the utility is defined a priori and based on external (covariate) information. In the proposed approach utility-weights represent the relative importance (utility) of discriminating between different types of normal and abnormal cases (pairs of the same type are combined in the classes termed utility-classes). We also describe a simple nonparametric procedure for comparing the proposed indices as computed from paired data. Computer simulations were conducted to evaluate the behavior of the type I error of the proposed test in the simple albeit important instance of two utility-classes. RESULTS: The new index provides an extension of the commonly used area under the ROC curve. It allows for incorporation of utility-weights into the analysis and reduces to the conventional AUC index when all assigned utility-weights are equal to unity. Computer simulations indicate that in the considered scenario of two utility-classes, the type I error of the proposed test is comparable to that of the conventional nonparametric test for equality of AUC indices. CONCLUSIONS: The proposed index and the statistical test provide a practical approach of incorporating utilities when comparing diagnostic systems.

Algorithms↗

Variability in observer performance studies experimental observations.

RATIONALE AND OBJECTIVES: The aim of the study is to assess variance components in observer performance studies and the possible impact on study results and conclusions. MATERIALS AND METHODS: Two previously performed retrospective receiver operating characteristic-type observer performance studies to evaluate the performance of seven radiologists in detecting interstitial disease on conventional posteroanterior chest films and nine radiologists in detecting interstitial disease on a high-resolution workstation were reanalyzed by using the Beiden, Wagner, and Campbell nine-component model to estimate the different variance components. We estimated case-, reader-, and mode-related components of the variance for the group as a whole and after excluding (round robin) each reader. Overall variance was evaluated, and the effect of individual readers on overall study conclusions was assessed. RESULTS: Overall results and conclusions of the reanalysis agreed with the original one in that, as a group, radiologists performed significantly better when using conventional films (P < .05) in both studies. Reader variability was large compared with all other components, and in one study, it was substantially larger for the workstation reading mode. Reader variability was affected substantially by one observer in each study, and in one study, reader-by-mode variability was affected by another reader who performed better on the workstation. CONCLUSION: Estimates of variance components can shed light on the appropriateness of study design, as well as the sensitivity of results to the inclusion (or exclusion) of individual observers.

Humans↗

A comparison of two data analyses from two observer performance studies using Jackknife ROC and JAFROC.

The authors compared two methodological approaches, Jackknife ROC and JAFROC, in analyzing data ascertained during FROC (free-response receiver operating characteristics) type studies. Observer rating data obtained from two observer performance studies were analyzed. During the first study, seven radiologists interpreted 120 mammography examinations depicting 57 masses under five different conditions with and without the results of computer-aided detection (CAD). In the second study, eight radiologists interpreted 110 examinations depicting 51 masses under six different display conditions with and without CAD results. Readers rated the detection task in a FROC type response. Jackknife ROC (using the software of LABMRMC with the highest rating per case) and JAFROC were used to compute differences, if any, in summary performance levels among all reading modes in each study as well as for all paired data sets. The results of the different analytical approaches are compared. The overall results for all modes were significantly different for the first study (p < 0.05) and not significant (p > 0.05) for the second study using either analytical approach. In the first study, the performance levels represented by three paired data sets were significantly different (p < 0.05) when computed using LABMRMC and four pairs were significantly different (p < 0.05) using JAFROC. In eight of ten pairs, JAFROC produced lower p values than LABMRMC. In the second study, LABMRMC showed no significant differences for any paired data sets and JAFROC showed a significant difference for one pair. In 15 of 16 pairs, p values computed by JAFROC were lower than those computed by LABMRMC.

Breast Neoplasms↗

Lung cancer screening: simulations of effects of imperfect detection on temporal dynamics.

PURPOSE: To use a mathematic model to demonstrate effects of imperfect detection on temporal dynamics of radiologic lung cancer screening. MATERIALS AND METHODS: Monte Carlo simulations of lung cancer screening programs were performed in subjects at high risk for developing cancer. The effects of detection probabilities, symptomatic presentation of tumors, tumor volume doubling time, and time between screenings were examined. Computed tomography (CT) and chest radiography models were used. RESULTS: For imperfect detection probabilities, the percentage of subjects with cancers detected with repeated screenings decreased to a steady-state value. The transition period was the period during which screenings were performed and detection rates decreased. At steady-state repeat screening, the proportion of subjects with cancers diagnosed at screening or by means of symptomatic presentation was determined by the annual probability of developing cancer and not by the sensitivity of the screening modality. The sensitivity of the screening technique did affect detected cancer size, number of interval cancers, and total number of cancers observed. CT was used to detect more total cancers over the course of the screening program and cancers with a smaller average size; moreover, fewer interval cancers were observed with CT screening than with chest radiography screening. CONCLUSION: Lung cancer screening with imperfect detection has a transition period between baseline screening and steady-state behavior of annual screenings. Advantages of CT screening include a decrease in the average cancer size at detection, a decrease in the number of observed interval cancers, and an increase in the total number of cancers observed. Steady-state behavior indicates that long-term trials of screening may not be necessary.

Humans↗

Is maximum positive predictive value a good indicator of an optimal screening mammography practice?

OBJECTIVE: Positive predictive value (PPV1) has been used as one important indicator of the quality of screening mammography programs. We show how the relationship between sensitivity and recall rate may affect the operating point at which optimal (maximum) PPV1 occurs. CONCLUSION: Optimal (maximum) PPV1 can occur at any sensitivity level and should not be used as the sole indicator for practice optimization because it does not take into account the number of cancers that would be missed at that sensitivity.

Breast Neoplasms↗

Performance and reproducibility of a computerized mass detection scheme for digitized mammography using rotated and resampled images: an assessment.

OBJECTIVE: Our objective was to compare the performance and reproducibility of a computer-aided detection (CAD) scheme that uses multiple rotated and resampled images with an in-house-developed CAD scheme (single-image-based) and a commercial CAD product in detecting masses depicted on digitized mammograms. MATERIALS AND METHODS: Ninety-two film mammograms (acquired from 23 patients) were selected. Forty-four mass regions associated with malignancy were visually identified. A commercial CAD system was used to scan and process each image four times, for a total of 368 digitized images depicting 176 mass regions. Images were processed using two CAD schemes developed in our laboratory. One uses the detection results generated from a single image, and the other averages five detection scores generated after processing the originally digitized image and four slightly rotated and resampled images. A region-based analysis was used to compare reproducibility and performance levels among the two in-house schemes and the commercial system. RESULTS: The commercial system detected a total of 98 mass regions (55.7% sensitivity) and 136 false-positive regions (an average of 0.37 per image). Among the detected mass regions, 76 represented 19 regions that were detected on all four scans and 22 represented 10 regions that were not fully reproducible. Eighty-eight false-positive detections represented 22 reproducible detections on all four scans. Our single-image-based scheme identified 87 mass regions and 160 false-positive regions. Seventeen mass regions and 28 false-positive regions were detected on all four scans. The multiple-image-based scheme identified 98 mass regions and 132 false-positive regions. Twenty-three mass regions were detected on all four scans. One hundred twelve of the 132 false-positive regions represented 28 reproducible detections. CONCLUSION: Averaging detection scores from multiple rotated and resampled images generated from a single digitization of a film can reduce variations in detection scores. Our multiple-image-based scheme improved both performance and reproducibility over the single-image-based scheme. The multiple-image-based scheme yielded an overall performance comparable to that of the commercial system but with improved reproducibility.

Breast Neoplasms↗

Computer-aided detection performance in mammographic examination of masses: assessment.

PURPOSE: To compare performance of two computer-aided detection (CAD) systems and an in-house scheme applied to five groups of sequentially acquired screening mammograms. MATERIALS AND METHODS: Two hundred nineteen film-based mammographic examinations, classified into five groups, were included in this study. Group 1 included 58 examinations in which verified malignant masses were detected during screening; group 2, 39 in which all available latest examinations were performed prior to diagnosis of these malignant masses (subset of 39 women from group 1); group 3, 22 in which findings were interpreted as negative but were verified as cancer within 1 year from the negative interpretation (missed cancers); group 4, 50 in which findings were negative and patients were not recalled for additional procedures; and group 5, 50 in which patients were recalled for additional procedures and findings were negative for cancer. In all examinations, images were processed with two Food and Drug Administration-approved commercially available CAD systems and an in-house scheme. Performance levels in terms of true-positive detection rates and number of false-positive identifications per image and per examination were compared. RESULTS: Mass detection rates in positive examinations (group 1) were 67%-72%. Detection rates among three systems were not significantly different (P > .05). In 50 negative screening examinations (group 4), false-positive rates ranged from 1.08 to 1.68 per four-view examination. Performance level differences among systems were significant for false-positive rates (P = .008). Performance of all systems was at levels lower than publicly suggested in some retrospective studies. False-positive CAD cueing rates were significantly higher for negative examinations in which patients were recalled (group 5) than they were for those in which patients were not recalled (group 4) (P < or = .002). CONCLUSION: Performance of CAD systems for mass detection at mammography varies significantly, depending on examination and system used. Actual performance of all systems in clinical environment can be improved.

Adult↗

Recall and detection rates in screening mammography.

BACKGROUND: The authors investigated the correlation between recall and detection rates in a group of 10 radiologists who had read a high volume of screening mammograms in an academic institution. METHODS: Practice-related and outcome-related databases of verified cases were used to compute recall rates and tumor detection rates for a group of 10 Mammography Quality Standard Act (MQSA)-certified radiologists who interpreted a total of 98,668 screening mammograms during the years 2000, 2001, and 2002. The relation between recall and detection rates for these individuals was investigated using parametric Pearson (r) and nonparametric Spearman (rho) correlation coefficients. The effect of the volume of mammograms interpreted by individual radiologists was assessed using partial correlations controlling for total reading volumes. RESULTS: A wide variability of recall rates (range, 7.7-17.2%) and detection rates (range, 2.6-5.4 per 1000 mammograms) was observed in the current study. A statistically significant correlation (P < 0.05) between recall and detection rates was observed in this group of 10 experienced radiologists. The results remained significant (P < 0.05) after accounting for the volume of mammograms interpreted by each radiologist. CONCLUSIONS: Optimal performance in screening mammography should be evaluated quantitatively. The general pressure to reduce recall rates through "practice guidelines" to below a fixed level for all radiologists should be assessed carefully.

Breast Neoplasms↗

Changes in breast cancer detection and mammography recall rates after the introduction of a computer-aided detection system.

BACKGROUND: Computer-aided mammography is rapidly gaining clinical acceptance, but few data demonstrate its actual benefit in the clinical environment. We assessed changes in mammography recall and cancer detection rates after the introduction of a computer-aided detection system into a clinical radiology practice in an academic setting. METHODS: We used verified practice- and outcome-related databases to compute recall rates and cancer detection rates for 24 Mammography Quality Standards Act-certified academic radiologists in our practice who interpreted 115,571 screening mammograms with (n = 59,139) or without (n = 56,432) the use of a computer-aided detection system. All statistical tests were two-sided. RESULTS: For the entire group of 24 radiologists, recall rates were similar for mammograms interpreted without and with computer-aided detection (11.39% versus 11.40%; percent difference = 0.09, 95% confidence interval [CI] = -11 to 11; P =.96) as were the breast cancer detection rates for mammograms interpreted without and with computer-aided detection (3.49% versus 3.55% per 1000 screening examinations; percent difference = 1.7, 95% CI = -11 to 19; P =.68). For the seven high-volume radiologists (i.e., those who interpreted more than 8000 screening mammograms each over a 3-year period), the recall rates were similar for mammograms interpreted without and with computer-aided detection (11.62% versus 11.05%; percent difference = -4.9, 95% CI = -21 to 4; P =.16), as were the breast cancer detection rates for mammograms interpreted without and with computer-aided detection (3.61% versus 3.49% per 1000 screening examinations; percent difference = -3.2, 95% CI = -15 to 9; P =.54). CONCLUSION: The introduction of computer-aided detection into this practice was not associated with statistically significant changes in recall and breast cancer detection rates, both for the entire group of radiologists and for the subset of radiologists who interpreted high volumes of mammograms.

Breast Neoplasms↗

Detection and classification performance levels of mammographic masses under different computer-aided detection cueing environments.

RATIONALE AND OBJECTIVES: The authors evaluated the impact of different computer-aided detection (CAD) cueing conditions on radiologists' performance levels in detecting and classifying masses depicted on mammograms. MATERIALS AND METHODS: In an observer performance study, eight radiologists interpreted 110 subtle cases six times under different display conditions to detect depicted masses and classify them as benign or malignant. Forty-five cases depicted biopsy-proven masses and 65 were negative. One mass-based cueing sensitivity of 80% and two false-positive cueing rates of 1.2 and 0.5 per image were used in this study. In one mode, radiologists first interpreted images without CAD results, followed by the display of cues and reinterpretation. In another mode, radiologists viewed CAD cues as images were presented and then interpreted images. Free-response receiver operating characteristic method was used to analyze and compare detection performance. The receiver operating characteristic method was used to evaluate classification performance. RESULTS: At these performance levels, providing cues after initial interpretation had little effect on the overall performance in detecting masses. However, in the mode with the highest false-positive cueing rate, viewing CAD cues immediately upon display of images significantly reduced average performance for both detection and classification tasks (P < .05). Viewing CAD cues during the initial display consistently resulted in fewer abnormalities being identified in noncued regions. CONCLUSION: CAD systems with low sensitivity (< or = 80% on mass-based detection) and high false-positive rate (> or = 0.5 per image) in a dataset with subtle abnormalities had little effect on radiologists' performance in the detection and classification of mammographic masses.

Area Under Curve↗

Assessment methodologies and statistical issues for computer-aided diagnosis of lung nodules in computed tomography: contemporary research topics relevant to the lung image database consortium.

Cancer of the lung and bronchus is the leading fatal malignancy in the United States. Five-year survival is low, but treatment of early stage disease considerably improves chances of survival. Advances in multidetector-row computed tomography technology provide detection of smaller lung nodules and offer a potentially effective screening tool. The large number of images per exam, however, requires considerable radiologist time for interpretation and is an impediment to clinical throughput. Thus, computer-aided diagnosis (CAD) methods are needed to assist radiologists with their decision making. To promote the development of CAD methods, the National Cancer Institute formed the Lung Image Database Consortium (LIDC). The LIDC is charged with developing the consensus and standards necessary to create an image database of multidetector-row computed tomography lung images as a resource for CAD researchers. To develop such a prospective database, its potential uses must be anticipated. The ultimate applications will influence the information that must be included along with the images, the relevant measures of algorithm performance, and the number of required images. In this article we outline assessment methodologies and statistical issues as they relate to several potential uses of the LIDC database. We review methods for performance assessment and discuss issues of defining "truth" as well as the complications that arise when truth information is not available. We also discuss issues about sizing and populating a database.

Algorithms↗

A method to test the reproducibility and to improve performance of computer-aided detection schemes for digitized mammograms.

The purpose of this study is to develop a new method for assessment of the reproducibility of computer-aided detection (CAD) schemes for digitized mammograms and to evaluate the possibility of using the implemented approach for improving CAD performance. Two thousand digitized mammograms (representing 500 cases) with 300 depicted verified masses were selected in the study. Series of images were generated for each digitized image by resampling after a series of slight image rotations. A CAD scheme developed in our laboratory was applied to all images to detect suspicious mass regions. We evaluated the reproducibility of the scheme using the detection sensitivity and false-positive rates for the original and resampled images. We also explored the possibility of improving CAD performance using three methods of combining results from the original and resampled images, including simple grouping, averaging output scores, and averaging output scores after grouping. The CAD scheme generated a detection score (from 0 to 1) for each identified suspicious region. A region with a detection score >0.5 was considered as positive. The CAD scheme detected 238 masses (79.3% case-based sensitivity) and identified 1093 false-positive regions (average 0.55 per image) in the original image dataset. In eleven repeated tests using original and ten sets of rotated and resampled images, the scheme detected a maximum of 271 masses and identified as many as 2359 false-positive regions. Two hundred and eighteen masses (80.4%) and 618 false-positive regions (26.2%) were detected in all 11 sets of images. Combining detection results improved reproducibility and the overall CAD performance. In the range of an average false-positive detection rate between 0.5 and 1 per image, the sensitivity of the scheme could be increased approximately 5% after averaging the scores of the regions detected in at least four images. At low false-positive rate (e.g., < or =average 0.3 per image), the grouping method alone could increase CAD sensitivity by 7%. The study demonstrated that reproducibility of a CAD scheme can be tested using a set of slightly rotated and resampled images. Because the reproducibility of true-positive detections is generally higher than that of false-positive detections, combining detection results generated from subsets of rotated and resampled images could improve both reproducibility and overall performance of CAD schemes.

Algorithms↗