Search PubMed⌕ Search

Biomedical subjects

Margaret S Pepe

Publications and source records attributed to Margaret S Pepe.

9 recordsLinked to original sources

Using the ROC curve for gauging treatment effect in clinical trials.

Non-parametric procedures such as the Wilcoxon rank-sum test, or equivalently the Mann-Whitney test, are often used to analyse data from clinical trials. These procedures enable testing for treatment effect, but traditionally do not account for covariates. We adapt recently developed methods for receiver operating characteristic (ROC) curve regression analysis to extend the Mann-Whitney test to accommodate covariate adjustment and evaluation of effect modification. Our approach naturally extends use of the Mann-Whitney statistic in a fashion that is analogous to how linear models extend the t-test. We illustrate the methodology with data from clinical trials of a therapy for Cystic Fibrosis.

Administration, Inhalation↗

Comparing the predictive values of diagnostic tests: sample size and analysis for paired study designs.

BACKGROUND: Although statistical methodology is well developed for comparing diagnostic tests in terms of their sensitivities and specificities, comparative inference about predictive values is not. PURPOSE: In this paper we consider the design and analysis of studies comparing the positive and negative predictive values of two diagnostic tests that are measured on all subjects. METHODS: We focus on comparing tests using the relative positive and negative predictive values. We discuss directly estimating these quantities from the data and derive analytic variance expressions. Sample size formulas for study design ensue. RESULTS: We analyze data on patients with cystic fibrosis to illustrate the methodology. This approach is compared and contrasted with an existing regression framework that can also be used for similar analysis purposes and yields similar results. CONCLUSIONS: We have developed a new approach for comparing the predictive values of two tests that gives rise to sample size formulas for study design.

Biometry↗

Inactivation of p16, RUNX3, and HPP1 occurs early in Barrett's-associated neoplastic progression and predicts progression risk.

Patients with Barrett's esophagus (BE) are at increased risk of developing esophageal adenocarcinoma (EAC). Clinical neoplastic progression risk factors, such as age and the length of the esophageal BE segment, have been identified. However, improved molecular biomarkers predicting increased progression risk are needed for improved risk assessment and stratification. Using real-time quantitative methylation-specific PCR, we screened 10 genes (HPP1, RUNX3, RIZ1, CRBP1, 3-OST-2, APC, TIMP3, p16, MGMT, p14) for promoter hypermethylation in 77 EAC, 93 BE, and 64 normal esophagus (NE) specimens. A subset of genes manifesting significant differences in methylation frequencies between BE and EAC was then analysed in 20 dysplastic specimens. All 10 genes except p14 were frequently methylated in EACs, with RUNX3, HPP1, CRBP1, RIZ1, and OST-2 representing novel methylation targets in EAC and/or BE. p16, RUNX3, and HPP1 displayed increasing methylation frequencies in BE vs EAC. Furthermore, these increases in methylation occurred early, at the interface between BE and low-grade dysplasia (LGD). To demonstrate the silencing effect of hypermethylation, we selected the EAC cells BIC1, in which the HPP1 promoter is natively methylated, and subjected them to 5-aza-2'-deoxycytidine (Aza-C) treatment. Real-time RT-PCR indicated increased HPP1 mRNA levels after 3 days of Aza-C treatment, as well as decreased levels of methylated HPP1 DNA. Hypermethylation of a subset of six genes (APC, TIMP3, CRBP1, p16, RUNX3, and HPP1) was then tested in a retrospective longitudinal study of 99 BE and nine LGD specimens obtained from 53 BE patients undergoing surveillance endoscopy. Only high-grade dysplasia (HGD) or EAC were defined as progression end points. Two patient groups were compared: eight progressors (P) and 45 nonprogressors (NP), using Cox proportional hazards models to determine the relative progression risks of age, BE segment length, and methylation events. Multivariate analyses revealed that only hypermethylation of p16 (odds ratio (OR) 1.74, 95% confidence interval (CI) 1.33-2.20), RUNX3 (OR 1.80, 95% CI 1.08-2.81), and HPP1 (OR 1.77, 95% CI 1.06-2.81) were independently associated with an increased risk of progression, whereas age, BE segment length, and hypermethylation of TIMP3, APC, or CRBP1 were not independent risk factors. In combined analyses, risk was detectable up to, but not earlier than, 2 years preceding neoplastic progression. Hypermethylation of p16, RUNX3, and HPP1 in BE or LGD may represent independent risk factors for the progression of BE to HGD or EAC. These findings have implications regarding risk stratification, early EAC detection, and the appropriate endoscopic surveillance interval for patients with BE.

Adenocarcinoma↗

Quantifying and comparing the accuracy of binary biomarkers when predicting a failure time outcome.

The positive and negative predictive values are standard measures used to quantify the predictive accuracy of binary biomarkers when the outcome being predicted is also binary. When the biomarkers are instead being used to predict a failure time outcome, there is no standard way of quantifying predictive accuracy. We propose a natural extension of the traditional predictive values to accommodate censored survival data. We discuss not only quantifying predictive accuracy using these extended predictive values, but also rigorously comparing the accuracy of two biomarkers in terms of their predictive values. Using a marginal regression framework, we describe how to estimate differences in predictive accuracy and how to test whether the observed difference is statistically significant.

Biomarkers↗

Quantifying and comparing the predictive accuracy of continuous prognostic factors for binary outcomes.

The positive and negative predictive values are standard ways of quantifying predictive accuracy when both the outcome and the prognostic factor are binary. Methods for comparing the predictive values of two or more binary factors have been discussed previously (Leisenring et al., 2000, Biometrics 56, 345-351). We propose extending the standard definitions of the predictive values to accommodate prognostic factors that are measured on a continuous scale and suggest a corresponding graphical method to summarize predictive accuracy. Drawing on the work of Leisenring et al. we make use of a marginal regression framework and discuss methods for estimating these predictive value functions and their differences within this framework. The methods presented in this paper have the potential to be useful in a number of areas including the design of clinical trials and health policy analysis.

Adolescent↗

Postprandial suppression of plasma ghrelin level is proportional to ingested caloric load but does not predict intermeal interval in humans.

Plasma ghrelin levels rise before meals and fall rapidly afterward. If ghrelin is a physiological meal-initiation signal, then a large oral caloric load should suppress ghrelin levels more than a small caloric load, and the request for a subsequent meal should be predicted by recovery of the plasma ghrelin level. To test this hypothesis, 10 volunteers were given, at three separate sessions, liquid meals (preloads) with widely varied caloric content (7.5%, 16%, or 33% of total daily energy expenditure) but equivalent volume. Preloads were consumed at 0900 h, and blood was sampled every 20 min from 0800 h until 80 min after subjects spontaneously requested a meal. The mean (+/- SE) intervals between ingestion of the 7.5%, 16%, and 33% preloads and the subsequent voluntary meal requests were 247 +/- 24, 286 +/- 20, and 321 +/- 27 min, respectively (P = 0.015), and the nadir plasma ghrelin levels were 80.2 +/- 2.8%, 72.7 +/- 2.7%, and 60.8 +/- 2.7% of baseline (the 0900 h value), respectively (P < 0.001). A Cox regression analysis failed to show a relationship between ghrelin profile and the spontaneous meal request. We conclude that the depth of postprandial ghrelin suppression is proportional to ingested caloric load but that recovery of plasma ghrelin is not a critical determinant of intermeal interval.

Adolescent↗

Endoscopic ultrasound, positron emission tomography, and computerized tomography for lung cancer.

Staging of patients with lung cancer to determine operability is intended to efficiently limit futile thoracotomies without denying possibly curative surgery. Currently available staging tests are imperfect alone and in combination. Imaged suspected metastases often require tissue confirmation before surgery can be denied. Endoscopic ultrasound (EUS) may help identify inoperable patients by providing tissue proof of inoperability in a single staging test, with similar sensitivity for identifying inoperable patients as other staging tests. Therefore, we compared computed tomography, positron emission tomography (PET), and EUS with fine-needle aspiration under conscious sedation, each test interpreted blinded with respect to the other tests, for identifying inoperable patients in a consecutive cohort of 79 potentially operable patients with suspected or proven lung cancer. An economic analysis was also performed. Thirty-nine patients were found inoperable (a 40th patient's inoperability was missed by all preoperative staging tests). The sensitivity of computerized tomography was 43%. PET and EUS each had similar sensitivities (68 and 63%, respectively) and similar negative predictive values (64 and 68%, respectively), but EUS's superior specificity (100 vs. 72% for PET) and considerably lower expense means it may be preferred to PET early in staging to identify inoperable patients.

Adult↗

Partial AUC estimation and regression.

Accurate diagnosis of disease is a critical part of health care. New diagnostic and screening tests must be evaluated based on their abilities to discriminate diseased from nondiseased states. The partial area under the receiver operating characteristic (ROC) curve is a measure of diagnostic test accuracy. We present an interpretation of the partial area under the curve (AUC), which gives rise to a nonparametric estimator. This estimator is more robust than existing estimators, which make parametric assumptions. We show that the robustness is gained with only a moderate loss in efficiency. We describe a regression modeling framework for making inference about covariate effects on the partial AUC. Such models can refine knowledge about test accuracy. Model parameters can be estimated using binary regression methods. We use the regression framework to compare two prostate-specific antigen biomarkers and to evaluate the dependence of biomarker accuracy on the time prior to clinical diagnosis of prostate cancer.

Area Under Curve↗

Sample size calculations for comparative studies of medical tests for detecting presence of disease.

Technologic advances give rise to new tests for detecting disease in many fields, including cancer and sexually transmitted disease. Before a new disease screening test is approved for public use, its accuracy should be shown to be better than or at least not inferior to an existing test. Standards do not yet exist for designing and analysing studies to address this issue. Established principles for the design of therapeutic studies can be adapted for studies of screening tests. In particular, drawing upon methods for superiority and non-inferiority studies of therapeutic agents, we propose that confidence intervals for the relative accuracy of dichotomous tests drive the design of comparative studies of disease screening tests. We derive sample size formulae for a variety of designs, including studies where patients undergo several tests and studies where patients receive only one of the tests under evaluation. Both cohort and case-control study designs are considered. Modifications to the confidence intervals and sample size formulae are discussed to accommodate studies where, because of the invasive nature of definitive testing, true disease status can only be obtained for subjects who are positive on one or more of the screening tests. The methods proposed are applied to a study comparing a modified pap test to the conventional pap for cervical cancer screening. The impact of error in the gold standard reference test on the design and evaluation of comparative screening test studies is also discussed.

Case-Control Studies↗