Search PubMed⌕ Search

Biomedical subjects

Holly Janes

Publications and source records attributed to Holly Janes.

8 recordsLinked to original sources

Insights into latent class analysis of diagnostic test performance.

Latent class analysis is used to assess diagnostic test accuracy when a gold standard assessment of disease is not available but results of multiple imperfect tests are. We consider the simplest setting, where 3 tests are observed and conditional independence (CI) is assumed. Closed-form expressions for maximum likelihood parameter estimates are derived. They show explicitly how observed 2- and 3-way associations between test results are used to infer disease prevalence and test true- and false-positive rates. Although interesting and reasonable under CI, the estimators clearly have no basis when it fails. Intuition for bias induced by conditional dependence follows from the analytic expressions. Further intuition derives from an Expectation Maximization (EM) approach to calculating the estimates. We discuss implications of our results and related work for settings where more than 3 tests are available. We conclude that careful justification of assumptions about the dependence between tests in diseased and nondiseased subjects is necessary in order to ensure unbiased estimates of prevalence and test operating characteristics and to provide these estimates clinical interpretations. Such justification must be based in part on a clear clinical definition of disease and biological knowledge about mechanisms giving rise to test results.

Algorithms↗

The optimal ratio of cases to controls for estimating the classification accuracy of a biomarker.

The case-control design is frequently used to study the discriminatory accuracy of a screening or diagnostic biomarker. Yet, the appropriate ratio in which to sample cases and controls has never been determined. It is common for researchers to sample equal numbers of cases and controls, a strategy that can be optimal for studies of association. However, considerations are quite different when the biomarker is to be used for classification. In this paper, we provide an expression for the optimal case-control ratio, when the accuracy of the biomarker is quantified by the receiver operating characteristic (ROC) curve. We show how it can be integrated with choosing the overall sample size to yield an efficient study design with specified power and type-I error. We also derive the optimal case-control ratios for estimating the area under the ROC curve and the area under part of the ROC curve. Our methods are applied to a study of a new marker for adenocarcinoma in patients with Barrett's esophagus.

Adenocarcinoma↗

Identifying target populations for screening or not screening using logic regression.

Colorectal cancer remains a significant public health concern despite the fact that effective screening procedures exist and that the disease is treatable when detected at early stages. Numerous risk factors for colon cancer have been identified, but none are very predictive alone. We sought to determine whether there are certain combinations of risk factors that distinguish well between cases and controls, and that could be used to identify subjects at particularly high or low risk of the disease to target screening. Using data from the Seattle site of the Colorectal Cancer Family Registry, we fit logic regression models to combine risk factor information. Logic regression is a methodology that identifies subsets of the population, described by Boolean combinations of binary coded risk factors. This method is well suited to situations in which interactions between many variables result in differences in disease risk. We found that neither the logic regression models nor stepwise logistic regression models fit for comparison resulted in criteria that could be used to direct subjects to screening. However, we believe that our novel statistical approach could be useful in settings where risk factors do discriminate between cases and controls, and illustrate this with a simulated data set.

Adult↗

Overlap bias in the case-crossover design, with application to air pollution exposures.

The case-crossover design uses cases only, and compares exposures just prior to the event times to exposures at comparable control, or 'referent' times, in order to assess the effect of short-term exposure on the risk of a rare event. It has commonly been used to study the effect of air pollution on the risk of various adverse health events. Proper selection of referents is crucial, especially with air pollution exposures, which are shared, highly seasonal, and often have a long-term time trend. Hence, careful referent selection is important to control for time-varying confounders, and in order to ensure that the distribution of exposure is constant across referent times, a key assumption of this method. Yet the referent strategy is important for a more basic reason: the conditional logistic regression estimating equations commonly used are biased when referents are not chosen a priori and are functions of the observed event times. We call this bias in the estimating equations overlap bias. In this paper, we propose a new taxonomy of referent selection strategies in order to emphasize their statistical properties. We give a derivation of overlap bias, explore its magnitude, and consider how the bias depends on properties of the exposure series. We conclude that the bias is usually small, though highly unpredictable, and easily avoided.

Air Pollution↗

Case-crossover analyses of air pollution exposure data: referent selection strategies and their implications for bias.

The case-crossover design has been widely used to study the association between short-term air pollution exposure and the risk of an acute adverse health event. The design uses cases only; for each individual case, exposure just before the event is compared with exposure at other control (or "referent") times. Time-invariant confounders are controlled by making within-subject comparisons. Even more important in the air pollution setting is that time-varying confounders can also be controlled by design by matching referents to the index time. The referent selection strategy is important for reasons in addition to control of confounding. The case-crossover design makes the implicit assumption that there is no trend in exposure across the referent times. In addition, the statistical method that is used-conditional logistic regression-is unbiased only with certain referent strategies. We review here the case-crossover literature in the air pollution context, focusing on key issues regarding referent selection. We conclude with a set of recommendations for choosing a referent strategy with air pollution exposure data. Specifically, we advocate the time-stratified approach to referent selection because it ensures unbiased conditional logistic regression estimates, avoids bias resulting from time trend in the exposure series, and can be tailored to match on specific time-varying confounders.

Acute Disease↗

Limitations of the odds ratio in gauging the performance of a diagnostic, prognostic, or screening marker.

A marker strongly associated with outcome (or disease) is often assumed to be effective for classifying persons according to their current or future outcome. However, for this assumption to be true, the associated odds ratio must be of a magnitude rarely seen in epidemiologic studies. In this paper, an illustration of the relation between odds ratios and receiver operating characteristic curves shows, for example, that a marker with an odds ratio of as high as 3 is in fact a very poor classification tool. If a marker identifies 10% of controls as positive (false positives) and has an odds ratio of 3, then it will correctly identify only 25% of cases as positive (true positives). The authors illustrate that a single measure of association such as an odds ratio does not meaningfully describe a marker's ability to classify subjects. Appropriate statistical methods for assessing and reporting the classification power of a marker are described. In addition, the serious pitfalls of using more traditional methods based on parameters in logistic regression models are illustrated.

Biomarkers↗