Search PubMed⌕ Search

Biomedical subjects

K J Lui

Publications and source records attributed to K J Lui.

At least 37 records · Page 2Linked to original sources

The performance of the O'Brien-Fleming multiple testing procedure in the presence of intraclass correlation.

Assuming that all subject responses were independent, O'Brien and Fleming (1979, Biometrics 35, 549-556) proposed a simple and useful multiple testing procedure for clinical trials for comparing two treatments with dichotomous data. Differences in the methods of evaluating subject responses at each evaluation time, however, may induce an intraclass correlation among these responses used in calculating the O'Brien-Fleming multiple testing procedure. On the basis of Monte Carlo simulations, we note that even a small intraclass correlation among subject responses in the same analysis can substantially inflate the Type I error of the O'Brien-Fleming multiple testing procedure. Furthermore, this inflation generally increases as either the number of analyses or the underlying response probability increases. We also have demonstrated that if we were able to maintain a uniform medical test procedure between the two treatments for each analysis, the actual Type I error of the O'Brien-Fleming multiple testing procedure may conversely become conservative.

Analysis of Variance↗

A note on the application of simple linear regression methods for trend detection at multiple sites and visits.

In comparing running median, tolerance, cusum, and regression methods for trend detection over a small number of visits, Yang et al. found that application of multiple Z-tests on the basis of a simple linear regression for each site separately was the most efficient for detection of trends at several sites simultaneously. Because the use of multiple Z-tests completely ignores the covariance among measurements taken from different sites, to improve the power we propose a global chi 2-test. Assuming the covariance matrix known, we have found that the proposed chi 2-test procedure is more powerful than multiple Z-tests for two-sided alternatives when both the correlation among measurements and the number of sites are small. We also have found that the former procedure can have power uniformly larger than the latter when ratios of slopes to standard deviations of measurements at different sites vary and the number of sites is large. In fact, in the latter situation, the proposed global chi 2-test procedure, usually used only for two-sided alternatives, can even have power larger than that of multiple Z-tests for one-sided alternatives. In the situation where the ratios of slopes to standard deviations of measurements are all equal, however, the proposed multivariate approach based on the chi 2-test distribution is the least efficient, especially when the number of sites and the correlation are moderate or large. Finally, to account for the effect of multiple tests over a series of visits on the overall alpha-level, on the basis of Monte Carlo simulations, we compute critical values for sequential use of the proposed multivariate test procedure.

Bias↗

A note on the effect of the intraclass correlation in the multiple reading procedure with a unanimity rule.

The use of multiple reading procedures to improve the performance of a diagnostic test occurs often in practice. Evaluation of the utility of multiple reading procedures, however, usually ignores the effect of the intraclass correlation. This paper provides a quantitative assessment of this effect in the multiple reading procedure with a unanimity rule with respect to sensitivity, specificity, positive and negative predictive values. We have found that when the disease prevalence is rare or moderate (less than or equal to 0.20), use of the multiple reading procedure with a unanimity rule is effective in increasing the positive predictive value of a single reading procedure for the situation in which the variation of responses among different subjects and the intraclass correlation among repeated tests are small. This is, however, not true for the situation in which the disease is rare and the variation of responses among different subjects is large, even when the intraclass correlation is small or 0. Furthermore, when the disease is rare and the variation of responses among subjects is small, a small or moderate intraclass correlation can substantially decrease the positive predictive value that one calculates under the assumption that the intraclass correlation is equal to 0. In general, when the disease is rare or moderate (less than or equal to 0.20), the intraclass correlation between repeated tests and the variation of responses among subjects have little effect on the negative predictive value.

Data Interpretation, Statistical↗

Sample size requirement for repeated measurements in continuous data.

In this paper we extend Bloch's discussion on the usefulness and the limitations in the application of repeated measurements per subject in study designs. We derive general sample size formulae for any finite number of comparison groups to calculate the required number of subjects with repeated measurements, that do not have to be conditionally independent. For fixed total cost, we discuss the optimal sample allocation for repeated measurements needed to maximize the power and the underestimation when using Bloch's sample size formula if in the hypothesis testing procedure the variance parameters are unknown. We have also included a quantitative investigation of the effectiveness of taking repeated measurements per subjects to reduced the required number of subjects for a given power at a given alpha-level.

Analysis of Variance↗

Sample size determination under an exponential model in the presence of a confounder and type I censoring.

In controlled clinical trials, random assignment of treatments to individuals is usually used to eliminate the effects of confounding variables. When there is censorship in data, however, confounding effects may not be automatically removed solely by random assignment of treatments to individuals under the exponential model. Therefore, it is important to incorporate the confounding effect into the sample size calculation even after randomization of treatments to individuals. In this paper, the discussion is restricted only to the situation where there are two comparison groups and one single Bernoulli confounding variable. Based on an exponential covariate model, an explicit sample size formula considering the confounding effect has been derived for the design of trials with type I censoring, in which an end time is fixed in advance and all responses occurring after that time are censored. The resulting sample size formula can also be applied to nonrandomized clinical trials. Finally, to provide insight into the influence of different factors on sample size calculation, a discussion on the effects of treatments, the confounder, the length of follow-up times for studied individuals, and the joint distribution of the treatment and the confounder has been included.

Clinical Trials as Topic↗

Sample sizes for repeated measurements in dichotomous data.

When the measurement of outcome varies within studied subjects and the cost of additional subjects is high, taking more than one measurement for each subject constitutes a useful alternative to increase the power or to reduce the total cost of the study. In this paper, I present sample size formulae for repeated measurements in dichotomous data under different situations. I also discuss optimal sample allocation for repeated measurements.

Markov Chains↗

A note on point estimation of the hazard ratio in exponential distributions.

The maximum likelihood estimator (MLE) of the ratio of the hazard rates in two exponential distributions is biased. This bias can be important when sample sizes are small or the ratio of these two hazard rates is large. When there is either no censoring or type II censoring, we propose using the uniformly minimum variance unbiased estimator (UMVUE). We show that using the UMVUE instead of the MLE reduces the mean-squared error (MSE). We have found that the UMVUE always has MSEs smaller than the MLE. We have also found that the UMVUE leads to important reductions in the MSE when the sample size used to calculate the hazard rate in the numerator of the hazard ratio is small (say less than or equal to 20) regardless of the sample size in the denominator. In the presence of type I censoring, the proposed estimator and the MLE are both biased. On the basis of a Monte Carlo study, however, we obtain similar reductions in the MSE using the UMVUE, as for no censoring or type II censoring.

Monte Carlo Method↗

Sample size determination for case-control studies: the influence of the joint distribution of exposure and confounder.

In case-control studies, the results about how the exposure distribution affects sample size are well known. This paper extends previous results by incorporating the effect of a confounder into the calculation of sample size for a desired size and power of a statistical test. The paper also includes a quantitative discussion on the influence of the joint distribution for exposure to a putative cause and a a confounder on required sample sizes. The results show that, to detect a specified alternative for a given size and power, the required sample size decreases as either the variance of exposure or the effect of exposure on disease increases. The required sample size, however, increases as either the variance of the confounder or the effect of the confounder on disease increases. Generally, the higher is the absolute value of the simple correlation between the exposure and the confounder, the larger is the required sample size.

Analysis of Variance↗

Risk of injury from resisting rape.

Women may resist rape by taking a variety of self-protective measures. To examine the association between a woman's use of self-protection during a rape incident and four injury outcomes, the authors analyzed data from the National Crime Survey, an ongoing survey of self-reported victimizations throughout the United States. The study population was 851 women greater than or equal to 12 years of age who reported being a victim of completed or attempted rape during 1973-1982. Logistic regression was used to control for eight covariates, including use of weapons by the offender and the nature of the victim-offender relationship. The use of self-protection during a rape incident was protective against completed rape. The odds ratios for completed rape were 0.2 for all measures of self-protection--nonforceful, forceful, and both forceful and nonforceful (all 95% confidence intervals between 0.1 and 0.4). After controlling for type of rape incident, the odds ratios for physical injury were greater than 1.0 for all measures of self-protection but, in the case of physical injury requiring medical attention, only the odds ratio for use of both forceful and nonforceful measures was statistically significant. Because of the limitations of the National Crime Survey, these findings should be interpreted cautiously. Further research is needed to help women respond in ways that will minimize injury should a rape incident occur.

Adolescent↗

An application of the empirical Bayes approach to directly adjusted rates: a note on suicide mapping in California.

A simple, reliable, and comparable measure for suicide mapping and other health problems is needed. Because standardized mortality ratios (SMRs) may not indicate the relative meaning of their magnitudes when compared with one another, and statistical significance levels of tests for SMRs overlook the areas that have small populations, neither of these approaches provides a satisfactory index. The results using directly adjusted rates can be ordered directly according to their magnitudes. However, because of the lack of reliable estimates of local age-specific rates, the usefulness of directly adjusted rates in mapping suicide is also limited. To extend the usefulness of directly adjusted rates, an empirical Bayes approach whereby information from other areas is borrowed to improve the precision of the estimates of local age-specific rates in calculating directly adjusted rates--especially in the areas with small population sizes--is proposed. When an empirical Bayes approach was applied to the 1983 suicide data for California counties, a more reasonable conclusion than could be obtained by using directly adjusted rates was reached.

Adolescent↗

Geographic distribution of heat-related deaths among elderly persons. Use of county-level dot maps for injury surveillance and epidemiologic research.

Mapping is a useful tool for initiating data analysis of relatively infrequent injury events and can lead to interesting hypotheses that can then be tested in further epidemiologic studies. From national death certificate data for the years 1979 through 1985, we made dot maps of fatalities due to excessive heat (International Classification of Diseases code E900) among persons 65 years or older. The maps show clusterings of deaths, particularly in the central, south central, and southeastern sections of the United States, to an extent not fully explained by the population density or temperature extremes. The counties principally affected were highly urbanized and, for races other than white, were relatively poor. Our maps identify counties in which heat-related health problems in the elderly are particularly severe. Public health officials in high-risk areas should undertake heat-wave contingency planning and physicians practicing in such areas should familiarize themselves with the treatment of the spectrum of heat-related illnesses.

Aged↗

Risk of developing AIDS in HIV-infected cohorts of hemophilic and homosexual men.

The latency period and/or incidence of the acquired immunodeficiency syndrome (AIDS) may differ in persons infected with the human immunodeficiency virus by different routes or having different "cofactors." We compared 79 hemophilic men in Pennsylvania and 117 homosexual and bisexual men in California, all having known dates of infection and long postinfection observation periods, to examine these hypotheses. By 1987, twenty-one percent of the hemophilic and 27% of the homosexual men had developed AIDS. However, seroconversion patterns differed for the two groups, and when this was taken into account, the conditional odds ratio for AIDS was 1.20. Kaplan-Meier survival analysis showed no significant difference in the cumulative proportion with AIDS, from time of infection. These results are limited by the small size and geographically localized nature of our study populations, but they suggest that currently the relative length of human immunodeficiency virus infection is of primary importance in comparing disease outcome for different populations.

Acquired Immunodeficiency Syndrome↗

An application of a mathematical model to adjust for time lag in case reporting.

In a dynamic, fluctuating surveillance system, the time lag of case reporting often causes an artificial plateau in an epidemic curve. Arbitrarily ignoring data reported in the most recent period to avoid this bias causes the loss of valuable information. In this report, we propose an application of a mathematical model to adjust for the underreporting bias owing to the time lag of the reporting process. We present an example using the acquired immunodeficiency syndrome incidence data for the homosexual group in the San Francisco surveillance system to illustrate and evaluate prospectively this proposed technique. The results show that the adjusted incidence obtained with the model agrees reasonably well with the true incidence, except for the last month of the period under consideration.

Acquired Immunodeficiency Syndrome↗

A model-based approach to the imputation of missing data: home injury incidences.

Missing or incomplete data cases are a problem in all types of statistical analyses. In disease surveillance, this problem inhibits determining the actual incidence of a disease event and monitoring the disease occurrence. Several statistical techniques have been developed to impute values for incomplete data cases. We present a model-based approach to the imputation of missing data elements as applied to determining the incidence of home injury deaths.

Accidents, Home↗

A discussion on the conventional estimator of sensitivity and specificity in multiple tests.

Both the variation of positive responses (negative responses) among individuals and the internal correlation among responses for the same individual affect the precision of the estimate of sensitivity (specificity). To estimate the sensitivity (specificity) of a medical diagnostic test, this paper proposes a Bayesian approach with a simple Markov model to evaluate the performance of the conventional estimator under different situations. On the basis of the assumed model, we derive a general formula for the variance of the conventional estimator regarding multiple tests and present a quantitative discussion on the limitations of this estimator.

Analysis of Variance↗

A model-based estimate of the mean incubation period for AIDS in homosexual men.

Because of the difficulty in identifying the date of exposure to type 1 of the human immunodeficiency virus (HIV-1) infection in persons other than transfusion recipients, studies of the incubation periods for acquired immunodeficiency syndrome (AIDS) have been limited. When data from a cohort of 84 homosexual and bisexual men that provided the information to determine the years of conversion of sera infected with HIV-1 were analyzed, a model for the proportion likely to develop AIDS and the incubation period for AIDS in homosexual men could be derived. The maximum likelihood estimate for the proportion of infected homosexual men developing AIDS is 0.99 (90% confidence interval ranging from 0.38 to 1). Furthermore, the maximum likelihood estimate for the mean incubation period for AIDS in homosexual men is 7.8 years (90% confidence interval ranging from 4.2 years to 15.0 years), which is close to the estimate of 8.2 years for adults developing transfusion-associated AIDS.

Acquired Immunodeficiency Syndrome↗