Search PubMed⌕ Search

Biomedical subjects

Sin-Ho Jung

Publications and source records attributed to Sin-Ho Jung.

At least 19 recordsLinked to original sources

How accurately can we control the FDR in analyzing microarray data?

SUMMARY: We want to evaluate the performance of two FDR-based multiple testing procedures by Benjamini and Hochberg (1995, J. R. Stat. Soc. Ser. B, 57, 289-300) and Storey (2002, J. R. Stat. Soc. Ser. B, 64, 479-498) in analyzing real microarray data. These procedures commonly require independence or weak dependence of the test statistics. However, expression levels of different genes from each array are usually correlated due to coexpressing genes and various sources of errors from experiment-specific and subject-specific conditions that are not adjusted for in data analysis. Because of high dimensionality of microarray data, it is usually impossible to check whether the weak dependence condition is met for a given dataset or not. We propose to generate a large number of test statistics from a simulation model which has asymptotically (in terms of the number of arrays) the same correlation structure as the test statistics that will be calculated from the given data and to investigate how accurately the FDR-based testing procedures control the FDR on the simulated data. Our approach is to directly check the performance of these procedures for a given dataset, rather than to check the weak dependency requirement. We illustrate the proposed method with real microarray datasets, one where the clinical endpoint is disease group and another where it is survival.

Algorithms↗

Rank tests for clustered survival data when dependent subunits are randomized.

In clustered survival data, subunits within each cluster share similar characteristics, so that observations made from them tend to be positively correlated. In clinical trials, the correlated subunits from the same cluster are often randomized to different treatment groups. In this case, the variance formulas of the standard rank tests such as the logrank, Gehan-Wilcoxon or Prentice-Wilcoxon, proposed for independent samples, need to be adjusted for intracluster correlations both within and between treatment groups for testing equality of marginal survival distributions. In this paper we derive a general form of simple variance formulas of the rank tests when subunits from the same cluster are randomized into different treatment groups. Extensive simulation studies are conducted to investigate small sample performance of the variance formulas. We compare our non-parametric rank tests based on the adjusted variances with one from a shared frailty model, which is an optimal semi-parametric testing procedure when the intracluster correlations within and between groups are the same.

Animals↗

A multiple testing procedure to associate gene expression levels with survival.

In many microarray studies the primary objective is to identify, from a large panel of genes, those which are prognostic markers of a censored survival endpoint such as time to disease recurrence or death. Often, these genes are considered prognostic in that their respective expressions are associated with the survival endpoint of interest. To assess this association requires specifying an appropriate measure of association, a suitable test statistic and, as the number of genes is large, proper handling of multiplicity issues. In this paper, we will address these issues by utilizing a general correlation measure, a non-parametric test statistic, and control of the family-wise error rate by employing permutation resampling. Comprehensive simulation studies are conducted to investigate the statistical properties of the proposed procedure. The proposed method is applied to a recently published data set on patients with lung cancer.

Adenocarcinoma↗

Sample size for a two-group comparison of repeated binary measurements using GEE.

Controlled clinical trials often randomize subjects to two treatment groups and repeatedly evaluate them at baseline and intervals across a treatment period of fixed duration. A popular primary objective in these trials is to compare the change rates in the repeated measurements between treatment groups. Repeated measurements usually involve missing data and a serial correlation within each subject. The generalized estimating equation (GEE) method has been widely used to fit the time trend in repeated measurements because of its robustness to random missing and mispecification of the true correlation structure. In this paper, we propose a closed form sample size formula for comparing the change rates of binary repeated measurements using GEE for a two-group comparison. The sample size formula is derived incorporating missing patterns, such as independent missing and monotone missing, and correlation structures, such as AR(1) model. We also propose an algorithm to generate correlated binary data with arbitrary marginal means and a Markov dependency and use it in simulation studies.

Algorithms↗

Sample size for FDR-control in microarray data analysis.

We consider identifying differentially expressing genes between two patient groups using microarray experiment. We propose a sample size calculation method for a specified number of true rejections while controlling the false discovery rate at a desired level. Input parameters for the sample size calculation include the allocation proportion in each group, the number of genes in each array, the number of differentially expressing genes and the effect sizes among the differentially expressing genes. We have a closed-form sample size formula if the projected effect sizes are equal among differentially expressing genes. Otherwise, our method requires a numerical method to solve an equation. Simulation studies are conducted to show that the calculated sample sizes are accurate in practical settings. The proposed method is demonstrated with a real study.

Algorithms↗

Sample size calculation for simulation-based multiple-testing procedures.

In this article, we present a simple method to calculate sample size and power for a simulation-based multiple testing procedure which gives a sharper critical value than the standard Bonferroni method. The method is especially useful when several highly correlated test statistics are involved in a multiple-testing procedure. The formula for sample size calculation will be useful in designing clinical trials with multiple endpoints or correlated outcomes. We illustrate our method with a quality-of-life study for patients with early stage prostate cancer. Our method can also be used for comparing multiple independent groups.

Computer Simulation↗

Sample size computation for two-sample noninferiority log-rank test.

When an experimental therapy is less extensive, less toxic, or less expensive than a standard therapy, we may want to prove that the former is not worse than the latter through a noninferiority trial. In this article, we discuss a modification of the log-rank test for noninferiority trials with survival endpoint and propose a sample size formula that can be used in designing such trials. Performance of our sample size formula is investigated through simulations. Our formula is applied to design a real clinical trial.

Clinical Trials as Topic↗

Effect of dropouts on sample size estimates for test on trends across repeated measurements.

Sample size calculation is an important component at the design stage of clinical trials. We investigate the implications of dropouts for the sample size estimates in testing differences in the rates of changes produced by two treatments in a randomized parallel-groups repeated measurement design. Statistical models for calculating sample sizes for repeated measurement designs often fail to take into account the impact of dropouts correctly. In this article, we examine the impact of dropouts on sample size estimate and compare the power with the approach of Jung and Ahn [Jung, S. H., Ahn, C. (2003). Sample size estimation for GEE method for comparing slopes in repeated measurements data. Stat. Med. 22: 1305-1315] with that suggested by Patel and Rowe [Patel, H., Rowe, E. (1999). Sample size for comparing linear growth curves. J. Biopharm. Stat. 9:339-350] through a simulation study.

Clinical Trials as Topic↗

Sample size calculation for multiple testing in microarray data analysis.

Microarray technology is rapidly emerging for genome-wide screening of differentially expressed genes between clinical subtypes or different conditions of human diseases. Traditional statistical testing approaches, such as the two-sample t-test or Wilcoxon test, are frequently used for evaluating statistical significance of informative expressions but require adjustment for large-scale multiplicity. Due to its simplicity, Bonferroni adjustment has been widely used to circumvent this problem. It is well known, however, that the standard Bonferroni test is often very conservative. In the present paper, we compare three multiple testing procedures in the microarray context: the original Bonferroni method, a Bonferroni-type improved single-step method and a step-down method. The latter two methods are based on nonparametric resampling, by which the null distribution can be derived with the dependency structure among gene expressions preserved and the family-wise error rate accurately controlled at the desired level. We also present a sample size calculation method for designing microarray studies. Through simulations and data analyses, we find that the proposed methods for testing and sample size calculation are computationally fast and control error and power precisely.

Computer Simulation↗

A phase II study of single agent gemcitabine in relapsed or refractory follicular or small lymphocytic non-Hodgkin lymphomas: a Hoosier Oncology Group Study.

Gemcitabine is a pyrimidine analog that is active in patients with aggressive lymphomas and Hodgkin disease. This study assessed tumor response in patients with previously treated follicular or small lymphocytic non-Hodgkin lymphoma. This was a 2-stage phase II trial with the first stage requiring 2 of 13 responses to proceed to the second stage. Gemcitabine was given as a single agent to patients with previously treated follicular or small lymphocytic lymphomas. Gemcitabine was administered at 1250 mg/m2 over 30 minutes on days 1 and 8 of a 21-day cycle for a maximum of 6 cycles. Thirteen patients were treated with 1 to 6 cycles of chemotherapy. Two patients experienced grade 4 toxicity with neutropenia. No grade 4 nonhematologic toxicity was seen. There was 1 partial response and 8 patients (61%) had either minimal response or stable disease. Single-agent gemcitabine administered at this dose and schedule produced 1 partial remission and half the patients had stable disease. However, the study had to be stopped early because of lack of meaningful response.

Aged↗

On the estimation of the binomial probability in multistage clinical trials.

Due to the optional sampling effect in a sequential design, the maximum likelihood estimator (MLE) following sequential tests is generally biased. In a typical two-stage design employed in a phase II clinical trial in cancer drug screening, a fixed number of patients are enrolled initially. The trial may be terminated for lack of clinical efficacy of treatment if the observed number of treatment responses after the first stage is too small. Otherwise, an additional fixed number of patients are enrolled to accumulate additional information on efficacy as well as on safety. There have been numerous suggestions for design of such two-stage studies. Here we establish that under the two-stage design the sufficient statistic, i.e. stopping stage and the number of treatment responses, for the parameter of the binomial distribution is also complete. Then, based on the Rao-Blackwell theorem, we derive the uniformly minimum variance unbiased estimator (UMVUE) as the conditional expectation of an unbiased estimator, which in this case is simply the maximum likelihood estimator based only on the first stage data, given the complete sufficient statistic. Our results generalize to a multistage design. We will illustrate features of the UMVUE based on two-stage phase II clinical trial design examples and present results of numerical studies on the properties of the UMVUE in comparison to the usual MLE.

Clinical Trials, Phase II as Topic↗

Admissible two-stage designs for phase II cancer clinical trials.

In a typical two-stage design for a phase II cancer clinical trial for efficacy screening of cytotoxic agents, a fixed number of patients are initially enrolled and treated. The trial may be terminated for lack of efficacy if the observed number of tumour responses after the first stage is too small, thus avoiding treatment of patient with inefficacious regimen. Otherwise, an additional fixed number of patients are enrolled and treated to accumulate additional information on efficacy as well as safety. The minimax and the so-called 'optimal' designs by Simon have been widely used, and other designs have largely been ignored in the past for such two-stage cancer clinical trials. Recently Jung et al. proposed a graphical method to search for compromise designs with features more favourable than either the minimax or the optimal design. In this paper, we develop a family of two-stage designs that are admissible according to a Bayesian decision-theoretic criterion based on an ethically justifiable loss function. We show that the admissible designs include as special cases the Simon's minimax and the optimal designs as well as the compromise designs introduced by Jung et al. We also present a Java program to search for admissible designs that are compromises between the minimax and the optimal designs.

Bayes Theorem↗

Characterization of gains, losses, and regional amplification in testicular germ cell tumor cell lines by comparative genomic hybridization.

We have performed comparative genomic hybridization on 12 testicular germ cell tumor (TGCT) cell lines and one paraffin-embedded surgical specimen to identify and characterize genome-wide gains and losses of chromosomes in these specimens. All specimens demonstrated overrepresentation of 12p. Other significant chromosomal gains, apart from 12p, included the X chromosome and chromosome arms 1q and 20q. Chromosomal losses were observed for chromosomes 4 and 18 and chromosome arms 2q, 9q, and 13q. Genomic differences were observed between an embryonal carcinoma component of a mixed tumor, 833K, and its cisplastin-resistant derivative line, 64CP, including losses of 6q23 approximately qter and 9p22 approximately q21. Five lines also demonstrated gain of 12p and additional 12p12 approximately p13 material. Similarly, two lines demonstrated gain of 12p and additional 12p11.2 approximately p12 material. The data supports the consistent gain of 12p in adult TGCT cell lines and additional regional amplification of 12p in some lines. This regional amplification has been observed in both primary tumor specimens and TGCT cell lines and may support a hypothesis that at least two different regions of 12p, one proximal and one distal, harbor genes important for the pathogenesis of testicular germ cell neoplasia.

Cell Line, Tumor↗

A phase I trial of olanzapine (Zyprexa) for the prevention of delayed emesis in cancer patients: a Hoosier Oncology Group study.

Chemotherapy-induced delayed emesis (DE) can affect up to 50% to 70% of patients receiving moderately and highly emetogenic chemotherapy, although rates are improving. DE most commonly occurs within the first 24 to 48 hours of chemotherapy administration and can persist for 2 to 5 days. Olanzapine, due to its activity at multiple dopaminergic, serotonergic, muscarinic, and histaminic receptor sites, has potential as antiemetic therapy. A phase I study was designed with olanzapine, using a four-cohort dose escalation of 3 to 6 patients per cohort, for the prevention of DE in cancer patients receiving their first cycle of chemotherapy consisting of cyclophosphamide, doxorubicin, platinum, and/or irinotecan. All patients received standard premedication. Olanzapine was administered on days -2 and -1 prior to chemotherapy and continued for 8 days (days 0-7). Episodes of vomiting as well as daily measurements of nausea, sedation, and toxicity were monitored at each dose level. Fifteen patients completed the protocol. No grade 4 toxicities were seen, and three patients experienced a dose-limiting toxicity (grade 3) of a depressed level of consciousness during the study. The maximum tolerated dose appeared to be 5 mg (for days -2 and -1) and 10 mg (for days 0-7). Four of six patients receiving highly emetogenic chemotherapy (cisplatin, > or = 70 mg/m2) and nine of nine patients receiving moderately emetogenic chemotherapy (doxorubicin, > or = 50 mg/m2) had complete control (no vomiting episodes) of DE. Therefore, olanzapine may be an effective agent for the prevention of chemotherapy-induced DE. A phase II trial is underway.

Adult↗

Assessment of quality of life in outpatients with advanced cancer: the accuracy of clinician estimations and the relevance of spiritual well-being--a Hoosier Oncology Group Study.

PURPOSE: To evaluate the association between quality-of-life (QOL) impairment as reported by patients and QOL impairment as judged by nurses or physicians, with and without consideration of spiritual well-being (SWB). PATIENTS AND METHODS: A total of 163 patients with advanced cancer were enrolled onto a therapeutic trial, and cross-sectional data were derived from clinical and demographic questionnaires obtained at baseline, including assessment of patient QOL and SWB. Clinicians rated the QOL impairment of their patients as mild, moderate, or severe. Clinician-estimated QOL impairment and patient-derived QOL categories were compared. Correlation coefficients were estimated to associate QOL scores using different instruments. The analysis of variance method was used to compare Functional Assessment of Cancer Therapy-General scores on categorical variables. RESULTS: There was no significant association between self-assessment scores and marital status, education level, performance status, or predicted life expectancy. However, a strong relationship between SWB and QOL was noted (P <.0001). Clinician-estimated QOL impairment matched the level of patient-derived QOL correctly in approximately 60% of cases, with only slight variation depending on the method of categorizing patient-derived QOL scores. The accuracy of clinician estimates was not associated with the level of SWB. Interestingly, a subset analysis of the inaccurate estimates revealed an association between lower SWB and clinician underestimation of QOL impairment (P =.0025). CONCLUSION: Clinician estimates of QOL impairment were accurate in more than 60% of patients. SWB is strongly associated with QOL, but it is not associated with the overall accuracy of clinicians' judgments about QOL impairment.

Activities of Daily Living↗

Fluoxetine versus placebo in advanced cancer outpatients: a double-blinded trial of the Hoosier Oncology Group.

PURPOSE: To determine whether fluoxetine improves overall quality of life (QOL) in advanced cancer patients with symptoms of depression revealed by a simple survey. PATIENTS AND METHODS: One hundred sixty-three patients with an advanced solid tumor and expected survival between 3 and 24 months were randomly assigned in a double-blinded fashion to receive either fluoxetine (20 mg daily) or placebo for 12 weeks. Patients were screened for at least minimal depressive symptoms and assessed every 3 to 6 weeks for QOL and depression. Patients with recent exposure to antidepressants were excluded. RESULTS: The groups were comparable at baseline in terms of age, sex, disease distribution, performance status, and level of depressive symptoms. One hundred twenty-nine patients (79%) completed at least one follow-up assessment. Analysis using generalized estimating equation modeling revealed that patients treated with fluoxetine exhibited a significant improvement in QOL as shown by the Functional Assessment of Cancer Therapy-General, compared with patients given placebo (P =.01). Specifically, the level of depressive symptoms expressed was lower in patients treated with fluoxetine (P =.0005), and the subgroup of patients showing higher levels of depressive symptoms on the two-question screening survey were the most likely to benefit from treatment. CONCLUSION: In this mix of patients with advanced cancer who had symptoms of depression as determined by a two-question bedside survey, use of fluoxetine was well tolerated, overall QOL was improved, and depressive symptoms were reduced.

Ambulatory Care↗

Sample size estimation for GEE method for comparing slopes in repeated measurements data.

Sample size calculation is an important component at the design stage of clinical trials. Controlled clinical trials often use a repeated measurement design in which individuals are randomly assigned to treatment groups and followed-up for measurements at intervals across a treatment period of fixed duration. In studies with repeated measurements, one of the popular primary interests is the comparison of the rates of change in a response variable between groups. Statistical models for calculating sample sizes for repeated measurement designs often fail to take into account the impact of missing data correctly. In this paper we propose to use the generalized estimating equation (GEE) method in comparing the rates of change in repeated measurements and introduce closed form formulae for sample size and power that can be calculated using a scientific calculator. Since the sample size formula is based on asymptotic theory, we investigate the performance of the estimated sample size in practical settings through simulations.

Data Interpretation, Statistical↗

A parametric model for long-term follow-up data from phase III breast cancer clinical trials.

We propose a parametric version of a univariate gamma frailty model. The proposed model is shown to be flexible enough to model long-term follow-up survival data from breast cancer clinical trials when the treatment effect diminishes as time progresses, a case for which neither the proportional hazards nor proportional odds assumptions are satisfied. The observed information matrix is computed to evaluate the variances of parameter estimates. A simple parametric test statistic to test proportional odds assumption is also constructed. The model is applied to a data set from a phase III clinical trial on breast cancer.

Antineoplastic Agents, Hormonal↗