Search PubMed⌕ Search

Biomedical subjects

HyungJun Cho

Publications and source records attributed to HyungJun Cho.

5 recordsLinked to original sources

Rank-invariant resampling based estimation of false discovery rate for analysis of small sample microarray data.

BACKGROUND: The evaluation of statistical significance has become a critical process in identifying differentially expressed genes in microarray studies. Classical p-value adjustment methods for multiple comparisons such as family-wise error rate (FWER) have been found to be too conservative in analyzing large-screening microarray data, and the False Discovery Rate (FDR), the expected proportion of false positives among all positives, has been recently suggested as an alternative for controlling false positives. Several statistical approaches have been used to estimate and control FDR, but these may not provide reliable FDR estimation when applied to microarray data sets with a small number of replicates. RESULTS: We propose a rank-invariant resampling (RIR) based approach to FDR evaluation. Our proposed method generates a biologically relevant null distribution, which maintains similar variability to observed microarray data. We compare the performance of our RIR-based FDR estimation with that of four other popular methods. Our approach outperforms the other methods both in simulated and real microarray data. CONCLUSION: We found that the SAM's random shuffling and SPLOSH approaches were liberal and the other two theoretical methods were too conservative while our RIR approach provided more accurate FDR estimation than the other approaches.

Algorithms↗

Robust classification modeling on microarray data using misclassification penalized posterior.

MOTIVATION: Genome-wide microarray data are often used in challenging classification problems of clinically relevant subtypes of human diseases. However, the identification of a parsimonious robust prediction model that performs consistently well on future independent data has not been successful due to the biased model selection from an extremely large number of candidate models during the classification model search and construction. Furthermore, common criteria of prediction model performance, such as classification error rates, do not provide a sensitive measure for evaluating performance of such astronomic competing models. Also, even though several different classification approaches have been utilized to tackle such classification problems, no direct comparison on these methods have been made. RESULTS: We introduce a novel measure for assessing the performance of a prediction model, the misclassification-penalized posterior (MiPP), the sum of the posterior classification probabilities penalized by the number of incorrectly classified samples. Using MiPP, we implement a forward step-wise cross-validated procedure to find our optimal prediction models with different numbers of features on a training set. Our final robust classification model and its dimension are determined based on a completely independent test dataset. This MiPP-based classification modeling approach enables us to identify the most parsimonious robust prediction models only with two or three features on well-known microarray datasets. These models show superior performance to other models in the literature that often have more than 40-100 features in their model construction. AVAILABILITY: Our MiPP software program is available at the Bioconductor website (http://www.bioconductor.org).

Algorithms↗

The number of metastatic lymph nodes in extrahepatic bile duct carcinoma as a prognostic factor.

The number of lymph nodes with metastases is known to be an important prognostic factor in carcinomas of many organs. The insufficient sampling of lymph nodes has also been associated with worse outcome in several types of carcinoma. However, the prognostic significance of lymph node dissection is not well characterized in extrahepatic bile duct (EBD) carcinomas. For 209 patients with EBD carcinoma, the total number of retrieved lymph nodes and the number of metastatic lymph nodes were evaluated, and other clinicopathologic variables were correlated with patient survival. The number of retrieved lymph nodes was not significantly correlated with survival in this study. The presence of metastasis to lymph nodes significantly decreased survival of patients with EBD carcinoma. The patients with 5 or more metastatic lymph nodes had significantly worse survival than those with 4 or less metastatic lymph nodes. To evaluate the prognosis of the patients with EBD carcinomas more precisely, the number of metastatic lymph nodes as well as the status of metastasis to lymph nodes should be examined and reported. Based on the present data, we propose that nodal classification should be divided into N1 (metastasis in 1 to 4 regional lymph nodes) and N2 (metastasis in 5 or more regional lymph nodes).

Adult↗

CDX2 and MUC2 protein expression in extrahepatic bile duct carcinoma.

Although CDX2-mediated intestinal metaplasia and its association with gastric and esophageal carcinoma have been well described, its function in extrahepatic bile duct (EBD) carcinoma remains unclear. CDX2 and MUC2 expression were examined in 193 EBD carcinomas, and observed in 37.3% and 42.0%, respectively. Both CDX2 and MUC2 were observed in 27.4%. CDX2 (P<.001) and MUC2 (P<.001) were correlated with histologic subtypes and present, respectively, in all intestinal-type adenocarcinomas and mucinous carcinomas, 12 (71%) and 13 (76%) of 17 papillary carcinomas, 2 (40%) and 2 (40%) of 5 adenosquamous carcinomas, and 28.4% and 33.5% of adenocarcinomas, not otherwise specified. CDX2 was observed more frequently in tumors with papillary growth (P=.03) and no vascular invasion (P=.04), whereas MUC2 was more common in cases with low stage (P=.01) and no vascular invasion (P<.001). Patients with CDX2+/MUC2+ tumors had significantly better overall survival in univariate but not multivariate analysis than patients with other tumors (P<.05). Expression of CDX2 and MUC2 supports that intestinal differentiation is present in specific subtypes of EBD carcinomas, and their expression status correlates with patients' overall survival.

Adult↗

Bayesian hierarchical error model for analysis of gene expression data.

MOTIVATION: Analysis of genome-wide microarray data requires the estimation of a large number of genetic parameters for individual genes and their interaction expression patterns under multiple biological conditions. The sources of microarray error variability comprises various biological and experimental factors, such as biological and individual replication, sample preparation, hybridization and image processing. Moreover, the same gene often shows quite heterogeneous error variability under different biological and experimental conditions, which must be estimated separately for evaluating the statistical significance of differential expression patterns. Widely used linear modeling approaches are limited because they do not allow simultaneous modeling and inference on the large number of these genetic parameters and heterogeneous error components on different genes, different biological and experimental conditions, and varying intensity ranges in microarray data. RESULTS: We propose a Bayesian hierarchical error model (HEM) to overcome the above restrictions. HEM accounts for heterogeneous error variability in an oligonucleotide microarray experiment. The error variability is decomposed into two components (experimental and biological errors) when both biological and experimental replicates are available. Our HEM inference is based on Markov chain Monte Carlo to estimate a large number of parameters from a single-likelihood function for all genes. An F-like summary statistic is proposed to identify differentially expressed genes under multiple conditions based on the HEM estimation. The performance of HEM and its F-like statistic was examined with simulated data and two published microarray datasets-primate brain data and mouse B-cell development data. HEM was also compared with ANOVA using simulated data. AVAILABILITY: The software for the HEM is available from the authors upon request.

Algorithms↗