Search PubMed⌕ Search

Biomedical subjects

Yeun-Jun Chung

Publications and source records attributed to Yeun-Jun Chung.

7 recordsLinked to original sources

Genomic signatures associated with epidemiologically defined high-risk pathogenic Escherichia coli isolates identified by interpretable machine learning.

Pathogenic Escherichia coli is a major cause of foodborne illness worldwide and includes strains capable of causing severe disease. To establish a genome-informed framework for foodborne outbreak surveillance, we analyzed 1,029 E. coli isolates from clinical, food, livestock, and environmental sources using whole-genome sequencing. Pathogenic isolates obtained from human clinical cases or linked to documented outbreaks were classified as epidemiologically defined high-risk (EpiHR), whereas the remaining pathogenic isolates were classified as non-EpiHR. Virulence-associated genomic features were extracted using a bioinformatics pipeline, and four machine learning (ML) algorithms, including gradient boosting machine, random forest (RF), and support vector machines with linear and radial basis function kernels, were evaluated. Among them, the RF model showed the best performance, achieving an area under the curve (AUC) of 0.98 and accuracy of 0.93 in 10-fold cross-validation. Additional leave-one-group-out validation showed retained discrimination across held-out sequence types and serotypes, although performance was reduced when isolates were grouped by isolation source. Evaluation using an independent test dataset of 1,908 publicly available pathogenic E. coli genomes showed an AUC of 0.97 and a sensitivity of 0.98. Feature importance analysis using Shapley additive explanations identified influential predictive features, including traT, etpB, and enterotoxin-associated genes. A reduced 10-feature model achieved an AUC of 0.79 in the independent test dataset, supporting its exploratory use for future simplified screening approaches. These results indicate that genome-based ML provides a sensitive framework for surveillance-oriented prioritization of EpiHR pathogenic E. coli isolates, with model predictions interpreted together with epidemiological information.

Escherichia coli↗

Recurrent genomic alterations with impact on survival in colorectal cancer identified by genome-wide array comparative genomic hybridization.

BACKGROUND & AIMS: Although genetic aspects of tumorigenesis in colorectal cancer (CRC) have been well studied, reliable biomarkers predicting prognosis are scarce. We aimed to identify recurrently altered genomic regions (RAR) in CRC with high resolution, to investigate their implications on survival and to explore novel cancer-related genes in prognosis-associated RARs. METHODS: A 1-Mb resolution microarray-based comparative genomic hybridization (array CGH) was applied to 59 CRCs. RARs, defined as genomic alterations, detected in more than 10 cases were identified and analyzed for their association with survival. Expression levels of genes in prognosis-associated RARs were examined by real-time quantitative polymerase chain reaction. RESULTS: Twenty-seven RARs were identified. Eleven high-level amplifications and 2 homozygous deletions also were detected, but they were not as common as RARs. Multivariate analysis revealed RAR-L1 (loss on 1p36; hazard ratio = 8.15, P = .002) and RAR-L20 (loss on 21q22; hazard ratio = 3.53, P = .034) are independent indicators of poor prognosis. Expression of CAMTA1, located in RAR-L1, was reduced frequently in CRCs, and low CAMTA1 expression was associated significantly with poor prognosis, which indicates that CAMTA1 may play a role as a tumor suppressor in CRC. Five pairs of RARs were correlated significantly to each other and 3 pairs share genes involved in the same biological functions, suggesting possible collaborative roles in tumorigenesis. CONCLUSIONS: We identified recurrent genomic changes in 59 CRCs. RARs could be more important in sporadic tumors where the effect of genomic changes on tumorigenesis is relatively smaller than in familial cancer. Our results and analysis strategy will be helpful to elucidate pathogenesis of CRCs or to develop biomarkers for predicting prognosis.

Adult↗

DNA copy number alterations and expression of relevant genes in mouse thymic lymphomas induced by gamma-irradiation and N-methyl-N-nitrosourea.

The genetic mechanism for the development and progression of a lymphoma is unclear. This study investigated the alterations in the DNA copy number and the expression profiles of the genes located in the altered regions in mouse thymic lymphomas that were induced by two mutagens, gamma-irradiation and N-methyl-N-nitrosourea (MNU). Microarray-based comparative genomic hybridization was used to precisely delineate the boundaries of the altered region. The copy number gains of chromosomes 4 and 5 were observed only in the radiation-induced lymphomas, and gains of chromosomes 10 and 14 were observed only in the MNU-induced lymphomas. Regional copy number losses in chromosomes 11, 16, and 19 appeared frequently in the radiation-induced lymphomas. The cancer-related genes Pten, Ikaros/Znfn1a1, Ercc4, and Top3b were located in the minimal deletion regions. In particular, the expression levels of the Pten, Top3b, and Ikaros genes were downregulated in both lymphoma groups, but the expression level of Ercc4 was downregulated only in the MNU group. This study also examined the expression levels of Sparc, Cxcl1, and Myc (synonym: c-Myc), which are located in the copy number gained chromosomes. Sparc was upregulated specifically in the radiation group, and Cxcl1 in the MNU group. c-Myc was upregulated in both groups. There was limited correlation between the DNA copy number profiles and the expression of the cancer-related genes in mouse lymphomagenesis. The chromosome aberrations and novel expression profiles of the cancer-related genes within the altered regions may provide important clues to the genetic mechanism for the development of lymphoma.

Alkylating Agents↗

Genome-wide screening of genomic alterations and their clinicopathologic implications in non-small cell lung cancers.

PURPOSE: Although many genomic alterations have been observed in lung cancer, their clinicopathologic significance has not been thoroughly investigated. This study screened the genomic aberrations across the whole genome of non-small cell lung cancer cells with high-resolution and investigated their clinicopathologic implications. EXPERIMENTAL DESIGN: One-megabase resolution array comparative genomic hybridization was applied to 29 squamous cell carcinomas and 21 adenocarcinomas of the lung. Tumor and normal tissues were microdissected and the extracted DNA was used directly for hybridization without genomic amplification. The recurrent genomic alterations were analyzed for their association with the clinicopathologic features of lung cancer. RESULTS: Overall, 36 amplicons, 3 homozygous deletions, and 17 minimally altered regions common to many lung cancers were identified. Among them, genomic changes on 13q21, 1p32, Xq, and Yp were found to be significantly associated with clinical features such as age, stage, and disease recurrence. Kaplan-Meier survival analysis revealed that genomic changes on 10p, 16q, 9p, 13q, 6p21, and 19q13 were associated with poor survival. Multivariate analysis showed that alterations on 6p21, 7p, 9q, and 9p remained as independent predictors of poor outcome. In addition, significant correlations were observed for three pairs of minimally altered regions (19q13 and 6p21, 19p13 and 19q13, and 8p12 and 8q11), which indicated their possible collaborative roles. CONCLUSIONS: These results show that our approach is robust for high-resolution mapping of genomic alterations. The novel genomic alterations identified in this study, along with their clinicopathologic implications, would be useful to elucidate the molecular mechanisms of lung cancer and to identify reliable biomarkers for clinical application.

Carcinoma, Non-Small-Cell Lung↗

ArrayCyGHt: a web application for analysis and visualization of array-CGH data.

UNLABELLED: ArrayCyGHt is a web-based application tool for analysis and visualization of microarray-comparative genomic hybridization (array-CGH) data. Full process of array-CGH data analysis, from normalization of raw data to the final visualization of copy number gain or loss, can be straightforwardly achieved on this arrayCyGHt system without the use of any further software. ArrayCyGHt, therefore, provides an easy and fast tool for the analysis of copy number aberrations in any kinds of data format. AVAILABILITY: ArrayCyGHt can be accessed at http://genomics.catholic.ac.kr/arrayCGH/

Algorithms↗

Regions syntenic to human 17q are gained in mouse and rat neuroblastoma.

Gain of chromosome arm 17q is the most frequent chromosomal change in human neuroblastoma and is a powerful predictor of adverse outcome of disease. This suggests that the region of gain includes a gene or genes critical for tumor pathogenesis. Analyses of breakpoint positions have revealed that the shortest region of gain (SRG) extends from MPO (17q23.1) to 17qter. Because this encompasses >300 genes, it precludes the identification of candidate genes from human breakpoint data alone. However, mouse chromosome 11, which is syntenic to human chromosome 17, is gained in up to 30% of neuroblastoma tumors developed in a murine MYCN transgenic model of this disease. To confirm that this key genetic change indicates the involvement of a molecular pathway conserved between mouse and man and is not occurring coincidentally in the transgenic model, we used fluorescence in situ hybridization to analyze sporadic cases of both mouse and rat neuroblastoma. Our results confirmed the presence of chromosome 11 gain in all three of the mouse cell lines we analyzed, with the SRG extending from Stat5b (101.6 Mb) to tel. In addition, the rat neuroblastoma cell line harbors an extra copy of distal chromosome 10, extending from 92.8 to 109.3 Mb, which is also syntenic to human 17q. Comparison of the regions gained in all three species has excluded 4.2 Mb from the previously defined region of 17q gain in humans as a likely location of the candidate gene or genes, and strongly suggests that the molecular etiology of neuroblastoma is similar in all three species.

Animals↗

A whole-genome mouse BAC microarray with 1-Mb resolution for analysis of DNA copy number changes by array comparative genomic hybridization.

Microarray-based comparative genomic hybridization (CGH) has become a powerful method for the genome-wide detection of chromosomal imbalances. Although BAC microarrays have been used for mouse CGH studies, the resolving power of these analyses was limited because high-density whole-genome mouse BAC microarrays were not available. We therefore developed a mouse BAC microarray containing 2803 unique BAC clones from mouse genomic libraries at 1-Mb intervals. For the general amplification of BAC clone DNA prior to spotting, we designed a set of three novel degenerate oligonucleotide-primed (DOP) PCR primers that preferentially amplify mouse genomic sequences while minimizing unwanted amplification of contaminating Escherichia coli DNA. The resulting 3K mouse BAC microarrays reproducibly identified DNA copy number alterations in cell lines and primary tumors, such as single-copy deletions, regional amplifications, and aneuploidy.

Animals↗