Search PubMed⌕ Search

Biomedical subjects

Yidong Chen

Publications and source records attributed to Yidong Chen.

16 recordsLinked to original sources

shinyDeepGxP: a user-friendly R shiny app for predicting surface protein abundance from scRNA-seq expression using deep learning in blood cells.

MOTIVATION: Understanding accurate immune cell heterogeneity and function in single-cell datasets requires access to protein-level information, which is often unavailable due to experimental limitations. RESULTS: We present shinyDeepGxP, an interactive web application featuring our deep learning model, DeepGxP, for predicting surface protein abundance from single-cell RNA-sequencing (scRNA-seq) data. This platform makes DeepGxP accessible to researchers without programming skills. Users can upload scRNA-seq count matrices and use "Predict Protein" to predict the abundance of 224 biologically relevant surface proteins. shinyDeepGxP provides visualizations to help identify distinct cell populations based on predicted protein profiles. Moreover, users can choose "Explore Model" to reveal key RNA predictors and their associated biological pathways for each protein. Overall, shinyDeepGxP is a user-friendly, freely available web tool that provides protein-level detail for RNA-only single-cell datasets, enabling multimodal discovery without additional experiments. AVAILABILITY AND IMPLEMENTATION: shinyDeepGxP can be launched on https://shiny.crc.pitt.edu/deepgxp/.

Journal Article↗

Antibody-Mediated Targeting of Secretory Protein SCUBE3 Suppresses Cancer Progression by Inhibiting Oncogenic Signaling and Inducing Antitumor Immunity.

UNLABELLED: Approaches targeting factors that simultaneously promote tumor growth and progression, induce therapy resistance, and inhibit antitumor immunity offer clear benefits over therapies targeting only one of these tumor-promoting processes. Through comprehensive loss-of-function genomic screening, we identified SCUBE3 as a pivotal factor that supports survival and therapy resistance and also orchestrates an immunosuppressive tumor microenvironment. Secretory SCUBE3 supported oncogenic activity through interactions with key oncogenic cell surface receptor proteins, including EGFR, mutant CALR, and TGFβRI/II. These interactions activated the transcription factors FOXR2 and c-Myc, promoting cancer cell proliferation and therapy resistance by enhancing DNA damage repair. Additionally, the SCUBE3-FOXR2 axis created an immunosuppressive tumor microenvironment by facilitating recruitment of the DNMT1 epigenetic repressor complex to the transcription regulator IRF1, thereby inhibiting the expression of MHC-I and MHC-II genes. A first-in-class neutralizing antibody targeting SCUBE3, which was developed using a sophisticated antibody discovery platform and engineered with specific mutations in the heavy chain for enhanced specificity and efficacy, demonstrated profound therapeutic potential across various cancer types in preclinical models, including patient-derived breast and ovarian cancer xenografts. This discovery marks an advancement toward developing a targeted therapy for cancers characterized by hyperactive SCUBE3-associated signaling pathways. SIGNIFICANCE: Targeting SCUBE3 with a neutralizing antibody inhibits tumor growth and metastasis by blocking oncogenic signaling through FOXR2 and c-Myc and by circumventing immunosuppression, providing a promising pan-cancer treatment approach.

Humans↗

Gene expression profile in multiple sclerosis patients and healthy controls: identifying pathways relevant to disease.

Multiple sclerosis (MS) and other T cell-mediated autoimmune diseases develop in individuals carrying a complex susceptibility trait, probably following exposure to various environmental triggers. Owing to the presumed weak influence of single genes on disease predisposition and the recognized genetic heterogeneity of autoimmune disorders in humans, candidate gene searches in MS have been difficult. In an attempt to identify molecular markers indicative of disease status rather than susceptibility genes for MS, we show that gene expression profiling of peripheral blood mononuclear cells by cDNA microarrays can distinguish MS patients from healthy controls. Our findings support the concept that the activation of autoreactive T cells is of primary importance for this complex organ-specific disorder and prompt further investigations on gene expression in peripheral blood cells aimed at characterizing disease phenotypes.

Adult↗

Predicting hepatitis B virus-positive metastatic hepatocellular carcinomas using gene expression profiling and supervised machine learning.

Hepatocellular carcinoma (HCC) is one of the most common and aggressive human malignancies. Its high mortality rate is mainly a result of intra-hepatic metastases. We analyzed the expression profiles of HCC samples without or with intra-hepatic metastases. Using a supervised machine-learning algorithm, we generated for the first time a molecular signature that can classify metastatic HCC patients and identified genes that were relevant to metastasis and patient survival. We found that the gene expression signature of primary HCCs with accompanying metastasis was very similar to that of their corresponding metastases, implying that genes favoring metastasis progression were initiated in the primary tumors. Osteopontin, which was identified as a lead gene in the signature, was over-expressed in metastatic HCC; an osteopontin-specific antibody effectively blocked HCC cell invasion in vitro and inhibited pulmonary metastasis of HCC cells in nude mice. Thus, osteopontin acts as both a diagnostic marker and a potential therapeutic target for metastatic HCC.

Algorithms↗

Molecular classification of familial non-BRCA1/BRCA2 breast cancer.

In the decade since their discovery, the two major breast cancer susceptibility genes BRCA1 and BRCA2, have been shown conclusively to be involved in a significant fraction of families segregating breast and ovarian cancer. However, it has become equally clear that a large proportion of families segregating breast cancer alone are not caused by mutations in BRCA1 or BRCA2. Unfortunately, despite intensive effort, the identification of additional breast cancer predisposition genes has so far been unsuccessful, presumably because of genetic heterogeneity, low penetrance, or recessive/polygenic mechanisms. These non-BRCA1/2 breast cancer families (termed BRCAx families) comprise a histopathologically heterogeneous group, further supporting their origin from multiple genetic events. Accordingly, the identification of a method to successfully subdivide BRCAx families into recognizable groups could be of considerable value to further genetic analysis. We have previously shown that global gene expression analysis can identify unique and distinct expression profiles in breast tumors from BRCA1 and BRCA2 mutation carriers. Here we show that gene expression profiling can discover novel classes among BRCAx tumors, and differentiate them from BRCA1 and BRCA2 tumors. Moreover, microarray-based comparative genomic hybridization (CGH) to cDNA arrays revealed specific somatic genetic alterations within the BRCAx subgroups. These findings illustrate that, when gene expression-based classifications are used, BRCAx families can be grouped into homogeneous subsets, thereby potentially increasing the power of conventional genetic analysis.

Adult↗

Effects of ligand and thyroid hormone receptor isoforms on hepatic gene expression profiles of thyroid hormone receptor knockout mice.

Little is known about the overall patterns of thyroid hormone (Th)-mediated gene regulation by the main Th receptor (Tr) isoforms, Tr-alpha and Tr-beta, in vivo. We used 48 complementary DNA microarrays to examine hepatic gene expression profiles of wild-type and Thra and Thrb knockout mice under different Th conditions: no treatment, treatment with 3,3',5-triiodothyronine (T(3)), Th-deprivation using propylthiouracil (PTU), and treatment with a combination of PTU and T(3). Hierarchical clustering analyses showed that positively regulated genes fit into three main expression patterns. In addition, only a subpopulation of target genes repressed basal transcription in the absence of ligand. Interestingly, Thra and Thrb knockout mice showed similar gene expression patterns to wild-type mice, suggesting that these isoforms co-regulate most hepatic target genes. Differences in the gene expression patterns of Thra/Thrb double-knockout mice and Th-deprived wild-type mice show that absence of receptor and of hormone can have different effects. This large-scale study of hormonal regulation reveals the functions of Th and of Tr isoforms in the regulation of gene expression patterns.

Animals↗

Gene expression after treatment with hydrogen peroxide, menadione, or t-butyl hydroperoxide in breast cancer cells.

Global gene expression patterns in breast cancer cells after treatment with oxidants (hydrogen peroxide, menadione, and t-butyl hydroperoxide) were investigated in three replicate experiments. RNA collected after treatment (at 1, 3, 7, and 24 h) rather than after a single time point, enabled an analysis of gene expression patterns. Using a 17,000 microarray, template-based clustering and multidimensional scaling analysis of the gene expression over the entire time course identified 421 genes as being either up- or down-regulated by the three oxidants. In contrast, only 127 genes were identified for any single time point and a 2-fold change criteria. Surprisingly, the patterns of gene induction were highly similar among the three oxidants; however, differences were observed, particularly with respect to p53, IL-6, and heat-shock related genes. Replicate experiments increased the statistical confidence of the study, whereas changes in gene expression patterns over a time course demonstrated significant additional information versus a single time point. Analyzing the three oxidants simultaneously by template cluster analysis identified genes that heretofore have not been associated with oxidative stress.

Breast Neoplasms↗

Gene expression signature of benign prostatic hyperplasia revealed by cDNA microarray analysis.

BACKGROUND: Despite the high prevalence of benign prostatic hyperplasia (BPH) in the aging male, little is known regarding the etiology of this disease. A better understanding of the molecular etiology of BPH would be facilitated by a comprehensive analysis of gene expression patterns that are characteristic of benign growth in the prostate gland. Since genes differentially expressed between BPH and normal prostate tissues are likely to reflect underlying pathogenic mechanisms involved in the development of BPH, we performed comparative gene expression analysis using cDNA microarray technology to identify candidate genes associated with BPH. METHODS: Total RNA was extracted from a set of 9 BPH specimens from men with extensive hyperplasia and a set of 12 histologically normal prostate tissues excised from radical prostatectomy specimens. Each of these 21 RNA samples was labeled with Cy3 in a reverse transcription reaction and cohybridized with a Cy5 labeled common reference sample to a cDNA microarray containing 6,500 human genes. Normalized fluorescent intensity ratios from each hybridization experiment were extracted to represent the relative mRNA abundance for each gene in each sample. Weighted gene and random permutation analyses were performed to generate a subset of genes with statistically significant differences in expression between BPH and normal prostate tissues. Semi-quantitative PCR analysis was performed to validate differential expression. RESULTS: A subset of 76 genes involved in a wide range of cellular functions was identified to be differentially expressed between BPH and normal prostate tissues. Semi-quantitative PCR was performed on 10 genes and 8 were validated. Genes consistently upregulated in BPH when compared to normal prostate tissues included: a restricted set of growth factors and their binding proteins (e.g. IGF-1 and -2, TGF-beta3, BMP5, latent TGF-beta binding protein 1 and -2); hydrolases, proteases, and protease inhibitors (e.g. neuropathy target esterase, MMP2, alpha-2-macroglobulin); stress response enzymes (e.g. COX2, GSTM5); and extracellular matrix molecules (e.g. laminin alpha 4 and beta 1, chondroitin sulfate proteoglycan 2, lumican). Genes consistently expressing less mRNA in BPH than in normal prostate tissues were less commonly observed and included the transcription factor KLF4, thrombospondin 4, nitric oxide synthase 2A, transglutaminase 3, and gastrin releasing peptide. CONCLUSIONS: We identified a diverse set of genes that are potentially related to benign prostatic hyperplasia, including genes both previously implicated in BPH pathogenesis as well as others not previously linked to this disease. Further targeted validation and investigations of these genes at the DNA, mRNA, and protein levels are warranted to determine the clinical relevance and possible therapeutic utility of these genes.

Adult↗

Mutation of melanosome protein RAB38 in chocolate mice.

Mutations of genes needed for melanocyte function can result in oculocutaneous albinism. Examination of similarities in human gene expression patterns by using microarray analysis reveals that RAB38, a small GTP binding protein, demonstrates a similar expression profile to melanocytic genes. Comparative genomic analysis localizes human RAB38 to the mouse chocolate (cht) locus. A G146T mutation occurs in the conserved GTP binding domain of RAB38 in cht mice. Rab38(cht)/Rab38(cht) mice exhibit a brown coat similar in color to mice with a mutation in tyrosinase-related protein 1 (Tyrp1), a mouse model for oculocutaneous albinism. The targeting of TYRP1 protein to the melanosome is impaired in Rab38(cht)/Rab38(cht) melanocytes. These observations, and the fact that green fluorescent protein-tagged RAB38 colocalizes with end-stage melanosomes in wild-type melanocytes, suggest that RAB38 plays a role in the sorting of TYRP1. This study demonstrates the utility of expression profile analysis to identify mammalian disease genes.

Animals↗

Expression profiling of synovial sarcoma by cDNA microarrays: association of ERBB2, IGFBP2, and ELF3 with epithelial differentiation.

Synovial sarcoma is an aggressive spindle cell sarcoma with two major histological subtypes, biphasic and monophasic, defined respectively by the presence or absence of areas of glandular epithelial differentiation. It is characterized by a specific chromosomal translocation, t(X;18)(p11.2;q11.2), which juxtaposes the SYT gene on chromosome 18 to either the SSX1 or the SSX2 gene on chromosome X. The chimeric SYT-SSX products are thought to function as transcriptional proteins that deregulate gene expression, thereby providing a putative oncogenic stimulus. We investigated the pattern of gene expression in synovial sarcoma using cDNA microarrays containing 6548 sequence-verified human cDNAs. A tissue microarray containing 37 synovial sarcoma samples verified to bear the SYT-SSX fusion was constructed for complementary analyses. Gene expression analyses were performed on individual tumor samples; 14 synovial sarcomas, 4 malignant fibrous histiocytomas, and 1 fibrosarcoma. Statistical analysis showed a distinct expression profile for the group of synovial sarcomas as compared to the other soft tissue sarcomas, which included variably high expression of ERBB2, IGFBP2, and IGF2 in the synovial sarcomas. Immunohistochemical analysis of protein expression in tissue microarrays of 37 synovial sarcomas demonstrated strong expression of ERBB2 and IGFBP2 in the glandular epithelial component of biphasic tumors and in solid epithelioid areas of some monophasic tumors. Fluorescence in situ hybridization analysis indicated that the ERBB2 overexpression was not because of gene amplification. Differentially expressed genes were also found in a comparison of the expression profiles of the biphasic and monophasic histological subgroups of synovial sarcoma, notably several keratin genes, and ELF3, an epithelial-specific transcription factor gene. Finally, we also noted differential overexpression of several neural- or neuroectodermal-associated genes in synovial sarcomas relative to the comparison sarcoma group, including OLFM1, TLE2, CNTNAP1, and DRPLA. Our high-throughput studies of gene expression patterns, complemented by tissue microarray studies, confirm the distinctive expression profile of synovial sarcoma, provide leads for the study of glandular morphogenesis in this tumor, and identify a new potential therapeutic target, ERBB2, in a subset of cases.

Cell Differentiation↗

Amine-modified random primers to label probes for DNA microarrays.

DNA microarrays have been used to study the expression of thousands of genes at the same time in a variety of cells and tissues. The methods most commonly used to label probes for microarray studies require a minimum of 20 microg of total RNA or 2 microg of poly(A) RNA. This has made it difficult to study small and rare tissue samples. RNA amplification techniques and improved labeling methods have recently been described. These new procedures and reagents allow the use of less input RNA, but they are relatively time-consuming and expensive. Here we introduce a technique for preparing fluorescent probes that can be used to label as little as 1 microg of total RNA. The method is based on priming cDNA synthesis with random hexamer oligonucleotides, on the 5' ends of which are bases with free amino groups. These amine-modified primers are incorporated into the cDNA along with aminoallyl nucleotides, and fluorescent dyes are then chemically added to the free amines. The method is simple to execute, and amine-reactive dyes are considerably less expensive than dye-labeled bases or dendrimers.

3T3 Cells↗

Inference from clustering with application to gene-expression microarrays.

There are many algorithms to cluster sample data points based on nearness or a similarity measure. Often the implication is that points in different clusters come from different underlying classes, whereas those in the same cluster come from the same class. Stochastically, the underlying classes represent different random processes. The inference is that clusters represent a partition of the sample points according to which process they belong. This paper discusses a model-based clustering toolbox that evaluates cluster accuracy. Each random process is modeled as its mean plus independent noise, sample points are generated, the points are clustered, and the clustering error is the number of points clustered incorrectly according to the generating random processes. Various clustering algorithms are evaluated based on process variance and the key issue of the rate at which algorithmic performance improves with increasing numbers of experimental replications. The model means can be selected by hand to test the separability of expected types of biological expression patterns. Alternatively, the model can be seeded by real data to test the expected precision of that output or the extent of improvement in precision that replication could provide. In the latter case, a clustering algorithm is used to form clusters, and the model is seeded with the means and variances of these clusters. Other algorithms are then tested relative to the seeding algorithm. Results are averaged over various seeds. Output includes error tables and graphs, confusion matrices, principal-component plots, and validation measures. Five algorithms are studied in detail: K-means, fuzzy C-means, self-organizing maps, hierarchical Euclidean-distance-based and correlation-based clustering. The toolbox is applied to gene-expression clustering based on cDNA microarrays using real data. Expression profile graphics are generated and error analysis is displayed within the context of these profile graphics. A large amount of generated output is available over the web.

Computational Biology↗

Strong feature sets from small samples.

For small samples, classifier design algorithms typically suffer from overfitting. Given a set of features, a classifier must be designed and its error estimated. For small samples, an error estimator may be unbiased but, owing to a large variance, often give very optimistic estimates. This paper proposes mitigating the small-sample problem by designing classifiers from a probability distribution resulting from spreading the mass of the sample points to make classification more difficult, while maintaining sample geometry. The algorithm is parameterized by the variance of the spreading distribution. By increasing the spread, the algorithm finds gene sets whose classification accuracy remains strong relative to greater spreading of the sample. The error gives a measure of the strength of the feature set as a function of the spread. The algorithm yields feature sets that can distinguish the two classes, not only for the sample data, but for distributions spread beyond the sample data. For linear classifiers, the topic of the present paper, the classifiers are derived analytically from the model, thereby providing an enormous savings in computation time. The algorithm is applied to cancer classification via cDNA microarrays. In particular, the genes BRCA1 and BRCA2 are associated with a hereditary disposition to breast cancer, and the algorithm is used to find gene sets whose expressions can be used to classify BRCA1 and BRCA2 tumors.

Breast Neoplasms↗

Assessing the significance of consistently mis-regulated genes in cancer associated gene expression matrices.

MOTIVATION: The simplest level of statistical analysis of cancer associated gene expression matrices is aimed at finding consistently up- or down-regulated genes within a given set of tumor samples. Considering the high level of gene expression diversity detected in cancer, one needs to assess the probability that the consistent mis-regulation of a given gene is due to chance. Furthermore, it is important to determine the required sample number that will ensure the meaningful statistical analysis of massively parallel gene expression measurements. RESULTS: The probability of consistent mis-regulation is calculated in this paper for binarized gene expression data, using combinatorial considerations. For practical purposes, we also provide a set of accurate approximate formulas for determining the same probability in a computationally less intensive way. When the pool of mis-regulatable genes is restricted, the probability of consistent mis-regulation can be overestimated. We show, however, that this effect has little practical consequences for cancer associated gene expression measurements published in the literature. Finally, in order to aid experimental design, we have provided estimates on the required sample number that will ensure that the detected consistent mis-regulation is not due to chance. Our results suggest that less than 20 sufficiently diverse tumor samples may be enough to identify consistently mis-regulated genes in a statistically significant manner. AVAILABILITY: An implementation using Mathematica (tm) of the main equation of the paper, (4), is available at www.me.chalmers.se/~mwahde/bioinfo.html.

Algorithms↗

Ratio statistics of gene expression levels and applications to microarray data analysis.

MOTIVATION: Expression-based analysis for large families of genes has recently become possible owing to the development of cDNA microarrays, which allow simultaneous measurement of transcript levels for thousands of genes. For each spot on a microarray, signals in two channels must be extracted from their backgrounds. This requires algorithms to extract signals arising from tagged mRNA hybridized to arrayed cDNA locations and algorithms to determine the significance of signal ratios. RESULTS: This paper focuses on estimation of signal ratios from the two channels, and the significance of those ratios. The key issue is the determination of whether a ratio is significantly high or low in order to conclude whether the gene is upregulated or downregulated. The paper builds on an earlier study that involved a hypothesis test based on a ratio statistic under the supposition that the measured fluorescent intensities subsequent to image processing can be assumed to reflect the signal intensities. Here, a refined hypothesis test is considered in which the measured intensities forming the ratio are assumed to be combinations of signal and background. The new method involves a signal-to-noise ratio, and for a high signal-to-noise ratio the new test reduces (with close approximation) to the original test. The effect of low signal-to-noise ratio on the ratio statistics constitutes the main theme of the paper. Finally, and in this vein, a quality metric is formulated for spots. This measure can be used to decide whether or not a spot ratio should be deleted, or to adjust various measurements to reflect confidence in the quality of the measurement. CONTACT: e-dougherty@tamu.edu

Cell Line↗

Simulation of cDNA microarrays via a parameterized random signal model.

cDNA microarrays provide simultaneous expression measurements for thousands of genes that are the result of processing images to recover the average signal intensity from a spot composed of pixels covering the area upon which the cDNA detector has been put down. The accuracy of the signal measurement depends on using an appropriate algorithm to process the images. This includes determining spot locations and processing the data in such a way as to take into account spot geometry, background noise, and various kinds of noise that degrade the signal. This paper presents a stochastic model for microarray images. There are over 20 model parameters, each governed by a probability distribution, that control the signal intensity, spot geometry, spot drift, background effects, and the many kinds of noise that affect microarray images owing to the manner in which they are formed. The model can be used to analyze the performance of image algorithms designed to measure the true signal intensity because the ground truth (signal intensity) for each spot is known. The levels of foreground noise, background noise, and spot distortion can be set, and algorithms can be evaluated under varying conditions.

Algorithms↗