Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “cancer subtype classification”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Multi-class cancer subtype classification based on gene expression signatures with reliability analysis.

Differential diagnosis among a group of histologically similar cancers poses a challenging problem in clinical medicine. Constructing a classifier based on gene expression signatures comprising multiple discriminatory molecular markers derived from microarray data analysis is an emerging trend for cancer diagnosis. To identify the best genes for classification using a small number of samples relative to the genome size remains the bottleneck of this approach, despite its promise. We have devised a new method of gene selection with reliability analysis, and demonstrated that this method can identify a more compact set of genes than other methods for constructing a classifier with optimum predictive performance for both small round blue cell tumors and leukemia. High consensus between our result and the results produced by methods based on artificial neural networks and statistical techniques confers additional evidence of the validity of our method. This study suggests a way for implementing a reliable molecular cancer classifier based on gene expression signatures.

Artificial Intelligence↗

New gene selection method for classification of cancer subtypes considering within-class variation.

In this work we propose a new method for finding gene subsets of microarray data that effectively discriminates subtypes of disease. We developed a new criterion for measuring the relevance of individual genes by using mean and standard deviation of distances from each sample to the class centroid in order to treat the well-known problem of gene selection, large within-class variation. Also this approach has the advantage that it is applicable not only to binary classification but also to multiple classification problems. We demonstrated the performance of the method by applying it to the publicly available microarray datasets, leukemia (two classes) and small round blue cell tumors (four classes). The proposed method provides a very small number of genes compared with the previous methods without loss of discriminating power and thus it can effectively facilitate further biological and clinical researches.

Acute Disease↗

Cross-platform array comparative genomic hybridization meta-analysis separates hematopoietic and mesenchymal from epithelial tumors.

A series of studies have been published that evaluate the chromosomal copy number changes of different tumor classes using array comparative genomic hybridization (array CGH); however, the chromosomal aberrations that distinguish the different tumor classes have not been fully characterized. Therefore, we performed a meta-analysis of different array CGH data sets in an attempt to classify samples tested across different platforms. As opposed to RNA expression, a common reference is used in dual channel CGH arrays: normal human DNA, theoretically facilitating cross-platform analysis. To this aim, cell line and primary cancer data sets from three different dual channel array CGH platforms obtained by four different institutes were integrated. The cell line data were used to develop preprocessing methods, which performed noise reduction and transformed samples into a common format. The transformed array CGH profiles allowed perfect clustering by cell line, but importantly not by platform or institute. The same preprocessing procedures used for the cell line data were applied to data from 373 primary tumors profiled by array CGH, including controls. Results indicated that there is no apparent feature related to the institute or platform and that array CGH allows for unambiguous cross-platform meta-analysis. Major clusters with common tissue origin were identified. Interestingly, tumors of hematopoietic and mesenchymal origins cluster separately from tumors of epithelial origin. Therefore, it can be concluded that chromosomal aberrations of tumors from hematopoietic and mesenchymal origin versus tumors of epithelial origin are distinct, and these differences can be picked up by meta-analysis of array CGH data. This suggests the possibility of prospectively using combined analysis of diverse copy number data sets for cancer subtype classification.

Chromosome Aberrations↗

Redefinition of Affymetrix probe sets by sequence overlap with cDNA microarray probes reduces cross-platform inconsistencies in cancer-associated gene expression measurements.

BACKGROUND: Comparison of data produced on different microarray platforms often shows surprising discordance. It is not clear whether this discrepancy is caused by noisy data or by improper probe matching between platforms. We investigated whether the significant level of inconsistency between results produced by alternative gene expression microarray platforms could be reduced by stringent sequence matching of microarray probes. We mapped the short oligo probes of the Affymetrix platform onto cDNA clones of the Stanford microarray platform. Affymetrix probes were reassigned to redefined probe sets if they mapped to the same cDNA clone sequence, regardless of the original manufacturer-defined grouping. The NCI-60 gene expression profiles produced by Affymetrix HuFL platform were recalculated using these redefined probe sets and compared to previously published cDNA measurements of the same panel of RNA samples. RESULTS: The redefined probe sets displayed a substantially higher level of cross-platform consistency at the level of gene correlation, cell line correlation and unsupervised hierarchical clustering. The same strategy allowed an almost complete correspondence of breast cancer subtype classification between Affymetrix gene chip and cDNA microarray derived gene expression data, and gave an increased level of similarity between normal lung derived gene expression profiles using the two technologies. In total, two Affymetrix gene-chip platforms were remapped to three cDNA platforms in the various cross-platform analyses, resulting in improved concordance in each case. CONCLUSION: We have shown that probes which target overlapping transcript sequence regions on cDNA microarrays and Affymetrix gene-chips exhibit a greater level of concordance than the corresponding Unigene or sequence matched features. This method will be useful for the integrated analysis of gene expression data generated by multiple disparate measurement platforms.

Breast Neoplasms↗

Bladder cancer outcome and subtype classification by gene expression.

Models of bladder tumor progression have suggested that genetic alterations may determine both phenotype and clinical course. We have applied expression microarray analysis to a divergent set of bladder tumors to further elucidate the course of disease progression and to classify tumors into more homogeneous and clinically relevant subgroups. cDNA microarrays containing 10,368 human gene elements were used to characterize the global gene expression patterns in 80 bladder tumors, 9 bladder cancer cell lines, and 3 normal bladder samples. Robust statistical approaches accounting for the multiple testing problem were used to identify differentially expressed genes. Unsupervised hierarchical clustering successfully separated the samples into two subgroups containing superficial (pT(a) and pT(1)) versus muscle-invasive (pT(2)-pT(4)) tumors. Supervised classification had a 90.5% success rate separating superficial from muscle-invasive tumors based on a limited subset of genes. Tumors could also be classified into transitional versus squamous subtypes (89% success rate) and good versus bad prognosis (78% success rate). The performance of our stage classifiers was confirmed in silico using data from an independent tumor set. Validation of differential expression was done using immunohistochemistry on tissue microarrays for cathepsin E, cyclin A2, and parathyroid hormone-related protein. Genes driving the separation between tumor subsets may prove to be important biomarkers for bladder cancer development and progression and eventually candidates for therapeutic targeting.

Aged↗

A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks.

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, DNA methylation, copy number variations, and gene expression, and were used to assess the model's generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

Humans↗

Optimal approach for classification of acute leukemia subtypes based on gene expression data.

The classification of cancer subtypes, which is critical for successful treatment, has been studied extensively with the use of gene expression profiles from oligonucleotide chips or cDNA microarrays. Various pattern recognition methods have been successfully applied to gene expression data. However, these methods are not optimal, rather they are high-performance classifiers that emphasize only classification accuracy. In this paper, we propose an approach for the construction of the optimal linear classifier using gene expression data. Two linear classification methods, linear discriminant analysis (LDA) and discriminant partial least-squares (DPLS), are applied to distinguish acute leukemia subtypes. These methods are shown to give satisfactory accuracy. Moreover, we determined optimally the number of genes participating in the classification (a remarkably small number compared to previous results) on the basis of the statistical significance test. Thus, the proposed method constructs the optimal classifier that is composed of a small size predictor and provides high accuracy.

Algorithms↗

Expert review of non-Hodgkin's lymphomas in a population-based cancer registry: reliability of diagnosis and subtype classifications.

Incidence rates of non-Hodgkin's lymphomas (NHLs) have nearly doubled in recent decades. Understanding the reasons behind these trends will require detailed surveillance and epidemiological study of NHL subtypes in large populations, using cancer registry or other multicenter data. However, little is known regarding the reliability of NHL diagnosis and subtype classification in such data, despite implications for the accuracy of incidence statistics and studies. Expert pathological re-review was completed for 1526 NHL patients who were reported to the Greater Bay Area Cancer Registry and who participated in a large population-based case-control study. Agreement of registry diagnosis with expert diagnosis and with International Classification of Diseases for Oncology-2 (Working Formulation) subtype classifications was measured with positive predictive values and kappa statistics. Agreement of registry and expert diagnoses was high (98%). Thirty patients were found on review not to have NHL; most of these had leukemia. For subtypes, agreement of registry and expert classification was more moderate (59%). Agreement varied substantially by subtype from 5% to 100% and was 77% for the most common subtype, diffuse large cell lymphoma. Seventy-seven percent of 128 registry-unclassified lymphomas were assigned a subtype on re-review. Our analyses suggest excellent diagnostic reliability but poorer subtype reliability of NHL in cancer registry data information that is critical to the interpretation of lymphoma time trends. Thus, overall NHL incidence and survival statistics from the early 1990s are probably accurate, but subtype-specific statistics could be substantially biased, especially because of high (15-20%) proportions of unclassified lymphomas.

Adult↗

Class discovery analysis of the lung cancer gene expression data.

Traditional histological classification of lung cancer subtypes is informative, but incomplete. Recent studies of gene expression suggest that molecular classification can be used for effective diagnostic and prediction of the treatment outcome. We attempt to build a molecular classification based on the public data available from a few independent sources. The data is reanalyzed with a new cluster analysis algorithm. This algorithm allows us to preserve the high dimensionality of data and produce the cluster structure without preliminary selection of significant genes or any other presumption about the relation between different cancer and normal tissue samples. The resulting clusters are generally consistent with the histological classification. However, our analysis reveals many additional details and subtypes of previously defined types of lung cancer. Large histological cancer types can be further divided into subclasses with different patterns of gene expression. These subtypes should be taken into account in diagnostics, drug testing, and treatment development for lung cancer patients.

Cluster Analysis↗

New developments in microarray technology.

Microarrays have emerged as indispensable research tools for gene expression profiling and mutation analysis. New classification of cancer subtypes, dissecting the yeast metabolism and large-scale genotyping of human single nucleotide polymorphisms are important results being obtained with this technique. Realizing the microsphere-based massively parallel signature sequencing technique as fluid microarrays, building new types of protein arrays and constructing miniaturized flow-through systems, which can potentially take this technology from the research bench into industrial, clinical and other routine applications, exemplify the intense developments that are now ongoing in this field.

Biotechnology↗

Large-scale gene expression analysis in molecular target discovery.

The evolution of simple arrays consisting of a few genes to ones composed of thousands of genes and/or ESTs has allowed investigators unprecedented views of the molecular mechanisms within cells. Due to the enormous quantities of information derived from microarray analysis, new types of problems have surfaced, such as where to store all of the data. The ability to solve database or statistical problems has required the bench biologist to collaborate with database developers, software designers and statisticians to determine solutions for storage, analysis and interpretation of microarray data. The collaborative effort between these extremely diverse disciplines has led to the development of creative database query and gene expression analysis tools, producing significant reductions in the time required by researchers to filter through the datasets and discover the key processes perturbed by the diseases of interest. Both unsupervised and supervised analysis methods have been applied to gene expression data leading to the discovery of novel therapeutic targets and diagnostic markers. Furthermore, tumor classification based on their respective molecular fingerprints has led to the classification of cancer subtypes and the discovery of novel molecular taxonomies that may eventually lead to improved patient stratification and superior therapeutic strategies.

Databases, Factual↗

Histological grading in gastric cancer by Ming classification: correlation with histopathological subtypes, metastasis, and prognosis.

The aim of this prospective study was to analyze Ming's classification in correlation with other currently used classification systems of gastric cancer. In addition, we wanted to define the prognostic significance of the Ming classification system. The present study analyzed material of 117 patients with gastric carcinoma who underwent D2-gastrectomy with curative intent. All specimens were categorized according to International Union Against Cancer (UICC) classification, World Health Organization (WHO) classification, Borrmann classification, Laurén classification, Goseki classification, Ming classification, and tumor differentiation. For analysis of correlation between the classification systems, the correlation coefficient according to Spearman was calculated. The survival curves have been calculated according to the Kaplan-Meier method. According to the Ming classification, 38.5% of the carcinomas exhibited an expanding growth pattern, and 61.5% of specimens showed an infiltrating growth pattern. The subtypes according to the Ming and Laurén classification correlated significantly (P < 0.001). WHO classification (P < 0.001), tumor differentiation (P < 0.001), and Goseki classification (P < 0.001), as well as the macroscopic classification of Borrmann (P < 0.001) and the pT and pN categories of the UICC classification exhibited a highly significant correlation with the Ming classification (P < 0.001 and 0.001, respectively). Median overall survival was 31.3 months. In Kaplan-Meier analysis, the 3-year survival rates were lower in the infiltrative tumor type when compared to the expansive tumor type according to Ming (P = 0.0847). In multivariate analysis, only the UICC system presented as an independent prognostic factor in multivariate analysis (P < 0.001). This study shows that the Ming classification correlates significantly with the currently used classification systems for gastric cancer and with the UICC staging system, especially, the pT and pN category. The 3-year survival rates were lower in the infiltrative tumor type than in the expansive tumor type according to Ming. However, the Ming classification is not an independent prognostic factor.

Adenocarcinoma↗

Histological grading in gastric cancer by Goseki classification: correlation with histopathological subtypes and prognosis.

BACKGROUND: Many different classification systems have been proposed for the histological classification and grading of gastric cancer. In 1992 Goseki described a novel classification system for gastric cancer based on tubular differentiation and mucus in the cytoplasm. The aim of the study was to compare the Goseki classification with the currently used classification systems and to define the prognostic significance of the Goseki classification system. PATIENTS AND METHODS: The present study analyzed material from 200 gastric carcinoma patients who underwent gastrectomy with curative intention. All specimens were categorized to UICC-classification, WHO-classification, Laurén classification, tumor differentiation and Goseki classification. The median follow-up for surviving patients was 3.75 years (range, 0.14-11.52). RESULTS: According to the Goseki classification 32% of patients were classified as group I, 11.5% as group II, 9.5% as group III and 48% as group IV. The Goseki classification was found to correlate with the WHO and Lauren classification as well as with conventional grading. Goseki classification as well as tumor differentiation, Lauren and WHO classification did not have prognostic value for survival. Only the UICC system presented as an independent prognostic factor in multivariate analysis (p < 0.000001). CONCLUSION: In our series Goseki classification correlated with conventional classification systems, but not with survival.

Adenocarcinoma↗

Quantitative chromatics analysis for computer imaging of cytologic subtypes of lung cancer stained by Papanicolaou stain.

OBJECTIVE: To determine the role of quantitative chromatics analysis in the classification of subtypes of lung cancer stained by Papanicolaou stain. STUDY DESIGN: By means of computer image analysis, 60 keratinized squamous carcinoma cells (KSCC), 88 nonkeratinized squamous carcinoma cells (NKSCC) and 150 adenocarcinoma cells (ACC) from lung cancer in sputum smears stained by Papanicolaou stain were analyzed and distinguished based on quantitative colorimetry. The features measured were the content of three primary colors, red (R), green (G) and blue (B) and the coefficients of R, G and B (r, g and b, respectively). Hue, saturation, brightness and gray level were also measured. A stepwise discriminant analysis was carried out. RESULTS: The values of R, G and B and r, g and b, hue and saturation in NKSCC and ACC were significantly different from those of KSCC, and the changes in the three primary colors were more sensitive than those in the gray level. Computer assessment based on three primary color coefficients, hue and saturation yielded accuracy of distinguishing KSCC from NKSCC and KSCC from ACC of 95.2% and 95%, respectively. CONCLUSION: Quantitative analyses of R, G and B and r, g, b and hue and saturation are valuable in distinguishing KSCC from NKSCC and ACC.

Adenocarcinoma↗

IGCN: integrative graph convolution networks for patient level insights and biomarker discovery in multi-omics integration.

MOTIVATION: Developing computational tools for integrative analysis across multiple types of omics data has been of immense importance in cancer molecular biology and precision medicine research. While recent advancements have yielded integrative prediction solutions for multi-omics data, these methods lack a comprehensive and cohesive understanding of the rationale behind their specific predictions. To shed light on personalized medicine and unravel previously unknown characteristics within integrative analysis of multi-omics data, we introduce a novel integrative neural network approach for cancer molecular subtype and biomedical classification applications, named Integrative Graph Convolutional Networks (IGCN). RESULTS: To demonstrate the superiority of IGCN, we compare its performance with other state-of-the-art approaches across different cancer subtype and biomedical classification tasks. Our experimental results show that our proposed model outperforms the state-of-the-art and baseline methods. IGCN identifies which types of omics data receive more emphasis for each patient when predicting a specific class. Additionally, IGCN has the capability to pinpoint significant biomarkers from a range of omics data types. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/bozdaglab/IGCN.

Humans↗

Breast cancer molecular subtypes respond differently to preoperative chemotherapy.

PURPOSE: Molecular classification of breast cancer has been proposed based on gene expression profiles of human tumors. Luminal, basal-like, normal-like, and erbB2+ subgroups were identified and were shown to have different prognoses. The goal of this research was to determine if these different molecular subtypes of breast cancer also respond differently to preoperative chemotherapy. EXPERIMENTAL DESIGN: Fine needle aspirations of 82 breast cancers were obtained before starting preoperative paclitaxel followed by 5-fluorouracil, doxorubicin, and cyclophosphamide chemotherapy. Gene expression profiling was done with Affymetrix U133A microarrays and the previously reported "breast intrinsic" gene set was used for hierarchical clustering and multidimensional scaling to assign molecular class. RESULTS: The basal-like and erbB2+ subgroups were associated with the highest rates of pathologic complete response (CR), 45% [95% confidence interval (95% CI), 24-68] and 45% (95% CI, 23-68), respectively, whereas the luminal tumors had a pathologic CR rate of 6% (95% CI, 1-21). No pathologic CR was observed among the normal-like cancers (95% CI, 0-31). Molecular class was not independent of conventional cliniocopathologic predictors of response such as estrogen receptor status and nuclear grade. None of the 61 genes associated with pathologic CR in the basal-like group were associated with pathologic CR in the erbB2+ group, suggesting that the molecular mechanisms of chemotherapy sensitivity may vary between these two estrogen receptor-negative subtypes. CONCLUSIONS: The basal-like and erbB2+ subtypes of breast cancer are more sensitive to paclitaxel- and doxorubicin-containing preoperative chemotherapy than the luminal and normal-like cancers.

Adult↗

Detection of keratin subtypes in routinely processed cervical tissue: implications for tumour classification and the study of cervix cancer aetiology.

We investigated the expression of keratin subtypes 7, 8, 10, 13, 14, 17, 18 and 19 in the normal cervix, in cervical intraepithelial neoplasia (CIN) lesions and in cervical carcinomas, using a selected panel of monoclonal keratin antibodies, reactive with routinely processed, formalin fixed paraffin embedded tissue fragments. The reaction patterns derived for each keratin antibody were compared with known expression patterns of the various epithelia, previously examined in frozen tissues. Although the reactivity of the antibodies was generally acceptable, considerable modifications to the manufacturers' staining instructions were often necessary. For some antibodies, which were previously thought to be reactive with fresh frozen tissue only, we developed staining protocols rendering them reactive with routinely processed material. As with previous findings in frozen sections we observed increasing expression of keratins 7, 8, 17, 18 and 19 with increasing grade of CIN. In cervical carcinomas the differences in keratin detectability between the main categories were more pronounced than in frozen sections, probably due to fixation and processing. For routine pathology, keratin phenotyping of cervical lesions may be of value in classification. The fact that keratin 7 was detected for the first time in reserve cells, and that this keratin was also found to be expressed in a considerable number of CIN lesions and cervical carcinomas supports the suggestion that reserve cells are a common progenitor cell for these lesions.

Adenocarcinoma↗

Flow cytometric analysis of DNA ploidy pattern from deparaffinized formalin-fixed gastric cancer tissue.

Histologically processed tissue from gastric cancers has been analyzed by flow cytometry in an attempt to correlate DNA ploidy pattern and behavior of the tumor. Of the mucosal and submucosal cancers (so-called early, all stage I in the present series), 62.7% show a diploid DNA pattern and 37.3% show a single aneuploid pattern. Of the deeply infiltrating (beyond the submucosa) cancers (stage II and III), 52.1% are single aneuploid and 47.9% are multiploid. While stage-I patients are all alive at the end of the follow-up period (6 years), in stage II and III cases Cox's regression model shows that the hazard function depends on DNA pattern: survival is negatively influenced by multiploidy. On this basis, it may be assumed that the DNA pattern is a useful prognostic indicator of gastric cancer. As expected, in Cox's regression model an even more important negative correlation exists between survival and stage: single aneuploid cases in stage II have a better prognosis than those in stage III. Instead, no correlation is found between histological cancer subtype (Laurén and WHO classifications), grade and DNA pattern.

Adult↗