Search PubMed⌕ Search

PubMed · 15806617

Gene selection for microarray data analysis using principal component analysis.

Abstract

Principal component analysis (PCA) has been widely used in multivariate data analysis to reduce the dimensionality of the data in order to simplify subsequent analysis and allow for summarization of the data in a parsimonious manner. It has become a useful tool in microarray data analysis. For a typical microarray data set, it is often difficult to compare the overall gene expression difference between observations from different groups or conduct the classification based on a very large number of genes. In this paper, we propose a gene selection method based on the strategy proposed by Krzanowski. We demonstrate the effectiveness of this procedure using a cancer gene expression data set and compare it with several other gene selection strategies. It turns out that the proposed method selects the best gene subset for preserving the original data structure.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Antai Wang, Edmund A Gehan. 2005-07-15. Gene selection for microarray data analysis using principal component analysis.. https://doi.org/10.1002/sim.2082

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Protocol to improve isoform-level quantification of low-abundance transcripts via STALARD pre-amplification.

STALARD (selective target amplification for low-abundance RNA detection) enables isoform-level quantification of low-abundance RNAs using conventional laboratory equipment. Here, we describe steps for RNA isolation, primer design, reverse transcription, selective target amplification, and downstream analysis. The protocol couples selective pre-amplification with a quantitative reverse-transcription PCR (RT-qPCR) readout and optional nanopore sequencing. Using 1 μg input RNA and 12 pre-amplification cycles, STALARD reduces Cq values by approximately 10-12 cycles, bringing the target into a reliably quantifiable range. For complete details on the use and execution of this protocol, please refer to Jeong et al.1.

Gene Expression↗

In silico comparison of gene expression levels in ten human tumor types reveals candidate genes associated with carcinogenesis.

Most human cancers are characterized by genomic instability. Changes associated with such may result in altered expression of numerous genes. The sequence information available in the public databases can be used to identify transcripts differentially expressed in cancers. Determining cancer-related genes that are commonly deregulated in different tumor types may facilitate identification of targets for cancer diagnoses and therapeutic treatments. Using a data-mining tool named Digital Differential Display (DDD) from the UniGene database at the NCBI web site, gene expression levels of ten different tumor types and their counterpart normal tissues were analyzed. Unigenes which showed transcriptional regulation in more than five tumor types with > or =2-fold differences from normal tissues were identified. The expression data of selected Unigenes were subjected to clustering analysis. 127 commonly up-regulated genes and 92 commonly down-regulated genes were identified. Clustering analysis using these genes showed that most tumor types can be clustered into a separate branch from most normal tissues. Nineteen genes that have been shown to be involved in carcinogenesis by experimental evidence were also identified. Present computational analyses revealed 219 candidate cancer-related genes that are commonly deregulated in ten human tumor types which may contribute to the progress of carcinogenesis.

Gene Expression↗