Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Gene selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Multi-class cancer classification using multinomial probit regression with Bayesian gene selection.

We consider the problems of multi-class cancer classification from gene expression data. After discussing the multinomial probit regression model with Bayesian gene selection, we propose two Bayesian gene selection schemes: one employs different strongest genes for different probit regressions; the other employs the same strongest genes for all regressions. Some fast implementation issues for Bayesian gene selection are discussed, including preselection of the strongest genes and recursive computation of the estimation errors using QR decomposition. The proposed gene selection techniques are applied to analyse real breast cancer data, small round blue-cell tumours, the national cancer institute's anti-cancer drug-screen data and acute leukaemia data. Compared with existing multi-class cancer classifications, our proposed methods can find which genes are the most important genes affecting which kind of cancer. Also, the strongest genes selected using our methods are consistent with the biological significance. The recognition accuracies are very high using our proposed methods.

Algorithms↗

A comparison of univariate and multivariate gene selection techniques for classification of cancer datasets.

BACKGROUND: Gene selection is an important step when building predictors of disease state based on gene expression data. Gene selection generally improves performance and identifies a relevant subset of genes. Many univariate and multivariate gene selection approaches have been proposed. Frequently the claim is made that genes are co-regulated (due to pathway dependencies) and that multivariate approaches are therefore per definition more desirable than univariate selection approaches. Based on the published performances of all these approaches a fair comparison of the available results can not be made. This mainly stems from two factors. First, the results are often biased, since the validation set is in one way or another involved in training the predictor, resulting in optimistically biased performance estimates. Second, the published results are often based on a small number of relatively simple datasets. Consequently no generally applicable conclusions can be drawn. RESULTS: In this study we adopted an unbiased protocol to perform a fair comparison of frequently used multivariate and univariate gene selection techniques, in combination with a ränge of classifiers. Our conclusions are based on seven gene expression datasets, across several cancer types. CONCLUSION: Our experiments illustrate that, contrary to several previous studies, in five of the seven datasets univariate selection approaches yield consistently better results than multivariate approaches. The simplest multivariate selection approach, the Top Scoring method, achieves the best results on the remaining two datasets. We conclude that the correlation structures, if present, are difficult to extract due to the small number of samples, and that consequently, overly-complex gene selection algorithms that attempt to extract these structures are prone to overtraining.

Algorithms↗

Generation of marker-free plastid transformants using a transiently cointegrated selection gene.

Genetic engineering of higher plant plastids typically involves stable introduction of antibiotic resistance genes as selection markers. Even though chloroplast genes are maternally inherited in most crops, the possibility of marker transfer to wild relatives or microorganisms cannot be completely excluded. Furthermore, marker expression can be a substantial metabolic drain. Therefore, efficient methods for complete marker removal from plastid transformants are necessary. One method to remove the selection gene from higher plant plastids is based on loop-out recombination, a process difficult to control because selection of homoplastomic transformants is unpredictable. Another method uses the CRE/lox system, but requires additional retransformation and sexual crossing for introduction and subsequent removal of the CRE recombinase. Here we describe the generation of marker-free chloroplast transformants in tobacco using the reconstitution of wild-type pigmentation in combination with plastid transformation vectors, which prevent stable integration of the kanamycin selection marker. One benefit of a procedure using mutants is that marker-free plastid transformants can be produced directly in the first generation (T0) without retransformation or crossing.

Base Sequence↗

Bayesian model averaging: development of an improved multi-class, gene selection and classification tool for microarray data.

MOTIVATION: Selecting a small number of relevant genes for accurate classification of samples is essential for the development of diagnostic tests. We present the Bayesian model averaging (BMA) method for gene selection and classification of microarray data. Typical gene selection and classification procedures ignore model uncertainty and use a single set of relevant genes (model) to predict the class. BMA accounts for the uncertainty about the best set to choose by averaging over multiple models (sets of potentially overlapping relevant genes). RESULTS: We have shown that BMA selects smaller numbers of relevant genes (compared with other methods) and achieves a high prediction accuracy on three microarray datasets. Our BMA algorithm is applicable to microarray datasets with any number of classes, and outputs posterior probabilities for the selected genes and models. Our selected models typically consist of only a few genes. The combination of high accuracy, small numbers of genes and posterior probabilities for the predictions should make BMA a powerful tool for developing diagnostics from expression data. AVAILABILITY: The source codes and datasets used are available from our Supplementary website.

Algorithms↗

A stable gene selection in microarray data analysis.

BACKGROUND: Microarray data analysis is notorious for involving a huge number of genes compared to a relatively small number of samples. Gene selection is to detect the most significantly differentially expressed genes under different conditions, and it has been a central research focus. In general, a better gene selection method can improve the performance of classification significantly. One of the difficulties in gene selection is that the numbers of samples under different conditions vary a lot. RESULTS: Two novel gene selection methods are proposed in this paper, which are not affected by the unbalanced sample class sizes and do not assume any explicit statistical model on the gene expression values. They were evaluated on eight publicly available microarray datasets, using leave-one-out cross-validation and 5-fold cross-validation. The performance is measured by the classification accuracies using the top ranked genes based on the training datasets. CONCLUSION: The experimental results showed that the proposed gene selection methods are efficient, effective, and robust in identifying differentially expressed genes. Adopting the existing SVM-based and KNN-based classifiers, the selected genes by our proposed methods in general give more accurate classification results, typically when the sample class sizes in the training dataset are unbalanced.

Algorithms↗

Characterization of gene regulatory elements for selective gene expression in human melanoma cells.

Since chemotherapy is not sufficiently effective, an alternative strategy for the treatment of advanced melanoma could be an in vivo gene therapy approach. For this purpose, a highly accurate delivery of the therapeutic gene and cell specific gene expression is essential. Since melanocytic cells are characterized by their pigmentation, and since tyrosinase is the key enzyme involved in melanogenesis, we studied the expression of a reporter gene which is under the control of the tyrosinase promoter or a combination of melanocyte-specific enhancer and tyrosinase promoter in ten human melanoma and four epithelial cell lines. Reporter gene expression was upregulated up to 21-fold using the tyrosinase promoter and up 154-fold using the enhancer/promoter construct compared to a control plasmid. Gene expression was strongly associated with capacity of cells for melanin synthesis. The results suggest that the use of tissue specific gene regulatory elements might provide a new opportunity for targeting therapeutic genes to melanoma cells.

Chloramphenicol O-Acetyltransferase↗

Gene selection for sample classification based on gene expression data: study of sensitivity to choice of parameters of the GA/KNN method.

MOTIVATION: We recently introduced a multivariate approach that selects a subset of predictive genes jointly for sample classification based on expression data. We tested the algorithm on colon and leukemia data sets. As an extension to our earlier work, we systematically examine the sensitivity, reproducibility and stability of gene selection/sample classification to the choice of parameters of the algorithm. METHODS: Our approach combines a Genetic Algorithm (GA) and the k-Nearest Neighbor (KNN) method to identify genes that can jointly discriminate between different classes of samples (e.g. normal versus tumor). The GA/KNN method is a stochastic supervised pattern recognition method. The genes identified are subsequently used to classify independent test set samples. RESULTS: The GA/KNN method is capable of selecting a subset of predictive genes from a large noisy data set for sample classification. It is a multivariate approach that can capture the correlated structure in the data. We find that for a given data set gene selection is highly repeatable in independent runs using the GA/KNN method. In general, however, gene selection may be less robust than classification. AVAILABILITY: The method is available at http://dir.niehs.nih.gov/microarray/datamining CONTACT: LI3@niehs.nih.gov

Algorithms↗

Cancer classification and prediction using logistic regression with Bayesian gene selection.

In microarray-based cancer classification and prediction, gene selection is an important research problem owing to the large number of genes and the small number of experimental conditions. In this paper, we propose a Bayesian approach to gene selection and classification using the logistic regression model. The basic idea of our approach is in conjunction with a logistic regression model to relate the gene expression with the class labels. We use Gibbs sampling and Markov chain Monte Carlo (MCMC) methods to discover important genes. To implement Gibbs Sampler and MCMC search, we derive a posterior distribution of selected genes given the observed data. After the important genes are identified, the same logistic regression model is then used for cancer classification and prediction. Issues for efficient implementation for the proposed method are discussed. The proposed method is evaluated against several large microarray data sets, including hereditary breast cancer, small round blue-cell tumors, and acute leukemia. The results show that the method can effectively identify important genes consistent with the known biological findings while the accuracy of the classification is also high. Finally, the robustness and sensitivity properties of the proposed method are also investigated.

Algorithms↗

Gene selection using support vector machines with non-convex penalty.

MOTIVATION: With the development of DNA microarray technology, scientists can now measure the expression levels of thousands of genes simultaneously in one single experiment. One current difficulty in interpreting microarray data comes from their innate nature of 'high-dimensional low sample size'. Therefore, robust and accurate gene selection methods are required to identify differentially expressed group of genes across different samples, e.g. between cancerous and normal cells. Successful gene selection will help to classify different cancer types, lead to a better understanding of genetic signatures in cancers and improve treatment strategies. Although gene selection and cancer classification are two closely related problems, most existing approaches handle them separately by selecting genes prior to classification. We provide a unified procedure for simultaneous gene selection and cancer classification, achieving high accuracy in both aspects. RESULTS: In this paper we develop a novel type of regularization in support vector machines (SVMs) to identify important genes for cancer classification. A special nonconvex penalty, called the smoothly clipped absolute deviation penalty, is imposed on the hinge loss function in the SVM. By systematically thresholding small estimates to zeros, the new procedure eliminates redundant genes automatically and yields a compact and accurate classifier. A successive quadratic algorithm is proposed to convert the non-differentiable and non-convex optimization problem into easily solved linear equation systems. The method is applied to two real datasets and has produced very promising results. AVAILABILITY: MATLAB codes are available upon request from the authors.

Algorithms↗

Identification of probabilistic transcriptional switches in the Ly49 gene cluster: a eukaryotic mechanism for selective gene activation.

Murine natural killer cells selectively express members of the Ly49 family of class I MHC receptors; however, the molecular mechanism controlling probabilistic expression of Ly49 proteins has not been defined. A pair of overlapping, divergent promoters discovered in the Ly49g gene functions as a molecular switch that can produce a forward transcript containing the coding region of the gene (on position) or a noncoding transcript in the opposite direction (off position), and this element maintains transcription in the chosen direction. Competition of C/EBP and TBP transcription factors for overlapping binding sites determines the relative strength of the competing promoters and the probability of transcription in a given direction. Similar elements precede all Ly49 family members, and the relative strength of the forward promoter in each inhibitory Ly49 gene correlates with the percentage of natural killer cells that express a given receptor, supporting a promoter competition model of selective gene activation.

Animals↗

The ties problem resulting from counting-based error estimators and its impact on gene selection algorithms.

MOTIVATION: Feature selection approaches, such as filter and wrapper, have been applied to address the gene selection problem in the literature of microarray data analysis. In wrapper methods, the classification error is usually used as the evaluation criterion of feature subsets. Due to the nature of high dimensionality and small sample size of microarray data, however, counting-based error estimation may not necessarily be an ideal criterion for gene selection problem. RESULTS: Our study reveals that evaluating genes in terms of counting-based error estimators such as resubstitution error, leave-one-out error, cross-validation error and bootstrap error may encounter severe ties problem, i.e. two or more gene subsets score equally, and this in turn results in uncertainty in gene selection. Our analysis finds that the ties problem is caused by the discrete nature of counting-based error estimators and could be avoided by using continuous evaluation criteria instead. Experiment results show that continuous evaluation criteria such as generalised the absolute value of w2 measure for support vector machines and modified Relief's measure for k-nearest neighbors produce improved gene selection compared with counting-based error estimators. AVAILABILITY: The companion website is at http://www.ntu.edu.sg/home5/pg02776030/wrappers/ The website contains (1) the source code of all the gene selection algorithms and (2) the complete set of tables and figures of experiments.

Algorithms↗

Borrelia burgdorferi genes selectively expressed in the infected host.

An immunological screening strategy was used to select microbial genes expressed only in the host. Differential screening of a Borrelia burgdorferi (the Lyme disease agent) expression library identified a gene (p21) encoding a 20.7-kDa antigen that reacted with antibodies in serum from actively infected mice but not serum from mice immunized with heat-killed B. burgdorferi. Selective expression of p21 in the infected host was confirmed by Northern blot analysis and RNA PCR. Further differential screening of the expression library identified at least five additional B. burgdorferi genes are selectively expressed in vivo. This screening method can be used to identify genes induced in vivo in a wide variety of pathogenic microorganisms for which a gene transfer system is not currently available.

Amino Acid Sequence↗

Gene selection in hemoglobin and in antibody-synthesizing cells.

Close linkage of mutually exclusive genes occurs in the non-alpha chain hemoglobin genes and in the immunoglobulin genes of man and other mammals. The expression of one gene in the cluster precludes the expression of any other linked gene. A simple, testable theory of gene selection called "looping-out excision"which was designed only to explain this mutual exclusivity in the hemoglobin system is described. The theory is closely concordant with a wide range of previously unexplained findings concerning hematopoiesis- including the developmental changes of hemoglobins, the increases in immature or fetal forms of hemoglobin that accompany anemia, and with the distribution of adult and fetal hemoglobins among erythrocytes during normal embryogenesis and in various pathological conditions. One corollary of this theory is that erythroid tissue in the normal adult bone marrow is constantly recapitulating the developmental stages of its embryogenesis. Another corollary is that the selection from among the linked globin genes occurs independently on the two chromosomes of the diploid organism. Both of these corollaries are supported by the available data. The same theory of gene selection is also remarkably consistent with known data for immunoglobulin synthesis; it could explain not only the mutually exclusive activation of linked variable genes but also the splicing which occurs between genetically linked variable and constant region genes for the immunoglobulin polypeptide chains. The agreement between these two different tissues is considered to be strong evidence that the proposed mechanism is correct at least in broad outline. Evidence from the genetics of maize and of drosophila also supports this theory of somatic tissue variegation. On the basis of these comparisons, I suggest that looping-out excision probably occurs also in other tissues and may be one means of gene selection and activation in differentiating cells.

Alleles↗

Ovary-selective genes I: the generation and characterization of an ovary-selective complementary deoxyribonucleic acid library.

The importance of several ovary-selective/specific genes, i.e. genes preferentially or exclusively expressed in the ovary, has been established. Indeed, null mutant female mice for the c-mos, growth and differentiation factor-9, alpha-inhibin, and zona pellucida-3 genes proved sterile. A loss of function mutation of the human FSH receptor gene established its critical role in ovarian function. These data support the hypothesis that genes expressed selectively or specifically in the ovary are probably essential for the normal functioning of this organ system. We have used the differential screening technique suppression subtractive hybridization to systematically isolate and clone genes that are expressed in an ovary-selective/specific manner. The resultant target complementary DNA (cDNA) library has been exhaustively screened to a point at which additional sequencing was increasingly unlikely (< or = 4%) to yield additional previously unencountered cDNAs. In toto, 844 clones were sequenced and analyzed for homology to known genes using the Basic Local Alignment Tool (BLAST). Of those, 342 were determined to be independent (nonredundant). One hundred and fifty-nine independent clones proved identical to previously characterized genes, whereas an additional 100 independent clones proved significantly homologous (but not identical) to previously characterized genes. Yet 83 other independent clones did not display significant homology to previously characterized genes now listed in the publicly accessible nonredundant databases. As such, these latter genes were deemed novel. Of these 83 novel genes, a total of 36 displayed ovary-specific/selective expression, as determined by probing mouse multitissue Northern blots with 32P-labeled/PCR-amplified cDNA inserts. Under these circumstances, the false positive rate was minimal, as only one novel clone was expressed at a higher level in nonovarian tissues relative to ovary. Of the 36 ovary-specific/ selective novel genes, 22 proved subject to hormonal regulation during a simulated estrous cycle. In this communication we focus on 2 such novel ovary-specific/hormonally-dependent genes, the full-length sequences of which were isolated using rapid amplification of 3'-cDNA ends technology. Taken together, the present study accomplished systematic identification of those genes that are restricted in their expression to the ovary. These ovary-selective genes may have significant implications for the understanding of ovarian function in molecular terms and for the development of innovative strategies for the promotion of fertility or its control.

17-Hydroxysteroid Dehydrogenases↗

Gene selection in cancer classification using sparse logistic regression with Bayesian regularization.

MOTIVATION: Gene selection algorithms for cancer classification, based on the expression of a small number of biomarker genes, have been the subject of considerable research in recent years. Shevade and Keerthi propose a gene selection algorithm based on sparse logistic regression (SLogReg) incorporating a Laplace prior to promote sparsity in the model parameters, and provide a simple but efficient training procedure. The degree of sparsity obtained is determined by the value of a regularization parameter, which must be carefully tuned in order to optimize performance. This normally involves a model selection stage, based on a computationally intensive search for the minimizer of the cross-validation error. In this paper, we demonstrate that a simple Bayesian approach can be taken to eliminate this regularization parameter entirely, by integrating it out analytically using an uninformative Jeffrey's prior. The improved algorithm (BLogReg) is then typically two or three orders of magnitude faster than the original algorithm, as there is no longer a need for a model selection step. The BLogReg algorithm is also free from selection bias in performance estimation, a common pitfall in the application of machine learning algorithms in cancer classification. RESULTS: The SLogReg, BLogReg and Relevance Vector Machine (RVM) gene selection algorithms are evaluated over the well-studied colon cancer and leukaemia benchmark datasets. The leave-one-out estimates of the probability of test error and cross-entropy of the BLogReg and SLogReg algorithms are very similar, however the BlogReg algorithm is found to be considerably faster than the original SLogReg algorithm. Using nested cross-validation to avoid selection bias, performance estimation for SLogReg on the leukaemia dataset takes almost 48 h, whereas the corresponding result for BLogReg is obtained in only 1 min 24 s, making BLogReg by far the more practical algorithm. BLogReg also demonstrates better estimates of conditional probability than the RVM, which are of great importance in medical applications, with similar computational expense. AVAILABILITY: A MATLAB implementation of the sparse logistic regression algorithm with Bayesian regularization (BLogReg) is available from http://theoval.cmp.uea.ac.uk/~gcc/cbl/blogreg/

Algorithms↗

Recursive gene selection based on maximum margin criterion: a comparison with SVM-RFE.

BACKGROUND: In class prediction problems using microarray data, gene selection is essential to improve the prediction accuracy and to identify potential marker genes for a disease. Among numerous existing methods for gene selection, support vector machine-based recursive feature elimination (SVM-RFE) has become one of the leading methods and is being widely used. The SVM-based approach performs gene selection using the weight vector of the hyperplane constructed by the samples on the margin. However, the performance can be easily affected by noise and outliers, when it is applied to noisy, small sample size microarray data. RESULTS: In this paper, we propose a recursive gene selection method using the discriminant vector of the maximum margin criterion (MMC), which is a variant of classical linear discriminant analysis (LDA). To overcome the computational drawback of classical LDA and the problem of high dimensionality, we present efficient and stable algorithms for MMC-based RFE (MMC-RFE). The MMC-RFE algorithms naturally extend to multi-class cases. The performance of MMC-RFE was extensively compared with that of SVM-RFE using nine cancer microarray datasets, including four multi-class datasets. CONCLUSION: Our extensive comparison has demonstrated that for binary-class datasets MMC-RFE tends to show intermediate performance between hard-margin SVM-RFE and SVM-RFE with a properly chosen soft-margin parameter. Notably, MMC-RFE achieves significantly better performance with a smaller number of genes than SVM-RFE for multi-class datasets. The results suggest that MMC-RFE is less sensitive to noise and outliers due to the use of average margin, and thus may be useful for biomarker discovery from noisy data.

Algorithms↗

Filter versus wrapper gene selection approaches in DNA microarray domains.

DNA microarray experiments generating thousands of gene expression measurements, are used to collect information from tissue and cell samples regarding gene expression differences that could be useful for diagnosis disease, distinction of the specific tumor type, etc. One important application of gene expression microarray data is the classification of samples into known categories. As DNA microarray technology measures the gene expression en masse, this has resulted in data with the number of features (genes) far exceeding the number of samples. As the predictive accuracy of supervised classifiers that try to discriminate between the classes of the problem decays with the existence of irrelevant and redundant features, the necessity of a dimensionality reduction process is essential. We propose the application of a gene selection process, which also enables the biology researcher to focus on promising gene candidates that actively contribute to classification in these large scale microarrays. Two basic approaches for feature selection appear in machine learning and pattern recognition literature: the filter and wrapper techniques. Filter procedures are used in most of the works in the area of DNA microarrays. In this work, a comparison between a group of different filter metrics and a wrapper sequential search procedure is carried out. The comparison is performed in two well-known DNA microarray datasets by the use of four classic supervised classifiers. The study is carried out over the original-continuous and three-intervals discretized gene expression data. While two well-known filter metrics are proposed for continuous data, four classic filter measures are used over discretized data. The same wrapper approach is used for both continuous and discretized data. The application of filter and wrapper gene selection procedures leads to considerably better accuracy results in comparison to the non-gene selection approach, coupled with interesting and notable dimensionality reductions. Although the wrapper approach mainly shows a more accurate behavior than filter metrics, this improvement is coupled with considerable computer-load necessities. We note that most of the genes selected by proposed filter and wrapper procedures in discrete and continuous microarray data appear in the lists of relevant-informative genes detected by previous studies over these datasets. The aim of this work is to make contributions in the field of the gene selection task in DNA microarray datasets. By an extensive comparison with more popular filter techniques, we would like to make contributions in the expansion and study of the wrapper approach in this type of domains.

Artificial Intelligence↗

Gene selection: a Bayesian variable selection approach.

UNLABELLED: Selection of significant genes via expression patterns is an important problem in microarray experiments. Owing to small sample size and the large number of variables (genes), the selection process can be unstable. This paper proposes a hierarchical Bayesian model for gene (variable) selection. We employ latent variables to specialize the model to a regression setting and uses a Bayesian mixture prior to perform the variable selection. We control the size of the model by assigning a prior distribution over the dimension (number of significant genes) of the model. The posterior distributions of the parameters are not in explicit form and we need to use a combination of truncated sampling and Markov Chain Monte Carlo (MCMC) based computation techniques to simulate the parameters from the posteriors. The Bayesian model is flexible enough to identify significant genes as well as to perform future predictions. The method is applied to cancer classification via cDNA microarrays where the genes BRCA1 and BRCA2 are associated with a hereditary disposition to breast cancer, and the method is used to identify a set of significant genes. The method is also applied successfully to the leukemia data. SUPPLEMENTARY INFORMATION: http://stat.tamu.edu/people/faculty/bmallick.html.

Algorithms↗