Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Cytology of ductal lavage fluid of the breast.

The cytologic evaluation of nipple aspirate fluids has been shown to identify women at increased risk for developing breast cancer. One limitation of this assay is the often scant cellularity of the specimen. An improved technique, ductal lavage, utilizes a microcatheter inserted into individual breast ducts to collect large numbers of cells for cytologic evaluation. Epithelial cells in ductal lavage fluids can be categorized as benign, malignant, or showing mildly or markedly atypical changes. The cell characteristics which were most helpful in identifying abnormal cells were related to cell arrangement, cell size, nuclear size, and size variation, nuclear membrane irregularity, chromatin granularity, and the presence of large nucleoli. Cell size, nuclear size variation, and large nucleoli were the most robust features, as determined by agreement between two pathologists. Moderate cell enlargement and the presence of large nucleoli were the features selected by structured tree analysis for classifying the specimens into the diagnostic groups. The similarity of the cytology of ductal lavage fluid to nipple aspirate fluid strongly suggests that these specimens will also be useful for predicting breast cancer risk.

Body Fluids↗

Prognostic classification of relapsing favorable histology Wilms tumor using cDNA microarray expression profiling and support vector machines.

Treatment of Wilms tumor has a high success rate, with some 85% of patients achieving long-term survival. However, late effects of treatment and management of relapse remain significant clinical problems. If accurate prognostic methods were available, effective risk-adapted therapies could be tailored to individual patients at diagnosis. Few molecular prognostic markers for Wilms tumor are currently defined, though previous studies have linked allele loss on 1p or 16q, genomic gain of 1q, and overexpression from 1q with an increased risk of relapse. To identify specific patterns of gene expression that are predictive of relapse, we used high-density (30 k) cDNA microarrays to analyze RNA samples from 27 favorable histology Wilms tumors taken from primary nephrectomies at the time of initial diagnosis. Thirteen of these tumors relapsed within 2 years. Genes differentially expressed between the relapsing and nonrelapsing tumor classes were identified by statistical scoring (t test). These genes encode proteins with diverse molecular functions, including transcription factors, developmental regulators, apoptotic factors, and signaling molecules. Use of a support vector machine classifier, feature selection, and test evaluation using cross-validation led to identification of a generalizable expression signature, a small subset of genes whose expression potentially can be used to predict tumor outcome in new samples. Similar methods were used to identify genes that are differentially expressed between tumors with and without genomic 1q gain. This set of discriminators was highly enriched in genes on 1q, indicating close agreement between data obtained from expression profiling with data from genomic copy number analyses.

Adolescent↗

Automated diagnosis of pigmented skin lesions.

Since advanced melanoma remains practically incurable, early detection is an important step toward a reduction in mortality. High expectations are entertained for a technique known as dermoscopy or epiluminescence light microscopy; however, evaluation of pigmented skin lesions by this method is often extremely complex and subjective. To obviate the problem of qualitative interpretation, methods based on mathematical analysis of pigmented skin lesions, such as digital dermoscopy analysis, have been developed. In the present study, we used a digital dermoscopy analyzer (DBDermo-Mips system) to evaluate a series of 588 excised, clinically atypical, flat pigmented skin lesions (371 benign, 217 malignant). The analyzer evaluated 48 parameters grouped into 4 categories (geometries, colors, textures and islands of color), which were used to train an artificial neural network. To evaluate the diagnostic performance of the neural network and to check it during the training process, we used the error area over the receiver operating characteristic curve. The discriminating power of the digital dermoscopy analyzer plus artificial neural network was compared with histologic diagnosis. A feature selection procedure indicated that as few as 13 of the variables were sufficient to discriminate the 2 groups of lesions, and this also ensured high generalization power. The artificial neural network designed with these variables enabled a diagnostic accuracy of about 94%. In conclusion, the good diagnostic performance and high speed in reading and analyzing lesions (real time) of our method constitute an important step in the direction of automated diagnosis of pigmented skin lesions.

Automation↗

Gene expression profiling of 30 cancer cell lines predicts resistance towards 11 anticancer drugs at clinically achieved concentrations.

Cancer patients with tumors of similar grading, staging and histogenesis can have markedly different treatment responses to different chemotherapy agents. So far, individual markers have failed to correctly predict resistance against anticancer agents. We tested 30 cancer cell lines for sensitivity to 5-fluorouracil, cisplatin, cyclophosphamide, doxorubicin, etoposide, methotrexate, mitomycin C, mitoxantrone, paclitaxel, topotecan and vinblastine at drug concentrations that can be systemically achieved in patients. The resistance index was determined to designate the cell lines as sensitive or resistant, and then, the subset of resistant vs. sensitive cell lines for each drug was compared. Gene expression signatures for all cell lines were obtained by interrogating Affymetrix U133A arrays. Prediction Analysis of Microarrays was applied for feature selection. An individual prediction profile for the resistance against each chemotherapy agent was constructed, containing 42-297 genes. The overall accuracy of the predictions in a leave-one-out cross validation was 86%. A list of the top 67 multidrug resistance candidate genes that were associated with the resistance against at least 4 anticancer agents was identified. Moreover, the differential expressions of 46 selected genes were also measured by quantitative RT-PCR using a TaqMan micro fluidic card system. As a single gene can be correlated with resistance against several agents, associations with resistance were detected all together for 76 genes and resistance phenotypes, respectively. This study focuses on the resistance at the in vivo concentrations, making future clinical cancer response prediction feasible. The TaqMan-validated gene expression patterns provide new gene candidates for multidrug resistance.

Antineoplastic Agents↗

Accurate prediction of the blood-brain partitioning of a large set of solutes using ab initio calculations and genetic neural network modeling.

A genetic algorithm-based artificial neural network model has been developed for the accurate prediction of the blood-brain barrier partitioning (in logBB scale) of chemicals. A data set of 123 logBB (115 old molecules and 8 new molecules) of a diverse set of chemicals was chosen in this study. The optimum 3D geometry of the molecules was estimated by the ab initio calculations at the level of RHF/STO-3G, and consequently, different electronic descriptors were calculated for each molecule. Indeed, logP as a measure of hydrophobicity and different topological indices were also calculated. A three-layered artificial neural network with backpropagation of an error-learning algorithm was employed to process the nonlinear relationship between the calculated descriptors and logBB data. Genetic algorithm was used as a feature selection method to select the most relevant set of descriptors as the input of the network. Modeling of the logBB data by the only quantum descriptors produced a 5:4:1 ANN structure with RMS error of validation and crossvalidation equal to 0.224 and 0.227, respectively. Better nonlinear model (RMS(V) and RMS(CV) equals to 0.097 and 0.099, respectively) was obtained by the incorporation of the logP and the principal components of the topological indices to electronic descriptors. The ultimate performances of the models were obtained by the application of the models to predict the logBB of 23 molecules that did not have contribution in the steps of model development. The best model produced RMS error of prediction 0.140, and could predict about 98% of variances in the logBB data.

Blood-Brain Barrier↗

Structure-activity studies of barbiturates using pattern recognition techniques.

The relationship between molecular structure and duration of depressant effect for barbiturates was investigated. A data set of 160 5,5'-disubstituted barbiturates with various acyclic substituents was coded using 47 numerical descriptors including fragments, substructures, environmental descriptors, and molecular connectivity indexes. All descriptors were derived directly from the connection tables of the barbiturates. Using an interactive error-correction feedback algorithm, linear discriminant functions were developed that could dichotomize the data set with respect to several thresholds separating longer from shorter acting compounds. Feature selection was used to focus on the relatively few structural descriptors sufficient to support linear separability. For three specific thresholds, nine, 11, and nine descriptors were sufficient. The importance of these descriptors and the utility of the technique are discussed. Predictive abilities of approximately 94% were obtained for known barbiturates of the same general molecular types.

Animals↗

A comparison of the chemical analyses of cell lipids with their complete proton NMR spectrum.

Whole cells are made up of molecules in different environments to which NMR spectroscopy is sensitive. In particular, malignant and transformed cells contain lipids not only in bilayers but in isotropically tumbling domains which give rise to high-resolution spectra. We have recently developed a technique for simultaneously analyzing broadline and high-resolution signals (M. Bloom, K. T. Holmes, C. E. Mountford, and P. G. Williams, J. Magn. Reson., in press) and we report here its application to a range of rat, mouse, and human cell lines. Some selected features of the NMR spectra were compared with the chemical analysis of the whole-cell lipid. We found that in general the proportion of protons in the narrow methylene resonance at 1.3 ppm increased with the neutral lipid content of the cells. This peak was chosen because its T2 relaxation behavior correlates with metastatic potential in a rat model system. This new technique could be applied to other high-resolution components both in healthy and in diseased states.

Animals↗

Eigenimage filtering in MR imaging: an application in the abnormal chest wall.

A postprocessing linear filter was applied to spin-echo images on 10 patients with known or suspected chest wall invasion due to bronchogenic carcinoma. This technique known as eigenimage filtering allows selective feature extraction of suspected abnormalities from conventional MR images. The final result is an image with marked increased contrast range through enhancement of a desired process (tumor) with suppression of an interfering process (e.g., normal surrounding tissue). This preliminary work demonstrates the ease with which the technique may be implemented, the contrast enhancement obtained between the desired and the interfering feature in the final eigenimage, and its ability to correct for partial volume averaging effects. Also demonstrated are artifacts that can interfere with the interpretation of the eigenimage and a method for minimizing these artifacts in the final eigenimage.

Carcinoma, Bronchogenic↗

Examination of 2-DE in the Human Proteome Organisation Brain Proteome Project pilot studies with the new RAIN gel matching technique.

The Human Proteome Organisation (HUPO) Brain Proteome Project (BPP) pilot studies have generated over 200 2-D gels from eight participating laboratories. This data includes 67 single-channel and 60 DIGE gels comparing 30 whole frozen C57/BL6 female mouse brains, ten each at embryonic day 16, postnatal day 7 (juvenile) and postnatal day 54-56 (adult); and ten single-channel and three DIGE gels comparing human epilepsy surgery of the temporal front lobe with a corresponding post-mortem specimen. The samples were generated centrally and distributed to the participating laboratories, but otherwise no restrictions were placed on sample preparation, running and staining protocols, nor on the 2-D gel analysis packages used. Spots were characterised by MS and the annotated gel images published on a ProteinScape web server. In order to examine the resultant differential expression and protein identifications, we have reprocessed a large subset of the gels using the newly developed RAIN (Robust Automated Image Normalisation) 2-D gel matching algorithm. Traditional approaches use symbolic representation of spots at the very early stages of the analysis, which introduces persistent errors due to inaccuracies in spot modelling and matching. With RAIN, image intensity distributions, rather than selected features, are used, where smooth geometric deformation and expression bias are modelled using multi-resolution image registration and bias-field correction. The method includes a new approach of volume-invariant warping which ensures the volume of protein expression under transformation is preserved. An image-based statistical expression analysis phase is then proposed, where small insignificant expression changes over one gel pair can be revealed when reinforced by the same consistent changes in others. Results of the proposed method as applied to the HUPO BPP data show significant intra-laboratory improvements in matching accuracy over a previous state-of-the-art technique, Multi-resolution Image Registration (MIR), and the commercial Progenesis PG240 package.

Algorithms↗

Normal mode analysis of macromolecular motions in a database framework: developing mode concentration as a useful classifying statistic.

We investigated protein motions using normal modes within a database framework, determining on a large sample the degree to which normal modes anticipate the direction of the observed motion and were useful for motions classification. As a starting point for our analysis, we identified a large number of examples of protein flexibility from a comprehensive set of structural alignments of the proteins in the PDB. Each example consisted of a pair of proteins that were considerably different in structure given their sequence similarity. On each pair, we performed geometric comparisons and adiabatic-mapping interpolations in a high-throughput pipeline, arriving at a final list of 3,814 putative motions and standardized statistics for each. We then computed the normal modes of each motion in this list, determining the linear combination of modes that best approximated the direction of the observed motion. We integrated our new motions and normal mode calculations in the Macromolecular Motions Database, through a new ranking interface at http://molmovdb.org. Based on the normal mode calculations and the interpolations, we identified a new statistic, mode concentration, related to the mathematical concept of information content, which describes the degree to which the direction of the observed motion can be summarized by a few modes. Using this statistic, we were able to determine the fraction of the 3,814 motions where one could anticipate the direction of the actual motion from only a few modes. We also investigated mode concentration in comparison to related statistics on combinations of normal modes and correlated it with quantities characterizing protein flexibility (e.g., maximum backbone displacement or number of mobile atoms). Finally, we evaluated the ability of mode concentration to automatically classify motions into a variety of simple categories (e.g., whether or not they are "fragment-like"), in comparison to motion statistics. This involved the application of decision trees and feature selection (particular machine-learning techniques) to training and testing sets derived from merging the "list" of motions with manually classified ones.

Databases, Protein↗

Protein classification based on text document classification techniques.

The need for accurate, automated protein classification methods continues to increase as advances in biotechnology uncover new proteins. G-protein coupled receptors (GPCRs) are a particularly difficult superfamily of proteins to classify due to extreme diversity among its members. Previous comparisons of BLAST, k-nearest neighbor (k-NN), hidden markov model (HMM) and support vector machine (SVM) using alignment-based features have suggested that classifiers at the complexity of SVM are needed to attain high accuracy. Here, analogous to document classification, we applied Decision Tree and Naive Bayes classifiers with chi-square feature selection on counts of n-grams (i.e. short peptide sequences of length n) to this classification task. Using the GPCR dataset and evaluation protocol from the previous study, the Naive Bayes classifier attained an accuracy of 93.0 and 92.4% in level I and level II subfamily classification respectively, while SVM has a reported accuracy of 88.4 and 86.3%. This is a 39.7 and 44.5% reduction in residual error for level I and level II subfamily classification, respectively. The Decision Tree, while inferior to SVM, outperforms HMM in both level I and level II subfamily classification. For those GPCR families whose profiles are stored in the Protein FAMilies database of alignments and HMMs (PFAM), our method performs comparably to a search against those profiles. Finally, our method can be generalized to other protein families by applying it to the superfamily of nuclear receptors with 94.5, 97.8 and 93.6% accuracy in family, level I and level II subfamily classification respectively.

Algorithms↗

Hierarchical Multi-Label Classification With Gene-Environment Interactions in Disease Modeling.

In biomedical studies, gene-environment (G-E) interactions have been demonstrated to have important implications for analyzing disease outcomes beyond the main G and main E effects. Many approaches have been developed for G-E interaction analysis, yielding important findings. However, hierarchical multi-label classification, which provides insightful information on disease outcomes, remains unexplored in G-E analysis literature. Moreover, unlabeled data are commonly observed in practical settings but omitted by many existing methods of hierarchical multi-label classification. In this study, we consider a semi-supervised scenario and develop a novel approach for the two-layer hierarchical response with G-E interactions. A two-step penalized estimation is then proposed using an efficient expectation-maximization (EM) algorithm. Simulation shows that it has superior performance in classification and feature selection. The analysis of The Cancer Genome Atlas (TCGA) data on lung cancer demonstrates the practical utility of the proposed method. Overall, this study can fill the important knowledge gap in G-E interaction analysis by providing a widely applicable framework for hierarchical multi-label classification of complex disease outcomes.

Humans↗

The relation of ERP components to complex memory processing.

The relation between various ERP components generated during encoding of a word and its subsequent recall were investigated using a "rote" serial-order and an "elaborative" category memory task. Words (flashed separately) were time-locked to EEG recordings from 21 cortical sites. ERP components from the five subjects having the highest recall scores were compared to the five lowest scoring subjects. Results based on the P200 peak amplitude data as well as the N400 and late positive component peak amplitude and latency data suggest that anterior and posterior distributional differences are elicited during encoding of words for rote and elaborative memory tasks. Furthermore, strong individual differences in these patterns were found as a function of task. A tentative argument was made that the obtained anterior and posterior differences may index different word feature selection and encoding processes, which are differentially utilized by high and low recallers.

Adult↗

Increased apolipoprotein E mRNA in the hippocampus in Alzheimer disease and in rats after entorhinal cortex lesioning.

The distribution of apolipoprotein E (ApoE) mRNA was characterized in the hippocampus of humans with Alzheimer disease (AD) and in rats with experimental lesions (unilateral ablation of the entorhinal cortex) that model selected features of AD. In both AD and the lesion model, we observed a shift in the location of astrocytes containing prevalent ApoE mRNA from the neuropil to regions with densely packed neurons. The increased abundance of ApoE mRNA in astrocytes close to neuron cell bodies could be indicative of lipid uptake in regions where neurons are degenerating or where synaptic remodeling is taking place.

Alzheimer Disease↗

Modular bacterial artificial chromosome vectors for transfer of large inserts into mammalian cells.

To facilitate the use of large-insert bacterial clones for functional analysis, we have constructed new bacterial artificial chromosome vectors, pPAC4 and pBACe4. These vectors contain two genetic elements that enable stable maintenance of the clones in mammalian cells: (1) The Epstein-Barr virus replicon, oriP, is included to ensure stable episomal propagation of the large insert clones upon transfection into mammalian cells. (2) The blasticidin deaminase gene is placed in a eukaryotic expression cassette to enable selection for the desired mammalian clones by using the nucleoside antibiotic blasticidin. Sequences important to select for loxP-specific genome targeting in mammalian chromosomes are also present. In addition, we demonstrate that the attTn7 sequence present on the vectors permits specific addition of selected features to the library clones. Unique sites have also been included in the vector to enable linearization of the large-insert clones, e. g., for optical mapping studies. The pPAC4 vector has been used to generate libraries from the human, mouse, and rat genomes. We believe that clones from these libraries would serve as an important reagent in functional experiments, including the identification or validation of candidate disease genes, by transferring a particular clone containing the relevant wildtype gene into mutant cells or transgenic or knock-out animals.

Animals↗

Vasopressin gene related products are markers of human breast cancer.

Immunohistochemical analysis for products of vasopressin and oxytocin gene expression was performed on acetone-fixed tissues from 19 breast cancers representing a variety of tumor sub-types. Studies employed the avidin-biotin complex (ABC) immunohistochemical procedure and utilized rabbit polyclonal antibodies to arginine vasopressin (VP), provasopressin (ProVP), vasopressin-associated human glycopeptide (VAG), oxytocin (OT), oxytocin-associated human neurophysin (OT-HNP), and a mouse monoclonal antibody to vasopressin-associated human neurophysin (VP-HNP). Western Blot analysis was performed on protein extracts of fresh-frozen tissues from 12 additional breast tumors. While VP gene related proteins were not detected in normal breast tissue, immunohistochemistry revealed the presence of VP, ProVP, and VAG in all neoplastic cells for all of the tumor tissues examined. Vasopressin-associated human neurophysin was evident in only one of 19 acetone-fixed tumor preparations. However, Western blot analysis for all 12 fresh-frozen tumor samples showed the presence of two proteins, 42,000 and 20,000 daltons, that were immunoreactive with antibodies to VP, VP-HNP, and VAG. Oxytocin and OT-HNP, by immunohistochemistry, were found to be common to cells of normal breast tissues. For tumors, positive staining for OT was observed in 8 of 18 tumors, while OT-HNP was not detected in any of the tumors examined. These findings indicate that VP gene expression is a selective feature of all breast cancers, and that products of this expression might therefore be useful as markers for early detection of this disease and as possible targets for immunotherapy.

Antibodies, Monoclonal↗

Prediction of the axillary lymph node status in mammary cancer on the basis of clinicopathological data and flow cytometry.

Axillary lymph node status is a major prognostic factor in mammary carcinoma. It is clinically desirable to predict the axillary lymph node status from data from the mammary cancer specimen. In the study, the axillary lymph node status, routine histological parameters and flow-cytometric data were retrospectively obtained from 1139 specimens of invasive mammary cancer. The ten variables: age, tumour type, tumour grade, tumour size, skin infiltration, lymphangiosis carcinomatosa, pT4 category, percentage of tumour cells in G2/M- and S-phases of the cell cycle, and ploidy index were considered as predictor variables, and the single variable lymph node metastasis pN (0 for pN0, or 1 for pN1 or pN2) was used as an output variable. A stepwise logistic regression analysis, with the axillary lymph node as a dependent variable, was used for feature selection. Only lymphangiosis carcinomatosa and tumour size proved to be significant as independent predictor variables; the other variables were non-contributory. Three paradigms with supervised learning rules (multilayer perceptron, learning vector quantisation and support vector machines) were used for the purpose of prediction. If any of these paradigms was used with the information from all ten input variables, 73% of cases could be correctly predicted, with specificity ranging from 82 to 84% and sensitivity ranging from 60 to 63%. If only the two significant input variables were used, lymphangiosis carcinomatosa and tumour diameter, the prediction accuracy was no worse. Nearly identical results were obtained by two different techniques of cross-validation (leave-one-out against ten-fold cross validation). It was concluded that: artificial neural networks can be used for risk stratification on the basis of routine data in individual cases of mammary cancer; and lymphangiosis carcinomatosa and tumour size are independent predictors of axillary lymph node metastasis in mammary cancer.

Algorithms↗

A study of the histological criteria for ulcerative colitis: retrospective evaluation of multiple colonic biopsies.

It is clinically important to distinguish idiopathic inflammatory bowel disease (IBD) from other colitides, and ulcerative colitis (UC) from Crohn's disease (CD); however only a few histological criteria based on colonic biopsies have been established. We investigated 209 consecutive series of biopsies taken from 38 patients with UC, 12 with CD, and 105 with other colitides, to evaluate whether combinations of histological features, selected on the basis of our experience, and listed below, could be useful criteria for the differential diagnosis of IBD, and, more specifically, of UC: (A) chronic inflammation with a predominant increase of plasma cells, (B) crypt distortion, (C) crypt atrophy, (D) diffuse chronic inflammation within a biopsy and between biopsies, and (E) diffuse mucin depletion within a biopsy and between biopsies. Findings that fulfilled all or two of A-C distinguished IBD from the other colitides with high sensitivity (94.3%) and specificity (95.8%). When the findings fulfilled the additional criteria of D and/or E, UC was differentiated from CD or the other colitides with high sensitivity (86.4%) and specificity (99.3%).

Adolescent↗