Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

New approach to generalized two-dimensional correlation spectroscopy. III: Eigenvalue manipulation transformation (EMT) for spectral selectivity enhancement.

This paper demonstrates the potential of eigenvalue manipulating transformation (EMT) of a data matrix for spectral selectivity enhancement, especially useful in 2D correlation analysis. The EMT operation aims at the accentuation of select features of the information content of the original data matrix. For example, by uniformly lowering the power of a set of eigenvalues associated with the original data, the smaller eigenvalues become more prominent and the contributions of secondary loadings become amplified. As a direct consequence of the minor factor accentuation by such EMT operations, 2D correlation spectra gain much stronger discriminating power. The selectivity enhancement effect of such manipulation of eigenvalues is much more noticeable on the synchronous 2D correlation spectrum. This improvement for the spectral selectivity of synchronous 2D correlation spectra is potentially very important, as we usually put more emphasis on the interpretation of asynchronous 2D spectra in 2D correlation analysis due to overlaps of synchronous peaks. Such EMT operations tend to exaggerate the information content of minor PCs and reduce that of major PCs. Thus, much more subtle difference of spectral behavior for each component is now highlighted. Surprisingly, asynchronous 2D correlation spectra are found to be much less sensitive to such EMT operations. The result indicates that the distinction of different band responses has already been accomplished effectively by the original asynchronous 2D correlation analysis.

Algorithms↗

Prediction of the effect of mobile-phase salt type on protein retention and selectivity in anion exchange systems.

This study examines the effect of different salt types on protein retention and selectivity in anion exchange systems. Particularly, linear retention data for various proteins were obtained on two structurally different anion exchange stationary-phase materials in the presence of three salts with different counterions. The data indicated that the effects are, for the most part, nonspecific, although various specific effects could also be observed. Quantitative structure retention relationship (QSRR) models based on support vector machine feature selection and regression models were developed using the experimental chromatographic data in conjunction with various molecular descriptors computed from protein crystal structure geometries. Star plots for each descriptor used in the final model were generated to aid in interpretation. The resulting QSRR models were predictive, with cross-validated r2 values of 0.9445, 0.9676, and 0.8897 for Source 15Q and 0.9561, 0.9876, and 0.9760 for Q Sepharose resins in the presence of three different salts. The predictive power of these models was validated using a set of test proteins that were not used in the generation of these models. Interpretation of the models revealed that particular trends for proteins and salts could be captured using QSRR techniques.

Algorithms↗

Support vector machines with selective kernel scaling for protein classification and identification of key amino acid positions.

MOTIVATION: Data that characterize primary and tertiary structures of proteins are now accumulating at a rapid and accelerating rate and require automated computational tools to extract critical information relating amino acid changes with the spectrum of functionally attributes exhibited by a protein. We propose that immunoglobulin-type beta-domains, which are found in approximate 400 functionally distinct forms in humans alone, provide the immense genetic variation within limited conformational changes that might facilitate the development of new computational tools. As an initial step, we describe here an approach based on Support Vector Machine (SVM) technology to identify amino acid variations that contribute to the functional attribute of pathological self-assembly by some human antibody light chains produced during plasma cell diseases. RESULTS: We demonstrate that SVMs with selective kernel scaling are an effective tool in discriminating between benign and pathologic human immunoglobulin light chains. Initial results compare favorably against manual classification performed by experts and indicate the capability of SVMs to capture the underlying structure of the data. The data set consists of 70 proteins of human antibody kappa1 light chains, each represented by aligned sequences of 120 amino acids. We perform feature selection based on a first-order adaptive scaling algorithm, which confirms the importance of changes in certain amino acid positions and identifies other positions that are key in the characterization of protein function.

Algorithms↗

GECKO: a complete large-scale gene expression analysis platform.

BACKGROUND: Gecko (Gene Expression: Computation and Knowledge Organization) is a complete, high-capacity centralized gene expression analysis system, developed in response to the needs of a distributed user community. RESULTS: Based on a client-server architecture, with a centralized repository of typically many tens of thousands of Affymetrix scans, Gecko includes automatic processing pipelines for uploading data from remote sites, a data base, a computational engine implementing approximately 50 different analysis tools, and a client application. Among available analysis tools are clustering methods, principal component analysis, supervised classification including feature selection and cross-validation, multi-factorial ANOVA, statistical contrast calculations, and various post-processing tools for extracting data at given error rates or significance levels. On account of its open architecture, Gecko also allows for the integration of new algorithms. The Gecko framework is very general: non-Affymetrix and non-gene expression data can be analyzed as well. A unique feature of the Gecko architecture is the concept of the Analysis Tree (actually, a directed acyclic graph), in which all successive results in ongoing analyses are saved. This approach has proven invaluable in allowing a large (approximately 100 users) and distributed community to share results, and to repeatedly return over a span of years to older and potentially very complex analyses of gene expression data. CONCLUSIONS: The Gecko system is being made publicly available as free software http://sourceforge.net/projects/geckoe. In totality or in parts, the Gecko framework should prove useful to users and system developers with a broad range of analysis needs.

Carcinoma↗

Multiple chemical sensitivity: discriminant validity of case definitions.

In this study, the authors used the University of Toronto's Health Survey self-administered questionnaire to determine discriminant validity of multiple chemical sensitivity definitions. The authors distributed a total of 4,126 questionnaires to adults who attended general, allergy, occupational, and environmental health practices. The authors then matched responses to features selected from existing case definitions posited by Thomson et al.; the National Research Council; Cullen; Ashford and Miller; Randolph; Nethercott et al.; and the 1999 Consensus (references 4-7, 2, 9, and 10, respectively, herein). The overall response rate was 61.7%. The prevalence of reported symptoms was lowest in general practices, was intermediate in occupational health and allergy practices, and was highest in environmental health practices. Features from the definitions presented by Nethercott et al. and the 1999 Consensus (references 9 and 10, respectively, herein) correctly identified more than 80% of environmental health practice patients and more than 70% of general practice patients. Combinations of 4 symptoms (i.e., having a stronger sense of smell than others, feeling dull/groggy, feeling "spacey," and having difficulty concentrating) also discriminated successfully. In summary, features from 2 of 7 case definitions assessed by the University of Toronto Health Survey achieved good discrimination and identified patients with an increased likelihood of multiple chemical sensitivity.

Adolescent↗

FREP: a database of functional repeats in mouse cDNAs.

The FREP database (http://facts.gsc.riken.go.jp/FREP/) contains 31 396 RepeatMasker-identified non-redundant variant repeat sequences derived from 16,527 mouse cDNAs with protein-coding potential. The repeats were computationally associated with potential effects on transcriptional variation, translation, protein function or involvement in disease to identify Functional REPeats (FREPs). FREPs are defined by the (i) occurrence of exon-exon boundaries in repeats, (ii) presence of polyadenylation sites in 3'UTR-located repeats, (iii) effect on translation, (iv) position in the protein- coding region or protein domains or (v) conditional association with disease MeSH terms. Currently the database contains 9261 (29.5%) inferred FREPs derived from 6861 (41.5%) mouse cDNAs. Integrated evidence of the functional assignments and dynamically generated sequence similarity search results support the exploration and annotation of functional, ancestral or taxon-specific repeats. Keyword and pre-selected feature searches (e.g. coding sequence-repeat or splice site-repeat relations) support intuitive database querying as well as the retrieval of repeat sequences. Integrated sequence search and alignment tools allow the analysis of known or identification of new functional repeat candidates. FREP is a unique resource for illuminating the role of transposons and repetitive sequences in shaping the coding part of the mouse transcriptome and for selecting the appropriate experimental model to study diseases with suspected repeat etiology contributions.

Animals↗

Selective adaptation to linear frequency-modulated sweeps: evidence for direction-specific FM channels?

Psychometric functions were obtained for detection of linear frequency-modulated pure tones which were preceded by either a pure tone or a linear FM pure-tone adaptor. The results of Gardner and Wilson [J. Acoust. Soc. Am. 66, 704-709(1979)] were generally confirmed: Thresholds were larger by about a factor of 1.7 when the adaptor and test sweeps rose in frequency. This increase in threshold corresponds to a change in performance from 75% to 65% correct. As an alternative to feature-selective channels, we propose that this small effect is due to nonsensory factors, specifically, the use of an adaptor-like reference in the "adapted" condition. Performance similar to that obtained in humans is shown by an ideal receiver that uses an inappropriate reference to match the signal in the detection task.

Acoustic Stimulation↗

Predicting the orientation of invisible stimuli from activity in human primary visual cortex.

Humans can experience aftereffects from oriented stimuli that are not consciously perceived, suggesting that such stimuli receive cortical processing. Determining the physiological substrate of such effects has proven elusive owing to the low spatial resolution of conventional human neuroimaging techniques compared to the size of orientation columns in visual cortex. Here we show that even at conventional resolutions it is possible to use fMRI to obtain a direct measure of orientation-selective processing in V1. We found that many parts of V1 show subtle but reproducible biases to oriented stimuli, and that we could accumulate this information across the whole of V1 using multivariate pattern recognition. Using this information, we could then successfully predict which one of two oriented stimuli a participant was viewing, even when masking rendered that stimulus invisible. Our findings show that conventional fMRI can be used to reveal feature-selective processing in human cortex, even for invisible stimuli.

Adult↗

What makes an Escherichia coli promoter sigma(S) dependent? Role of the -13/-14 nucleotide promoter positions and region 2.5 of sigma(S).

The sigmaS and sigma70 subunits of Escherichia coli RNA polymerase recognize very similar promoter sequences. Therefore, many promoters can be activated by both holoenzymes in vitro. The same promoters, however, often exhibit distinct sigma factor selectivity in vivo. It has been shown that high salt conditions, reduced negative supercoiling and the formation of complex nucleoprotein structures in a promoter region can contribute to or even generate sigmaS selectivity. Here, we characterize the first positively acting sigmaS-selective feature in the promoter sequence itself. Using the sigmaS-dependent csiD promoter as a model system, we demonstrate that C and T at the -13 and -14 positions, respectively, result in strongest expression. We provide allele-specific suppression data indicating that these nucleotides are contacted by K173 in region 2.5 of sigmaS. In contrast, sigma70, which features a glutamate at the corresponding position (E458), as well as the sigmaS(K173E) variant, exhibit a preference for a G(-13). C(-13) is highly conserved in sigmaS-dependent promoters, and additional data with the osmY promoter demonstrate that the K173/C(-13) interaction is of general importance. In conclusion, our data demonstrate an important role for region 2.5 in sigmaS in transcription initiation. Moreover, we propose a consensus sequence for a sigmaS-selective promoter and discuss its emergence and functional properties from an evolutionary point of view.

Amino Acid Sequence↗

FeatureMap3D--a tool to map protein features and sequence conservation onto homologous structures in the PDB.

FeatureMap3D is a web-based tool that maps protein features onto 3D structures. The user provides sequences annotated with any feature of interest, such as post-translational modifications, protease cleavage sites or exonic structure and FeatureMap3D will then search the Protein Data Bank (PDB) for structures of homologous proteins. The results are displayed both as an annotated sequence alignment, where the user-provided annotations as well as the sequence conservation between the query and the target sequence are displayed, and also as a publication-quality image of the 3D protein structure with the selected features and sequence conservation enhanced. The results are also returned in a readily parsable text format as well as a PyMol (http://pymol.sourceforge.net/) script file, which allows the user to easily modify the protein structure image to suit a specific purpose. FeatureMap3D can also be used without sequence annotation, to evaluate the quality of the alignment of the input sequences to the most homologous structures in the PDB, through the sequence conservation colored 3D structure visualization tool. FeatureMap3D is available at: http://www.cbs.dtu.dk/services/FeatureMap3D/.

Amino Acid Sequence↗

Using in vitro prediction models instead of the rabbit eye irritation test to classify and label new chemicals: a post hoc data analysis of the international EC/HO validation study.

The international validation study on alternative methods to replace the Draize rabbit eye irritation test, funded by the European Commission (EC) and the British Home Office (HO), took place during 1992-1994, and the results were published in 1995. The results of this EC/HO study are analysed by employing discriminant analysis, taking into account the classification of the in vivo data into eye irritation classes A (risk of serious damage to eyes), B (irritating to eyes) and NI (non-irritant). A data set for 59 test items was analysed, together with three subsets: surfactants, water-soluble chemicals, and water-insoluble chemicals. The new statistical methods of feature selection and estimation of the discriminant functions classification error were used. Normal distributed random numbers were added to the mean values of each in vitro endpoint, depending on the observed standard deviations. Thereafter, the reclassification error of the random observations was estimated by applying the fixed function of the mean values. Moreover, the leaving-one-out cross-classification method was applied to this random data set. Subsequently, random data were generated r times (for example, r = 1000) for a feature combination. Eighteen features were investigated in nine in vitro test systems to predict the effects of a chemical in the rabbit eye. 72.5% of the chemicals in the undivided sample were correctly classified when applying the in vitro endpoints lgNRU of the neutral red uptake test and lgBCOPo5 of the bovine opacity and permeability test. The accuracy increased to 80.9% when six in vitro features were used, and the sample was subdivided. The subset of surfactants was correctly classified in more than 90% of cases, which is an excellent performance.

Animal Testing Alternatives↗

Prediction of genomewide conserved epitope profiles of HIV-1: classifier choice and peptide representation.

Identification of peptides binding to Major Histocompatibility Complex (MHC) molecules is important for accelerating vaccine development and improving immunotherapy. Accordingly, a wide variety of prediction methods have been applied in this context. In this paper, we introduce (tree-based) ensemble classifiers for such problems and contrast their predictive performance with forefront existing methods for both MHC class I and class II molecules. In addition, we investigate the impact of differing peptide representation schemes on performance. Finally, classifier predictions are used to conduct genomewide scans of a diverse collection of HIV-1 strains, enabling assessment of epitope conservation. We investigated all combinations of six classification methods (classification trees, artificial neural networks, support vector machines, as well as the more recently devised ensemble methods (bagging, random forests, boosting) with four peptide representation schemes (amino acid sequence, select biophysical properties, select quantitative structure-activity relationship (QSAR) descriptors, and the combination of the latter two) in predicting peptide binding to an MHC class I molecule (HLA-A2) and MHC class II molecule (HLA-DR4). Our results show that the ensemble methods are consistently more accurate than the other three alternatives. Furthermore, they are robust with respect to parameter tuning. Among the four representation schemes, the amino acid sequence representation gave consistently (across classifiers) best results. This finding obviates the need for feature selection strategies incurred by use of biophysical and/or QSAR properties. We obtained, and aligned, a diverse set of 32 HIV-1 genomes and pursued genomewide HLA-DR4 epitope profiling by querying with respect to classifier predictions, as obtained under each of the four peptide representation schemes. We validated those epitopes conserved across strains against known T-cell epitopes. Once again, amino acid sequence representation was at least as effective as using properties. Assessment of novel epitope predictions awaits experimental verification.

Journal Article↗

In silico prediction of pregnane X receptor activators by machine learning approaches.

Pregnane X receptor (PXR) regulates drug metabolism and is involved in drug-drug interactions. Prediction of PXR activators is important for evaluating drug metabolism and toxicity. Computational pharmacophore and quantitative structure-activity relationship models have been developed for predicting PXR activators. Because of the structural diversity of PXR activators, more efforts are needed for exploring methods applicable to a broader spectrum of compounds. We explored three machine learning methods (MLMs) for predicting PXR activators, which were trained and tested by using significantly higher number of compounds, 128 PXR activators (98 human) and 77 PXR non-activators, than those of previous studies. The recursive feature-selection method was used to select molecular descriptors relevant to PXR activator prediction, which are consistent with conclusions from other computational and structural studies. In a 10-fold cross-validation test, our MLM systems correctly predicted 81.2 to 84.0% of PXR activators, 80.8 to 85.0% of hPXR activators, 61.2 to 70.3% of PXR nonactivators, and 67.7 to 73.6% of hPXR nonactivators. Our systems also correctly predicted 73.3 to 86.7% of 15 newly published hPXR activators. MLMs seem to be useful for predicting PXR activators and for providing clues to physicochemical features of PXR activation.

Artificial Intelligence↗

Event-related potentials: a critical review of methods for single-trial detection.

The analysis of ERP data has followed several lines over the last 20 years. The most prevalent method is simply to average ERPs for a given class of stimuli. The ERPs are compared for differences across classes of stimuli. Little other special data processing is used. The ERP comparisons are usually performed using visual examination of the wave-shapes. Sometimes statistics are calculated such as means, variances, and confidence limits. Linear filtering is used to reduce interference. Another approach is to model or analyze the ERP as a sequence of vectors or frames of data samples. These samples may be of the ERP time waveform or they may be of the frequency transform of the ERP waveform. The frames of data vary in length from the entire ERP waveform (500 to 1000 msec) to frames as short as ten sample points (100 msec). Recognition of an event in the ERP is achieved by computing a distance measure between parameter vectors for one class of stimuli and corresponding parameter vectors for another class of stimuli. Recognition is achieved by selecting the ERP with the lowest distance score. This approach is "pattern matching" and relies on two assumptions: adjacent frames of data are uncorrelated, and the variability of the data can be accounted for by the distance measured for all stimuli in the classes presented. Subject variability is generally not accounted for, other than to assume it is the same for all classes of stimuli. The data are clustered into a variety of reference patterns that represent particular manifestations of a particular stimulus. Another approach is "feature-based" recognition. The idea is to identify and automatically extract features of the data that can provide a characterization of stimuli. The features selected may be abstract. They are calculated from the data or transforms of the data.

Biometry↗

Animal models of ALS.

Animal models of amyotrophic lateral sclerosis (ALS) provide a unique opportunity to study this incurable and fatal human disease both clinically and pathologically. This is particularly true for certain pathological and therapeutic studies that are impractical or impossible to perform in human patients. Nonetheless, postmortem ALS tissue remains the "gold standard" against which pathologic findings in animal models must be compared. Four natural disease models have been most extensively studied, including three mouse models: motor neuron degeneration (Mnd), progressive motor neuronopathy (pmn), wobbler, and one canine model: hereditary canine spinal muscular atrophy (HCSMA). The wobbler mouse has been the most extensively studied of these models with analyses of clinical, pathological (perikaryon, axon, muscle), and biochemical features. Experimentally induced ALS animal models have allowed controlled testing of various neurotoxic, viral and immune-mediated mechanisms. Molecular techniques have recently generated mouse models in which genes relevant to the human disease or motor neuron biology have been manipulated. The most clinically relevant of these is a transgenic mouse overexpressing the mutated SOD1 gene of FALS patients, which has already provided significant insights into mechanisms of motor neuron degeneration in this disease. Because no single animal model perfectly reflects all the clinical and pathological characteristics of ALS, study of selected features from the most relevant models will contribute to a better understanding of the pathogenesis and/or etiology of this disease.

Amyotrophic Lateral Sclerosis↗

Quantitative analysis of synovial membrane inflammation: a comparison between automated and conventional microscopic measurements.

OBJECTIVE: The objective of this study was to quantify selected features of chronic synovial tissue inflammation by computerised image analysis and to validate the results by comparison with conventional microscopic measurements. METHODS: Synovial biopsy samples were obtained from the knee joints of patients with chronic arthritis and prepared for immunohistochemical analysis using standard techniques. Following the development of special software, four parameters of chronic synovial inflammation were evaluated: intimal layer thickness, CD3+ cell infiltration, CD8+ cell infiltration and vascularity. Intimal layer thickness was expressed in microns. The intensity of CD3+ and CD8+ cell infiltration was expressed as the percentage area of the tissue section occupied by positively stained cells. Vascularity was expressed as the percentage area occupied by blood vessels. Conventional quantitative microscopic analysis was also undertaken and the results from both methods compared. RESULTS: Seventy eight tissue sections were selected for study. Measurements of intimal layer thickness by both techniques correlated strongly: r = 0.85, p = 0.0006. Measurements of CD8+ cell infiltration, usually widely dispersed, also correlated well: r = 0.64, p = 0.005. Measurements of CD3+ cell infiltration, often densely aggregated, correlated less well: r = 0.55, p = 0.02. Measurements of vascularity demonstrated no statistically significant correlation: r = 0.41, p = 0.07. Proficiency in the use of computerised image analysis was readily acquired. CONCLUSION: Computerised image analysis was successfully applied to the measurement of some features of synovial tissue inflammation. Further software development is required to validate measurement of blood vessels of variable size.

Arthritis↗

Prediction of lymph node metastasis by analysis of gene expression profiles in non-small cell lung cancer.

OBJECTIVE: Non-small cell lung carcinoma (NSCLC) is one of the leading causes of death in the world. Lymph node metastasis is not only an important factor in estimating the extent and the metastatic potential of an NSCLC but also in prognosticating the patient outcome. Preoperative prediction of lymph node metastasis might greatly facilitate the choice of appropriate surgical and medical options in patients with NSCLC. METHODS AND RESULTS: Using a cDNA array, we analyzed the expression profiles of 1,289 genes in 92 cancer tissues of NSCLC (37 squamous cell carcinomas and 55 adenocarcinomas). We divided the patients into two groups (classes) for each of various pathological factors, such as lymph node metastasis and pT-stage. For each pair of classes, we searched for an optimal combination of genes to classify the cases using a sequential forward selection algorithm starting from a gene set that showed significant difference in expression between the classes. We used the leave-one-out error cross-validation on a k-nearest neighbor classifier to sequentially choose the gene. Using the optimized set of genes, it was possible to stratify the patients for lymph node metastasis (pN-stage) and pT-stage at, respectively, 100% (23 genes) and 100% (55 genes) for cases with squamous cell carcinomas and 94% (43 genes) and 92% (35 genes) for those with adenocarcinomas. CONCLUSION: We conclude that expression profiling using feature selection provides a powerful means of stratification (personalization) of NSCLC patients and choice in treatment options, particularly for factors such as lymph node metastasis whose radiological diagnosis is presently incomplete.

Adenocarcinoma↗

Automatic detection of calcifications in the aorta from CT scans of the abdomen. 3D computer-aided diagnosis.

RATIONALE AND OBJECTIVES: Automated detection and quantification of arterial calcifications can facilitate epidemiologic research and, eventually, the use of full-body calcium scoring in clinical practice. An automatic computerized method to detect calcifications in CT scans is presented. MATERIALS AND METHODS: Forty abdominal CT scans have been randomly selected from clinical practice. They all contained contrast material and belonged to one of four categories: containing "no," "small," "moderate," or "large" amounts of arterial calcification. There were ten scans in each category. The experiments were restricted to the vertical range from the point where the superior mesenteric artery branches off of the descending aorta until the first bifurcation of the iliac arteries. The automatic method starts by extracting all connected objects above 220 Hounsfield units (HU) from the scan. These objects include all calcifications, as well as bony structures and contrast material. To distinguish calcifications from non-calcifications, a number of features are calculated for each object. These features are based on the object's size, location, shape characteristics, and surrounding structures. Subsequently a classification of each object is performed in two stages. First the probability that an object represents a calcification is computed assuming a multivariate Gaussian distribution for the calcifications. Objects with low probability are discarded. The remaining objects are then classified into calcifications and non-calcifications using a 5-nearest-neighbor classifier and sequential forward feature selection. Based on the total volume of calcifications determined by the system, the scan is assigned to one of the four categories mentioned above. RESULTS: The 40 scans contained a total of 249 calcifications as determined by a human observer. The method detected 209 calcifications (sensitivity 83.9%) at the expense of on average 1.0 false-positive object per scan. The correct category label was assigned to 30 scans and only 2 scans were off by more than one category. Most incorrect classifications can be attributed to the presence of contrast material in the scans. CONCLUSION: It is possible to identify the majority of arterial calcifications in abdominal CT scans in a completely automatic fashion with few false positive objects, even if the scans contain contrast material.

Aorta, Abdominal↗