Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

A pattern classification approach to characterizing solitary pulmonary nodules imaged on high resolution CT: preliminary results.

The purpose of this research is to characterize solitary pulmonary nodules as benign or malignant based on quantitative measures extracted from high resolution CT (HRCT) images. High resolution CT images of 31 patients with solitary pulmonary nodules and definitive diagnoses were obtained. The diagnoses of these 31 cases (14 benign and 17 malignant) were determined from either radiologic follow-up or pathological specimens. Software tools were developed to perform the classification task. On the HRCT images, solitary nodules were identified using semiautomated contouring techniques. From the resulting contours, several quantitative measures were extracted related to each nodule's size, shape, attenuation, distribution of attenuation, and texture. A stepwise discriminant analysis was performed to determine which combination of measures were best able to discriminate between the benign and malignant nodules. A linear discriminant analysis was then performed using selected features to evaluate the ability of these features to predict the classification for each nodule. A jackknifed procedure was performed to provide a less biased estimate of the linear discriminator's performance. The preliminary discriminant analysis identified two different texture measures--correlation and difference entropy--as the top features in discriminating between benign and malignant nodules. The linear discriminant analysis using these features correctly classified 28/31 cases (90.3%) of the training set. A less biased estimate, using jackknifed training and testing, yielded the same results (90.3% correct). The preliminary results of this approach are very promising in characterizing solitary nodules using quantitative measures extracted from HRCT images. Future work involves including contrast enhancement and three-dimensional measures extracted from volumetric CT scans, as well as the use of several pattern classifiers.

Biophysical Phenomena↗

Use of cephalexin-aztreonam-arabinose agar for selective isolation of Enterococcus faecium.

Cephalexin-aztreonam-arabinose agar (CAA), a new selective agar, was examined in comparison with nalidixic acid-colistin agar for the differentiation of Enterococcus faecium from other enterococci and the ability to isolate the organism from feces. Two hundred sixteen enterococcus isolates and a variety of gram-positive and gram-negative control strains were inoculated onto both media. All control strains of E. faecium were easily differentiated from Enterococcus faecalis and Enterococcus durans on the basis of arabinose fermentation on CAA. Differentiation of E. faecium from other enterococci or Streptococcus bovis was not possible on nalidixic acid-colistin agar. Increased isolation of E. faecium was demonstrated on CAA when both media were compared for the isolation of the organism from feces. CAA has been shown to possess excellent differential and selective features allowing the simple and effective isolation of E. faecium from heavily contaminated sites.

Agar↗

Data mining of inputs: analysing magnitude and functional measures.

The problem of data encoding and feature selection for training back-propagation neural networks is well known. The basic principles are to avoid encrypting the underlying structure of the data, and to avoid using irrelevant inputs. This is not easy in the real world, where we often receive data which has been processed by at least one previous user. The data may contain too many instances of some class, and too few instances of other classes. Real data sets often include many irrelevant or redundant input fields. This paper examines the use of weight matrix analysis techniques and functional measures using two real (and hence noisy) data sets. The first part of this paper examines the use of the weight matrix of the trained neural network itself to determine which inputs are significant. A new technique is introduced and compared with two other techniques from the literature. We present our experience and results on some satellite data augmented by a terrain model. The task was to predict the forest supra-type based on the available information. A brute force technique eliminating randomly selected inputs was used to validate our approach. The second part of this paper examines the use of measures to determine the functional contribution of inputs to outputs. Inputs which include minor but unique information to the network are more significant than inputs with higher magnitude contribution but providing redundant information, which is also provided by another input. A comparison is made to sensitivity analysis, where the sensitivity of outputs to input perturbation is used as a measure of the significance of inputs. This paper presents a novel functional analysis of the weight matrix based on a technique developed for determining the behavioral significance of hidden neurons. This is compared with the application of the same technique to the training and test data. Finally, a novel aggregation technique is introduced.

Algorithms↗

Direct prediction of T-cell epitopes using support vector machines with novel sequence encoding schemes.

New peptide encoding schemes are proposed to use with support vector machines for the direct recognition of T cell epitopes. The methods enable the presentation of information on (1) amino acid positions in peptides, (2) neighboring side chain interactions, and (3) the similarity between amino acids through a BLOSUM matrix. A procedure of feature selection is also introduced to strengthen the prediction. The computational results demonstrate competitive performance over previous techniques.

Amino Acid Sequence↗

Organizing principles for single joint movements. III. Speed-insensitive strategy as a default.

1. Human subjects made discrete elbow flexions in a horizontal plane over different distances, from a stationary initial position to a visually defined stationary target 9 degrees wide. We measured joint angle, acceleration, and electromyograms (EMGs) from two agonist and two antagonist muscles. 2. Subjects made movements over four different distances following one of four different instructions. The first instructed the subject simply to choose a comfortable speed. The other three explicitly emphasized either speed, accuracy, or maintenance of the "same" speed over different distances. These instructions produced a wide range of movement velocities. 3. The initial rises of the acceleration (and therefore of the inertial torque), as well as the initial slope of the agonist EMG, were all invariant over changes in the target distance for any single instruction but were all sensitive to the given instruction. 4. Our results demonstrate that the speed-insensitive strategy is a standard or default pattern for performing movements that may be carried out for different instructions over a wide range of speeds. A uniform intensity of excitation pulse is not a byproduct of moving at maximal speed. Submaximal intensities are associated with submaximal speeds and are a selected feature of the pattern of movement control.

Acceleration↗

Identifying marker genes in transcription profiling data using a mixture of feature relevance experts.

Transcription profiling experiments permit the expression levels of many genes to be measured simultaneously. Given profiling data from two types of samples, genes that most distinguish the samples (marker genes) are good candidates for subsequent in-depth experimental studies and developing decision support systems for diagnosis, prognosis, and monitoring. This work proposes a mixture of feature relevance experts as a method for identifying marker genes and illustrates the idea using published data from samples labeled as acute lymphoblastic and myeloid leukemia (ALL, AML). A feature relevance expert implements an algorithm that calculates how well a gene distinguishes samples, reorders genes according to this relevance measure, and uses a supervised learning method [here, support vector machines (SVMs)] to determine the generalization performances of different nested gene subsets. The mixture of three feature relevance experts examined implement two existing and one novel feature relevance measures. For each expert, a gene subset consisting of the top 50 genes distinguished ALL from AML samples as completely as all 7,070 genes. The 125 genes at the union of the top 50s are plausible markers for a prototype decision support system. Chromosomal aberration and other data support the prediction that the three genes at the intersection of the top 50s, cystatin C, azurocidin, and adipsin, are good targets for investigating the basic biology of ALL/AML. The same data were employed to identify markers that distinguish samples based on their labels of T cell/B cell, peripheral blood/bone marrow, and male/female. Selenoprotein W may discriminate T cells from B cells. Results from analysis of transcription profiling data from tumor/nontumor colon adenocarcinoma samples support the general utility of the aforementioned approach. Theoretical issues such as choosing SVM kernels and their parameters, training and evaluating feature relevance experts, and the impact of potentially mislabeled samples on marker identification (feature selection) are discussed.

Acute Disease↗

Two-dimensional transcriptome profiling: identification of messenger RNA isoform signatures in prostate cancer from archived paraffin-embedded cancer specimens.

The expression of specific mRNA isoforms may uniquely reflect the biological state of a cell because it reflects the integrated outcome of both transcriptional and posttranscriptional regulation. In this study, we constructed a splicing array to examine approximately 1,500 mRNA isoforms from a panel of genes previously implicated in prostate cancer and identified a large number of cell type-specific mRNA isoforms. We also developed a novel "two-dimensional" profiling strategy to simultaneously quantify changes in splicing and transcript abundance; the results revealed extensive covariation between transcription and splicing in prostate cancer cells. Taking advantage of the ability of our technology to analyze RNA from formalin-fixed, paraffin-embedded tissues, we derived a specific set of mRNA isoform biomarkers for prostate cancer using independent panels of tissue samples for feature selection and cross-analysis. A number of cancer-specific splicing switch events were further validated by laser capture microdissection. Quantitative changes in transcription/RNA stability and qualitative differences in splicing ratio may thus be combined to characterize tumorigenic programs and signature mRNA isoforms may serve as unique biomarkers for tumor diagnosis and prognosis.

Aged↗

A systems approach to model metastatic progression.

Proteomic profiling of human disease has seen much early activity with the accessibility of the newest generation of high-throughput platforms and technologies. Nevertheless, the nature of the dynamic physiologic milieu and high dimensionality of the data has complicated major diagnostic and prognostic breakthroughs. Our recent article in Cancer Cell delineates an integrative model for culling a molecular signature of metastatic progression in prostate cancer from proteomic and transcriptomic analyses and shows its facility as a predictor of prognosis. The study leveraged direct proteomic analysis of tumor tissue extracts, differential feature selection characterizing the proteomic alterations of prostate cancer subclasses, and integration with public and study-derived genomic data to construct a multiplex gene signature representing progression of indolent cancer to aggressive disease. This further predicted clinical outcome in a variety of solid tumors. This review describes the context of the work, the framework for the analysis itself, and a look forward to the promise of this systems approach to human disease.

Computational Biology↗

Tumors and malformations of the caudal spinal axis.

The early development of the neural tube has been well studied in animals and humans. After axial determinants have been accomplished the processes of primary and secondary neurulation take place. Successful completion results in a spinal cord that has arisen from primary neurulation and a lower sacro-coccygeal portion from secondary neurulation. The latter region is the site of numerous skin-covered clinical lesions, which include tumors and malformations. A listing of selected features in 764 cases of skin-covered sacrococcygeal lesions is presented. The manner in which these lesions arise and the potential for genetic factors being responsible is discussed.

Diagnosis, Differential↗

Categorization of voice disorders with six perceptual dimensions.

To obtain a perceptual reference for acoustic feature selection, 94 male and 124 female voices were categorized using the ratings of 6 clinicians on visual analog scales for pathology, roughness, breathiness, strain, asthenia, and pitch. Partial correlations showed that breathiness and roughness were the main determinants of pathology. The six-dimensional ratings (the six median scores for each voice) were categorized with the aid of the Sammon map and the self-organizing map. The five categories created differed with respect to the breathiness/roughness ratio and the degree of pathology.

Adult↗

An autopsy study of the incidence of lacunes in relation to age, hypertension, and arteriosclerosis.

We investigated selected features of lacunes in 1,086 necropsy cases. Lacunes were found in brains from patients above the age of 40 years and were most common in brains from persons in their sixties but decreased in number in brains from older persons. The most common site of lacunes was the frontal lobe white matter, followed by the putamen, pons, parietal lobe white matter, thalamus, and caudate nucleus in descending order of frequency. By dividing the 1,086 cases into three groups according to blood pressure, we found more lacunes in the hypertensive and borderline hypertensive groups than in the normotensive group; the average number of lacunes per brain in each group was 3.61, 2.77, and 1.15, respectively. Diastolic hypertension was more closely related to the number of lacunes than was systolic hypertension. The extent of arteriolosclerosis of the medullary arteries in the frontal lobe white matter was measured and compared with the number of lacunes. There was a close correlation between lacunes and arterioloslerosis in all age groups.

Adult↗

RSPOP: rough set-based pseudo outer-product fuzzy rule identification algorithm.

System modeling with neuro-fuzzy systems involves two contradictory requirements: interpretability verses accuracy. The pseudo outer-product (POP) rule identification algorithm used in the family of pseudo outer-product-based fuzzy neural networks (POPFNN) suffered from an exponential increase in the number of identified fuzzy rules and computational complexity arising from high-dimensional data. This decreases the interpretability of the POPFNN in linguistic fuzzy modeling. This article proposes a novel rough set-based pseudo outer-product (RSPOP) algorithm that integrates the sound concept of knowledge reduction from rough set theory with the POP algorithm. The proposed algorithm not only performs feature selection through the reduction of attributes but also extends the reduction to rules without redundant attributes. As many possible reducts exist in a given rule set, an objective measure is developed for POPFNN to correctly identify the reducts that improve the inferred consequence. Experimental results are presented using published data sets and real-world application involving highway traffic flow prediction to evaluate the effectiveness of using the proposed algorithm to identify fuzzy rules in the POPFNN using compositional rule of inference and singleton fuzzifier (POPFNN-CRI(S)) architecture. Results showed that the proposed rough set-based pseudo outer-product algorithm reduces computational complexity, improves the interpretability of neuro-fuzzy systems by identifying significantly fewer fuzzy rules, and improves the accuracy of the POPFNN.

Algorithms↗

Optical coherence tomography machine learning classifiers for glaucoma detection: a preliminary study.

PURPOSE: Machine-learning classifiers are trained computerized systems with the ability to detect the relationship between multiple input parameters and a diagnosis. The present study investigated whether the use of machine-learning classifiers improves optical coherence tomography (OCT) glaucoma detection. METHODS: Forty-seven patients with glaucoma (47 eyes) and 42 healthy subjects (42 eyes) were included in this cross-sectional study. Of the glaucoma patients, 27 had early disease (visual field mean deviation [MD] > or = -6 dB) and 20 had advanced glaucoma (MD < -6 dB). Machine-learning classifiers were trained to discriminate between glaucomatous and healthy eyes using parameters derived from OCT output. The classifiers were trained with all 38 parameters as well as with only 8 parameters that correlated best with the visual field MD. Five classifiers were tested: linear discriminant analysis, support vector machine, recursive partitioning and regression tree, generalized linear model, and generalized additive model. For the last two classifiers, a backward feature selection was used to find the minimal number of parameters that resulted in the best and most simple prediction. The cross-validated receiver operating characteristic (ROC) curve and accuracies were calculated. RESULTS: The largest area under the ROC curve (AROC) for glaucoma detection was achieved with the support vector machine using eight parameters (0.981). The sensitivity at 80% and 95% specificity was 97.9% and 92.5%, respectively. This classifier also performed best when judged by cross-validated accuracy (0.966). The best classification between early glaucoma and advanced glaucoma was obtained with the generalized additive model using only three parameters (AROC = 0.854). CONCLUSIONS: Automated machine classifiers of OCT data might be useful for enhancing the utility of this technology for detecting glaucomatous abnormality.

Adult↗

Alkaline phosphatase: placental and tissue-nonspecific isoenzymes hydrolyze phosphoethanolamine, inorganic pyrophosphate, and pyridoxal 5'-phosphate. Substrate accumulation in carriers of hypophosphatasia corrects during pregnancy.

Hypophosphatasia features selective deficiency of activity of the tissue-nonspecific (liver/bone/kidney) alkaline phosphatase (ALP) isoenzyme (TNSALP); placental and intestinal ALP isoenzyme (PALP and IALP, respectively) activity is not reduced. Three phosphocompounds (phosphoethanolamine [PEA], inorganic pyrophosphate [PPi], and pyridoxal 5'-phosphate [PLP]) accumulate endogenously and appear, therefore, to be natural substrates for TNSALP. Carriers for hypophosphatasia may have decreased serum ALP activity and elevated substrate levels. To test whether human PALP and TNSALP are physiologically active toward the same substrates, we studied PEA, PPi, and PLP levels during and after pregnancy in three women who are carriers for hypophosphatasia. Hypophosphatasemia corrected during the third trimester because of PALP in maternal blood. Blood or urine concentrations of PEA, PPi, and PLP diminished substantially during that time. After childbirth, maternal circulating levels of PALP decreased, and PEA, PPi, and PLP levels abruptly increased. In serum, unremarkable concentrations of IALP and low levels of TNSALP did not change during the study period. We conclude that PALP, like TNSALP, is physiologically active toward PEA, PPi, and PLP in humans. We speculate from molecular/crystallographic information, indicating significant similarity of structure of the substrate-binding site of ALPs throughout nature, that all ALP isoenzymes recognize these same three phosphocompound substrates.

Alkaline Phosphatase↗

Psychiatry residency programs: trends in psychotherapy supervision.

The evolving dominance of psychobiologic over psychodynamic theoretical influences on education and practice presents new challenges for psychiatry. This article features selected data from the 1989 American Association of Directors of Psychiatric Residency Training annual survey (n = 215) that describe current teaching activities related to psychodynamic psychiatry, mainly psychotherapy. Results are based on a 50 percent return rate (107/215 questionnaires). Responses confirm the emergence of psychobiological (48%) over psychodynamic (40%) departmental orientations and report that the psychodynamic orientation has maintained strength as a secondary emphasis. Residents generally gain experience in a range of psychotherapy theories and modalities, including psychodynamic, cognitive, behavioral, individual, couples, family, and group therapies. Training in brief and short-term individual psychodynamic psychotherapy predominates, however. Use of video- and audiotaping in supervision is limited. Full-time faculty provide the bulk of psychotherapy instruction. This is carried out in both individual and group sessions, which are organized primarily around case reviews. Supervision-related problems include faculty availability, skill diversity, competence, theoretical flexibility, and attitudes, as well as program structure and standards.

Female↗

Breast tissue classification using diagnostic ultrasound and pattern recognition techniques: I. Methods of pattern recognition.

This paper discusses the application of statistical pattern recognition techniques to problems in diagnostic ultrasound. Using our own system as an example, we describe the concepts and specific methods that we have applied to a problem involving the computer-aided classification of breast tissue in vivo. Topics include feature generation, feature selection and classification, as well as a method which estimates the probability of error on classifying future data. An accompanying paper applies these methods to the classification of backscattered RF signals from normal and diseased breast tissue.

Breast Diseases↗

Anti-inflammatory pre-treatment and the resultant effects of interleukin-10: adjuncts to multi-therapeutical strategies.

With the advent of off-pump coronary bypass surgery, there is increasing demand for research in attenuating the deleterious effects of cardiopulmonary bypass (CPB). An improved understanding of the systemic inflammatory response syndrome (SIRS) has distinguished which areas of components have the most adverse effects and which are, in fact, anti-inflammatory. This classification of inflammatory components allows strategic treatment for those likely to cause the most clinically significant 'effect', suitably termed 'effectors'. This article will identify current methods in treating 'effectors', as well as those components having anti-inflammatory effects. This article selectively features certain inflammatory components by: (1) grouping them as being 'mediators' or 'effectors'; (2) relating them to interleukin-10 (IL-10) and treatments potentiating anti-inflammatory effects; (3) summarizing their mechanisms of action; (4) recognizing the time periods during bypass exhibiting peak levels; and (5) investigating current treatment. methods and identifying their significance to 'effectors'. A literature search in MEDLINE was performed, featuring articles of the English-language within the past 5 years. Because of the characteristic of having interlinked multi-component cascades, it is evident that treating SIRS with a one-dimensional method would be inadequate. This article not only confirms the importance of a multi-factorial therapeutic approach, but also targets the inflammatory components having the highest potential for causing direct tissue damage, known as 'effectors'. In addition, previous studies have found IL-10 to have 'regulatory effects' during periods of excessive pro-inflammatory stimuli. These findings may arouse new ideas in exploring the area of anti-inflammatory cytokines. In fact, future treatments may suggest a new classification featuring 'mediators', 'effectors', and 'regulators'.

Cardiopulmonary Bypass↗

Serum proteomic pattern analysis for early cancer detection.

The ability of physicians to effectively treat and cure cancer is directly dependent on their ability to detect cancers at their early stages. The early detection of cancer has the potential to dramatically reduce mortality. Recently, the use of mass spectrometry to develop profiles of patient serum proteins has been reported as a promising method to achieve this goal. In this paper, we analyzed the ovarian cancer and prostate cancer data sets using support vector machine (SVM) to detect cancer at the early stages based on serum proteomic pattern. The results showed that SVM, in general, performed well on these two data sets, as measured by sensitivity, specificity, positive predictive value, negative predictive value, and accuracy. Linear kernel worked the best on ovarian cancer data with a sensitivity of 0.99 and an accuracy of 0.97, while polynomial kernel worked the best on prostate cancer data with a sensitivity of 0.79 and an accuracy of 0.82. When redial kernel was applied to either of the two data sets, all the samples were predicted as cancer samples, with a sensitivity of 1 and a specificity of 0. Furthermore, feature selection did not improve SVM performance.

Artificial Intelligence↗