Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Biomarker identification by feature wrappers.

Gene expression studies bridge the gap between DNA information and trait information by dissecting biochemical pathways into intermediate components between genotype and phenotype. These studies open new avenues for identifying complex disease genes and biomarkers for disease diagnosis and for assessing drug efficacy and toxicity. However, the majority of analytical methods applied to gene expression data are not efficient for biomarker identification and disease diagnosis. In this paper, we propose a general framework to incorporate feature (gene) selection into pattern recognition in the process to identify biomarkers. Using this framework, we develop three feature wrappers that search through the space of feature subsets using the classification error as measure of goodness for a particular feature subset being "wrapped around": linear discriminant analysis, logistic regression, and support vector machines. To effectively carry out this computationally intensive search process, we employ sequential forward search and sequential forward floating search algorithms. To evaluate the performance of feature selection for biomarker identification we have applied the proposed methods to three data sets. The preliminary results demonstrate that very high classification accuracy can be attained by identified composite classifiers with several biomarkers.

Algorithms↗

Active shape model segmentation with optimal features.

An active shape model segmentation scheme is presented that is steered by optimal local features, contrary to normalized first order derivative profiles, as in the original formulation [Cootes and Taylor, 1995, 1999, and 2001]. A nonlinear kNN-classifier is used, instead of the linear Mahalanobis distance, to find optimal displacements for landmarks. For each of the landmarks that describe the shape, at each resolution level taken into account during the segmentation optimization procedure, a distinct set of optimal features is determined. The selection of features is automatic, using the training images and sequential feature forward and backward selection. The new approach is tested on synthetic data and in four medical segmentation tasks: segmenting the right and left lung fields in a database of 230 chest radiographs, and segmenting the cerebellum and corpus callosum in a database of 90 slices from MRI brain images. In all cases, the new method produces significantly better results in terms of an overlap error measure (p < 0.001 using a paired T-test) than the original active shape model scheme.

Adolescent↗

QSAR with few compounds and many features.

Fitting quantitative structure-activity relationships (QSAR) requires different statistical methodologies and, to some degree, philosophies depending on the "shape" of the data matrix. When few features are used and there are many compounds, it is a reasonable expectation that good feature subset selection may be made and that nonlinearities and nonadditivities can be detected and diagnosed. Where there are many features and few compounds, this is unrealistic. Methods such as ridge regression RR, PLS, and principal component regression PCR, which abjure feature selection and rely on linearity may provide good predictions and fair understanding. We report a development of ridge regression for the underdetermined case by using generalized cross-validation to choose the ridge constant and perform F-tests for additional information. Conventional regression diagnostics can be used in followup to identify nonlinearities and other departures from model. We illustrate the approach with QSAR models of four data sets using calculated molecular descriptors.

Algorithms↗

Integrated approach using protein and ligand information to analyze selectivity- and affinity-determining features of carbonic anhydrase isozymes.

The application and comparison of selected protein- and ligand-based approaches to elucidate factors important for affinity and selectivity towards the carbonic anhydrase isozymes I, II, and IV are described. Carbonic anhydrases are abundant in pro- and eukaryotes. These enzymes catalyze the reversible hydration of carbon dioxide to bicarbonate and H(+) ions and are thus involved in many important physiological and pathophysiological processes. Due to the fact that the human carbonic anhydrase family consists of 16 closely related isozymes, the design of selective inhibitors is a special challenge for medicinal chemists. In order to extract selectivity-determining features, we applied purely ligand-based 3D QSAR techniques as well as qualitative comparative molecular field analyses of the targets' binding sites using consensus principal component analysis (CPCA). The dataset for the QSAR studies was deliberately compiled from 1,748 inhibitors and comprises about 140 ligands, mainly of the sulfonamide type. Additionally, we employed the novel AFMoC approach, which intrinsically combines protein and ligand information. The simultaneous use of these different techniques gives deeper insight into selectivity and affinity-determining features and provides quantitative models for prediction.

Carbonic Anhydrase Inhibitors↗

Tumour grading from magnetic resonance spectroscopy: a comparison of feature extraction with variable selection.

Magnetic resonance spectroscopy (MRS) provides a non-invasive measurement of the biochemistry of living tissue. However, signal variation due to tissue heterogeneity causes considerable mixing between different disease categories, making accurate class assignments difficult. This paper compares a systematic methodology for classifier design using multivariate bayesian variable selection (MBVS), with one based on feature extraction using independent component analysis (ICA). We illustrate the methodology and assess the classification performance using a data set comprising 41 magnetic resonance spectra acquired in vivo from two grades of brain tumour, namely low- and medium-grade astrocytic tumours, labelled astrocytomas (AST), and high-grade gliomas and glioblastomas labelled glioblastomas (GL). The aim of this study is threefold. First, to describe the application of the alternative methodologies to MRS, then to benchmark their classification performance, and finally to interpret the classification models in terms of biologically relevant signals derived from the spectra. The classification performance is assessed using the bootstrap method and by application to a test sample in a retrospective study.

Astrocytoma↗

[Population dynamics of the additive polygenic system under truncation selection].

Common features of the equations describing dynamics of the additive polygenic system under truncation selection are summarized. A combination of parameters playing the role of the effective selective pressure on the ith polygenic locus was revealed. The product of mean relative fitnesses of the individual polygenic loci, [formula: see text], was shown to play the role of relative mean fitness of the polygenic population. This value depends on the measurable parameters of the character distribution in the population: [formula: see text]. It was shown that under the constant population number during truncation selection, the characteristic of the best genotype increases, [formula: see text]; which is also a product of the frequencies of preferable genotypes at individual polygenic loci. This value plays the role of the proportion of the number of the best ("champion") genotype in the population. In fact, this is the champion genotype polygene consensus pattern frequency, which a priori indicates the possibility of the champion pattern fixation. The analogue of Haldane's dilemma for the polygenic system which restrict the number of polygenes simultaneously subjected to adaptive evolution [formula: see text] was obtained for the case of constant effective population number (Ne = const).

Genotype↗

A multivariate classification study of attentional orienting in patients with right hemisphere lesions.

OBJECTIVE: To characterize patterns of orienting dysfunction that were typical to patients with lesions in the right hemisphere (RH). BACKGROUND: Brain lesions in the right hemisphere are commonly associated with dysfunction of visual orienting (e.g., visual neglect and extinction). In a clinical study of these symptoms, the multicomponent nature of attentional orienting calls upon a multivariate statistical design with careful selection of neurocognitive variables. METHODS: Thirty-eight patients with verified brain lesions and four patients with peripheral motor dysfunction after poliomyelitis were included in the study. Cognitive function was evaluated in all patients. Three reaction-time (RT) measures derived from the cue-target paradigm were selected as features in a data-driven multivariate classification scheme to generate natural subgroups of patients. RESULTS: Four subgroups were generated, with only RH patients allocated to two of them. All patients within one of the RH subgroups were characterized by a pattern of impairment earlier described as a "disengage failure" by Posner et al. ( 5, 6), signs of visual neglect on Behavioral Inattention Test, and a severely impaired cognitive function. Results for the other RH subgroup were heterogeneous on both experimental and clinical variables. Age and cognitive function were found to strongly influence two of the features, but could not predict the clusters generated from the RT measures. CONCLUSIONS: The present findings show that the cue-target paradigm together with multivariate clustering of selected feature variables can serve as a useful tool to explore and characterize patients with right hemisphere lesions. Although the concept of "disengage failure" was suited to describe the typical RT pattern in patients with neglect, other neurocognitive models and concepts can be applied to the results in the current study.

Adolescent↗

Right- and left-sided colorectal cancers display distinct expression profiles and the anatomical stratification allows a high accuracy prediction of lymph node metastasis.

Accurate preoperative prediction of lymph node metastasis and degree of tumor invasion would facilitate an appropriate decision of the extent of surgical resection of cancers to reduce unnecessary complication or to minimize the risk of recurrence in patients. We analyzed gene expression profiles characteristic of the invasiveness of colorectal carcinoma in a total of 89 cases, using a cDNA array and pattern classification algorithms. We set binary classes for a panel of clinicopathologic parameters, each of which was divided at different levels for categories (discrete) or values (continuous). We searched an optimal combination of genes to discriminate the classes by using of a feature subset selection algorithm, which was applied to a set of genes preselected on the basis of statistical difference in expression (two-sided t test, P < or = 0.05). We used a sequential forward feature selection which additively searched a combination of genes, giving a minimal leave-one-out classification error rate of a k-nearest neighbor classifier. In the process of gene preselection, we found a remarkable difference in the expression pattern of genes according to the anatomical location of cancers. The difference was most prominent when the classes were set for cecum, ascending colon, transverse colon, and descending colon (CATD) versus sigmoid colon and rectum (SR). By stratifying these two locations, we were able to extract gene expression profiles characteristic of the classes of the presence versus absence of lymph node metastasis, lymphatic invasion, vascular invasion and degree of mural invasion, and pathological stages, with an accuracy of more than 90%. These results suggest that colorectal cancers harbor distinct molecular pathophysiological statuses according to their right-to-left locations, of which stratification is important for pattern classification of cDNA array data.

Adult↗

Evolving neural networks to identify bent-double galaxies in the FIRST survey.

The FIRST (Faint Images of the Radio Sky at Twenty-cm) survey is an ambitious project scheduled to cover 10,000 square degrees of the northern and southern galactic caps. Until recently, astronomers associated with FIRST identified radio-emitting galaxies with a bent-double morphology through a visual inspection of images. Besides being subjective, prone to error and tedious, this manual approach is becoming increasingly infeasible: upon completion, FIRST will include almost a million galaxies. This paper describes the application of six methods of evolving neural networks (NNs) with genetic algorithms (GAs) to the identification of bent-double galaxies. The objective is to demonstrate that GAs can successfully address some common problems in the application of NNs to classification problems, such as training the networks, choosing appropriate network topologies, and selecting relevant features. We measured the overall accuracy of the networks using the arithmetic and geometric means of the accuracies on bent and non-bent galaxies. Most of the combinations of GAs and NNs perform equally well on our data, but using GAs to select feature subsets produces the best results, reaching accuracies of 90% using the arithmetic mean and 87% with the geometric mean. The networks found by the GAs were more accurate than hand-designed networks and decision trees.

Astronomy↗

A fuzzy-classifier system to distinguish respiratory patterns evolving after diaphragm paralysis in the cat.

We applied the fuzzy "k-nearest neighbor" (k-NN) classifier of the pattern recognition theory to fathom the abnormal way of breathing resulting from diaphragm paralysis and to distinguish the dominant component, tidal or frequency, of the breathing pattern on which ventilatory compensation relies in such a pathological state. We addressed this issue in the experimental model of diaphragm paralysis as a result of bilateral phrenicotomy in anesthetized, spontaneously breathing cats. Of several variables recorded, we selected two features, minute ventilation and arterial CO(2) tension, that were used for the k-NN analysis. The results demonstrate that the ability to maintain ventilation critically depended on the increase in frequency of breathing. Other breathing pattern strategies were ineffective. The k-NN evaluation with the two selected features discerned the prevailing pattern of breathing with sufficient probability. Such an evaluation may be a useful tool in predicting the development of compensatory strategies in disordered patterns of breathing.

Animals↗

Enhancement of contrast regions in suboptimal ultrasound images with application to echocardiography.

In this paper we propose a novel feature-based contrast enhancement approach to enhance the quality of noisy ultrasound (US) images. Our approach uses a phase-based feature detection algorithm, followed by sparse surface interpolation and subsequent nonlinear postprocessing. We first exploited the intensity-invariant property of phase-based acoustic feature detection to select a set of relevant image features in the data. Then, an approximation to the low-frequency components of the sparse set of selected features was obtained using a fast surface interpolation algorithm. Finally, a nonlinear postprocessing step was applied. Results of applying the method to echocardiographic sequences (2-D + T) are presented. The results demonstrate that the method can successfully enhance the intensity of the interesting features in the image. Better balanced contrasted images are obtained, which is important and useful both for manual processing and assessment by a clinician, and for computer analysis of the sequence. An evaluation protocol is proposed in the case of echocardiographic data and quantitative results are presented. We show that the correction is consistent over time and does not introduce any temporal artefacts.

Acoustics↗

Implicit attentional selection of bound visual features.

Traditionally, research on visual attention has been focused on the processes involved in conscious, explicit selection of task-relevant sensory input. Recently, however, it has been shown that attending to a specific feature of an object automatically increases neural sensitivity to this feature throughout the visual field. Here we show that directing attention to a specific color of an object results in attentional modulation of the processing of task-irrelevant and not consciously perceived motion signals that are spatiotemporally associated with this color throughout the visual field. Such implicit cross-feature spreading of attention takes place according to the veridical physical associations between the color and motion signals, even under special circumstances when they are perceptually misbound. These results imply that the units of implicit attentional selection are spatiotemporally colocalized feature clusters that are automatically bound throughout the visual field.

Acoustic Stimulation↗

Molecular differentiation of ischemic and valvular heart disease by liquid chromatography/fourier transform ion cyclotron resonance mass spectrometry.

Proteomic patterns of myocardial tissue in different etiologies of heart failure were investigated using a direct analytical approach with High Performance Liquid Chromatography (HPLC)/Fourier Transform Ion Cyclotron Resonance Mass Spectrometry (FT-ICR MS). Right atrial appendages from 20 patients, 10 with hemodynamically significant isolated aortic valve disease and 10 with symptomatic coronary artery disease were collected during elective cardiac surgery. After preparation of tissue samples and tryptic digestion of proteins, the peptide mixture was HPLC-separated and on-line analyzed by electrospray FT-ICR MS. Data obtained from HPLC / FT-ICR MS runs were compared for classification. To extract the classification features, the selection of best individual features was applied and the "nearest mean classifier" was used for the classification of test samples and the sample projection onto classification patterns. The pattern distribution characteristics of aortic and coronary diseases were clearly different. No interference between samples of both disease categories was registered, even if the distribution of unsupervised classified test samples were closer. Samples representing aortic valve disease showed a closer accumulation pattern of spots compared to the samples representing coronary disease, which indicated a more specific protein classification. Through selective identification of specific peptides and protein patterns with FTMS, valvular and coronary heart disease is for the first time clearly distinguished at molecular level. The described methodology could also be feasible in search for specific biomarkers in plasma or serum for diagnostic purposes.

Adult↗

Analysis of nonlinear systems to estimate intraocular lens position after cataract surgery.

PURPOSE: To compare the performance of neural networks with that of linear regression to predict the postoperative effective lens position (ELP) from preoperative biometry measurements. SETTING: Departments of Ophthalmology, Medical Cybernetics and Artificial Intelligence, and Medical Physics, Medical University of Vienna, Vienna, Austria. METHODS: The neural-network-type multilayer perceptron (MLP) and a linear regression technique were used to predict ELP. Suitable MLP models and variable input combinations were selected by extended-feature subset selection. Apart from the usual preoperative biometric variables, anterior chamber depth and lens thickness were measured with partial coherence interferometry and white-to-white measurements were used as input variables. RESULTS: Prediction of ELP could be improved from a correlation coefficient (Pearson) of 0.54 for linear regression to a coefficient of 0.68 for the MLP; however, this difference was not statistically significant. CONCLUSION: The prediction of postoperative ACD with the MLP was not significantly better than the prediction using linear regression.

Adult↗

Relation between quantitative description of ultrasonographic image and clinical and laboratory findings in lymphocytic thyroiditis.

OBJECTIVE: Relations between measurable properties of B-mode ultrasound images of thyroid gland and clinical and laboratory findings in patients with chronic inflammation of thyroid gland were studied. METHODS: Data from 65 patients with lymphocytic thyroiditis (LT) and 38 control subjects were analysed. Raw values of individual B-mode image pixels and standard co-occurrence second-order texture features were selected as quantitative image features. Thyroid antibodies, thyrotropin level, thyroxine replacement therapy, and body mass index were used as clinical variables. RESULTS: In the LT group, significant differences (t-tests, p<0.05) in image features were found for body mass indices (BMI) under and over 25 kg.m(-2), for thyroxine replacement therapy, and for the presence and absence of thyroid antibodies. Forward stepwise multiple regression was performed for the clinical or laboratory values as dependent variables and image features as independent variables. The following correlations were found: 1. between BMI and four image features in the normal group; 2. between the dose of thyroxine replacement therapy and two of image features in the LT group; and 3. for the level of thyroid antibodies in the LT group: five image features have correlated with the level of anti-thyroglobulin and three image features with level of anti-thyroperoxidase. CONCLUSION: These findings suggest the possibility of using quantitative indicators of ultrasound image of thyroid gland as predictors of the presence or absence of thyroid antibodies in patient's blood or as an auxiliary tool for dose recommendation of thyroxine replacement therapy.

Adolescent↗

Multiple pharmacophores for the investigation of human UDP-glucuronosyltransferase isoform substrate selectivity.

The UDP-glucuronosyltransferase (UGT) enzyme 'superfamily' contributes to the metabolism of a myriad of drugs, nondrug xenobiotic agents, and endogenous compounds. Although the individual UGT isoforms exhibit distinct but overlapping substrate selectivities, structural features of substrates that confer selectivity remain largely unknown. Using methods developed for pharmacophore fingerprinting combined with optimization and pattern recognition techniques, subsets of pharmacophores associated with the substrates and nonsubstrates of 12 human UGT isoforms were selected to generate predictive models of substrate selectivity and to elucidate the chemical and structural features associated with substrates and nonsubstrates. For all 12 UGT isoforms, the pharmacophore model generated showed predictive ability, as determined by a test set comprising 30% of the available data for each isoform. Models for UGT1A6, -1A7, -1A9, and -2B4 displayed the best predictive ability (>75% of test set predicted correctly) and were further analyzed to interpret the pharmacophores selected as important. The individual pharmacophores differed among isoforms but generally represented relatively simple structural and chemical features. For example, an aromatic ring attached to the nucleophilic group was found to increase the likelihood of glucuronidation by UGT1A6, UGT1A7 and UGT1A9. A large hydrophobic region close to the nucleophile and a hydrogen bond acceptor 10 A from the nucleophile were found to be common to most UGT2B4 substrates. The pharmacophores further suggest that the environment immediately adjacent to the nucleophilic site of conjugation is an important determinant of metabolism by a particular UGT.

Glucuronosyltransferase↗

Selected socio-economic features and the prevalence of peptic ulcer among Polish rural population.

The aim of the study was to determine the relationship between the prevalence of peptic ulcer and the occurrence of selected socio-economic features among Polish rural population. The study was conducted based on the all- Polish representative study of the state of health of rural population, and covered a group of 6,512 rural inhabitants aged 20-64 -- 3,107 males and 3,405 females selected by two-stage stratified sampling. Peptic ulcer was diagnosed in 348 people in the study (5.3%): 250 males (8.0%) and 98 females (2.9%). Duodenal ulcer occurred in 3.2% of people examined, followed by gastric ulcer -- 1.2%, duodenal and gastric ulcer -- 0.2%, and 0.9% of patients underwent surgical procedures due to peptic ulcer. Peptic ulcer occurred more frequently among people with a lower education level (lack of education -- 7.8%, elementary school education -- 5.8%), compared to those with higher education categories (elementary vocational -- 4.9%, secondary school and college -- 3.7%). The disease was more often diagnosed among respondents who described their material standard as poor (7.7%), compared to those who described this standard as good (4.0%). Among people who considered their material standard as poor, gastric ulcer was noted more frequently than duodenal ulcer. A correlation was observed between the prevalence of peptic ulcer and such socio-economic features of Polish rural population as the level of education and material standard.

Adolescent↗