Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Methodological aspects of using decision trees to characterise leiomyomatous tumors.

The aim of the present work is to present the potential uses of a classification technique labeled the "decision tree" for tumor characterisation when faced with a large number of features. The decision tree technique enables multifeature logical classification rules to be produced by determining discriminatory values for each feature selected. In this report, we propose a methodology that used decision trees to compare and evaluate the information contributed by different types of features for tumor characterisation. This methodology is able to produce a set of hypotheses related to a diagnosis and or prognosis problem. For example, hypotheses can be producted (on the basis of a set of descriptive features) to explain why tumor cases belong to a given histopathological group. To illustrate our purpose, this methodology was applied to the difficult problem of leiomyomatous tumour diagnosis. The aim was to illustrate what kind of diagnostic information can be extracted from a sample data set including 23 smooth muscle tumors (14 benign leiomyomas and 9 malignant leiomyosarcomas) described by a large set of computer-assisted, microscope-generated features. Three groups of features were used relating to: (1) ploidy level determination (10 features), (2) quantitative chromatin pattern description (15 features), and (3) immunohistochemically related antigen specificities (6 features). All these features were quantified by digital cell image analysis. The results suggest that an objective distinction between leiomyomas and leiomyosarcomas can be established by means of simple logical rules depending on only a few features among which the immunohistochemically revealed antigen expression of desmin plays a preponderant part. One of the combinations of features proposed by the methodology is interesting for pathologists, because it includes two features describing the appearance of a nucleus in terms of chromatin distribution homogeneity and density, two features widely used by pathologists in tumor-grading systems.

Adolescent↗

Automated computerized classification of malignant and benign masses on digitized mammograms.

RATIONALE AND OBJECTIVES: To develop a method for differentiating malignant from benign masses in which a computer automatically extracts lesion features and merges them into an estimated likelihood of malignancy. MATERIALS AND METHODS: Ninety-five mammograms depicting masses in 65 patients were digitized. Various features related to the margin and density of each mass were extracted automatically from the neighborhoods of the computer-identified mass regions. Selected features were merged into an estimated likelihood of malignancy by, using three different automated classifiers. The performance of the three classifiers in distinguishing between benign and malignant masses was evaluated by receiver operating characteristic analysis and compared with the performance of an experienced mammographer and that of five less experienced mammographers. RESULTS: Our computer classification scheme yielded an area under the receiver operating characteristic curve (Az) value of 0.94, which was similar to that for an experienced mammographer (Az = 0.91) and was statistically significantly higher than the average performance of the radiologists with less mammographic experience (Az = 0.81) (P = .013). With the database used, the computer scheme achieved, at 100% sensitivity, a positive predictive value of 83%, which was 12% higher than that for the performance of the experienced mammographer and 21% higher than that for the average performance of the less experienced mammographers (P < .0001). CONCLUSION: Automated computerized classification schemes may be useful in helping radiologists distinguish between benign and malignant masses and thus reducing the number of unnecessary biopsies.

Breast Neoplasms↗

Feature genes of hepatitis B virus-positive hepatocellular carcinoma, established by its molecular discrimination approach using prediction analysis of microarray.

Recent introduction of a learning algorithm for cDNA microarray analysis has permitted to select feature set to accurately distinguish human cancers according to their pathological judgments. Here, we demonstrate that hepatitis B virus-positive hepatocellular carcinoma (HCC) could successfully be identified from non-tumor liver tissues by supervised learning analysis of gene expression profiling. Through learning and cross-validating HCC sample set, we could identify an optimized set of 44 genes to discriminate the status of HCC from non-tumor liver tissues. In an analysis of other blind-tested HCC sample sets, this feature set was found to be statistically significant, indicating the reproducibility of our molecular discrimination approach with the defined genes. One prominent finding was an asymmetrical distribution pattern of expression profiling in HCC, in which the number of down-regulated genes was greater than that of up-regulated genes. In conclusion, the present findings indicate that application of learning algorithm to HCC may establish a reliable feature set of genes to be useful for therapeutic target of HCC, and that the asymmetric expression pattern may emphasize the importance of suppressed genes in HCC.

Carcinoma, Hepatocellular↗

Computer-assisted bone age assessment: graphical user interface for image processing and comparison.

The current study is part of a project resulting in a computer-assisted analysis of a hand radiograph yielding an assessment of skeletal maturity. The image analysis is based on features selected from six regions of interest. At various stages of skeletal development different image processing problems have to be addressed. At the early stage, feature extraction is based on Lee filtering followed by the random Gibbs fields and mathematical morphology. Once the fusion starts, wavelet decomposition methods are implemented. The user interface displays the closest neighbors to each image under consideration. Results show the sensitivity of different regions to both stages of development and certain feature sensitivity within each region. At the early stage of development, the distal features are more reliable indicators, whereas at the stage of epiphyseal fusion, a larger dynamic range of middle features makes them more sensitive. In the current study, a graphical user interface has been designed and implemented for testing the image processing routines and comparing the results of quantitative image analysis with the visual interpretation of extracted regions of interest. The user interface may also serve as a teaching tool. At the later stage of the project it will be used as a classification tool.

Adolescent↗

Searching for biomarkers of heart failure in the mass spectra of blood plasma.

We have developed a technique for analysing blood plasma using MALDI-MS with subsequent data analysis to identify significant and specific differences between heart failure (HF) patients and healthy individuals. A training dataset comprising 100 HF patients and 100 healthy individuals was used to search for biomarkers (m/z range 1000-10,000). EWP cartridges when used in tandem with microcon centrifugal filters were found to give the best results. A data management chain including event binning, background subtraction and feature extraction was developed to reduce the data, and statistical analysis was used to map feature intensities on to a common scale. Various mathematical approaches including a simple cumulative score, support vector machines (SVM) and genetic algorithms (GAs) were then used to combine the results from individual features and provide a robust classification algorithm. The SVM gave the most promising results (accuracy 95%, receiver operating characteristic (ROC) score of 0.997 using 18 selected features). Finally, a test dataset comprising a further 32 HF patients and 20 controls was used to verify that the 18 putative biomarkers and classification algorithms gave reliable predictions (accuracy 88.5%, ROC score 0.998).

Biomarkers↗

Quantitative structure-activity relationship studies of progesterone receptor binding steroids.

The selection of appropriate descriptors is an important step in the successful formulation of quantitative structure-activity relationships (QSARs). This paper compares a number of feature selection routines and mapping methods that are in current use. They include forward stepping regression (FSR), genetic function approximation (GFA), generalized simulated annealing (GSA), and genetic neural network (GNN). On the basis of a data set of steroids of known in vitro binding affinity to the progsterone receptor, a number of QSAR models are constructed. A comparison of the predictive qualities for both training and test compounds demonstrates that the GNN protocol achieves the best results among the 2D QSAR that are considered. Analysis of the choice of descriptors by the GNN method shows that the results are consistent with established SARs on this series of compounds.

Neural Networks, Computer↗

Weighting features to recognize 3D patterns of electron density in X-ray protein crystallography.

Feature selection and weighting are central problems in pattern recognition and instance-based learning. In this work, we discuss the challenges of constructing and weighting features to recognize 3D patterns of electron density to determine protein structures. We present SLIDER, a feature-weighting algorithm that adjusts weights iteratively such that patterns that match query instances are better ranked than mismatching ones. Moreover, SLIDER makes judicious choices of weight values to be considered in each iteration, by examining specific weights at which matching and mismatching patterns switch as nearest neighbors to query instances. This approach reduces the space of weight vectors to be searched. We make the following two main observations: (1) SLIDER efficiently generates weights that contribute significantly in the retrieval of matching electron density patterns; (2) the optimum weight vector is sensitive to the distance metric i.e. feature relevance can be, to a certain extent, sensitive to the underlying metric used to compare patterns.

Absorptiometry, Photon↗

Extraction of features from ultrasound acoustic emissions: a tool to assess the hydraulic vulnerability of Norway spruce trunkwood?

The aim of this study was to assess the hydraulic vulnerability of Norway spruce (Picea abies) trunkwood by extraction of selected features of acoustic emissions (AEs) detected during dehydration of standard size samples. The hydraulic method was used as the reference method to assess the hydraulic vulnerability of trunkwood of different cambial ages. Vulnerability curves were constructed by plotting the percentage loss of conductivity vs an overpressure of compressed air. Differences in hydraulic vulnerability were very pronounced between juvenile and mature wood samples; therefore, useful AE features, such as peak amplitude, duration and relative energy, could be filtered out. The AE rates of signals clustered by amplitude and duration ranges and the AE energies differed greatly between juvenile and mature wood at identical relative water losses. Vulnerability curves could be constructed by relating the cumulated amount of relative AE energy to the relative loss of water and to xylem tension. AE testing in combination with feature extraction offers a readily automated and easy to use alternative to the hydraulic method.

Botany↗

Plasma free fatty acid levels and the risk of ischemic heart disease in men: prospective results from the Québec Cardiovascular Study.

Insulin resistance, through numerous related disturbances in glucose and lipoprotein-lipid metabolism, is associated with an increased risk of ischemic heart disease (IHD). The purpose of the present study was to examine the relationship between increased plasma free fatty acid (FFA) concentrations, as a feature of the insulin resistance syndrome, and the risk of IHD in men. Analyses were carried out in a nested, case-control sample of men selected from a population of 2103 individuals without IHD at baseline among whom 114 developed IHD during a 5-year follow-up period. Incident IHD cases were matched with controls for age, body mass index, smoking habits and alcohol intake. Analyses were performed while excluding (88 cases and 98 controls) and including (103 cases and 99 controls) patients with type 2 diabetes. Among non-diabetic individuals, elevated plasma FFA concentrations (3rd tertile of the distribution) yielded a twofold increase in the risk of IHD (odds ratio [OR] 2.1, P=0.05) compared with lower plasma FFA levels (lowest tertile) after adjusting for non-lipid risk factors. Further adjustment for insulin, triglycerides, apolipoprotein B, HDL cholesterol and small dense LDL attenuated significantly the relationship between plasma FFA concentrations and the risk of IHD. High plasma FFA levels showed no synergism with selected features of the insulin resistance syndrome in determining the risk of IHD. Inclusion of diabetic subjects in the study did not improve FFA independent prognostic value to the risk of IHD. These results suggest that elevated plasma FFA concentrations are associated with an increased risk of IHD. However, a single fasting measurement of plasma FFA levels does not appear to improve our ability to predict IHD onset in men when information on other risk factors is considered.

Adult↗

Oriented principal component analysis for large margin classifiers.

Large margin classifiers (such as MLPs) are designed to assign training samples with high confidence (or margin) to one of the classes. Recent theoretical results of these systems show why the use of regularisation terms and feature extractor techniques can enhance their generalisation properties. Since the optimal subset of features selected depends on the classification problem, but also on the particular classifier with which they are used, global learning algorithms for large margin classifiers that use feature extractor techniques are desired. A direct approach is to optimise a cost function based on the margin error, which also incorporates regularisation terms for controlling capacity. These terms must penalise a classifier with the largest margin for the problem at hand. Our work shows that the inclusion of a PCA term can be employed for this purpose. Since PCA only achieves an optimal discriminatory projection for some particular distribution of data, the margin of the classifier can then be effectively controlled. We also propose a simple constrained search for the global algorithm in which the feature extractor and the classifier are trained separately. This allows a degree of flexibility for including heuristics that can enhance the search and the performance of the computed solution. Experimental results demonstrate the potential of the proposed method.

Algorithms↗

Interval change analysis to improve computer aided detection in mammography.

We are developing computer aided diagnosis (CAD) techniques to study interval changes between two consecutive mammographic screening rounds. We have previously developed methods for the detection of malignant masses based on features extracted from single mammographic views. The goal of the present work was to improve our detection method by including temporal information in the CAD program. Toward this goal, we have developed a regional registration technique. This technique links a suspicious location on the current mammogram with a corresponding location on the prior mammogram. The novelty of our method is that the search for correspondence is done in feature space. This has the advantage that very small lesions and architectural distortions may be found as well. Following the linking process several features are calculated for the current and prior region. Temporal features are obtained by combining the feature values from both regions. We evaluated the detection performance with and without the use of temporal features on a data set containing 2873 temporal film pairs from 938 patients. There were 589 cases in which the current mammogram contained exactly one malignant mass. Cross validation was used to partition the data set into a train set and a test set. The train set was used for feature selection and classifier training, the test set for classifier evaluation. FROC (free response operating characteristic) analysis showed an improvement in detection performance with the use of temporal features.

Aged↗

Selection of electrode positions for an EEG-based brain computer interface (BCI).

One major question in designing an EEG-based Brain Computer Interface to bypass the normal motor pathways is the selection of proper electrode positions. This study investigates electrode selection with a Distinction Sensitive Learning Vector Quantizer (DSLVQ). DSLVQ is an extended Learning Vector Quantizer (LVQ) which employs a weighted distance function for dynamical scaling and feature selection. The data analysed and classified were 56-channel EEG recordings over sensorimotor areas during preparation for discrete left or right index finger flexions. Data from 3 subjects are reported. It was found by DSLVQ that the most important electrode positions for differentiation between planning of left and right finger movement overlie cortical finger/hand areas over both hemispheres.

Adult↗

Physicochemical descriptors to discriminate protein-protein interactions in permanent and transient complexes selected by means of machine learning algorithms.

Analyzing protein-protein interactions at the atomic level is critical for our understanding of the principles governing the interactions involved in protein-protein recognition. For this purpose, descriptors explaining the nature of different protein-protein complexes are desirable. In this work, the authors introduced Epic Protein Interface Classification as a framework handling the preparation, processing, and analysis of protein-protein complexes for classification with machine learning algorithms. We applied four different machine learning algorithms: Support Vector Machines, C4.5 Decision Trees, K Nearest Neighbors, and Naïve Bayes algorithm in combination with three feature selection methods, Filter (Relief F), Wrapper, and Genetic Algorithms, to extract discriminating features from the protein-protein complexes. To compare protein-protein complexes to each other, the authors represented the physicochemical characteristics of their interfaces in four different ways, using two different atomic contact vectors, DrugScore pair potential vectors and SFCscore descriptor vectors. We classified two different datasets: (A) 172 protein-protein complexes comprising 96 monomers, forming contacts enforced by the crystallographic packing environment (crystal contacts), and 76 biologically functional homodimer complexes; (B) 345 protein-protein complexes containing 147 permanent complexes and 198 transient complexes. We were able to classify up to 94.8% of the packing enforced/functional and up to 93.6% of the permanent/transient complexes correctly. Furthermore, we were able to extract relevant features from the different protein-protein complexes and introduce an approach for scoring the importance of the extracted features.

Algorithms↗

CLIFF: clustering of high-dimensional microarray data via iterative feature filtering using normalized cuts.

We present CLIFF, an algorithm for clustering biological samples using gene expression microarray data. This clustering problem is difficult for several reasons, in particular the sparsity of the data, the high dimensionality of the feature (gene) space, and the fact that many features are irrelevant or redundant. Our algorithm iterates between two computational processes, feature filtering and clustering. Given a reference partition that approximates the correct clustering of the samples, our feature filtering procedure ranks the features according to their intrinsic discriminability, relevance to the reference partition, and irredundancy to other relevant features, and uses this ranking to select the features to be used in the following round of clustering. Our clustering algorithm, which is based on the concept of a normalized cut, clusters the samples into a new reference partition on the basis of the selected features. On a well-studied problem involving 72 leukemia samples and 7130 genes, we demonstrate that CLIFF outperforms standard clustering approaches that do not consider the feature selection issue, and produces a result that is very close to the original expert labeling of the sample set.

Algorithms↗

A protocol for the assessment of 3D movements of the head in persons with cervical dystonia.

OBJECTIVE: To design and test a protocol for the assessment of neck movements in patients affected by cervical dystonia by using an electromagnetic system. This approach could overcome the limits of the current assessment scales in this specific field. BACKGROUND: Initial assessment and function recovery during treatments are diagnosed by the clinician using outcome scales which present many drawbacks in terms of easiness of use, sensitivity, and reliability. DESIGN: A three-dimensional motion analysis system was used to record six different head movements. METHODS: Six able-bodied subjects and 10 subjects affected by cervical dystonia participated in this study. For the different head movements three kinematic parameters (a symmetry index and two indexes related to the reduction of the range of motion) have been extracted in order to compare the performance of able-bodied and disabled persons. RESULTS: The features selected allowed highlighting of the differences between able-bodied and disabled subjects for the degrees of freedom of the neck. CONCLUSIONS: Using a motion analysis system, three kinematic features were extracted from head movements. They seem to allow a more objective assessment of the disability and a more appropriated strategy for the management of patients affected by cervical dystonia.

Adult↗

An intelligent framework for the classification of the 12-lead ECG.

An intelligent framework has been proposed to classify an unknown 12-Lead electrocardiogram into one of a possible number of mutually exclusive and combined diagnostic classes. The framework segregates the classification problem into a number of bi-dimensional classification problems, requiring individual bi-group classifiers for each individual diagnostic class. The bi-group classifiers were generated employing Neural Networks (NN), combined with a combination framework containing an Evidential Reasoning framework to accommodate for any conflicting situations between the bi-group classifiers. A number of different feature selection techniques were investigated with the aim of generating the most appropriate input vector for the bi-group classifiers. It was found that by reducing the original input feature vector, the generalisation ability of the classifiers, when exposed to unseen data, was enhanced and subsequently this reduced the computational requirements of the network itself. The entire framework was compared with a conventional approach to NN classification and a rule based classification approach. The framework attained a significantly higher level of classification in comparison with the other methods; 80.0% compared with 66.7% for the rule based technique and 68.00% for the conventional neural approach.

Computer Simulation↗

Graph neural network-based risk stratification of prostate cancer using gene expression and SHAP interpretability.

Accurate risk stratification is essential for guiding treatment decisions and preventing over treatment of prostate cancer, which remains one of the most prevalent cancers among adult men. While the Gleason score, obtained from prostate biopsies, is routinely used to assess tumor aggressiveness, the biopsy procedure carries risks such as pain, infection, and, in some cases, serious complications such as sepsis. In this study, we proposed an artificial intelligence-based framework that integrates mRNA expression profiles with functional interaction networks to classify prostate cancer patients into low-, medium-, and high-risk groups defined by Gleason scores. The pipeline comprised five steps: (1) data collection from The Cancer Genome Atlas (TCGA), (2) preprocessing of gene expression data, (3) two-stage feature selection to identify informative biomarkers, (4) risk classification using a dual-branch graph neural network (GNN) that combines gene-gene interaction graphs with sample-level expression features, and (5) model interpretation using SHAP to quantify feature contributions. Differentially expressed genes were identified in the High (ASPN, GMNN, PEBP4, C2, KNCK17), Medium (C2, IGSF1, ASPN, CDKN3, AMH), and Low (TNMD, VWA5B2, ST6GALNAC5, CYP3A5, PHGR1) risk groups, underscoring the molecular heterogeneity of disease progression. On an independent held-out test set, the model achieved AUCs of 0.86, 0.88, and 0.95 for the low-, medium-, and high-risk groups, respectively, with an overall accuracy of 80%. These results suggest that combining GNN-based modeling with explainable AI can capture both global and local molecular patterns relevant to tumor aggressiveness. However, as the model was developed and evaluated solely on the TCGA cohort, the findings should be regarded as exploratory, and external validation will be required to establish generalizability. Within these limitations, the proposed framework highlights the potential of molecular profiling and graph-based deep learning to support more precise, potentially less invasive, risk assessment and individualized treatment planning in prostate cancer.

Prostatic Neoplasms↗

The selective value of bacterial shape.

Why do bacteria have shape? Is morphology valuable or just a trivial secondary characteristic? Why should bacteria have one shape instead of another? Three broad considerations suggest that bacterial shapes are not accidental but are biologically important: cells adopt uniform morphologies from among a wide variety of possibilities, some cells modify their shape as conditions demand, and morphology can be tracked through evolutionary lineages. All of these imply that shape is a selectable feature that aids survival. The aim of this review is to spell out the physical, environmental, and biological forces that favor different bacterial morphologies and which, therefore, contribute to natural selection. Specifically, cell shape is driven by eight general considerations: nutrient access, cell division and segregation, attachment to surfaces, passive dispersal, active motility, polar differentiation, the need to escape predators, and the advantages of cellular differentiation. Bacteria respond to these forces by performing a type of calculus, integrating over a number of environmental and behavioral factors to produce a size and shape that are optimal for the circumstances in which they live. Just as we are beginning to answer how bacteria create their shapes, it seems reasonable and essential that we expand our efforts to understand why they do so.

Bacteria↗