Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Oligosaccharide microarrays fabricated on aminooxyacetyl functionalized glass surface for characterization of carbohydrate-protein interaction.

Carbohydrate-protein interactions play important biological roles in biological processes. But there is a lack of high-throughput methods to elucidate recognition events between carbohydrates and proteins. This paper reported a convenient and efficient method for preparing oligosaccharide microarrays, wherein the underivatized oligosaccharide probes were efficiently immobilized on aminooxyacetyl functionalized glass surface by formation of oxime bonding with the carbonyl group at the reducing end of the suitable carbohydrates via irreversible condensation. Prototypes of carbohydrate microarrays containing 10 oligosaccharides were fabricated on aminooxyacetyl functionalized glass by robotic arrayer. Utilization of the prepared carbohydrate microarrays for the characterization of carbohydrate-protein interaction reveals that carbohydrates with different structural features selectively bound to the corresponding lectins with relative binding affinities that correlated with those obtained from solution-based assays. The limit of detection (LOD) for lectin ConA on the fabricated carbohydrate microarrays was determined to be approximately 0.008 microg/mL. Inhibition experiment with soluble carbohydrates also demonstrated that the binding affinities of lectins to different carbohydrates could be analyzed quantitatively by determining IC(50) values of the soluble carbohydrates with the carbohydrate microarrays. This work provides a simple procedure to prepare carbohydrate microarray for high-throughput parallel characterization of carbohydrate-protein interaction.

Aminooxyacetic Acid↗

Prediction of gas chromatographic retention indices of a diverse set of toxicologically relevant compounds.

For a set of 846 organic compounds, relevant in forensic analytical chemistry, with highly diverse chemical structures, the gas chromatographic Kovats retention indices have been quantitatively modeled by using a large set of molecular descriptors generated by software Dragon. Best and very similar performances for prediction have been obtained by a partial least squares regression (PLS) model using all considered 529 descriptors, and a multiple linear regression (MLR) model using only 15 descriptors obtained by a stepwise feature selection. The standard deviations of the prediction errors (SEP), were estimated in four experiments with differently distributed training and prediction sets. For the best models SEP is about 80 retention index units, corresponding to 2.1-7.2% of the covered retention index interval of 1110-3870. The molecular properties known to be relevant for GC retention data, such as molecular size, branching and polar functional groups are well covered by the selected 15 descriptors. The developed models support the identification of substances in forensic analytical work by GC-MS in cases the retention data for candidate structures are not available.

Calibration↗

Multivariate adaptive regression splines (MARS) in chromatographic quantitative structure-retention relationship studies.

The multivariate adaptive regression splines (MARS) methodology was applied to build quantitative structure-retention relationships (QSRRs). The response (dependent variable) in the MARS models consisted of the logarithms of the extrapolated retention factors (log k(w)) of 83 structurally diverse drugs on a Unisphere PBD column, using isocratic elutions at pH 11.7. A set of 266 molecular descriptors was used as predictor (independent) variables in the MARS model building. The optimal MARS model uses 34 basis functions to describe the retention and has acceptable predictive properties for new objects. The molecular descriptors included in the model describe hydrophobicity, molecular size, complexity, shape and polarisability. Some additional MARS models were created using alternative strategies. These include models with log P as the single predictor and models obtained with only the three most important molecular descriptors. The use of classification and regression trees (CART) as feature selection technique for predictor variables used in the MARS model was also investigated. Further, it is also studied whether allowing quadratic terms instead of interaction terms might lead to better MARS models.

Chromatography, Liquid↗

Optimisation of a new headspace mass spectrometry instrument. Discrimination of different geographical origin olive oils.

A fast head-space analysis instrument, constituted by an automatic sample introduction system directly coupled to a mass detector without performing any chromatographic separation, was assembled. A suitable and original response was computed to optimise, by experimental design, the measured signals for discrimination purposes. The volatile fractions of 105 extra virgin olive oils coming from five different Mediterranean areas were analysed. The rough information collected by this system was unravelled and explained by well-known chemometrical techniques of display (principal component analysis), feature selection (stepwise linear discriminant analysis) and classification (linear discriminant analysis). The 93.4% of samples resulted to be correctly classified and the 90.5% correctly predicted by cross-validation procedure, whilst the 80.0% of an external test set, created to full validate the classification rule, were correctly assigned.

Mass Spectrometry↗

Identification of Africanized honeybees.

Gas chromatography and pattern recognition methods were used to develop a potential method for differentiating European honeybees from Africanized honeybees. The test data consisted of 237 gas chromatograms of hydrocarbon extracts obtained from the wax glands, cuticle, and exocrine glands of European and Africanized honeybees. Each gas chromatogram contained 65 peaks corresponding to a set of standardized retention time windows. A genetic algorithm (GA) for pattern recognition was used to identify features in the gas chromatograms characteristic of the genotype. The pattern recognition GA searched for features in the chromatograms that optimized the separation of the European and Africanized honeybees in a plot of the two or three largest principal components of the data. Because the largest principal components capture the bulk of the variance in the data, the peaks identified by the pattern recognition GA primarily contained information about differences between gas chromatograms of European and Africanized honeybees. The principal component analysis routine embedded in the fitness function of the pattern recognition GA acted as an information filter, significantly reducing the size of the search space since it restricted the search to feature sets whose principal component plots showed clustering on the basis of the bees' genotype. In addition, the algorithm focused on those classes and/or samples that were difficult to classify as it trained using a form of boosting. Samples that consistently classify correctly are not as heavily weighted as samples that are difficult to classify. Over time, the algorithm learns its optimal parameters in a manner similar to a neural network. The pattern recognition GA integrates aspects of artificial intelligence and evolutionary computations to yield a "smart" one-pass procedure for feature selection and classification.

Animals↗

Applications of quantum AI in brain disorder diagnosis: A systematic review.

BACKGROUND AND OBJECTIVE: Brain disorder diagnosis and prediction remain challenging because neuroimaging, electrophysiological, behavioral, and multimodal data are high-dimensional, noisy, heterogeneous, and limited by small clinical cohorts. This systematic review synthesised applications of quantum artificial intelligence (QAI) for brain disorder diagnosis, prediction, detection, and monitoring. METHODS: Following PRISMA guidelines, studies published from 2016 to 13 January 2026 were retrieved from Scopus, Web of Science, and IEEE Xplore. After screening, 36 studies met the eligibility criteria and were qualitatively analysed according to disorder category, data modality, QAI method, implementation setting, validation strategy, and performance. RESULTS: At the broader disease-group level, neurodegenerative disorders were the most frequently investigated, followed by mental health and psychiatric disorders. At the individual level, Parkinson's disease and schizophrenia were the leading applications, followed by depression, anxiety, Alzheimer's disease, and stress-related tasks. MRI-based modalities were the most frequently used data source, followed by multimodal data and EEG. Methodologically, primary QAI approaches were dominated by quantum neural and QDL architectures, followed by quantum-inspired optimization or feature-selection methods and quantum-kernel/conventional QML classifiers. Qiskit/IBM Quantum and PennyLane were the most frequently reported quantum software frameworks. However, most studies relied on simulators, classical quantum-inspired implementations, or unclear implementation settings, with limited real-hardware evaluation. CONCLUSIONS: QAI shows emerging potential for brain disorder analysis, particularly through hybrid quantum-classical learning, quantum neural architectures, quantum-kernel methods, and quantum-inspired optimization. Nevertheless, current evidence remains preliminary and requires larger datasets, subject-level and external validation, fair classical benchmarking, noise-resilient circuits, real quantum hardware evaluation, explainability, and clinical validation.

Humans↗

Degree prediction of malignancy in brain glioma using support vector machines.

The degree of malignancy in brain glioma needs to be assessed by MRI findings and clinical data before operations. There have been previous attempts to solve this problem with a fuzzy rule extraction algorithm based on fuzzy min-max neural networks. We utilize support vector machines with floating search method to select relevant features and to predict the degree of malignancy. Computation results show that the feature subset selected by our techniques can yield better classification performance. In contrast with the base line method, which generated two rules and obtained 83.21% accuracy on the whole data set, our method generates one rule to yield 88.21% accuracy.

Algorithms↗

A method based on multispectral imaging technique for white blood cell segmentation.

Aiming at the problem of image analysis of White Blood Cells in bone marrow microscopic images, multispectral imaging techniques with spectral calibration method was proposed to acquire device-independent images, which is almost impossible in conventional color imaging method. For image segmentation, Support Vector Machine (SVM) was applied directly to the spectrum of each pixel, and using sequential minimal optimization (SMO) algorithm for feature selection to reduce the time of training SVM classifier. Mass of experiments showed that the method is robust, effective and insensitive to smear staining and illumination condition.

Algorithms↗

A wavelet-based data pre-processing analysis approach in mass spectrometry.

Recently, mass spectrometry analysis has a become an effective and rapid approach in detecting early-stage cancer. To identify proteomic patterns in serum to discriminate cancer patients from normal individuals, machine-learning methods, such as feature selection and classification, have already been involved in the analysis of mass spectrometry (MS) data with some success. However, the performance of existing machine learning methods for MS data analysis still needs improving. The study in this paper proposes a wavelet-based pre-processing approach to MS data analysis. The approach applies wavelet-based transforms to MS data with the aim of de-noising the data that are potentially contaminated in acquisition. The effects of the selection of wavelet function and decomposition level on the de-noising performance have also been investigated in this study. Our comparative experimental results demonstrate that the proposed de-noising pre-processing approach has potentials to remove possible noise embedded in MS data, which can lead to improved performance for existing machine learning methods in cancer detection.

Algorithms↗

Neuroimaging: seeing the trees for the forest.

New functional imaging studies demonstrate that it is possible to decode a sensory visual pattern, and even an internal perceptual state, by combining seemingly insignificant feature selective signal biases present in a large number of voxels.

Attention↗

A machine learning perspective on the development of clinical decision support systems utilizing mass spectra of blood samples.

Currently, the best way to reduce the mortality of cancer is to detect and treat it in the earliest stages. Technological advances in genomics and proteomics have opened a new realm of methods for early detection that show potential to overcome the drawbacks of current strategies. In particular, pattern analysis of mass spectra of blood samples has attracted attention as an approach to early detection of cancer. Mass spectrometry provides rapid and precise measurements of the sizes and relative abundances of the proteins present in a complex biological/chemical mixture. This article presents a review of the development of clinical decision support systems using mass spectrometry from a machine learning perspective. The literature is reviewed in an explicit machine learning framework, the components of which are preprocessing, feature extraction, feature selection, classifier training, and evaluation.

Artificial Intelligence↗

Prediction of estrogen receptor agonists and characterization of associated molecular descriptors by statistical learning methods.

Specific estrogen receptor (ER) agonists have been used for hormone replacement therapy, contraception, osteoporosis prevention, and prostate cancer treatment. Some ER agonists and partial-agonists induce cancer and endocrine function disruption. Methods for predicting ER agonists are useful for facilitating drug discovery and chemical safety evaluation. Structure-activity relationships and rule-based decision forest models have been derived for predicting ER binders at impressive accuracies of 87.1-97.6% for ER binders and 80.2-96.0% for ER non-binders. However, these are not designed for identifying ER agonists and they were developed from a subset of known ER binders. This work explored several statistical learning methods (support vector machines, k-nearest neighbor, probabilistic neural network and C4.5 decision tree) for predicting ER agonists from comprehensive set of known ER agonists and other compounds. The corresponding prediction systems were developed and tested by using 243 ER agonists and 463 ER non-agonists, respectively, which are significantly larger in number and structural diversity than those in previous studies. A feature selection method was used for selecting molecular descriptors responsible for distinguishing ER agonists from non-agonists, some of which are consistent with those used in other studies and the findings from X-ray crystallography data. The prediction accuracies of these methods are comparable to those of earlier studies despite the use of significantly more diverse range of compounds. SVM gives the best accuracy of 88.9% for ER agonists and 98.1% for non-agonists. Our study suggests that statistical learning methods such as SVM are potentially useful for facilitating the prediction of ER agonists and for characterizing the molecular descriptors associated with ER agonists.

Forecasting↗

Development and validation of a comprehensive prognostic model for 28-day ICU mortality in non-traumatic subarachnoid hemorrhage: an analysis based on the MIMIC-IV database.

BACKGROUND: Due to the complex pathophysiology of non-traumatic subarachnoid hemorrhage (SAH), accurate risk prediction remains a challenge. Our aim is to develop and validate a comprehensive prognostic model that integrates demographic characteristics, vital signs, laboratory parameters, and more, to provide clinical decision-making support in real-world practice. METHODS: We conducted a retrospective cohort study of 785 Non-traumatic subarachnoid hemorrhage patients. The cohort was randomly divided into a training set (n = 549) and a validation set (n = 236). Feature selection was performed using LASSO regression, followed by backward stepwise Cox regression for optimization. A nomogram was constructed based on independent predictive factors, and model performance was assessed using discrimination, calibration, and decision curve analysis. To prevent immortal-time bias, all predictors were anchored to a fixed early (first-24-hour) measurement window, treatment variables were modelled as binary indicators rather than cumulative exposures, and a five-model sensitivity analysis with baseline-severity adjustment was performed. RESULTS: The development of our model followed a systematic approach: first, 15 potential predictive factors were selected via LASSO regression, which were then refined to 12 independent predictors using backward stepwise Cox regression. The final predictive factors included: Ventilation, AHT, Nimodipine 60 mg, Age, SAPS.II, Input amount, Calcium total, Platelet count, White blood cells, Anion gap, pH, and Chloride. The integrated model demonstrated excellent predictive ability for 7-day, 14-day, and 21-day mortality in both the training set (AUC: 0.972, 0.934, 0.898) and the validation set (AUC: 0.968, 0.948, 0.911). Calibration curves and decision curve analysis confirmed the model's reliability and clinical utility across different time points. We constructed a nomogram for individualized risk prediction. Univariate Kaplan-Meier survival analysis demonstrated significant stratification of survival outcomes by each predictor, while restricted cubic spline analysis revealed non-linear relationships between continuous variables and mortality risk. Random survival forest analysis identified the top three predictive factors (Nimodipine 60 mg, Ventilation, AHT) and compared them with our full 12-variable model, confirming superior performance of the integrated model at all time points. At the 28-day primary endpoint, the model achieved a time-dependent AUC of 0.898 (training) and 0.904 (validation); after restricting predictors to the early baseline window, the leakage-controlled model retained good discrimination (validation C-index 0.803). CONCLUSIONS: Our ICU 28-day mortality prognosis model demonstrated robust performance in predicting ICU 28-day mortality in non-traumatic subarachnoid hemorrhage. The model, through the nomogram, provides individualized risk assessment, aiding clinical decision-making and patient stratification.

Humans↗

Integrated salivary proteomic and metabolomic analyses reveal molecular characterization and novel biomarker panels of chronic obstructive pulmonary disease.

Chronic obstructive pulmonary disease (COPD) is a respiratory disorder characterized by chronic inflammation, oxidative stress, and metabolic dysregulation. The lack of convenient and easily-accessible non-invasive diagnostic approaches remains a major clinical challenge. This study applied an integrated saliva-based proteomic and untargeted metabolomic strategy to identify potential biomarkers for COPD classification. Comprehensive multi-omics analyses identified 225 differentially abundant proteins and 60 differentially abundant metabolites between patients with COPD and healthy controls, including 24 biologically relevant endogenous metabolites. Functional enrichment analyses revealed pronounced dysregulation of mitochondrial energy metabolism, redox homeostasis, lipid remodeling, and inflammatory-related pathways in COPD. By integrating salivary proteomic and metabolomic biomarkers, a stepwise feature selection combined with LASSO logistic regression was used to construct diagnostic models, yielding an optimized biomarker panel consisting of 11 proteins and 2 endogenous metabolites. This integrated model achieved excellent diagnostic performance, with an area under the ROC curve of 0.96. Collectively, these findings demonstrate that integrated salivary proteomic and metabolomic profiling provides a robust, non-invasive approach for COPD classification and offers a promising foundation for the development of biosensor-based diagnostic platforms and early disease detection. SIGNIFICANCE: Chronic obstructive pulmonary disease (COPD) remains a major global health burden. Current diagnostic approaches rely largely on spirometry and clinical assessment, which are limited in sensitivity for early-stage disease and unsuitable for large-scale screening. This study employs an integrated saliva-based proteomic and metabolomic strategy to identify non-invasive biomarkers for COPD classification. Our findings reveal coordinated dysregulation of mitochondrial energy metabolism, redox homeostasis, and lipid remodeling in COPD, highlighting the interconnected roles of metabolic reprogramming, oxidative stress, and inflammation in disease pathophysiology. Notably, a robust diagnostic panel comprising 11 proteins and 2 endogenous metabolites was established, achieving excellent classification performance (AUC of 0.96). To our knowledge, the integrated application of salivary proteomics and metabolomics for COPD diagnosis remains largely unexplored, underscoring the significance and translational potential of our findings.

Humans↗

Multiomics Integration Identifies a Molecular Subtype of Intrahepatic Cholangiocarcinoma With Enhanced Benefit From Adjuvant Therapy.

Intrahepatic cholangiocarcinoma (iCCA) is a molecularly heterogeneous liver cancer with a poor prognosis. Improved stratification is needed to guide postoperative therapy. In this study, we applied integrative multiomics analysis to classify iCCA and identify biomarkers predictive of adjuvant treatment benefit. Using publicly available datasets (including whole exome sequencing, RNA sequencing, proteomics, and phosphoproteomics from FU-iCCA cohort and a transcriptomic cohort GSE244807), we defined 3 robust molecular subtypes of iCCA. These subtypes exhibited distinct genomic alterations, pathway activation, and immune microenvironments, with significant differences in overall survival (OS). Through protein-protein interaction network analysis and consensus feature selection using 10 clustering algorithms, we prioritized 8 marker genes distinguishing the subtypes. A Cox proportional-hazards model constructed from these markers stratified patients into high- and low-risk groups. High-risk iCCA, characterized by elevated expression of markers such as CLDN18, MUC1, and MUC5AC, had significantly worse OS in the absence of adjuvant therapy. Notably, in an independent validation of 174 patients with iCCA who underwent resection (single-center cohort), high expression of any of these 3 markers were associated with markedly prolonged OS in patients who received adjuvant chemotherapy or chemoembolization, compared with those who did not. In contrast, marker-negative patients showed no clear benefit from adjuvant therapy. In conclusion, our multiomics approach identified a high-risk, mucin-enriched subtype of iCCA. CLDN18, MUC1, and MUC5AC emerge as candidate predictive biomarkers for adjuvant chemotherapy benefit in iCCA, warranting prospective validation to improve personalized postoperative management.

Humans↗

Principal component and linear discriminant analysis of T1 histograms of white and grey matter in multiple sclerosis.

Twenty-three relapsing remitting multiple sclerosis (RRMS) patients and 14 controls were imaged to produce normal-appearing white and grey matter T1 histograms. These were used to assess whether histogram measures from principal component analysis (PCA) and linear discriminant analysis (LDA) out-perform traditional histogram metrics in classification of T1 histograms into control and RRMS subject groups and in correlation with the expanded disability status score (EDSS). The histograms were classified into one of two groups using a leave-one-out analysis. In addition, the patients were scanned serially, and the calculated parameters correlated with the EDSS. The classification results showed that the more complex techniques were at least as good at classifying the subjects as histogram mean, peak height and peak location, with PCA/LDA having success rates of 76% for white matter and 68%/65% for grey matter. No significant correlations were found with EDSS for any histogram parameter. These results indicate that there is much information contained within the grey matter as well as the white matter histograms. Although in these histograms PCA and LDA did not add greatly to the discriminatory power of traditional histogram parameters, they provide marginally better performance, while relying only on data-driven feature selection.

Brain↗

Predictive neural networks for gene expression data analysis.

Gene expression data generated by DNA microarray experiments have provided a vast resource for medical diagnosis and disease understanding. Most prior work in analyzing gene expression data, however, focuses on predictive performance but not so much on deriving human understandable knowledge. This paper presents a systematic approach for learning and extracting rule-based knowledge from gene expression data. A class of predictive self-organizing networks known as Adaptive Resonance Associative Map (ARAM) is used for modelling gene expression data, whose learned knowledge can be transformed into a set of symbolic IF-THEN rules for interpretation. For dimensionality reduction, we illustrate how the system can work with a variety of feature selection methods. Benchmark experiments conducted on two gene expression data sets from acute leukemia and colon tumor patients show that the proposed system consistently produces predictive performance comparable, if not superior, to all previously published results. More importantly, very simple rules can be discovered that have extremely high diagnostic power. The proposed methodology, consisting of dimensionality reduction, predictive modelling, and rule extraction, provides a promising approach to gene expression analysis and disease understanding.

Animals↗

Support vector machines for temporal classification of block design fMRI data.

This paper treats support vector machine (SVM) classification applied to block design fMRI, extending our previous work with linear discriminant analysis [LaConte, S., Anderson, J., Muley, S., Ashe, J., Frutiger, S., Rehm, K., Hansen, L.K., Yacoub, E., Hu, X., Rottenberg, D., Strother S., 2003a. The evaluation of preprocessing choices in single-subject BOLD fMRI using NPAIRS performance metrics. NeuroImage 18, 10-27; Strother, S.C., Anderson, J., Hansen, L.K., Kjems, U., Kustra, R., Siditis, J., Frutiger, S., Muley, S., LaConte, S., Rottenberg, D. 2002. The quantitative evaluation of functional neuroimaging experiments: the NPAIRS data analysis framework. NeuroImage 15, 747-771]. We compare SVM to canonical variates analysis (CVA) by examining the relative sensitivity of each method to ten combinations of preprocessing choices consisting of spatial smoothing, temporal detrending, and motion correction. Important to the discussion are the issues of classification performance, model interpretation, and validation in the context of fMRI. As the SVM has many unique properties, we examine the interpretation of support vector models with respect to neuroimaging data. We propose four methods for extracting activation maps from SVM models, and we examine one of these in detail. For both CVA and SVM, we have classified individual time samples of whole brain data, with TRs of roughly 4 s, thirty slices, and nearly 30,000 brain voxels, with no averaging of scans or prior feature selection.

Algorithms↗