Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

An Epicurean learning approach to gene-expression data classification.

We investigate the use of perceptrons for classification of microarray data where we use two datasets that were published in [Nat. Med. 7 (6) (2001) 673] and [Science 286 (1999) 531]. The classification problem studied by Khan et al. is related to the diagnosis of small round blue cell tumours (SRBCT) of childhood which are difficult to classify both clinically and via routine histology. Golub et al. study acute myeloid leukemia (AML) and acute lymphoblastic leukemia (ALL). We used a simulated annealing-based method in learning a system of perceptrons, each obtained by resampling of the training set. Our results are comparable to those of Khan et al. and Golub et al., indicating that there is a role for perceptrons in the classification of tumours based on gene-expression data. We also show that it is critical to perform feature selection in this type of models, i.e. we propose a method for identifying genes that might be significant for the particular tumour types. For SRBCTs, zero error on test data has been obtained for only 13 out of 2308 genes; for the ALL/AML problem, we have zero error for 9 out of 7129 genes that are used for the classification procedure. Furthermore, we provide evidence that Epicurean-style learning and simulated annealing-based search are both essential for obtaining the best classification results.

Algorithms↗

Knowledge-based approach to septic shock patient data using a neural network with trapezoidal activation functions.

In this contribution we present an application of a knowledge-based neural network technique in the domain of medical research. We consider the crucial problem of intensive care patients developing a septic shock during their stay at the intensive care unit. Septic shock is of prime importance in intensive care medicine due to its high mortality rate. Our analysis of the patient data is embedded in a medical data analysis cycle, including preprocessing, classification, rule generation and interpretation. For classification and rule generation we chose an improved architecture based on a growing trapezoidal basis function network for our metric variables. Our results extend those of a black box classification and give a deeper insight in our patient data. We evaluate our results with classification and rule performance measures. For feature selection we introduce a new importance measure.

Abdomen↗

A non-calcemic sulfone version of the vitamin D(3) analogue seocalcitol (EB 1089): chemical synthesis, biological evaluation and potency enhancement of the anticancer drug adriamycin.

Novel side-chain diene sulfones 5, analogues of the natural hormone 1alpha,25-dihydroxyvitamin D(3) (calcitriol, 1), were designed to incorporate some of the therapeutically most favorable structural features of the Leo Pharmaceutical Company's drug candidate diene EB 1089 (seocalcitol, 4) and of the Hopkins' non-calcemic side-chain sulfone analogues 2 and 3. Synthesis of diene sulfones 5 features selective Swern oxidation of a primary silyl ether in the presence of a secondary silyl ether (9-->10) and Horner-Wadsworth-Emmons aldehyde addition by a 1-phosphonyl-3-sulfonyl stabilized carbanion regiospecifically at the 1-position to form E,E-diene sulfone 11. Sulfone diene analogue 5a with natural 1alpha,3beta-diol functionality, but not its diastereomer 5b with unnatural A-ring stereochemistry, is antiproliferative in vitro toward murine keratinocytes and malignant melanoma cells, as well as toward MCF-7 human breast cancer cells. Combining diene sulfone 5a with the currently used anticancer drug adriamycin (ADR) caused a noteworthy 3-fold enhancement of ADR antiproliferative potency in MCF-7 cells. Sulfone diene analogue 5a is weakly active transcriptionally in MCF-7 and ROS 17/2.8 cells, binds poorly but measurably to the vitamin D receptor (VDR), and desirably is non-calcemic in vivo at a daily dose (7 days) of 10 microg/kg of rat body weight.

Animals↗

Design, construction and evaluation of systems to predict risk in obstetrics.

We present a systematic, practical approach to developing risk prediction systems, suitable for use with large databases of medical information. An important part of this approach is a novel feature selection algorithm which uses the area under the receiver operating characteristic (ROC) curve to measure the expected discriminative power of different sets of predictor variables. We describe this algorithm and use it to select variables to predict risk of a specific adverse pregnancy outcome: failure to progress in labour. Neural network, logistic regression and hierarchical Bayesian risk prediction models are constructed, all of which achieve close to the limit of performance attainable on this prediction task. We show that better prediction performance requires more discriminative clinical information rather than improved modelling techniques. It is also shown that better diagnostic criteria in clinical records would greatly assist the development of systems to predict risk in pregnancy.

Algorithms↗

Preprocessing of tandem mass spectrometric data based on decision tree classification.

In this study, we present a preprocessing method for quadrupole time-of-flight (Q-TOF) tandem mass spectra to increase the accuracy of database searching for peptide (protein) identification. Based on the natural isotopic information inherent in tandem mass spectra, we construct a decision tree after feature selection to classify the noise and ion peaks in tandem spectra. Furthermore, we recognize overlapping peaks to find the monoisotopic masses of ions for the following identification process. The experimental results show that this preprocessing method increases the search speed and the reliability of peptide identification.

Amino Acid Sequence↗

Prediction and classification of human G-protein coupled receptors based on support vector machines.

A computational system for the prediction and classification of human G-protein coupled receptors (GPCRs) has been developed based on the support vector machine (SVM) method and protein sequence information. The feature vectors used to develop the SVM prediction models consist of statistically significant features selected from single amino acid, dipeptide, and tripeptide compositions of protein sequences. Furthermore, the length distribution difference between GPCRs and non-GPCRs has also been exploited to improve the prediction performance. The testing results with annotated human protein sequences demonstrate that this system can get good performance for both prediction and classification of human GPCRs.

Amino Acids↗

A statewide characterization of hospital infection control practices and practitioners.

Selected features of infection control programs among the 163 general hospitals in Tennessee were surveyed in 1976 and 1979. Each hospital but one had a designated infection control practitioner. Three-fourths of the hospitals had fewer than 200 beds and most were in rural areas. The practitioners in these small hospitals worked in an isolated professional milieu: few (4%) had attended a basic training course or were members of a national (11%) or local (16%) infection control association. They also had significantly less access to standard infection control resource publications than did practitioners in large hospitals. Use of aqueous quaternary ammonium compounds for disinfection was reported by 37% of all hospitals in 1979; 68% of hospitals routinely performed bacteriologic cultures of personnel or the environment. In contrast, only 3% of hospitals did not have a policy specifying the use of sterile closed-system drainage of indwelling bladder catheters. Although these practices varied somewhat by hospital size, the differences were not statistically significant. Modest improvement in each parameter was noted since 1976. Pathology was the most common medical specialty (34%) among chairman of infection control committees; internal medicine and pediatrics accounted for only 13%. The practice of routine microbiologic monitoring was significantly more common among hospitals with chairmen who were pathologists. The implications of these findings for national priorities in hospital infection control are discussed.

Bacteria↗

Spectral-Proteomic Integration Analysis (SPIA) Deciphers Molecular Trajectories of Breast Cancer and Enables Multitarget Therapeutic Assessment.

Raman spectroscopy and mass spectrometry-based proteomics offer deeply complementary yet largely disconnected views of cancer biology: the former provides a label-free, real-time biochemical phenotype, while the latter delivers a quantitative inventory of specific protein effectors. Bridging this gap remains a fundamental challenge in analytical biomedicine. Here, we introduce Spectral-Proteomic Integration Analysis (SPIA)─a novel, data-driven integrative framework that systematically links Raman spectroscopic phenotypes with quantitative proteomic profiles through machine learning and statistical correlation. Using a DMBA-induced rat breast cancer model with and without Toremifene (TOR) intervention, SPIA dynamically maps tumor microenvironment remodeling, capturing progressive collagen deposition and lipid metabolic reprogramming. An SVM classifier trained on Raman spectra achieves exceptional diagnostic accuracy (AUC ≥ 99.0%) and successfully predicts TOR therapeutic response. Proteomic analysis identifies 1,350 differentially expressed proteins, with convergent machine learning feature selection (LASSO, Random Forest, XGBoost) pinpointing core regulators including Luc7l2, Nucb1, Cbx3, and Csnk2a1. Crucially, Spearman correlation analysis between key Raman bands and core DEPs reveals strong, statistically robust associations (median ρ ∼ 0.75 in the 1533-1669 cm-1 region), empirically validating SPIA's core integrative logic. Leveraging this multimodal map, we elucidate a multitarget mechanism for TOR involving concurrent suppression of collagen deposition and correction of aberrant lipid metabolism. SPIA establishes a powerful, generalizable paradigm for integrating phenotypic and molecular data, with broad implications for biomarker discovery, drug mechanism elucidation, and precision oncology.

Animals↗

Dual-Matrix Platform for Highly Specific Multi-Omics Profiling of Renal Cell Carcinoma.

Multiomics interrogation provides complementary information beyond single-omics approaches for improved disease characterization. To enable such multilayer profiling, we expanded the rapid functionalized mesoporous nanoparticle-coupled laser desorption/ionization mass spectrometry (fMNPLDI-MS) platform by designing two structurally homologous but functionally tailored fMNPs. This design enables efficient acquisition of both serum metabolic and peptide fingerprints from a total of only 2.05 μL of serum, with an LDI MS analysis time of approximately 90 s per sample, while addressing the limitation of single-matrix systems in simultaneously optimizing analytical performance for different biomolecular species. Through statistical analysis and machine learning-based feature selection, an integrated multiomics biomarker panel was established, comprising 5 peptides and 4 metabolites. Notably, this integrated panel outperformed both single-omics panels across all evaluation metrics in the validation set, improving the area under curve from 0.985 to 1.000 and increasing the classification accuracy from 0.947 (metabolites) and 0.930 (peptides) to 0.965, while showing consistent improvements in F1-score, precision, and recall. Collectively, these results demonstrate the robust performance of the dual-matrix design and multiomics integration for renal cell carcinoma classification, with potential relevance for broader applications in complex disease profiling.

Carcinoma, Renal Cell↗

From Variability to Consensus: Rescoring Harmonizes Peptide Identification across Diverse Search Engines and Data Sets.

Peptide-spectrum match (PSM) rescoring has become standard in proteomics workflows, improving peptide identification accuracy across diverse search engines. Despite the availability of multiple rescoring strategies, systematic comparisons spanning several search engines, data sets, and database configurations remain limited. Here, we benchmarked seven publicly available search engines, evaluating standard target-decoy-based false discovery rate (FDR) estimation alongside Percolator, MS2Rescore, and Oktoberfest across four data sets acquired on different mass spectrometry platforms in data-dependent mode and searched against protein databases of varying size and composition. Rescoring substantially increased identification consensus and reduced variability between search engines, with prediction-based approaches yielding the largest gains. While database size had limited impact for human data sets, it significantly affected identification rates on a metaproteomic data set. Entrapment-based evaluation indicated generally adequate FDR control across methods, although prediction-based rescoring exhibited a higher tendency toward FDR underestimation in specific configurations. Overall, advanced rescoring strategies harmonize peptide identification outcomes across search engines, thereby enhancing robustness and comparability in proteomics analyses. However, careful feature selection and appropriate database choice remain essential to ensure reliable FDR control and optimal performance across diverse experimental settings.

Search Engine↗

A novel method for building regression tree models for QSAR based on artificial ant colony systems.

Among the multitude of learning algorithms that can be employed for deriving quantitative structure-activity relationships, regression trees have the advantage of being able to handle large data sets, dynamically perform the key feature selection, and yield readily interpretable models. A conventional method of building a regression tree model is recursive partitioning, a fast greedy algorithm that works well in many, but not all, cases. This work introduces a novel method of data partitioning based on artificial ants. This method is shown to perform better than recursive partitioning on three well-studied data sets.

Algorithms↗

Prediction of aqueous solubility of heteroatom-containing organic compounds from molecular structure.

The use of quantitative structure-property relationships (QSPRs) to predict aqueous solubilities (log S) of heteroatom-containing organic compounds from their molecular structure is presented. Three data sets are examined. Data set 1 contains 176 compounds having one or more nitrogen atoms with some oxygen (log S[mol/L] range is -7.41 to 0.96). Data set 2 contains 223 compounds having one or more oxygen atoms, with no nitrogen (log S[mol/L] range is -8.77 to 1.57). Data set 3 contains all 399 compounds from sets 1 and 2 (log S/mol/L] range is -8.77 to 1.57). After descriptor generation and feature selection, multiple linear regression (MLR) and computational neural network (CNN) models are developed for aqueous solubility prediction. The best results were obtained with nonlinear CNN models. Root-mean-square (rms) errors for training with the three data sets ranged from 0.3 to 0.6 log units. All models were validated with external prediction sets, with the rms errors ranging from 0.6 log units to 1.5 log units.

Journal Article↗

Use of computer-assisted methods for the modeling of the retention time of a variety of volatile organic compounds: a PCA-MLR-ANN approach.

A hybrid method consisting of principal component analysis (PCA), multiple linear regressions (MLR), and artificial neural network (ANN) was developed to predict the retention time of 149 C(3)-C(12) volatile organic compounds for a DB-1 stationary phase. PCA and MLR methods were used as feature-selection tools, and a neural network was employed for predicting the retention times. The regression method was also used as a calibration model for calculating the retention time of VOCs and investigating their linear characteristics. The descriptors of the total information index of atomic composition, IAC, Wiener number, W, solvation connectivity index, X1sol, and number of substituted aromatic C(sp(2)), nCaR, appeared in the MLR model and were used as inputs for the ANN generation. Appearance of these parameters shows the importance of the dispersion interactions in the mechanism of retention. Comparison of the MLR and 5-2-1 ANN models indicates the superiority of the ANN over that of the MLR model. The values of 0.913 and 0.738 were obtained for the standard error of prediction set of MLR and ANN models, respectively.

Journal Article↗

Predicting protein-ligand binding affinities using novel geometrical descriptors and machine-learning methods.

Inspired by the concept of knowledge-based scoring functions, a new quantitative structure-activity relationship (QSAR) approach is introduced for scoring protein-ligand interactions. This approach considers that the strength of ligand binding is correlated with the nature of specific ligand/binding site atom pairs in a distance-dependent manner. In this technique, atom pair occurrence and distance-dependent atom pair features are used to generate an interaction score. Scoring and pattern recognition results obtained using Kernel PLS (partial least squares) modeling and a genetic algorithm-based feature selection method are discussed.

Algorithms↗

Diagnostic pattern recognition on gene-expression profile data by using one-class classification.

In this paper, we perform diagnostic pattern recognition on a gene-expression profile data set by using one-class classification. Unlike conventional multiclass classifiers, the one-class (OC) classifier is built on one class only. For optimal performance, it accepts samples coming from the class used for training and rejects all samples from other classes. We evaluate six OC classifiers: the Gaussian model, Parzen windows, support vector data description (with two types of kernels: inner product and Gaussian), nearest neighbor data description, K-means, and PCA on three gene-expression profile data sets, those being an SRBCT data set, a Colon data set, and a Leukemia data set. Providing there is a good splitting of training and test samples and feature selection, most OC classifiers can produce high quality results. Parzen windows and support vector data description are "over-strict" in most cases, while nearest neighbor data description is "over-loose". Other classifiers are intermediate between these two extremes. The main difficulty for the OC classifier is it is difficult to obtain an optimum decision threshold if there are a limited number of training samples.

Colon↗

Modeling of cyclin-dependent kinase inhibition by 1H-pyrazolo[3,4-d]pyrimidine derivatives using artificial neural network ensembles.

Artificial neural network ensembles were used for modeling the cyclin-dependent kinase inhibition of 1H-pyrazolo[3,4-d]pyrimidine derivatives. The structural characteristics of these inhibitors were encoded in relevant 3D-spatial descriptors extracted by genetic algorithm feature selection. Bayesian-regularized multilayer neural networks, trained by the back-propagation algorithm, were developed using these variables as inputs. The predictive power of the model was tested by leave-one-out cross validation. In addition, for a more rigorous measure of the predictive capacity, multiple validation sets were randomly generated as members of neural network ensembles, which makes doing averaged predictions feasible. In this way, the predictive power was analyzed accounting for the averaged test set R values and test set mean-square errors. Otherwise, Kohonen self-organizing maps were used as an additional tool for the same modeling. The location of the inhibitors in a map facilitates the analysis of the connection between compounds and serves as a useful tool for qualitative predictions.

Algorithms↗

Development and evaluation of an in silico model for hERG binding.

It has been recognized that drug-induced QT prolongation is related to blockage of the human ether-a-go-go-related gene (hERG) ion channel. Therefore, it is prudent to evaluate the hERG binding of active compounds in early stages of drug discovery. In silico approaches provide an economic and quick method to screen for potential hERG liability. A diverse set of 90 compounds with hERG IC(50) inhibition data was collected from literature references. Fragment-based QSAR descriptors and three different statistical methods, support vector regression, partial least squares, and random forests, were employed to construct QSAR models for hERG binding affinity. Important fragment descriptors relevant to hERG binding affinity were identified through an efficient feature selection method based on sparse linear support vector regression. The support vector regression predictive model built upon selected fragment descriptors outperforms the other two statistical methods in this study, resulting in an r(2) of 0.912 and 0.848 for the training and testing data sets, respectively. The support vector regression model was applied to predict hERG binding affinities of 20 in-house compounds belonging to three different series. The model predicted the relative binding affinity well for two out of three compound series. The hierarchical clustering and dendrogram results show that the compound series with the best prediction has much higher structural similarity and more neighbors of training compounds than the other two compound series, demonstrating the predictive scope of the model. The combination of a QSAR model and postprocessing analysis, such as clustering and visualization, provides a way to assess the confidence level of QSAR prediction results on the basis of similarity to the training set.

Cell Line↗

Quantitative structure-property relationships for the prediction of vapor pressures of organic compounds from molecular structures

A quantitative structure-property relationship (QSPR) is developed to relate the molecular structures of 420 diverse organic compounds to their vapor pressures at 25 degrees C expressed as log(vp), where vp is in pascals. The log(vp) values range over 8 orders of magnitude from -1.34 to 6.68 log units. The compounds are encoded with topological, electronic, geometrical, and hybrid descriptors. Statistical and computational neural network (CNN) models are built using subsets of the descriptors chosen by simulated annealing and genetic algorithm feature selection routines. An 8-descriptor CNN model, which contains only topological descriptors, is presented which has a root-mean-square (rms) error of 0.37 log unit for a 65-member external prediction set. A 10-descriptor CNN model containing a larger selection of descriptor types gives an improved rms error of 0.33 log unit for the external prediction set.

Journal Article↗