Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Semisupervised learning for molecular profiling.

Class prediction and feature selection are two learning tasks that are strictly paired in the search of molecular profiles from microarray data. Researchers have become aware how easy it is to incur a selection bias effect, and complex validation setups are required to avoid overly optimistic estimates of the predictive accuracy of the models and incorrect gene selections. This paper describes a semisupervised pattern discovery approach that uses the by-products of complete validation studies on experimental setups for gene profiling. In particular, we introduce the study of the patterns of single sample responses (sample-tracking profiles) to the gene selection process induced by typical supervised learning tasks in microarray studies. We originate sample-tracking profiles as the aggregated off-training evaluation of SVM models of increasing gene panel sizes. Genes are ranked by E-RFE, an entropy-based variant of the recursive feature elimination for support vector machines (RFE-SVM). A Dynamic Time Warping (DTW) algorithm is then applied to define a metric between sample-tracking profiles. An unsupervised clustering based on the DTW metric allows automating the discovery of outliers and of subtypes of different molecular profiles. Applications are described on synthetic data and in two gene expression studies.

Algorithms↗

GMDH-based feature ranking and selection for improved classification of medical data.

Medical applications are often characterized by a large number of disease markers and a relatively small number of data records. We demonstrate that complete feature ranking followed by selection can lead to appreciable reductions in data dimensionality, with significant improvements in the implementation and performance of classifiers for medical diagnosis. We describe a novel approach for ranking all features according to their predictive quality using properties unique to learning algorithms based on the group method of data handling (GMDH). An abductive network training algorithm is repeatedly used to select groups of optimum predictors from the feature set at gradually increasing levels of model complexity specified by the user. Groups selected earlier are better predictors. The process is then repeated to rank features within individual groups. The resulting full feature ranking can be used to determine the optimum feature subset by starting at the top of the list and progressively including more features until the classification error rate on an out-of-sample evaluation set starts to increase due to overfitting. The approach is demonstrated on two medical diagnosis datasets (breast cancer and heart disease) and comparisons are made with other feature ranking and selection methods. Receiver operating characteristics (ROC) analysis is used to compare classifier performance. At default model complexity, dimensionality reduction of 22 and 54% could be achieved for the breast cancer and heart disease data, respectively, leading to improvements in the overall classification performance. For both datasets, considerable dimensionality reduction introduced no significant reduction in the area under the ROC curve. GMDH-based feature selection results have also proved effective with neural network classifiers.

Breast Neoplasms↗

Classification of organelle trajectories using region-based curve analysis.

A method based on analysis of the region of movement and the functioning of the acto-myosin cytoskeleton has been elaborated to quantify and classify patterns of organelle movement in tobacco pollen tubes. The trajectory was dilated to the region of movement, which was then reduced to give a one-pixel-wide skeleton, represented by a graph structure. The longest line in this skeleton was hypothesized to represent the basic track of the organelle along a single actin filament. Quantitative features were derived from the graph structure, direction of movement on the longest skeletal line, and distance between skeletal line and particle. These features corresponded to biological events like the amount of linear movement or the probability of attachment of an organelle to the actin filament. From 81 analyzed organelle trajectories, 17 had completely linear, 17 had completely non-linear, and 47 had alternating linear and non-linear movement. Selected features were employed for classification and ranking of the movement patterns of a representative sample of the population of organelles moving in the cell tip. The presented methods can be applied to any field where analysis and classification of particle motion are intended.

Actin Cytoskeleton↗

Microbial carbohydrate specific antibodies distinguish between different stages of differentiating mouse cerebellum.

High titered anticarbohydrate antibodies were used to identify cell surface carbohydrates during different stages in histogenesis of mouse cerebellum in a micro tissue-culture system which mimics selected features of in vivo cerebellum development. Blockage of fiber formation within the first few days in vitro and inhibition of cell migrations by carbohydrate-specific antibodies served as an assay system for possible contributions of surface carbohydrates to the behavior of developing cerebellar cells. Microbial strains were selected on the basis of carbohydrate structures of their cell wall antigens, and anticarbohydrate antibodies were raised against treated whole bacteria and yeast in rabbits. We found that antibodies to mannan were active at all stages of development tested (embryonic day 13, E13; the day of birth, PO; and postnatal day 7, P7). Antibodies to sialic acids prepared against strains B and C of Neisseria meningitidis distinguish different subterminal structures: anti-B reacted with E13 and PO cerebellar cells, and anti-C mostly with cells older than P7. Antifetuin antibody recognized E13 and PO but not P7 cell populations. Pneumococcus C strain R36A-specific antibodies were effective only after coating cells to C type carbohydrate before application of the antibody. The results demonstrate that antimicrobiol carbohydrate antibodies cross-react with mammalian cell surface carbohydrate structures and therefore can be used as a powerful tool in tissue culture to analyse those structures which might control cell behaviors pertinent to cerebellar development.

Animals↗

Some practical conclusions following a longitudinal study of common macular lesions.

In a project which lasted 18 years, the natural history of five common macular lesions was followed in 268 patients, and the most significant clinical features selected in each. Methodology consisted of trichromatic analysis with direct ophthalmoscopy, contact lens fundoscopy including the triple mirror, refraction, and fundus photography. In several subjects in each group repeated fluorescence angiographies were carried out. Single eye records were kept and scattergrams of corrected visual acuities plotted. The latter--besides key clinical features--were then used to predict the probable course. Examples are related both to the biographies and to the clinical features to make the method more comprehensible. The five lesions studied were: cystoid macular degeneration, retinal pigment epithelial detachments, choroidal sclerosis, hyaline degenerations, and macular (foveal) degenerations. A new paramacular 'window' defect called a 'target' lesion is described and attention drawn to its close association with migraine. A feature of the study throughout was the frequency with which more than one of the lesions appeared in the same eye, although there was no set sequence. Thus a patient suffering from one macular lesion not infrequently ended by being blinded by another.

Adolescent↗

Prenatal detection of neuroblastoma by fetal ultrasonography.

PURPOSE: We report three cases of neuroblastoma diagnosed by prenatal ultrasound examination and examine the biologic features of tumors diagnosed prenatally. PATIENTS AND METHODS: Neuroblastoma is the most common tumor detected in the newborn period. Thus, some of these tumors develop prenatally and should be detectable by maternal ultrasound. Here we report three cases in which a neuroblastoma was suspected on prenatal ultrasonography. In addition, we review selected features of 17 additional cases reported in the literature. RESULTS AND CONCLUSIONS: These data indicate that, although the majority of patients have favorable clinical and biological features and do well, some patients do not, and the DNA index may be the most important predictor of outcome.

Female↗

Opportunities for machine learning to predict cross-neutralization in FMDV serotype O.

Accurately estimating cross-neutralization between serotype O foot-and-mouth disease viruses (FMDVs) is critical for guiding vaccine selection and disease management. In this study, we developed a machine learning approach to estimate r1 values-an established measure of antigenic similarity-using VP1 sequence data and published virus neutralization titer (VNT) results. Our dataset comprised 108 serum-virus pairs representing 73 distinct FMDV strains. We applied Boruta feature selection and random forest classifiers, optimizing model performance through tenfold cross-validation and sub-sampling to address class imbalance. Predictors included pairwise amino acid distances, site-specific polymorphisms, and differences in potential N-glycosylation sites. Using a 0.3 r1 threshold to define cross-neutralization, the final model achieved high accuracy (0.96), sensitivity (0.93), and specificity (0.96) in training, and performed robustly on independent test sets - accuracy was 0.75 (95% CI 0.60 and 0.90), F1 score 0.86% and PPV 0.77. Importantly, key VP1 residues-positions 48, 100, 135, 150, and 151-emerged as strong predictors of antigenic relationships. Our results demonstrate the utility of integrating routinely generated genomic data with machine learning to inform vaccine candidate selection and anticipate immune interactions among circulating FMDV strains. This approach offers a practical tool for accelerating vaccine decision-making and can be adapted to other FMDV serotypes. The latest version of the r1 predictive model is available for access via a Shiny dashboard (https://dmakau.shinyapps.io/PredImmune-FMD/).

Foot-and-Mouth Disease Virus↗

Computerized evaluation of mammographic lesions: what diagnostic role does the shape of the individual microcalcifications play compared with the geometry of the cluster?

OBJECTIVE: The objective of this study was to compare the diagnostic role of features reflecting the geometry of clusters with features reflecting the shape of the individual microcalcification in a mammographic computer-aided diagnosis system. MATERIALS AND METHODS: Three hundred twenty-four cases of clustered microcalcifications with biopsy-proven results were digitized at 42-microm resolution and analyzed on a computerized system. The shape factor and number of neighbors were computed for each microcalcification, and the eccentricity of the cluster was computed as well. The shape factor is related to the individual microcalcification; the average number of neighbors and the cluster eccentricity reflect the cluster geometry. Stepwise discriminant analysis was used to evaluate the contribution of the extracted features in predicting malignancy. The performance of a classifier based on the features selected by stepwise discriminant analysis was evaluated by receiver operating characteristic (ROC) analysis. RESULTS: To obtain the best discrimination model, we used stepwise discriminant analysis to select the average number of neighbors and the shape of the individual microcalcification, but excluded cluster eccentricity. A classification scheme assigned the average number of neighbors a weighting factor, which was 1.49 times greater than that assigned to the shape factor of the individual microcalcification. A scheme based only on these two features yielded an ROC curve with an area under the curve (A(z)) of 0.87, indicating a positive predictive value of 61% for 98% sensitivity. CONCLUSION: Computerized analysis permitted calculations reflecting the shape of individual microcalcification and the geometry of clusters of microcalcifications. For the computerized classification scheme studied, the cluster geometry was more effective in differentiating benign from malignant clusters than was the shape of individual microcalcification.

Adult↗

Direct identification of pure Penicillium species using image analysis.

This paper presents a method for direct identification of fungal species solely by means of digital image analysis of colonies as seen after growth on a standard medium. The method described is completely automated and hence objective once digital images of the reference fungi have been established. Using a digital image it is possible to extract precise information from the surface of the fungal colony. This includes color distribution, colony dimensions and texture measurements. For fungal identification, this is normally done by visual observation that often results in a very subjective data recording. Isolates of nine different species of the genus Penicillium have been selected for the purpose. After incubation for 7 days, the fungal colonies are digitized using a very accurate digital camera. Prior to the image analysis each image is corrected for self-illumination, thereby gaining a set of directly corresponding images with respect to illumination. A Windows application has been developed to locate the position and size of up to three colonies in the digitized image. Using the estimated positions and sizes of the colonies, a number of relevant features can be extracted for further analysis. The method used to determine the position of the colonies will be covered as well as the feature selection. The texture measurements of colonies of the nine species were analyzed and a clustering of the data into the correct species was confirmed. This indicates that it is indeed possible to identify a given colony merely by macromorphological features. A classifier (in the normal distribution) based on measurements of 151 colonies incubated on yeast extract sucrose agar (YES) was used to discriminate between the species. This resulted in a correct classification rate of 100% when used on the training set and 96% using cross-validation. The same methods applied to 194 colonies incubated on Czapek yeast extract agar (CYA) resulted in a correct classification rate of 98% on the training set and 71% using cross-validation.

Color↗

Bursting as an effective relay mode in a minimal thalamic model.

In recent years, accumulating evidence indicates that thalamic bursts are present during wakefulness and participate in information transmission as an effective relay mode with distinctive properties from the tonic activity. Thalamic bursts originate from activation of the low threshold calcium cannels via a local feedback inhibition, exerted by the thalamic reticular neurons upon the relay neurons. This article, examines if this simple mechanism is sufficient to explain the distinctive properties of thalamic bursting as an effective relay mode. A minimal model of thalamic circuit composed of a retinal spike train, a relay neuron and a reticular neuron is simulated to generate the tonic and burst firing modes. The integrate-and-fire-or-burst model is used to simulate the neurons. After discriminating the burst events with criteria based on inter-spike-intervals, statistical indices show that the bursts of the minimal model are stereotypic events. The relation between the rate of bursts and the parameters of the input spike train demonstrates marked nonlinearities. Burst response is shown to be selective to spike-silence-spike sequences in the input spike train. Moreover, burst events represent the input more reliably than the tonic spike in a considerable range of the parameters of the model. In conclusion, many of the distinctive properties of thalamic bursts such as stereotypy, nonlinear dependence on the sensory stimulus, feature selectivity and reliability are reproducible in the minimal model. Furthermore, the minimal model predicts that while the bursts are more frequent in the spike train of the off-center X relay neurons (corresponding to off-center X retinal ganglion cells), they are more reliable when generated by the on-center ones (corresponding to on-center X ganglion cells).

Action Potentials↗

The compressed feature matrix--a fast method for feature based substructure search.

The compressed feature matrix (CFM) is a feature based molecular descriptor for the fast processing of pharmacochemical applications such as adaptive similarity search, pharmacophore development and substructure search. Depending on the particular purpose, the descriptor may be generated upon either topological or Euclidean molecular data. To assure a variable utilizability, the assignment of the structural patterns to feature types is arbitrarily determined by the user. This step is based on a graph algorithm for substructure search, which resembles the common substructure descriptors. While these merely allow a screening for the predefined patterns, the CFM permits a real substructure/subgraph search, presuming that all desired elements of the query substructure are described by the selected feature set. In this work, the CFM based substructure search is evaluated with regard to both the different outputs resulting from varying feature sets and the search speed. As a benchmark we use the programmable atom typer (PATTY) graph algorithm. When comparing the two methods, the CFM based matrix algorithm is up to several hundred times faster than PATTY and when using the CFM as a basis for substructure screening, the search speed is accelerated by three orders of magnitude. Thus, the CFM based substructure search complies with the requirements for interactive usage, even for the evaluation of several hundred thousand compounds. The concept of the CFM is implemented in the software COFEA. FIGURE CFM based substructure search using the compounds dopamine and benzene-1,2-diol

Algorithms↗

Random forest: a classification and regression tool for compound classification and QSAR modeling.

A new classification and regression tool, Random Forest, is introduced and investigated for predicting a compound's quantitative or categorical biological activity based on a quantitative description of the compound's molecular structure. Random Forest is an ensemble of unpruned classification or regression trees created by using bootstrap samples of the training data and random feature selection in tree induction. Prediction is made by aggregating (majority vote or averaging) the predictions of the ensemble. We built predictive models for six cheminformatics data sets. Our analysis demonstrates that Random Forest is a powerful tool capable of delivering performance that is among the most accurate methods to date. We also present three additional features of Random Forest: built-in performance assessment, a measure of relative importance of descriptors, and a measure of compound similarity that is weighted by the relative importance of descriptors. It is the combination of relatively high prediction accuracy and its collection of desired features that makes Random Forest uniquely suited for modeling in cheminformatics.

Journal Article↗

Automated classification of clustered microcalcifications into malignant and benign types.

The objectives in this study were to design and test a fully automated method for classification of microcalcification clusters into malignant and benign types, and to compare the method's performance with that of radiologists. A novel aspect of the approach is that the relative location and orientation of clusters inside the breast was taken into account for feature calculation. Furthermore, correspondence of location of clusters in mediolateral oblique (MLO) and cranio-caudal (CC) views, was used in feature calculation and in final classification. Initially, microcalcifications were automatically detected by using a statistical method based on Bayesian techniques and a Markov random field model. To determine malignancy or benignancy of a cluster, a method based on two classification steps was developed. In the first step, classification of clusters was performed and in the second step a patient based classification was done. A total of 16 features was used in the study. To identify meaningful features, a feature selection was applied, using the area under the receiver operating characteristic (ROC) curve (Az value) as a criterion. For classification the k-nearest-neighbor method was used in a leave-one-patient-out procedure. A database of 192 mammograms with 280 true positive detected microcalcification clusters was used for evaluation of the method. The set consisted of cases that were selected for diagnostic work up during a 4 year period of screening in the Nijmegen region (The Netherlands). Because of the high positive predictive value in the screening program (50%), this set did not contain obvious benign cases. The method's best patient-based performance on this set corresponded with Az = 0.83, using nine features. A subset of the data set, containing mammograms from 90 patients, was used for comparing the computer results to radiologists' performance. Ten radiologists read these cases on a light-box and assessed the probability of malignancy for each patient. All participants had experience in clinical mammography and participated in our observer study during the last 2 days of a 2-week training session leading to screening mammography certification. Results on the subset showed that the method's performance (Az = 0.83) was considerably higher than that of the radiologists (Az = 0.63).

Breast Neoplasms↗

Regional structural characterization of the brain of schizophrenia patients.

RATIONALE AND OBJECTIVES: We study morphologic characteristics and age-related changes in patients with schizophrenia to investigate whether abnormal neurodevelopment and brain structure have a role in the pathophysiological course of this disease. MATERIALS AND METHODS: Our data consist of a set of cranial magnetic resonance images of 46 patients with schizophrenia and age- and sex-matched healthy controls. We deformed a template brain image to our set of subject images. Jacobian fields of these deformations were reduced to sets of 52 normalized region volumes for each subject by using a neuroanatomic atlas. Normalized regional volumes of the control and patient groups were compared by using Student t-test, and age correlation of each region volume was calculated for the two groups. All results were corrected for multiple comparisons by using permutation testing. We used a classifier based on support vector machines and a feature selection method to determine our ability to discriminate brains of controls from those of patients. RESULTS: Analysis of normalized region volumes shows enlargement of the third ventricle in patients. The age-correlation study showed a significant positive correlation in the third ventricle and right thalamus of controls, but not patients. Using an average of 6.5 features, our classifier was able to correctly identify 72% of patients and 70% of controls. CONCLUSION: In addition to enlargement of the third ventricle, brains of patients with schizophrenia show a different pattern of age-related changes.

Adult↗

Predicting telomerase reverse transcriptase promoter mutation status in glioblastoma by whole-tumor multi-sequence magnetic resonance texture analysis.

OBJECTIVE: This study aimed to determine the feasibility of preoperative multi-sequence magnetic resonance texture analysis (MRTA) for predicting TERT promoter mutation status in IDH-wildtype glioblastoma (IDHwt GB). METHODS: The clinical and imaging data of 111 patients with IDHwt GB at our hospital between November 2018 and June 2023 were retrospectively analyzed as the training set, and those of 23 patients with IDHwt GB between July 2023 and November 2023 were interpreted as the validation set. We used molecular sequencing results to classify the training set into TERT promoter mutation and wildtype groups. Textural features of the whole-tumor volume were extracted, including T2-weighted imaging (T2WI), T2-fluid-attenuated inversion recovery, apparent diffusion coefficient (ADC) map, and contrast-enhanced T1-weighted imaging (CE-T1). All textural features were obtained using open-source pyradiomics. After feature selection, logistic regression was used to build prediction models, and a nomogram was generated. Finally, the model was validated using validation cohort. RESULTS: The CE-T1_Model (AUC 0.704) had a better predictive ability than the T2_Model (AUC 0.684) and ADC_Model (AUC 0.624). The MRI_Combined_Model (CE-T1, T2, and ADC texture features) (AUC 0.780) had a better predictive ability than the Clinical_Model (AUC 0.758). The Combined_Model (CE-T1, T2, ADC texture features, and clinical features) had the best predictive performance (AUC 0.871), with a sensitivity, specificity, and accuracy of 82.60 %, 83.30 %, and 80.18 %, respectively. The AUC, sensitivity, specificity, and accuracy in the validation cohort were 0.775, 86.70 %, 75.00 %, and 69.57 %, respectively. CONCLUSIONS: Whole-tumor multi-sequence MRTA can be used as non-invasive quantitative parameters to assist in the preoperative clinical prediction of TERT promoter mutation status in IDHwt GB.

Humans↗

Inherited and inducible chromosomal instability: a fragile bridge between genome integrity mechanisms and tumourigenesis.

Cancer is a multi-step process evolving as the result of the accumulation of a number of mutational events. The growing body of evidence implicating genetic instability as a key feature of this evolutionary process and the risk of malignancy associated with chromosomal instability syndromes highlight the importance of understanding the mechanisms that cells use to maintain the integrity of their genomes. Classic examples of inherited chromosomal instability with cancer predisposition are Bloom's syndrome, ataxia telangiectasia, and Fanconi anaemia, although the mechanisms involved are far from understood. Selected features of these inherited disorders are reviewed to provide a background to the more recently discovered inducible chromosomal instability, a phenotype in which apparently normal cells that have survived ionizing radiation and certain chemical insults may produce descendants exhibiting a high frequency of de novo chromosome aberrations and gene mutations. The phenotype is induced at frequencies considerably greater than conventional mutation frequencies but little is understood of the underlying mechanism(s). To date, chromosomal instability induced by ionizing radiation has been the most extensively studied phenotype and it is evident that the expression of inducible instability has a strong dependence on the type of radiation exposure, the cell type irradiated, and the genetic 'predisposition' of the irradiated cell.

Animals↗

Automated identification of cancerous smears using various competitive intelligent techniques.

In this study the performance of various intelligent methodologies is compared in the task of pap-smear diagnosis. The selected intelligent methodologies are briefly described and explained, and then, the acquired results are presented and discussed for their comprehensibility and usefulness to medical staff, either for fault diagnosis tasks, or for the construction of automated computer-assisted classification of smears. The intelligent methodologies used for the construction of pap-smear classifiers, are different clustering approaches, feature selection, neuro-fuzzy systems, inductive machine learning, genetic programming, and second order neural networks. Acquired results reveal the power of most intelligent techniques to obtain high quality solutions in this difficult problem of medical diagnosis. Some of the methods obtain almost perfect diagnostic accuracy in test data, but the outcome lacks comprehensibility. On the other hand, results scoring high in terms of comprehensibility are acquired from some methods, but with the drawback of achieving lower diagnostic accuracy. The experimental data used in this study were collected at a previous stage, for the purpose of combining intelligent diagnostic methodologies with other existing computer imaging technologies towards the construction of an automated smear cell classification device.

Artificial Intelligence↗

Significance analysis of qualitative mammographic features, using linear classifiers, neural networks and support vector machines.

Advances in modern technologies and computers have enabled digital image processing to become a vital tool in conventional clinical practice, including mammography. However, the core problem of the clinical evaluation of mammographic tumors remains a highly demanding cognitive task. In order for these automated diagnostic systems to perform in levels of sensitivity and specificity similar to that of human experts, it is essential that a robust framework on problem-specific design parameters is formulated. This study is focused on identifying a robust set of clinical features that can be used as the base for designing the input of any computer-aided diagnosis system for automatic mammographic tumor evaluation. A thorough list of clinical features was constructed and the diagnostic value of each feature was verified against current clinical practices by an expert physician. These features were directly or indirectly related to the overall morphological properties of the mammographic tumor or the texture of the fine-scale tissue structures as they appear in the digitized image, while others contained external clinical data of outmost importance, like the patient's age. The entire feature set was used as an annotation list for describing the clinical properties of mammographic tumor cases in a quantitative way, such that subsequent objective analyses were possible. For the purposes of this study, a mammographic image database was created, with complete clinical evaluation descriptions and positive histological verification for each case. All tumors contained in the database were characterized according to the identified clinical features' set and the resulting dataset was used as input for discrimination and diagnostic value analysis for each one of these features. Specifically, several standard methodologies of statistical significance analysis were employed to create feature rankings according to their discriminating power. Moreover, three different classification models, namely linear classifiers, neural networks and support vector machines, were employed to investigate the true efficiency of each one of them, as well as the overall complexity of the diagnostic task of mammographic tumor characterization. Both the statistical and the classification results have proven the explicit correlation of all the selected features with the final diagnosis, qualifying them as an adequate input base for any type of similar automated diagnosis system. The underlying complexity of the diagnostic task has justified the high value of sophisticated pattern recognition architectures.

Analysis of Variance↗