Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Transfer Learning across Material Properties Using Center-Environment Features: From Energetics to Mechanical Properties in Multicomponent Mo Alloys.

Transfer learning (TL) provides a viable approach to mitigate data scarcity in materials informatics. While conventional TL focuses on predicting identical properties across different systems, this work demonstrates a cross-property extension of TL from energy to mechanical properties via end-to-end model weight pre-training and fine-tuning: knowledge learned from predicting substitution energies is transferred to predict distinctly different mechanical properties, substantially improving computational efficiency given the typically higher cost of acquiring target-domain data. To accelerate computational alloy design, machine learning models using center-environment (CE) features were first developed to predict substitution energies of alloying elements in molybdenum (Mo)-based alloys. The Random Forest models achieved the optimal performance and transferability-R2 = 0.97, 〈MAE〉 = 0.11 eV, and 〈RMSE〉 = 0.16 eV-against the density functional theory (DFT) benchmark. The model dependency of feature selection and importance analysis was discussed. The transferability of the energy models was validated on unknown systems with new elements. Subsequently, the energy models were fine-tuned using limited mechanical property data to construct energy-to-property (E2P) TL models capable of predicting elastic properties, including bulk modulus, Young's modulus, shear modulus, and elastic constants, achieving an improved accuracy over the non-transferred ML by ∼10-30%, with its transferability verified by additional DFT calculations. This cross-property E2P transfer learning framework opens a new avenue for accelerating computational materials discovery and may be extended to other multiproperty predictions governed by similar physical principles.

center-environment feature↗

Differential distribution of Fos-like immunoreactivity in the spinal trigeminal nucleus after noxious and innocuous thermal and chemical stimulation of rat cornea.

Corneal afferent nerves project to two spatially distinct sites within the spinal trigeminal nucleus: the subnucleus interpolaris/caudalis transition and the subnucleus caudalis/upper cervical spinal cord transition. The role of these two regions in processing corneal input is uncertain. To determine if neurons in these regions encode different features of an applied corneal stimulus, immunoreactivity for the immediate early gene protein product, Fos, was quantified in barbiturate-anesthetized rats. Intensity was varied across thermal (thermal probe 5, 35, 42, 52 degrees C; radiant heat of approximately 45 degrees C) stimuli and compared with that seen after mustard oil (5 microliters, 20%) or mineral oil application. All stimuli increased the number of Fos-positive neurons located at the ventrolateral pole of the subnucleus interpolaris/caudalis transition compared with unstimulated controls. By contrast, only 52 degrees C thermal probe and mustard oil produced an additional peak of Fos-positive neurons within the superficial laminae at the subnucleus caudalis/cervical cord transition. Further, the magnitudes of the bimodal peaks of Fos produced by 52 degrees C thermal probe and mustard oil stimuli were different quantitatively. Mustard oil caused a greater Fos response at the subnucleus interpolaris/caudalis transition than 52 degrees C thermal probe stimulation, whereas the opposite was true at the subnucleus caudalis/cervical cord transition. Double-labeling revealed that Fos immunoreactive neurons within the spinal trigeminal nucleus were restricted to regions densely labeled for calcitonin gene-related peptide. These results indicate that select features of corneal stimuli such as modality are encoded differently by neurons in the trigeminal subnucleus interpolaris/caudalis transition compared with those located in the subnucleus caudalis/cervical cord transition. It is likely that neurons in these two brainstem regions subserve different aspects of corneal sensation.

Animals↗

Detecting genomic features under weak selective pressure: the example of codon usage in animals and plants.

Large scale experiments of gene inactivation in yeast have shown that 50% of genes have no detectable impact on the phenotype, and similar observations have been made in other model organisms. This apparent paradox is probably due to the fact that many genes only have a marginal contribution to the fitness of organisms. Because of the size of populations and the number of generations that can be studied in laboratories, experimental approaches only permit to detect functional elements that have a strong phenotypic impact. Comparative sequence analysis can help to solve this problem: the analysis of sequences evolution permits to detect the action of selection, and hence to reveal functional features of genomes. This approach will be illustrated by the study of synonymous codon usage in animals and plants.

Animals↗

Three-dimensional reconstruction of temporal bone from computed tomographic scans on a personal computer.

The advantages of computer reconstruction of anatomical structures from computed tomographic scans are common knowledge by now. Unfortunately, to date most reconstructions have required the use of large computers and/or have entailed tedious manual contour tracing. The system described here allows largely automatic detection of surfaces in computed tomographic scans plus the usual display capabilities including feature selection, magnification, rotation, shading, and slicing as well as measurement of lengths and angles. It runs on a normal International Business Machines AT-compatible computer with a medium-resolution video card.

Child↗

Discontinuing therapy in childhood acute lymphocytic leukemia. A multicentric survey in Italy.

The results of discontinuing therapy in children with acute lymphocytic leukemia observed at four associated institutions are presented. Of the 247 patients who achieved complete remission, 122 (49.3%) reached the point of discontinuing therapy after 2-4 years of continuous remission. The median period off therapy was 13 months with a range of 1-69 months. Of the 122 children removed from therapy, 27 (22.1%) relapsed, mainly in the bone marrow; relapses occurred 1-32 months after cessation of therapy (median ten months) with only two relapses occurring later than two years. By actuarial analysis, 57% of the patients are projected in continuous remission after five years from cessation of therapy. Neither selected features at diagnosis nor single modalities of treatment were found to predict whether relapse would occur after discontinuing therapy. Long-term remission and possibly cure can be expected in over one-third of newly diagnosed children with ALL after 2-4 years of antileukemic treatment.

Adolescent↗

Multiresolution image registration for two-dimensional gel electrophoresis.

In proteomic research, two-dimensional electrophoresis (2-D) is an important tool for investigating differential patterns of qualitative and quantitative protein expression. The strength of the technique is due to its unrivalled power of being able to separate simultaneously thousands of proteins. The key to the comparison of 2-D protein profiles, however, lies in the use of a fast and robust image matching process which is essential to the subsequent quantification procedure. To satisfy the growing demand for a robust and fully automatic method of matching 2-D gel protein separation profiles, we describe in this paper a novel registration technique based on image intensity distribution rather than selected features. The method uses a multiresolution representation of the gel profiles and exploits the fact that coarse approximations to the optimal matching can be extracted efficiently from low-resolution images. This permits the removal of misalignments at different scales in a systematic manner and the strength of the new method has been confirmed by a double blind trial of 111 2-D gel pairs. The proposed method requires neither landmarks nor an a priori image alignment, and takes about five seconds for processing a typical gel pair on a standard personal computer.

Algorithms↗

The functional significance of tumour-associated cell surface alterations of embryonic and unknown origin.

The study of the phenotype of tumours aims to elucidate cell surface alterations that could be used for diagnostic, prognostic or therapeutic purposes. As tumours tend to escape the homeostatic growth control mechanisms of the host, it can be assumed that plasma membrane alterations are also responsible for the antisocial behaviour of tumour cells. Selected features of the transformed phenotype, of fetal or unknown origin, namely tumour-associated antigens, isozymes and growth factors, are discussed in relation to the altered growth pattern of the tumour cell. It is concluded that definitive structure-function relationships have not yet been established, but areas for future investigation are suggested.

Animals↗

Classifying "kinase inhibitor-likeness" by using machine-learning methods.

By using an in-house data set of small-molecule structures, encoded by Ghose-Crippen parameters, several machine learning techniques were applied to distinguish between kinase inhibitors and other molecules with no reported activity on any protein kinase. All four approaches pursued--support-vector machines (SVM), artificial neural networks (ANN), k nearest neighbor classification with GA-optimized feature selection (GA/kNN), and recursive partitioning (RP)--proved capable of providing a reasonable discrimination. Nevertheless, substantial differences in performance among the methods were observed. For all techniques tested, the use of a consensus vote of the 13 different models derived improved the quality of the predictions in terms of accuracy, precision, recall, and F1 value. Support-vector machines, followed by the GA/kNN combination, outperformed the other techniques when comparing the average of individual models. By using the respective majority votes, the prediction of neural networks yielded the highest F1 value, followed by SVMs.

Algorithms↗

On fully automatic feature measurement for banded chromosome classification.

Procedures for fully automatic location of chromosome axis and centromere in metaphase chromosomes are described for a practical interactive chromosome analysis system that omits the usual stages of interactive axis and centromere correction. Accuracy of centromere finding and consequential determination of a chromosome's polarity, i.e., which end is which, is measured experimentally. The saving in interaction by not correcting centromeres is compared to the increase in errors at the classification stage and the consequent increase in interaction needed to correct these errors. Some previously unreported features for banded chromosome classification are described, and in particular a set of global shape features is introduced. The discrimination capability of the feature measurements is evaluated by use of simple statistics and by reference to the performance of classifiers trained with various feature subsets. Class discrimination capability of the global shape feature set is shown to be comparable to that of centromere position, a widely used local shape feature. The variability of feature measurements that might occur in data from different laboratories on account of differing tissue, preparation methods, and digitiser hardware is assessed using three data bases of G-banded human metaphase cells. It is shown that the differences can be considerable and that appropriate feature selection and classifier training substantially improve classification performance.

Chromosomes↗

Cytology of ductal lavage fluid of the breast.

The cytologic evaluation of nipple aspirate fluids has been shown to identify women at increased risk for developing breast cancer. One limitation of this assay is the often scant cellularity of the specimen. An improved technique, ductal lavage, utilizes a microcatheter inserted into individual breast ducts to collect large numbers of cells for cytologic evaluation. Epithelial cells in ductal lavage fluids can be categorized as benign, malignant, or showing mildly or markedly atypical changes. The cell characteristics which were most helpful in identifying abnormal cells were related to cell arrangement, cell size, nuclear size, and size variation, nuclear membrane irregularity, chromatin granularity, and the presence of large nucleoli. Cell size, nuclear size variation, and large nucleoli were the most robust features, as determined by agreement between two pathologists. Moderate cell enlargement and the presence of large nucleoli were the features selected by structured tree analysis for classifying the specimens into the diagnostic groups. The similarity of the cytology of ductal lavage fluid to nipple aspirate fluid strongly suggests that these specimens will also be useful for predicting breast cancer risk.

Body Fluids↗

Prognostic classification of relapsing favorable histology Wilms tumor using cDNA microarray expression profiling and support vector machines.

Treatment of Wilms tumor has a high success rate, with some 85% of patients achieving long-term survival. However, late effects of treatment and management of relapse remain significant clinical problems. If accurate prognostic methods were available, effective risk-adapted therapies could be tailored to individual patients at diagnosis. Few molecular prognostic markers for Wilms tumor are currently defined, though previous studies have linked allele loss on 1p or 16q, genomic gain of 1q, and overexpression from 1q with an increased risk of relapse. To identify specific patterns of gene expression that are predictive of relapse, we used high-density (30 k) cDNA microarrays to analyze RNA samples from 27 favorable histology Wilms tumors taken from primary nephrectomies at the time of initial diagnosis. Thirteen of these tumors relapsed within 2 years. Genes differentially expressed between the relapsing and nonrelapsing tumor classes were identified by statistical scoring (t test). These genes encode proteins with diverse molecular functions, including transcription factors, developmental regulators, apoptotic factors, and signaling molecules. Use of a support vector machine classifier, feature selection, and test evaluation using cross-validation led to identification of a generalizable expression signature, a small subset of genes whose expression potentially can be used to predict tumor outcome in new samples. Similar methods were used to identify genes that are differentially expressed between tumors with and without genomic 1q gain. This set of discriminators was highly enriched in genes on 1q, indicating close agreement between data obtained from expression profiling with data from genomic copy number analyses.

Adolescent↗

Automated diagnosis of pigmented skin lesions.

Since advanced melanoma remains practically incurable, early detection is an important step toward a reduction in mortality. High expectations are entertained for a technique known as dermoscopy or epiluminescence light microscopy; however, evaluation of pigmented skin lesions by this method is often extremely complex and subjective. To obviate the problem of qualitative interpretation, methods based on mathematical analysis of pigmented skin lesions, such as digital dermoscopy analysis, have been developed. In the present study, we used a digital dermoscopy analyzer (DBDermo-Mips system) to evaluate a series of 588 excised, clinically atypical, flat pigmented skin lesions (371 benign, 217 malignant). The analyzer evaluated 48 parameters grouped into 4 categories (geometries, colors, textures and islands of color), which were used to train an artificial neural network. To evaluate the diagnostic performance of the neural network and to check it during the training process, we used the error area over the receiver operating characteristic curve. The discriminating power of the digital dermoscopy analyzer plus artificial neural network was compared with histologic diagnosis. A feature selection procedure indicated that as few as 13 of the variables were sufficient to discriminate the 2 groups of lesions, and this also ensured high generalization power. The artificial neural network designed with these variables enabled a diagnostic accuracy of about 94%. In conclusion, the good diagnostic performance and high speed in reading and analyzing lesions (real time) of our method constitute an important step in the direction of automated diagnosis of pigmented skin lesions.

Automation↗

Structure-activity studies of barbiturates using pattern recognition techniques.

The relationship between molecular structure and duration of depressant effect for barbiturates was investigated. A data set of 160 5,5'-disubstituted barbiturates with various acyclic substituents was coded using 47 numerical descriptors including fragments, substructures, environmental descriptors, and molecular connectivity indexes. All descriptors were derived directly from the connection tables of the barbiturates. Using an interactive error-correction feedback algorithm, linear discriminant functions were developed that could dichotomize the data set with respect to several thresholds separating longer from shorter acting compounds. Feature selection was used to focus on the relatively few structural descriptors sufficient to support linear separability. For three specific thresholds, nine, 11, and nine descriptors were sufficient. The importance of these descriptors and the utility of the technique are discussed. Predictive abilities of approximately 94% were obtained for known barbiturates of the same general molecular types.

Animals↗

A comparison of the chemical analyses of cell lipids with their complete proton NMR spectrum.

Whole cells are made up of molecules in different environments to which NMR spectroscopy is sensitive. In particular, malignant and transformed cells contain lipids not only in bilayers but in isotropically tumbling domains which give rise to high-resolution spectra. We have recently developed a technique for simultaneously analyzing broadline and high-resolution signals (M. Bloom, K. T. Holmes, C. E. Mountford, and P. G. Williams, J. Magn. Reson., in press) and we report here its application to a range of rat, mouse, and human cell lines. Some selected features of the NMR spectra were compared with the chemical analysis of the whole-cell lipid. We found that in general the proportion of protons in the narrow methylene resonance at 1.3 ppm increased with the neutral lipid content of the cells. This peak was chosen because its T2 relaxation behavior correlates with metastatic potential in a rat model system. This new technique could be applied to other high-resolution components both in healthy and in diseased states.

Animals↗

Eigenimage filtering in MR imaging: an application in the abnormal chest wall.

A postprocessing linear filter was applied to spin-echo images on 10 patients with known or suspected chest wall invasion due to bronchogenic carcinoma. This technique known as eigenimage filtering allows selective feature extraction of suspected abnormalities from conventional MR images. The final result is an image with marked increased contrast range through enhancement of a desired process (tumor) with suppression of an interfering process (e.g., normal surrounding tissue). This preliminary work demonstrates the ease with which the technique may be implemented, the contrast enhancement obtained between the desired and the interfering feature in the final eigenimage, and its ability to correct for partial volume averaging effects. Also demonstrated are artifacts that can interfere with the interpretation of the eigenimage and a method for minimizing these artifacts in the final eigenimage.

Carcinoma, Bronchogenic↗

Normal mode analysis of macromolecular motions in a database framework: developing mode concentration as a useful classifying statistic.

We investigated protein motions using normal modes within a database framework, determining on a large sample the degree to which normal modes anticipate the direction of the observed motion and were useful for motions classification. As a starting point for our analysis, we identified a large number of examples of protein flexibility from a comprehensive set of structural alignments of the proteins in the PDB. Each example consisted of a pair of proteins that were considerably different in structure given their sequence similarity. On each pair, we performed geometric comparisons and adiabatic-mapping interpolations in a high-throughput pipeline, arriving at a final list of 3,814 putative motions and standardized statistics for each. We then computed the normal modes of each motion in this list, determining the linear combination of modes that best approximated the direction of the observed motion. We integrated our new motions and normal mode calculations in the Macromolecular Motions Database, through a new ranking interface at http://molmovdb.org. Based on the normal mode calculations and the interpolations, we identified a new statistic, mode concentration, related to the mathematical concept of information content, which describes the degree to which the direction of the observed motion can be summarized by a few modes. Using this statistic, we were able to determine the fraction of the 3,814 motions where one could anticipate the direction of the actual motion from only a few modes. We also investigated mode concentration in comparison to related statistics on combinations of normal modes and correlated it with quantities characterizing protein flexibility (e.g., maximum backbone displacement or number of mobile atoms). Finally, we evaluated the ability of mode concentration to automatically classify motions into a variety of simple categories (e.g., whether or not they are "fragment-like"), in comparison to motion statistics. This involved the application of decision trees and feature selection (particular machine-learning techniques) to training and testing sets derived from merging the "list" of motions with manually classified ones.

Databases, Protein↗

Protein classification based on text document classification techniques.

The need for accurate, automated protein classification methods continues to increase as advances in biotechnology uncover new proteins. G-protein coupled receptors (GPCRs) are a particularly difficult superfamily of proteins to classify due to extreme diversity among its members. Previous comparisons of BLAST, k-nearest neighbor (k-NN), hidden markov model (HMM) and support vector machine (SVM) using alignment-based features have suggested that classifiers at the complexity of SVM are needed to attain high accuracy. Here, analogous to document classification, we applied Decision Tree and Naive Bayes classifiers with chi-square feature selection on counts of n-grams (i.e. short peptide sequences of length n) to this classification task. Using the GPCR dataset and evaluation protocol from the previous study, the Naive Bayes classifier attained an accuracy of 93.0 and 92.4% in level I and level II subfamily classification respectively, while SVM has a reported accuracy of 88.4 and 86.3%. This is a 39.7 and 44.5% reduction in residual error for level I and level II subfamily classification, respectively. The Decision Tree, while inferior to SVM, outperforms HMM in both level I and level II subfamily classification. For those GPCR families whose profiles are stored in the Protein FAMilies database of alignments and HMMs (PFAM), our method performs comparably to a search against those profiles. Finally, our method can be generalized to other protein families by applying it to the superfamily of nuclear receptors with 94.5, 97.8 and 93.6% accuracy in family, level I and level II subfamily classification respectively.

Algorithms↗