Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Transfer Learning across Material Properties Using Center-Environment Features: From Energetics to Mechanical Properties in Multicomponent Mo Alloys.

Transfer learning (TL) provides a viable approach to mitigate data scarcity in materials informatics. While conventional TL focuses on predicting identical properties across different systems, this work demonstrates a cross-property extension of TL from energy to mechanical properties via end-to-end model weight pre-training and fine-tuning: knowledge learned from predicting substitution energies is transferred to predict distinctly different mechanical properties, substantially improving computational efficiency given the typically higher cost of acquiring target-domain data. To accelerate computational alloy design, machine learning models using center-environment (CE) features were first developed to predict substitution energies of alloying elements in molybdenum (Mo)-based alloys. The Random Forest models achieved the optimal performance and transferability-R2 = 0.97, 〈MAE〉 = 0.11 eV, and 〈RMSE〉 = 0.16 eV-against the density functional theory (DFT) benchmark. The model dependency of feature selection and importance analysis was discussed. The transferability of the energy models was validated on unknown systems with new elements. Subsequently, the energy models were fine-tuned using limited mechanical property data to construct energy-to-property (E2P) TL models capable of predicting elastic properties, including bulk modulus, Young's modulus, shear modulus, and elastic constants, achieving an improved accuracy over the non-transferred ML by ∼10-30%, with its transferability verified by additional DFT calculations. This cross-property E2P transfer learning framework opens a new avenue for accelerating computational materials discovery and may be extended to other multiproperty predictions governed by similar physical principles.

center-environment feature↗

Differential distribution of Fos-like immunoreactivity in the spinal trigeminal nucleus after noxious and innocuous thermal and chemical stimulation of rat cornea.

Corneal afferent nerves project to two spatially distinct sites within the spinal trigeminal nucleus: the subnucleus interpolaris/caudalis transition and the subnucleus caudalis/upper cervical spinal cord transition. The role of these two regions in processing corneal input is uncertain. To determine if neurons in these regions encode different features of an applied corneal stimulus, immunoreactivity for the immediate early gene protein product, Fos, was quantified in barbiturate-anesthetized rats. Intensity was varied across thermal (thermal probe 5, 35, 42, 52 degrees C; radiant heat of approximately 45 degrees C) stimuli and compared with that seen after mustard oil (5 microliters, 20%) or mineral oil application. All stimuli increased the number of Fos-positive neurons located at the ventrolateral pole of the subnucleus interpolaris/caudalis transition compared with unstimulated controls. By contrast, only 52 degrees C thermal probe and mustard oil produced an additional peak of Fos-positive neurons within the superficial laminae at the subnucleus caudalis/cervical cord transition. Further, the magnitudes of the bimodal peaks of Fos produced by 52 degrees C thermal probe and mustard oil stimuli were different quantitatively. Mustard oil caused a greater Fos response at the subnucleus interpolaris/caudalis transition than 52 degrees C thermal probe stimulation, whereas the opposite was true at the subnucleus caudalis/cervical cord transition. Double-labeling revealed that Fos immunoreactive neurons within the spinal trigeminal nucleus were restricted to regions densely labeled for calcitonin gene-related peptide. These results indicate that select features of corneal stimuli such as modality are encoded differently by neurons in the trigeminal subnucleus interpolaris/caudalis transition compared with those located in the subnucleus caudalis/cervical cord transition. It is likely that neurons in these two brainstem regions subserve different aspects of corneal sensation.

Animals↗

Detecting genomic features under weak selective pressure: the example of codon usage in animals and plants.

Large scale experiments of gene inactivation in yeast have shown that 50% of genes have no detectable impact on the phenotype, and similar observations have been made in other model organisms. This apparent paradox is probably due to the fact that many genes only have a marginal contribution to the fitness of organisms. Because of the size of populations and the number of generations that can be studied in laboratories, experimental approaches only permit to detect functional elements that have a strong phenotypic impact. Comparative sequence analysis can help to solve this problem: the analysis of sequences evolution permits to detect the action of selection, and hence to reveal functional features of genomes. This approach will be illustrated by the study of synonymous codon usage in animals and plants.

Animals↗

Three-dimensional reconstruction of temporal bone from computed tomographic scans on a personal computer.

The advantages of computer reconstruction of anatomical structures from computed tomographic scans are common knowledge by now. Unfortunately, to date most reconstructions have required the use of large computers and/or have entailed tedious manual contour tracing. The system described here allows largely automatic detection of surfaces in computed tomographic scans plus the usual display capabilities including feature selection, magnification, rotation, shading, and slicing as well as measurement of lengths and angles. It runs on a normal International Business Machines AT-compatible computer with a medium-resolution video card.

Child↗

Discontinuing therapy in childhood acute lymphocytic leukemia. A multicentric survey in Italy.

The results of discontinuing therapy in children with acute lymphocytic leukemia observed at four associated institutions are presented. Of the 247 patients who achieved complete remission, 122 (49.3%) reached the point of discontinuing therapy after 2-4 years of continuous remission. The median period off therapy was 13 months with a range of 1-69 months. Of the 122 children removed from therapy, 27 (22.1%) relapsed, mainly in the bone marrow; relapses occurred 1-32 months after cessation of therapy (median ten months) with only two relapses occurring later than two years. By actuarial analysis, 57% of the patients are projected in continuous remission after five years from cessation of therapy. Neither selected features at diagnosis nor single modalities of treatment were found to predict whether relapse would occur after discontinuing therapy. Long-term remission and possibly cure can be expected in over one-third of newly diagnosed children with ALL after 2-4 years of antileukemic treatment.

Adolescent↗

Multiresolution image registration for two-dimensional gel electrophoresis.

In proteomic research, two-dimensional electrophoresis (2-D) is an important tool for investigating differential patterns of qualitative and quantitative protein expression. The strength of the technique is due to its unrivalled power of being able to separate simultaneously thousands of proteins. The key to the comparison of 2-D protein profiles, however, lies in the use of a fast and robust image matching process which is essential to the subsequent quantification procedure. To satisfy the growing demand for a robust and fully automatic method of matching 2-D gel protein separation profiles, we describe in this paper a novel registration technique based on image intensity distribution rather than selected features. The method uses a multiresolution representation of the gel profiles and exploits the fact that coarse approximations to the optimal matching can be extracted efficiently from low-resolution images. This permits the removal of misalignments at different scales in a systematic manner and the strength of the new method has been confirmed by a double blind trial of 111 2-D gel pairs. The proposed method requires neither landmarks nor an a priori image alignment, and takes about five seconds for processing a typical gel pair on a standard personal computer.

Algorithms↗

The functional significance of tumour-associated cell surface alterations of embryonic and unknown origin.

The study of the phenotype of tumours aims to elucidate cell surface alterations that could be used for diagnostic, prognostic or therapeutic purposes. As tumours tend to escape the homeostatic growth control mechanisms of the host, it can be assumed that plasma membrane alterations are also responsible for the antisocial behaviour of tumour cells. Selected features of the transformed phenotype, of fetal or unknown origin, namely tumour-associated antigens, isozymes and growth factors, are discussed in relation to the altered growth pattern of the tumour cell. It is concluded that definitive structure-function relationships have not yet been established, but areas for future investigation are suggested.

Animals↗

On fully automatic feature measurement for banded chromosome classification.

Procedures for fully automatic location of chromosome axis and centromere in metaphase chromosomes are described for a practical interactive chromosome analysis system that omits the usual stages of interactive axis and centromere correction. Accuracy of centromere finding and consequential determination of a chromosome's polarity, i.e., which end is which, is measured experimentally. The saving in interaction by not correcting centromeres is compared to the increase in errors at the classification stage and the consequent increase in interaction needed to correct these errors. Some previously unreported features for banded chromosome classification are described, and in particular a set of global shape features is introduced. The discrimination capability of the feature measurements is evaluated by use of simple statistics and by reference to the performance of classifiers trained with various feature subsets. Class discrimination capability of the global shape feature set is shown to be comparable to that of centromere position, a widely used local shape feature. The variability of feature measurements that might occur in data from different laboratories on account of differing tissue, preparation methods, and digitiser hardware is assessed using three data bases of G-banded human metaphase cells. It is shown that the differences can be considerable and that appropriate feature selection and classifier training substantially improve classification performance.

Chromosomes↗

Automated diagnosis of pigmented skin lesions.

Since advanced melanoma remains practically incurable, early detection is an important step toward a reduction in mortality. High expectations are entertained for a technique known as dermoscopy or epiluminescence light microscopy; however, evaluation of pigmented skin lesions by this method is often extremely complex and subjective. To obviate the problem of qualitative interpretation, methods based on mathematical analysis of pigmented skin lesions, such as digital dermoscopy analysis, have been developed. In the present study, we used a digital dermoscopy analyzer (DBDermo-Mips system) to evaluate a series of 588 excised, clinically atypical, flat pigmented skin lesions (371 benign, 217 malignant). The analyzer evaluated 48 parameters grouped into 4 categories (geometries, colors, textures and islands of color), which were used to train an artificial neural network. To evaluate the diagnostic performance of the neural network and to check it during the training process, we used the error area over the receiver operating characteristic curve. The discriminating power of the digital dermoscopy analyzer plus artificial neural network was compared with histologic diagnosis. A feature selection procedure indicated that as few as 13 of the variables were sufficient to discriminate the 2 groups of lesions, and this also ensured high generalization power. The artificial neural network designed with these variables enabled a diagnostic accuracy of about 94%. In conclusion, the good diagnostic performance and high speed in reading and analyzing lesions (real time) of our method constitute an important step in the direction of automated diagnosis of pigmented skin lesions.

Automation↗

Structure-activity studies of barbiturates using pattern recognition techniques.

The relationship between molecular structure and duration of depressant effect for barbiturates was investigated. A data set of 160 5,5'-disubstituted barbiturates with various acyclic substituents was coded using 47 numerical descriptors including fragments, substructures, environmental descriptors, and molecular connectivity indexes. All descriptors were derived directly from the connection tables of the barbiturates. Using an interactive error-correction feedback algorithm, linear discriminant functions were developed that could dichotomize the data set with respect to several thresholds separating longer from shorter acting compounds. Feature selection was used to focus on the relatively few structural descriptors sufficient to support linear separability. For three specific thresholds, nine, 11, and nine descriptors were sufficient. The importance of these descriptors and the utility of the technique are discussed. Predictive abilities of approximately 94% were obtained for known barbiturates of the same general molecular types.

Animals↗

A comparison of the chemical analyses of cell lipids with their complete proton NMR spectrum.

Whole cells are made up of molecules in different environments to which NMR spectroscopy is sensitive. In particular, malignant and transformed cells contain lipids not only in bilayers but in isotropically tumbling domains which give rise to high-resolution spectra. We have recently developed a technique for simultaneously analyzing broadline and high-resolution signals (M. Bloom, K. T. Holmes, C. E. Mountford, and P. G. Williams, J. Magn. Reson., in press) and we report here its application to a range of rat, mouse, and human cell lines. Some selected features of the NMR spectra were compared with the chemical analysis of the whole-cell lipid. We found that in general the proportion of protons in the narrow methylene resonance at 1.3 ppm increased with the neutral lipid content of the cells. This peak was chosen because its T2 relaxation behavior correlates with metastatic potential in a rat model system. This new technique could be applied to other high-resolution components both in healthy and in diseased states.

Animals↗

Eigenimage filtering in MR imaging: an application in the abnormal chest wall.

A postprocessing linear filter was applied to spin-echo images on 10 patients with known or suspected chest wall invasion due to bronchogenic carcinoma. This technique known as eigenimage filtering allows selective feature extraction of suspected abnormalities from conventional MR images. The final result is an image with marked increased contrast range through enhancement of a desired process (tumor) with suppression of an interfering process (e.g., normal surrounding tissue). This preliminary work demonstrates the ease with which the technique may be implemented, the contrast enhancement obtained between the desired and the interfering feature in the final eigenimage, and its ability to correct for partial volume averaging effects. Also demonstrated are artifacts that can interfere with the interpretation of the eigenimage and a method for minimizing these artifacts in the final eigenimage.

Carcinoma, Bronchogenic↗

Normal mode analysis of macromolecular motions in a database framework: developing mode concentration as a useful classifying statistic.

We investigated protein motions using normal modes within a database framework, determining on a large sample the degree to which normal modes anticipate the direction of the observed motion and were useful for motions classification. As a starting point for our analysis, we identified a large number of examples of protein flexibility from a comprehensive set of structural alignments of the proteins in the PDB. Each example consisted of a pair of proteins that were considerably different in structure given their sequence similarity. On each pair, we performed geometric comparisons and adiabatic-mapping interpolations in a high-throughput pipeline, arriving at a final list of 3,814 putative motions and standardized statistics for each. We then computed the normal modes of each motion in this list, determining the linear combination of modes that best approximated the direction of the observed motion. We integrated our new motions and normal mode calculations in the Macromolecular Motions Database, through a new ranking interface at http://molmovdb.org. Based on the normal mode calculations and the interpolations, we identified a new statistic, mode concentration, related to the mathematical concept of information content, which describes the degree to which the direction of the observed motion can be summarized by a few modes. Using this statistic, we were able to determine the fraction of the 3,814 motions where one could anticipate the direction of the actual motion from only a few modes. We also investigated mode concentration in comparison to related statistics on combinations of normal modes and correlated it with quantities characterizing protein flexibility (e.g., maximum backbone displacement or number of mobile atoms). Finally, we evaluated the ability of mode concentration to automatically classify motions into a variety of simple categories (e.g., whether or not they are "fragment-like"), in comparison to motion statistics. This involved the application of decision trees and feature selection (particular machine-learning techniques) to training and testing sets derived from merging the "list" of motions with manually classified ones.

Databases, Protein↗

Hierarchical Multi-Label Classification With Gene-Environment Interactions in Disease Modeling.

In biomedical studies, gene-environment (G-E) interactions have been demonstrated to have important implications for analyzing disease outcomes beyond the main G and main E effects. Many approaches have been developed for G-E interaction analysis, yielding important findings. However, hierarchical multi-label classification, which provides insightful information on disease outcomes, remains unexplored in G-E analysis literature. Moreover, unlabeled data are commonly observed in practical settings but omitted by many existing methods of hierarchical multi-label classification. In this study, we consider a semi-supervised scenario and develop a novel approach for the two-layer hierarchical response with G-E interactions. A two-step penalized estimation is then proposed using an efficient expectation-maximization (EM) algorithm. Simulation shows that it has superior performance in classification and feature selection. The analysis of The Cancer Genome Atlas (TCGA) data on lung cancer demonstrates the practical utility of the proposed method. Overall, this study can fill the important knowledge gap in G-E interaction analysis by providing a widely applicable framework for hierarchical multi-label classification of complex disease outcomes.

Humans↗

The relation of ERP components to complex memory processing.

The relation between various ERP components generated during encoding of a word and its subsequent recall were investigated using a "rote" serial-order and an "elaborative" category memory task. Words (flashed separately) were time-locked to EEG recordings from 21 cortical sites. ERP components from the five subjects having the highest recall scores were compared to the five lowest scoring subjects. Results based on the P200 peak amplitude data as well as the N400 and late positive component peak amplitude and latency data suggest that anterior and posterior distributional differences are elicited during encoding of words for rote and elaborative memory tasks. Furthermore, strong individual differences in these patterns were found as a function of task. A tentative argument was made that the obtained anterior and posterior differences may index different word feature selection and encoding processes, which are differentially utilized by high and low recallers.

Adult↗

Increased apolipoprotein E mRNA in the hippocampus in Alzheimer disease and in rats after entorhinal cortex lesioning.

The distribution of apolipoprotein E (ApoE) mRNA was characterized in the hippocampus of humans with Alzheimer disease (AD) and in rats with experimental lesions (unilateral ablation of the entorhinal cortex) that model selected features of AD. In both AD and the lesion model, we observed a shift in the location of astrocytes containing prevalent ApoE mRNA from the neuropil to regions with densely packed neurons. The increased abundance of ApoE mRNA in astrocytes close to neuron cell bodies could be indicative of lipid uptake in regions where neurons are degenerating or where synaptic remodeling is taking place.

Alzheimer Disease↗

Modular bacterial artificial chromosome vectors for transfer of large inserts into mammalian cells.

To facilitate the use of large-insert bacterial clones for functional analysis, we have constructed new bacterial artificial chromosome vectors, pPAC4 and pBACe4. These vectors contain two genetic elements that enable stable maintenance of the clones in mammalian cells: (1) The Epstein-Barr virus replicon, oriP, is included to ensure stable episomal propagation of the large insert clones upon transfection into mammalian cells. (2) The blasticidin deaminase gene is placed in a eukaryotic expression cassette to enable selection for the desired mammalian clones by using the nucleoside antibiotic blasticidin. Sequences important to select for loxP-specific genome targeting in mammalian chromosomes are also present. In addition, we demonstrate that the attTn7 sequence present on the vectors permits specific addition of selected features to the library clones. Unique sites have also been included in the vector to enable linearization of the large-insert clones, e. g., for optical mapping studies. The pPAC4 vector has been used to generate libraries from the human, mouse, and rat genomes. We believe that clones from these libraries would serve as an important reagent in functional experiments, including the identification or validation of candidate disease genes, by transferring a particular clone containing the relevant wildtype gene into mutant cells or transgenic or knock-out animals.

Animals↗