Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Acoustic-phonetic contrasts and intelligibility in the dysarthria associated with mixed cerebral palsy.

This study evaluated the relationship between specific acoustic features of speech and perceptual judgments of word intelligibility of adults with cerebral palsy-dysarthria. Use of a contrasting word task allowed for intelligibility analysis and correlated acoustic analysis according to specified spectral and temporal features. Selected phonemic contrasts included syllable-initial voicing; syllable-final voicing; stop-nasal; fricative-affricate; front-back, high-low, and tense-lax vowels. Speech materials included a set of CVC stimulus words. Acoustic data are reported on vowel duration, formant frequency locations, voice onset times, amplitude rise times, and frication durations. Listeners' perceptual assessment of intelligibility of the 16 dysarthric adults by transcription and rating tasks is also presented. All but one acoustic contrast was successfully made as evidenced by measured acoustic differences between contrast pairs. However, the generally successful acoustic contrasts stood in marked contrast to the poorly rated intelligibility scores and high error percentages that were ascribed to the opposite pair members. A second analysis examined the contribution of these acoustic features towards estimates and prediction of intelligibility deficits in speakers with dysarthria. The scaled intelligibility was predicted by multiple regression analysis with 62.6% accuracy by acoustic measures related to one consonant contrast (fricative-affricate) and three vowel contrasts (front-back, high-low, and tense-lax). Other measured contrasts, such as those related to contrast voicing effects and stop-nasal distinctions, did not seem to contribute in a significant way to variability in the intelligibility estimates. These findings are discussed in relation to specific areas of production deficiency that are consistent across different types of dysarthria with cerebral palsy as the etiology.

Adult↗

Assessing sequence comparison methods with the average precision criterion.

MOTIVATION: Comprehensive performance assessment is important for improving sequence database search methods. Sensitivity, selectivity and speed are three major yet usually conflicting evaluation criteria. The average precision (AP) measure aims to combine the sensitivity and selectivity features of a search algorithm. It can be easily visualized and extended to analyze results from a set of queries. Finally, the time-AP plot can clearly show the overall performance of different search methods. RESULTS: Experiments are performed based on the SCOP database. Popular sequence comparison algorithms, namely Smith-Waterman (SSEARCH), FASTA, BLAST and PSI-BLAST are evaluated. We find that (1) the low-complexity segment filtration procedure in BLAST actually harms its overall search quality; (2) AP scores of different search methods are approximately in proportion of the logarithm of search time; and (3) homologs in protein families with many members tend to be more obscure than those in small families. This measure may be helpful for developing new search algorithms and can guide researchers in selecting most suitable search methods. AVAILABILITY: Test sets and source code of this evaluation tool are available upon request.

Algorithms↗

Evaluation of a statistically derived decision tree for the cytodiagnosis of fine needle aspirates of the breast (FNAB).

A decision tree for the diagnosis of FNAB was derived from defined human observations using a rule induction method, C4.5 (a derivative of the ID3 algorithm). This algorithm is an implementation of the top-down induction method where the tree is determined iteratively by adding those nodes and branches which maximize the information gain at each step. The tree was derived from a training set of 200 FNAB with known outcome using 10 defined features (from one observer) and patient age. The tree contained a total of seven nodes (six observable features and patient age) with eight endpoints (four benign, four malignant). The tree was applied to a test set of 400 further FNAB with observations from the training observer and produced a sensitivity of 95%, specificity of 93% and a positive predictive value (PPV) of a malignant result of 89%. Four trainee pathologists were given a training session on the observable features and then used the tree to determine outcome in a further 50 FNAB. The observers were blind to clinical details apart from age and the endpoints were coded with letters and not labelled benign or malignant. The results from these observers produced ranges of sensitivity 80-96%, specificity 64-92%, PPV 73-92% and kappa statistics (with known outcome) 0.6-0.8. Reported difficulties in using the tree included estimation of nuclear size. These results were worse than the performance of the observers on a further 50 cases without using the decision tree (sensitivity 80-100%, specificity 72-100%, PPV 78-100%, kappa 0.72-0.92). The original 50 case test set was rerandomized and the four trainee observers made all 10 defined observations on each specimen without using the decision tree; these observations were then used to derive decisions from the tree. The performance from this method was similar to that using selected features from the tree, suggesting that observation of all features together does not improve the reliability of each specific observation. The poor performance of this tree suggests that this methodology may be unsuitable for producing decision support aids for diagnostic or training purposes in this domain.

Biopsy, Needle↗

Association of prognosis in surgically treated lung cancer patients with cytometric, histometric and ligand histochemical properties: with an emphasis on structural entropy.

OBJECTIVE: To explore new tumor features for refined category formation that permits the tailoring of individualized treatment schemes in lung cancer. STUDY DESIGN: Survival data on patients from six independent studies on cases with surgically treated lung cancer, primary lung carcinoids or metastasizing breast carcinoma (including data on primary breast carcinoma) were analyzed by nonhierarchic multivariant discriminant analysis with respect to a set of cytometric/histometric and immunohistochemical/ligand histochemical parameters. The number of stem lines, S-phase-related tumor cell fraction and the extent of structural entropy and its current were measured. In addition, the expression of binding capacities for histo-blood group trisacharides, galectins, the alpha/beta-interferon antagonist sarcolectin, the lymphokine macrophage migration inhibitory factor and a monoclonal antibody to the Le(y) epitope was monitored for insight into aspects of immunologic and biologic behavior. RESULTS: In all studies, a correlation between tumor parameters, according to TNM stage and survival, was seen. In order to refine this category formation, at least certain selected features should provide an even more stringent association than TNM stages. Indeed, statistical correlation of the cytometric and histometric parameters as well as the expression of receptors for the two histo-blood group trisaccharides, ligands for the galectins (CL-16, CL-14) and macrophage migration inhibitory factor was stronger than that of TNM stage. A large amount of the current of structural entropy was especially highly significantly associated with poor survival. This observation could be verified in each of the different studies. CONCLUSION: The obtained data strongly support the notion that thermodynamic evaluation of tumor growth focusing on the "entropy distance" of the tumor from its environment is a promising perspective warranting extended studies. Additionally, glycohistochemical features, including binding capacities for histo-blood group trisaccharides, have the potential to aid in establishment of a biologic marker set for tumor staging.

Biomarkers, Tumor↗

Transfer Learning across Material Properties Using Center-Environment Features: From Energetics to Mechanical Properties in Multicomponent Mo Alloys.

Transfer learning (TL) provides a viable approach to mitigate data scarcity in materials informatics. While conventional TL focuses on predicting identical properties across different systems, this work demonstrates a cross-property extension of TL from energy to mechanical properties via end-to-end model weight pre-training and fine-tuning: knowledge learned from predicting substitution energies is transferred to predict distinctly different mechanical properties, substantially improving computational efficiency given the typically higher cost of acquiring target-domain data. To accelerate computational alloy design, machine learning models using center-environment (CE) features were first developed to predict substitution energies of alloying elements in molybdenum (Mo)-based alloys. The Random Forest models achieved the optimal performance and transferability-R2 = 0.97, 〈MAE〉 = 0.11 eV, and 〈RMSE〉 = 0.16 eV-against the density functional theory (DFT) benchmark. The model dependency of feature selection and importance analysis was discussed. The transferability of the energy models was validated on unknown systems with new elements. Subsequently, the energy models were fine-tuned using limited mechanical property data to construct energy-to-property (E2P) TL models capable of predicting elastic properties, including bulk modulus, Young's modulus, shear modulus, and elastic constants, achieving an improved accuracy over the non-transferred ML by ∼10-30%, with its transferability verified by additional DFT calculations. This cross-property E2P transfer learning framework opens a new avenue for accelerating computational materials discovery and may be extended to other multiproperty predictions governed by similar physical principles.

center-environment feature↗

Differential distribution of Fos-like immunoreactivity in the spinal trigeminal nucleus after noxious and innocuous thermal and chemical stimulation of rat cornea.

Corneal afferent nerves project to two spatially distinct sites within the spinal trigeminal nucleus: the subnucleus interpolaris/caudalis transition and the subnucleus caudalis/upper cervical spinal cord transition. The role of these two regions in processing corneal input is uncertain. To determine if neurons in these regions encode different features of an applied corneal stimulus, immunoreactivity for the immediate early gene protein product, Fos, was quantified in barbiturate-anesthetized rats. Intensity was varied across thermal (thermal probe 5, 35, 42, 52 degrees C; radiant heat of approximately 45 degrees C) stimuli and compared with that seen after mustard oil (5 microliters, 20%) or mineral oil application. All stimuli increased the number of Fos-positive neurons located at the ventrolateral pole of the subnucleus interpolaris/caudalis transition compared with unstimulated controls. By contrast, only 52 degrees C thermal probe and mustard oil produced an additional peak of Fos-positive neurons within the superficial laminae at the subnucleus caudalis/cervical cord transition. Further, the magnitudes of the bimodal peaks of Fos produced by 52 degrees C thermal probe and mustard oil stimuli were different quantitatively. Mustard oil caused a greater Fos response at the subnucleus interpolaris/caudalis transition than 52 degrees C thermal probe stimulation, whereas the opposite was true at the subnucleus caudalis/cervical cord transition. Double-labeling revealed that Fos immunoreactive neurons within the spinal trigeminal nucleus were restricted to regions densely labeled for calcitonin gene-related peptide. These results indicate that select features of corneal stimuli such as modality are encoded differently by neurons in the trigeminal subnucleus interpolaris/caudalis transition compared with those located in the subnucleus caudalis/cervical cord transition. It is likely that neurons in these two brainstem regions subserve different aspects of corneal sensation.

Animals↗

Detecting genomic features under weak selective pressure: the example of codon usage in animals and plants.

Large scale experiments of gene inactivation in yeast have shown that 50% of genes have no detectable impact on the phenotype, and similar observations have been made in other model organisms. This apparent paradox is probably due to the fact that many genes only have a marginal contribution to the fitness of organisms. Because of the size of populations and the number of generations that can be studied in laboratories, experimental approaches only permit to detect functional elements that have a strong phenotypic impact. Comparative sequence analysis can help to solve this problem: the analysis of sequences evolution permits to detect the action of selection, and hence to reveal functional features of genomes. This approach will be illustrated by the study of synonymous codon usage in animals and plants.

Animals↗

Three-dimensional reconstruction of temporal bone from computed tomographic scans on a personal computer.

The advantages of computer reconstruction of anatomical structures from computed tomographic scans are common knowledge by now. Unfortunately, to date most reconstructions have required the use of large computers and/or have entailed tedious manual contour tracing. The system described here allows largely automatic detection of surfaces in computed tomographic scans plus the usual display capabilities including feature selection, magnification, rotation, shading, and slicing as well as measurement of lengths and angles. It runs on a normal International Business Machines AT-compatible computer with a medium-resolution video card.

Child↗

Discontinuing therapy in childhood acute lymphocytic leukemia. A multicentric survey in Italy.

The results of discontinuing therapy in children with acute lymphocytic leukemia observed at four associated institutions are presented. Of the 247 patients who achieved complete remission, 122 (49.3%) reached the point of discontinuing therapy after 2-4 years of continuous remission. The median period off therapy was 13 months with a range of 1-69 months. Of the 122 children removed from therapy, 27 (22.1%) relapsed, mainly in the bone marrow; relapses occurred 1-32 months after cessation of therapy (median ten months) with only two relapses occurring later than two years. By actuarial analysis, 57% of the patients are projected in continuous remission after five years from cessation of therapy. Neither selected features at diagnosis nor single modalities of treatment were found to predict whether relapse would occur after discontinuing therapy. Long-term remission and possibly cure can be expected in over one-third of newly diagnosed children with ALL after 2-4 years of antileukemic treatment.

Adolescent↗

Multiresolution image registration for two-dimensional gel electrophoresis.

In proteomic research, two-dimensional electrophoresis (2-D) is an important tool for investigating differential patterns of qualitative and quantitative protein expression. The strength of the technique is due to its unrivalled power of being able to separate simultaneously thousands of proteins. The key to the comparison of 2-D protein profiles, however, lies in the use of a fast and robust image matching process which is essential to the subsequent quantification procedure. To satisfy the growing demand for a robust and fully automatic method of matching 2-D gel protein separation profiles, we describe in this paper a novel registration technique based on image intensity distribution rather than selected features. The method uses a multiresolution representation of the gel profiles and exploits the fact that coarse approximations to the optimal matching can be extracted efficiently from low-resolution images. This permits the removal of misalignments at different scales in a systematic manner and the strength of the new method has been confirmed by a double blind trial of 111 2-D gel pairs. The proposed method requires neither landmarks nor an a priori image alignment, and takes about five seconds for processing a typical gel pair on a standard personal computer.

Algorithms↗

The functional significance of tumour-associated cell surface alterations of embryonic and unknown origin.

The study of the phenotype of tumours aims to elucidate cell surface alterations that could be used for diagnostic, prognostic or therapeutic purposes. As tumours tend to escape the homeostatic growth control mechanisms of the host, it can be assumed that plasma membrane alterations are also responsible for the antisocial behaviour of tumour cells. Selected features of the transformed phenotype, of fetal or unknown origin, namely tumour-associated antigens, isozymes and growth factors, are discussed in relation to the altered growth pattern of the tumour cell. It is concluded that definitive structure-function relationships have not yet been established, but areas for future investigation are suggested.

Animals↗

Classifying "kinase inhibitor-likeness" by using machine-learning methods.

By using an in-house data set of small-molecule structures, encoded by Ghose-Crippen parameters, several machine learning techniques were applied to distinguish between kinase inhibitors and other molecules with no reported activity on any protein kinase. All four approaches pursued--support-vector machines (SVM), artificial neural networks (ANN), k nearest neighbor classification with GA-optimized feature selection (GA/kNN), and recursive partitioning (RP)--proved capable of providing a reasonable discrimination. Nevertheless, substantial differences in performance among the methods were observed. For all techniques tested, the use of a consensus vote of the 13 different models derived improved the quality of the predictions in terms of accuracy, precision, recall, and F1 value. Support-vector machines, followed by the GA/kNN combination, outperformed the other techniques when comparing the average of individual models. By using the respective majority votes, the prediction of neural networks yielded the highest F1 value, followed by SVMs.

Algorithms↗

On fully automatic feature measurement for banded chromosome classification.

Procedures for fully automatic location of chromosome axis and centromere in metaphase chromosomes are described for a practical interactive chromosome analysis system that omits the usual stages of interactive axis and centromere correction. Accuracy of centromere finding and consequential determination of a chromosome's polarity, i.e., which end is which, is measured experimentally. The saving in interaction by not correcting centromeres is compared to the increase in errors at the classification stage and the consequent increase in interaction needed to correct these errors. Some previously unreported features for banded chromosome classification are described, and in particular a set of global shape features is introduced. The discrimination capability of the feature measurements is evaluated by use of simple statistics and by reference to the performance of classifiers trained with various feature subsets. Class discrimination capability of the global shape feature set is shown to be comparable to that of centromere position, a widely used local shape feature. The variability of feature measurements that might occur in data from different laboratories on account of differing tissue, preparation methods, and digitiser hardware is assessed using three data bases of G-banded human metaphase cells. It is shown that the differences can be considerable and that appropriate feature selection and classifier training substantially improve classification performance.

Chromosomes↗

Cytology of ductal lavage fluid of the breast.

The cytologic evaluation of nipple aspirate fluids has been shown to identify women at increased risk for developing breast cancer. One limitation of this assay is the often scant cellularity of the specimen. An improved technique, ductal lavage, utilizes a microcatheter inserted into individual breast ducts to collect large numbers of cells for cytologic evaluation. Epithelial cells in ductal lavage fluids can be categorized as benign, malignant, or showing mildly or markedly atypical changes. The cell characteristics which were most helpful in identifying abnormal cells were related to cell arrangement, cell size, nuclear size, and size variation, nuclear membrane irregularity, chromatin granularity, and the presence of large nucleoli. Cell size, nuclear size variation, and large nucleoli were the most robust features, as determined by agreement between two pathologists. Moderate cell enlargement and the presence of large nucleoli were the features selected by structured tree analysis for classifying the specimens into the diagnostic groups. The similarity of the cytology of ductal lavage fluid to nipple aspirate fluid strongly suggests that these specimens will also be useful for predicting breast cancer risk.

Body Fluids↗

Prognostic classification of relapsing favorable histology Wilms tumor using cDNA microarray expression profiling and support vector machines.

Treatment of Wilms tumor has a high success rate, with some 85% of patients achieving long-term survival. However, late effects of treatment and management of relapse remain significant clinical problems. If accurate prognostic methods were available, effective risk-adapted therapies could be tailored to individual patients at diagnosis. Few molecular prognostic markers for Wilms tumor are currently defined, though previous studies have linked allele loss on 1p or 16q, genomic gain of 1q, and overexpression from 1q with an increased risk of relapse. To identify specific patterns of gene expression that are predictive of relapse, we used high-density (30 k) cDNA microarrays to analyze RNA samples from 27 favorable histology Wilms tumors taken from primary nephrectomies at the time of initial diagnosis. Thirteen of these tumors relapsed within 2 years. Genes differentially expressed between the relapsing and nonrelapsing tumor classes were identified by statistical scoring (t test). These genes encode proteins with diverse molecular functions, including transcription factors, developmental regulators, apoptotic factors, and signaling molecules. Use of a support vector machine classifier, feature selection, and test evaluation using cross-validation led to identification of a generalizable expression signature, a small subset of genes whose expression potentially can be used to predict tumor outcome in new samples. Similar methods were used to identify genes that are differentially expressed between tumors with and without genomic 1q gain. This set of discriminators was highly enriched in genes on 1q, indicating close agreement between data obtained from expression profiling with data from genomic copy number analyses.

Adolescent↗

Automated diagnosis of pigmented skin lesions.

Since advanced melanoma remains practically incurable, early detection is an important step toward a reduction in mortality. High expectations are entertained for a technique known as dermoscopy or epiluminescence light microscopy; however, evaluation of pigmented skin lesions by this method is often extremely complex and subjective. To obviate the problem of qualitative interpretation, methods based on mathematical analysis of pigmented skin lesions, such as digital dermoscopy analysis, have been developed. In the present study, we used a digital dermoscopy analyzer (DBDermo-Mips system) to evaluate a series of 588 excised, clinically atypical, flat pigmented skin lesions (371 benign, 217 malignant). The analyzer evaluated 48 parameters grouped into 4 categories (geometries, colors, textures and islands of color), which were used to train an artificial neural network. To evaluate the diagnostic performance of the neural network and to check it during the training process, we used the error area over the receiver operating characteristic curve. The discriminating power of the digital dermoscopy analyzer plus artificial neural network was compared with histologic diagnosis. A feature selection procedure indicated that as few as 13 of the variables were sufficient to discriminate the 2 groups of lesions, and this also ensured high generalization power. The artificial neural network designed with these variables enabled a diagnostic accuracy of about 94%. In conclusion, the good diagnostic performance and high speed in reading and analyzing lesions (real time) of our method constitute an important step in the direction of automated diagnosis of pigmented skin lesions.

Automation↗

Structure-activity studies of barbiturates using pattern recognition techniques.

The relationship between molecular structure and duration of depressant effect for barbiturates was investigated. A data set of 160 5,5'-disubstituted barbiturates with various acyclic substituents was coded using 47 numerical descriptors including fragments, substructures, environmental descriptors, and molecular connectivity indexes. All descriptors were derived directly from the connection tables of the barbiturates. Using an interactive error-correction feedback algorithm, linear discriminant functions were developed that could dichotomize the data set with respect to several thresholds separating longer from shorter acting compounds. Feature selection was used to focus on the relatively few structural descriptors sufficient to support linear separability. For three specific thresholds, nine, 11, and nine descriptors were sufficient. The importance of these descriptors and the utility of the technique are discussed. Predictive abilities of approximately 94% were obtained for known barbiturates of the same general molecular types.

Animals↗