Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Automatic detection of calcifications in the aorta from CT scans of the abdomen. 3D computer-aided diagnosis.

RATIONALE AND OBJECTIVES: Automated detection and quantification of arterial calcifications can facilitate epidemiologic research and, eventually, the use of full-body calcium scoring in clinical practice. An automatic computerized method to detect calcifications in CT scans is presented. MATERIALS AND METHODS: Forty abdominal CT scans have been randomly selected from clinical practice. They all contained contrast material and belonged to one of four categories: containing "no," "small," "moderate," or "large" amounts of arterial calcification. There were ten scans in each category. The experiments were restricted to the vertical range from the point where the superior mesenteric artery branches off of the descending aorta until the first bifurcation of the iliac arteries. The automatic method starts by extracting all connected objects above 220 Hounsfield units (HU) from the scan. These objects include all calcifications, as well as bony structures and contrast material. To distinguish calcifications from non-calcifications, a number of features are calculated for each object. These features are based on the object's size, location, shape characteristics, and surrounding structures. Subsequently a classification of each object is performed in two stages. First the probability that an object represents a calcification is computed assuming a multivariate Gaussian distribution for the calcifications. Objects with low probability are discarded. The remaining objects are then classified into calcifications and non-calcifications using a 5-nearest-neighbor classifier and sequential forward feature selection. Based on the total volume of calcifications determined by the system, the scan is assigned to one of the four categories mentioned above. RESULTS: The 40 scans contained a total of 249 calcifications as determined by a human observer. The method detected 209 calcifications (sensitivity 83.9%) at the expense of on average 1.0 false-positive object per scan. The correct category label was assigned to 30 scans and only 2 scans were off by more than one category. Most incorrect classifications can be attributed to the presence of contrast material in the scans. CONCLUSION: It is possible to identify the majority of arterial calcifications in abdominal CT scans in a completely automatic fashion with few false positive objects, even if the scans contain contrast material.

Aorta, Abdominal↗

CoxKAN: Kolmogorov-Arnold networks for interpretable, high-performance survival analysis.

MOTIVATION: Survival analysis is a branch of statistics that is crucial in medicine for modeling the time to critical events such as death or relapse, in order to improve treatment strategies and patient outcomes. Selecting survival models often involves a trade-off between performance and interpretability; deep learning models offer high performance but lack the transparency of more traditional approaches. This poses a significant issue in medicine, where practitioners are reluctant to use black-box models for critical patient decisions. RESULTS: We introduce CoxKAN, a Cox proportional hazards Kolmogorov-Arnold Network for interpretable, high-performance survival analysis. Kolmogorov-Arnold Networks (KANs) were recently proposed as an interpretable and accurate alternative to multi-layer perceptrons. We evaluated CoxKAN on four synthetic and nine real datasets, including five cohorts with clinical data and four with genomics biomarkers. In synthetic experiments, CoxKAN accurately recovered interpretable hazard function formulae and excelled in automatic feature selection. Evaluations on real datasets showed that CoxKAN consistently outperformed the traditional Cox proportional hazards model (by up to 4% in C-index) and matched or surpassed the performance of deep learning-based models. Importantly, CoxKAN revealed complex interactions between predictor variables and uncovered symbolic formulae, which are key capabilities that other survival analysis methods lack, to provide clear insights into the impact of key biomarkers on patient risk. AVAILABILITY AND IMPLEMENTATION: CoxKAN is available at GitHub and Zenodo.

Humans↗

Identification of amino acid residues lining the pore of a gap junction channel.

Gap junctions represent a ubiquitous and integral part of multicellular organisms, providing the only conduit for direct exchange of nutrients, messengers and ions between neighboring cells. However, at the molecular level we have limited knowledge of their endogenous permeants and selectivity features. By probing the accessibility of systematically substituted cysteine residues to thiol blockers (a technique called SCAM), we have identified the pore-lining residues of a gap junction channel composed of Cx32. Analysis of 45 sites in perfused Xenopus oocyte pairs defined M3 as the major pore-lining helix, with M2 (open state) or M1 (closed state) also contributing to the wider cytoplasmic opening of the channel. Additional mapping of a close association between M3 and M4 allowed the helices of the low resolution map (Unger et al., 1999. Science. 283:1176-1180) to be tentatively assigned to the connexin transmembrane domains. Contrary to previous conceptions of the gap junction channel, the residues lining the pore are largely hydrophobic. This indicates that the selective permeabilities of this unique channel class may result from novel mechanisms, including complex van der Waals interactions of permeants with the pore wall, rather than mechanisms involving fixed charges or chelation chemistry as reported for other ion channels.

Amino Acid Sequence↗

Biological factors predisposing to traumatic posterior dislocation of the hip. A selection process in the mechanism of injury.

The factors involved in the mechanism leading to traumatic posterior dislocation of the hip are examined. In 47 adult patients who had previously suffered such a dislocation, ultrasound scans were used to measure femoral anteversion on both the affected and the uninjured side. In 36 normal adult volunteers, used as controls, similar measurements were made. Femoral anteversion on both the injured and uninjured side was significantly reduced in the patients compared with the volunteers. These findings are discussed in the light of previous work which indicates that medial rotation is a factor in the mechanism of posterior dislocation of the hip. It is suggested that reduced anteversion acts like medial rotation to make the hip more susceptible to posterior dislocation, and that the less the anteversion the more likely is the injury to be a dislocation rather than a fracture-dislocation. It is concluded that patients who suffer such dislocated hips belong at one extreme of the normal population, having either reduced femoral anteversion or even retroversion, and that this anatomical feature selects towards hip dislocation rather than to injury of the femoral shaft, knee or tibia during the appropriate type of accident.

Adult↗

Acoustic-phonetic contrasts and intelligibility in the dysarthria associated with mixed cerebral palsy.

This study evaluated the relationship between specific acoustic features of speech and perceptual judgments of word intelligibility of adults with cerebral palsy-dysarthria. Use of a contrasting word task allowed for intelligibility analysis and correlated acoustic analysis according to specified spectral and temporal features. Selected phonemic contrasts included syllable-initial voicing; syllable-final voicing; stop-nasal; fricative-affricate; front-back, high-low, and tense-lax vowels. Speech materials included a set of CVC stimulus words. Acoustic data are reported on vowel duration, formant frequency locations, voice onset times, amplitude rise times, and frication durations. Listeners' perceptual assessment of intelligibility of the 16 dysarthric adults by transcription and rating tasks is also presented. All but one acoustic contrast was successfully made as evidenced by measured acoustic differences between contrast pairs. However, the generally successful acoustic contrasts stood in marked contrast to the poorly rated intelligibility scores and high error percentages that were ascribed to the opposite pair members. A second analysis examined the contribution of these acoustic features towards estimates and prediction of intelligibility deficits in speakers with dysarthria. The scaled intelligibility was predicted by multiple regression analysis with 62.6% accuracy by acoustic measures related to one consonant contrast (fricative-affricate) and three vowel contrasts (front-back, high-low, and tense-lax). Other measured contrasts, such as those related to contrast voicing effects and stop-nasal distinctions, did not seem to contribute in a significant way to variability in the intelligibility estimates. These findings are discussed in relation to specific areas of production deficiency that are consistent across different types of dysarthria with cerebral palsy as the etiology.

Adult↗

Assessing sequence comparison methods with the average precision criterion.

MOTIVATION: Comprehensive performance assessment is important for improving sequence database search methods. Sensitivity, selectivity and speed are three major yet usually conflicting evaluation criteria. The average precision (AP) measure aims to combine the sensitivity and selectivity features of a search algorithm. It can be easily visualized and extended to analyze results from a set of queries. Finally, the time-AP plot can clearly show the overall performance of different search methods. RESULTS: Experiments are performed based on the SCOP database. Popular sequence comparison algorithms, namely Smith-Waterman (SSEARCH), FASTA, BLAST and PSI-BLAST are evaluated. We find that (1) the low-complexity segment filtration procedure in BLAST actually harms its overall search quality; (2) AP scores of different search methods are approximately in proportion of the logarithm of search time; and (3) homologs in protein families with many members tend to be more obscure than those in small families. This measure may be helpful for developing new search algorithms and can guide researchers in selecting most suitable search methods. AVAILABILITY: Test sets and source code of this evaluation tool are available upon request.

Algorithms↗

Evaluation of a statistically derived decision tree for the cytodiagnosis of fine needle aspirates of the breast (FNAB).

A decision tree for the diagnosis of FNAB was derived from defined human observations using a rule induction method, C4.5 (a derivative of the ID3 algorithm). This algorithm is an implementation of the top-down induction method where the tree is determined iteratively by adding those nodes and branches which maximize the information gain at each step. The tree was derived from a training set of 200 FNAB with known outcome using 10 defined features (from one observer) and patient age. The tree contained a total of seven nodes (six observable features and patient age) with eight endpoints (four benign, four malignant). The tree was applied to a test set of 400 further FNAB with observations from the training observer and produced a sensitivity of 95%, specificity of 93% and a positive predictive value (PPV) of a malignant result of 89%. Four trainee pathologists were given a training session on the observable features and then used the tree to determine outcome in a further 50 FNAB. The observers were blind to clinical details apart from age and the endpoints were coded with letters and not labelled benign or malignant. The results from these observers produced ranges of sensitivity 80-96%, specificity 64-92%, PPV 73-92% and kappa statistics (with known outcome) 0.6-0.8. Reported difficulties in using the tree included estimation of nuclear size. These results were worse than the performance of the observers on a further 50 cases without using the decision tree (sensitivity 80-100%, specificity 72-100%, PPV 78-100%, kappa 0.72-0.92). The original 50 case test set was rerandomized and the four trainee observers made all 10 defined observations on each specimen without using the decision tree; these observations were then used to derive decisions from the tree. The performance from this method was similar to that using selected features from the tree, suggesting that observation of all features together does not improve the reliability of each specific observation. The poor performance of this tree suggests that this methodology may be unsuitable for producing decision support aids for diagnostic or training purposes in this domain.

Biopsy, Needle↗

Association of prognosis in surgically treated lung cancer patients with cytometric, histometric and ligand histochemical properties: with an emphasis on structural entropy.

OBJECTIVE: To explore new tumor features for refined category formation that permits the tailoring of individualized treatment schemes in lung cancer. STUDY DESIGN: Survival data on patients from six independent studies on cases with surgically treated lung cancer, primary lung carcinoids or metastasizing breast carcinoma (including data on primary breast carcinoma) were analyzed by nonhierarchic multivariant discriminant analysis with respect to a set of cytometric/histometric and immunohistochemical/ligand histochemical parameters. The number of stem lines, S-phase-related tumor cell fraction and the extent of structural entropy and its current were measured. In addition, the expression of binding capacities for histo-blood group trisacharides, galectins, the alpha/beta-interferon antagonist sarcolectin, the lymphokine macrophage migration inhibitory factor and a monoclonal antibody to the Le(y) epitope was monitored for insight into aspects of immunologic and biologic behavior. RESULTS: In all studies, a correlation between tumor parameters, according to TNM stage and survival, was seen. In order to refine this category formation, at least certain selected features should provide an even more stringent association than TNM stages. Indeed, statistical correlation of the cytometric and histometric parameters as well as the expression of receptors for the two histo-blood group trisaccharides, ligands for the galectins (CL-16, CL-14) and macrophage migration inhibitory factor was stronger than that of TNM stage. A large amount of the current of structural entropy was especially highly significantly associated with poor survival. This observation could be verified in each of the different studies. CONCLUSION: The obtained data strongly support the notion that thermodynamic evaluation of tumor growth focusing on the "entropy distance" of the tumor from its environment is a promising perspective warranting extended studies. Additionally, glycohistochemical features, including binding capacities for histo-blood group trisaccharides, have the potential to aid in establishment of a biologic marker set for tumor staging.

Biomarkers, Tumor↗

Transfer Learning across Material Properties Using Center-Environment Features: From Energetics to Mechanical Properties in Multicomponent Mo Alloys.

Transfer learning (TL) provides a viable approach to mitigate data scarcity in materials informatics. While conventional TL focuses on predicting identical properties across different systems, this work demonstrates a cross-property extension of TL from energy to mechanical properties via end-to-end model weight pre-training and fine-tuning: knowledge learned from predicting substitution energies is transferred to predict distinctly different mechanical properties, substantially improving computational efficiency given the typically higher cost of acquiring target-domain data. To accelerate computational alloy design, machine learning models using center-environment (CE) features were first developed to predict substitution energies of alloying elements in molybdenum (Mo)-based alloys. The Random Forest models achieved the optimal performance and transferability-R2 = 0.97, 〈MAE〉 = 0.11 eV, and 〈RMSE〉 = 0.16 eV-against the density functional theory (DFT) benchmark. The model dependency of feature selection and importance analysis was discussed. The transferability of the energy models was validated on unknown systems with new elements. Subsequently, the energy models were fine-tuned using limited mechanical property data to construct energy-to-property (E2P) TL models capable of predicting elastic properties, including bulk modulus, Young's modulus, shear modulus, and elastic constants, achieving an improved accuracy over the non-transferred ML by ∼10-30%, with its transferability verified by additional DFT calculations. This cross-property E2P transfer learning framework opens a new avenue for accelerating computational materials discovery and may be extended to other multiproperty predictions governed by similar physical principles.

center-environment feature↗

Differential distribution of Fos-like immunoreactivity in the spinal trigeminal nucleus after noxious and innocuous thermal and chemical stimulation of rat cornea.

Corneal afferent nerves project to two spatially distinct sites within the spinal trigeminal nucleus: the subnucleus interpolaris/caudalis transition and the subnucleus caudalis/upper cervical spinal cord transition. The role of these two regions in processing corneal input is uncertain. To determine if neurons in these regions encode different features of an applied corneal stimulus, immunoreactivity for the immediate early gene protein product, Fos, was quantified in barbiturate-anesthetized rats. Intensity was varied across thermal (thermal probe 5, 35, 42, 52 degrees C; radiant heat of approximately 45 degrees C) stimuli and compared with that seen after mustard oil (5 microliters, 20%) or mineral oil application. All stimuli increased the number of Fos-positive neurons located at the ventrolateral pole of the subnucleus interpolaris/caudalis transition compared with unstimulated controls. By contrast, only 52 degrees C thermal probe and mustard oil produced an additional peak of Fos-positive neurons within the superficial laminae at the subnucleus caudalis/cervical cord transition. Further, the magnitudes of the bimodal peaks of Fos produced by 52 degrees C thermal probe and mustard oil stimuli were different quantitatively. Mustard oil caused a greater Fos response at the subnucleus interpolaris/caudalis transition than 52 degrees C thermal probe stimulation, whereas the opposite was true at the subnucleus caudalis/cervical cord transition. Double-labeling revealed that Fos immunoreactive neurons within the spinal trigeminal nucleus were restricted to regions densely labeled for calcitonin gene-related peptide. These results indicate that select features of corneal stimuli such as modality are encoded differently by neurons in the trigeminal subnucleus interpolaris/caudalis transition compared with those located in the subnucleus caudalis/cervical cord transition. It is likely that neurons in these two brainstem regions subserve different aspects of corneal sensation.

Animals↗

Detecting genomic features under weak selective pressure: the example of codon usage in animals and plants.

Large scale experiments of gene inactivation in yeast have shown that 50% of genes have no detectable impact on the phenotype, and similar observations have been made in other model organisms. This apparent paradox is probably due to the fact that many genes only have a marginal contribution to the fitness of organisms. Because of the size of populations and the number of generations that can be studied in laboratories, experimental approaches only permit to detect functional elements that have a strong phenotypic impact. Comparative sequence analysis can help to solve this problem: the analysis of sequences evolution permits to detect the action of selection, and hence to reveal functional features of genomes. This approach will be illustrated by the study of synonymous codon usage in animals and plants.

Animals↗

Three-dimensional reconstruction of temporal bone from computed tomographic scans on a personal computer.

The advantages of computer reconstruction of anatomical structures from computed tomographic scans are common knowledge by now. Unfortunately, to date most reconstructions have required the use of large computers and/or have entailed tedious manual contour tracing. The system described here allows largely automatic detection of surfaces in computed tomographic scans plus the usual display capabilities including feature selection, magnification, rotation, shading, and slicing as well as measurement of lengths and angles. It runs on a normal International Business Machines AT-compatible computer with a medium-resolution video card.

Child↗

Discontinuing therapy in childhood acute lymphocytic leukemia. A multicentric survey in Italy.

The results of discontinuing therapy in children with acute lymphocytic leukemia observed at four associated institutions are presented. Of the 247 patients who achieved complete remission, 122 (49.3%) reached the point of discontinuing therapy after 2-4 years of continuous remission. The median period off therapy was 13 months with a range of 1-69 months. Of the 122 children removed from therapy, 27 (22.1%) relapsed, mainly in the bone marrow; relapses occurred 1-32 months after cessation of therapy (median ten months) with only two relapses occurring later than two years. By actuarial analysis, 57% of the patients are projected in continuous remission after five years from cessation of therapy. Neither selected features at diagnosis nor single modalities of treatment were found to predict whether relapse would occur after discontinuing therapy. Long-term remission and possibly cure can be expected in over one-third of newly diagnosed children with ALL after 2-4 years of antileukemic treatment.

Adolescent↗

Multiresolution image registration for two-dimensional gel electrophoresis.

In proteomic research, two-dimensional electrophoresis (2-D) is an important tool for investigating differential patterns of qualitative and quantitative protein expression. The strength of the technique is due to its unrivalled power of being able to separate simultaneously thousands of proteins. The key to the comparison of 2-D protein profiles, however, lies in the use of a fast and robust image matching process which is essential to the subsequent quantification procedure. To satisfy the growing demand for a robust and fully automatic method of matching 2-D gel protein separation profiles, we describe in this paper a novel registration technique based on image intensity distribution rather than selected features. The method uses a multiresolution representation of the gel profiles and exploits the fact that coarse approximations to the optimal matching can be extracted efficiently from low-resolution images. This permits the removal of misalignments at different scales in a systematic manner and the strength of the new method has been confirmed by a double blind trial of 111 2-D gel pairs. The proposed method requires neither landmarks nor an a priori image alignment, and takes about five seconds for processing a typical gel pair on a standard personal computer.

Algorithms↗

The functional significance of tumour-associated cell surface alterations of embryonic and unknown origin.

The study of the phenotype of tumours aims to elucidate cell surface alterations that could be used for diagnostic, prognostic or therapeutic purposes. As tumours tend to escape the homeostatic growth control mechanisms of the host, it can be assumed that plasma membrane alterations are also responsible for the antisocial behaviour of tumour cells. Selected features of the transformed phenotype, of fetal or unknown origin, namely tumour-associated antigens, isozymes and growth factors, are discussed in relation to the altered growth pattern of the tumour cell. It is concluded that definitive structure-function relationships have not yet been established, but areas for future investigation are suggested.

Animals↗

Classifying "kinase inhibitor-likeness" by using machine-learning methods.

By using an in-house data set of small-molecule structures, encoded by Ghose-Crippen parameters, several machine learning techniques were applied to distinguish between kinase inhibitors and other molecules with no reported activity on any protein kinase. All four approaches pursued--support-vector machines (SVM), artificial neural networks (ANN), k nearest neighbor classification with GA-optimized feature selection (GA/kNN), and recursive partitioning (RP)--proved capable of providing a reasonable discrimination. Nevertheless, substantial differences in performance among the methods were observed. For all techniques tested, the use of a consensus vote of the 13 different models derived improved the quality of the predictions in terms of accuracy, precision, recall, and F1 value. Support-vector machines, followed by the GA/kNN combination, outperformed the other techniques when comparing the average of individual models. By using the respective majority votes, the prediction of neural networks yielded the highest F1 value, followed by SVMs.

Algorithms↗

On fully automatic feature measurement for banded chromosome classification.

Procedures for fully automatic location of chromosome axis and centromere in metaphase chromosomes are described for a practical interactive chromosome analysis system that omits the usual stages of interactive axis and centromere correction. Accuracy of centromere finding and consequential determination of a chromosome's polarity, i.e., which end is which, is measured experimentally. The saving in interaction by not correcting centromeres is compared to the increase in errors at the classification stage and the consequent increase in interaction needed to correct these errors. Some previously unreported features for banded chromosome classification are described, and in particular a set of global shape features is introduced. The discrimination capability of the feature measurements is evaluated by use of simple statistics and by reference to the performance of classifiers trained with various feature subsets. Class discrimination capability of the global shape feature set is shown to be comparable to that of centromere position, a widely used local shape feature. The variability of feature measurements that might occur in data from different laboratories on account of differing tissue, preparation methods, and digitiser hardware is assessed using three data bases of G-banded human metaphase cells. It is shown that the differences can be considerable and that appropriate feature selection and classifier training substantially improve classification performance.

Chromosomes↗