Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Using in vitro prediction models instead of the rabbit eye irritation test to classify and label new chemicals: a post hoc data analysis of the international EC/HO validation study.

The international validation study on alternative methods to replace the Draize rabbit eye irritation test, funded by the European Commission (EC) and the British Home Office (HO), took place during 1992-1994, and the results were published in 1995. The results of this EC/HO study are analysed by employing discriminant analysis, taking into account the classification of the in vivo data into eye irritation classes A (risk of serious damage to eyes), B (irritating to eyes) and NI (non-irritant). A data set for 59 test items was analysed, together with three subsets: surfactants, water-soluble chemicals, and water-insoluble chemicals. The new statistical methods of feature selection and estimation of the discriminant functions classification error were used. Normal distributed random numbers were added to the mean values of each in vitro endpoint, depending on the observed standard deviations. Thereafter, the reclassification error of the random observations was estimated by applying the fixed function of the mean values. Moreover, the leaving-one-out cross-classification method was applied to this random data set. Subsequently, random data were generated r times (for example, r = 1000) for a feature combination. Eighteen features were investigated in nine in vitro test systems to predict the effects of a chemical in the rabbit eye. 72.5% of the chemicals in the undivided sample were correctly classified when applying the in vitro endpoints lgNRU of the neutral red uptake test and lgBCOPo5 of the bovine opacity and permeability test. The accuracy increased to 80.9% when six in vitro features were used, and the sample was subdivided. The subset of surfactants was correctly classified in more than 90% of cases, which is an excellent performance.

Animal Testing Alternatives↗

Event-related potentials: a critical review of methods for single-trial detection.

The analysis of ERP data has followed several lines over the last 20 years. The most prevalent method is simply to average ERPs for a given class of stimuli. The ERPs are compared for differences across classes of stimuli. Little other special data processing is used. The ERP comparisons are usually performed using visual examination of the wave-shapes. Sometimes statistics are calculated such as means, variances, and confidence limits. Linear filtering is used to reduce interference. Another approach is to model or analyze the ERP as a sequence of vectors or frames of data samples. These samples may be of the ERP time waveform or they may be of the frequency transform of the ERP waveform. The frames of data vary in length from the entire ERP waveform (500 to 1000 msec) to frames as short as ten sample points (100 msec). Recognition of an event in the ERP is achieved by computing a distance measure between parameter vectors for one class of stimuli and corresponding parameter vectors for another class of stimuli. Recognition is achieved by selecting the ERP with the lowest distance score. This approach is "pattern matching" and relies on two assumptions: adjacent frames of data are uncorrelated, and the variability of the data can be accounted for by the distance measured for all stimuli in the classes presented. Subject variability is generally not accounted for, other than to assume it is the same for all classes of stimuli. The data are clustered into a variety of reference patterns that represent particular manifestations of a particular stimulus. Another approach is "feature-based" recognition. The idea is to identify and automatically extract features of the data that can provide a characterization of stimuli. The features selected may be abstract. They are calculated from the data or transforms of the data.

Biometry↗

Animal models of ALS.

Animal models of amyotrophic lateral sclerosis (ALS) provide a unique opportunity to study this incurable and fatal human disease both clinically and pathologically. This is particularly true for certain pathological and therapeutic studies that are impractical or impossible to perform in human patients. Nonetheless, postmortem ALS tissue remains the "gold standard" against which pathologic findings in animal models must be compared. Four natural disease models have been most extensively studied, including three mouse models: motor neuron degeneration (Mnd), progressive motor neuronopathy (pmn), wobbler, and one canine model: hereditary canine spinal muscular atrophy (HCSMA). The wobbler mouse has been the most extensively studied of these models with analyses of clinical, pathological (perikaryon, axon, muscle), and biochemical features. Experimentally induced ALS animal models have allowed controlled testing of various neurotoxic, viral and immune-mediated mechanisms. Molecular techniques have recently generated mouse models in which genes relevant to the human disease or motor neuron biology have been manipulated. The most clinically relevant of these is a transgenic mouse overexpressing the mutated SOD1 gene of FALS patients, which has already provided significant insights into mechanisms of motor neuron degeneration in this disease. Because no single animal model perfectly reflects all the clinical and pathological characteristics of ALS, study of selected features from the most relevant models will contribute to a better understanding of the pathogenesis and/or etiology of this disease.

Amyotrophic Lateral Sclerosis↗

Quantitative analysis of synovial membrane inflammation: a comparison between automated and conventional microscopic measurements.

OBJECTIVE: The objective of this study was to quantify selected features of chronic synovial tissue inflammation by computerised image analysis and to validate the results by comparison with conventional microscopic measurements. METHODS: Synovial biopsy samples were obtained from the knee joints of patients with chronic arthritis and prepared for immunohistochemical analysis using standard techniques. Following the development of special software, four parameters of chronic synovial inflammation were evaluated: intimal layer thickness, CD3+ cell infiltration, CD8+ cell infiltration and vascularity. Intimal layer thickness was expressed in microns. The intensity of CD3+ and CD8+ cell infiltration was expressed as the percentage area of the tissue section occupied by positively stained cells. Vascularity was expressed as the percentage area occupied by blood vessels. Conventional quantitative microscopic analysis was also undertaken and the results from both methods compared. RESULTS: Seventy eight tissue sections were selected for study. Measurements of intimal layer thickness by both techniques correlated strongly: r = 0.85, p = 0.0006. Measurements of CD8+ cell infiltration, usually widely dispersed, also correlated well: r = 0.64, p = 0.005. Measurements of CD3+ cell infiltration, often densely aggregated, correlated less well: r = 0.55, p = 0.02. Measurements of vascularity demonstrated no statistically significant correlation: r = 0.41, p = 0.07. Proficiency in the use of computerised image analysis was readily acquired. CONCLUSION: Computerised image analysis was successfully applied to the measurement of some features of synovial tissue inflammation. Further software development is required to validate measurement of blood vessels of variable size.

Arthritis↗

Prediction of lymph node metastasis by analysis of gene expression profiles in non-small cell lung cancer.

OBJECTIVE: Non-small cell lung carcinoma (NSCLC) is one of the leading causes of death in the world. Lymph node metastasis is not only an important factor in estimating the extent and the metastatic potential of an NSCLC but also in prognosticating the patient outcome. Preoperative prediction of lymph node metastasis might greatly facilitate the choice of appropriate surgical and medical options in patients with NSCLC. METHODS AND RESULTS: Using a cDNA array, we analyzed the expression profiles of 1,289 genes in 92 cancer tissues of NSCLC (37 squamous cell carcinomas and 55 adenocarcinomas). We divided the patients into two groups (classes) for each of various pathological factors, such as lymph node metastasis and pT-stage. For each pair of classes, we searched for an optimal combination of genes to classify the cases using a sequential forward selection algorithm starting from a gene set that showed significant difference in expression between the classes. We used the leave-one-out error cross-validation on a k-nearest neighbor classifier to sequentially choose the gene. Using the optimized set of genes, it was possible to stratify the patients for lymph node metastasis (pN-stage) and pT-stage at, respectively, 100% (23 genes) and 100% (55 genes) for cases with squamous cell carcinomas and 94% (43 genes) and 92% (35 genes) for those with adenocarcinomas. CONCLUSION: We conclude that expression profiling using feature selection provides a powerful means of stratification (personalization) of NSCLC patients and choice in treatment options, particularly for factors such as lymph node metastasis whose radiological diagnosis is presently incomplete.

Adenocarcinoma↗

Automatic detection of calcifications in the aorta from CT scans of the abdomen. 3D computer-aided diagnosis.

RATIONALE AND OBJECTIVES: Automated detection and quantification of arterial calcifications can facilitate epidemiologic research and, eventually, the use of full-body calcium scoring in clinical practice. An automatic computerized method to detect calcifications in CT scans is presented. MATERIALS AND METHODS: Forty abdominal CT scans have been randomly selected from clinical practice. They all contained contrast material and belonged to one of four categories: containing "no," "small," "moderate," or "large" amounts of arterial calcification. There were ten scans in each category. The experiments were restricted to the vertical range from the point where the superior mesenteric artery branches off of the descending aorta until the first bifurcation of the iliac arteries. The automatic method starts by extracting all connected objects above 220 Hounsfield units (HU) from the scan. These objects include all calcifications, as well as bony structures and contrast material. To distinguish calcifications from non-calcifications, a number of features are calculated for each object. These features are based on the object's size, location, shape characteristics, and surrounding structures. Subsequently a classification of each object is performed in two stages. First the probability that an object represents a calcification is computed assuming a multivariate Gaussian distribution for the calcifications. Objects with low probability are discarded. The remaining objects are then classified into calcifications and non-calcifications using a 5-nearest-neighbor classifier and sequential forward feature selection. Based on the total volume of calcifications determined by the system, the scan is assigned to one of the four categories mentioned above. RESULTS: The 40 scans contained a total of 249 calcifications as determined by a human observer. The method detected 209 calcifications (sensitivity 83.9%) at the expense of on average 1.0 false-positive object per scan. The correct category label was assigned to 30 scans and only 2 scans were off by more than one category. Most incorrect classifications can be attributed to the presence of contrast material in the scans. CONCLUSION: It is possible to identify the majority of arterial calcifications in abdominal CT scans in a completely automatic fashion with few false positive objects, even if the scans contain contrast material.

Aorta, Abdominal↗

CoxKAN: Kolmogorov-Arnold networks for interpretable, high-performance survival analysis.

MOTIVATION: Survival analysis is a branch of statistics that is crucial in medicine for modeling the time to critical events such as death or relapse, in order to improve treatment strategies and patient outcomes. Selecting survival models often involves a trade-off between performance and interpretability; deep learning models offer high performance but lack the transparency of more traditional approaches. This poses a significant issue in medicine, where practitioners are reluctant to use black-box models for critical patient decisions. RESULTS: We introduce CoxKAN, a Cox proportional hazards Kolmogorov-Arnold Network for interpretable, high-performance survival analysis. Kolmogorov-Arnold Networks (KANs) were recently proposed as an interpretable and accurate alternative to multi-layer perceptrons. We evaluated CoxKAN on four synthetic and nine real datasets, including five cohorts with clinical data and four with genomics biomarkers. In synthetic experiments, CoxKAN accurately recovered interpretable hazard function formulae and excelled in automatic feature selection. Evaluations on real datasets showed that CoxKAN consistently outperformed the traditional Cox proportional hazards model (by up to 4% in C-index) and matched or surpassed the performance of deep learning-based models. Importantly, CoxKAN revealed complex interactions between predictor variables and uncovered symbolic formulae, which are key capabilities that other survival analysis methods lack, to provide clear insights into the impact of key biomarkers on patient risk. AVAILABILITY AND IMPLEMENTATION: CoxKAN is available at GitHub and Zenodo.

Humans↗

Identification of amino acid residues lining the pore of a gap junction channel.

Gap junctions represent a ubiquitous and integral part of multicellular organisms, providing the only conduit for direct exchange of nutrients, messengers and ions between neighboring cells. However, at the molecular level we have limited knowledge of their endogenous permeants and selectivity features. By probing the accessibility of systematically substituted cysteine residues to thiol blockers (a technique called SCAM), we have identified the pore-lining residues of a gap junction channel composed of Cx32. Analysis of 45 sites in perfused Xenopus oocyte pairs defined M3 as the major pore-lining helix, with M2 (open state) or M1 (closed state) also contributing to the wider cytoplasmic opening of the channel. Additional mapping of a close association between M3 and M4 allowed the helices of the low resolution map (Unger et al., 1999. Science. 283:1176-1180) to be tentatively assigned to the connexin transmembrane domains. Contrary to previous conceptions of the gap junction channel, the residues lining the pore are largely hydrophobic. This indicates that the selective permeabilities of this unique channel class may result from novel mechanisms, including complex van der Waals interactions of permeants with the pore wall, rather than mechanisms involving fixed charges or chelation chemistry as reported for other ion channels.

Amino Acid Sequence↗

Biological factors predisposing to traumatic posterior dislocation of the hip. A selection process in the mechanism of injury.

The factors involved in the mechanism leading to traumatic posterior dislocation of the hip are examined. In 47 adult patients who had previously suffered such a dislocation, ultrasound scans were used to measure femoral anteversion on both the affected and the uninjured side. In 36 normal adult volunteers, used as controls, similar measurements were made. Femoral anteversion on both the injured and uninjured side was significantly reduced in the patients compared with the volunteers. These findings are discussed in the light of previous work which indicates that medial rotation is a factor in the mechanism of posterior dislocation of the hip. It is suggested that reduced anteversion acts like medial rotation to make the hip more susceptible to posterior dislocation, and that the less the anteversion the more likely is the injury to be a dislocation rather than a fracture-dislocation. It is concluded that patients who suffer such dislocated hips belong at one extreme of the normal population, having either reduced femoral anteversion or even retroversion, and that this anatomical feature selects towards hip dislocation rather than to injury of the femoral shaft, knee or tibia during the appropriate type of accident.

Adult↗

Acoustic-phonetic contrasts and intelligibility in the dysarthria associated with mixed cerebral palsy.

This study evaluated the relationship between specific acoustic features of speech and perceptual judgments of word intelligibility of adults with cerebral palsy-dysarthria. Use of a contrasting word task allowed for intelligibility analysis and correlated acoustic analysis according to specified spectral and temporal features. Selected phonemic contrasts included syllable-initial voicing; syllable-final voicing; stop-nasal; fricative-affricate; front-back, high-low, and tense-lax vowels. Speech materials included a set of CVC stimulus words. Acoustic data are reported on vowel duration, formant frequency locations, voice onset times, amplitude rise times, and frication durations. Listeners' perceptual assessment of intelligibility of the 16 dysarthric adults by transcription and rating tasks is also presented. All but one acoustic contrast was successfully made as evidenced by measured acoustic differences between contrast pairs. However, the generally successful acoustic contrasts stood in marked contrast to the poorly rated intelligibility scores and high error percentages that were ascribed to the opposite pair members. A second analysis examined the contribution of these acoustic features towards estimates and prediction of intelligibility deficits in speakers with dysarthria. The scaled intelligibility was predicted by multiple regression analysis with 62.6% accuracy by acoustic measures related to one consonant contrast (fricative-affricate) and three vowel contrasts (front-back, high-low, and tense-lax). Other measured contrasts, such as those related to contrast voicing effects and stop-nasal distinctions, did not seem to contribute in a significant way to variability in the intelligibility estimates. These findings are discussed in relation to specific areas of production deficiency that are consistent across different types of dysarthria with cerebral palsy as the etiology.

Adult↗

Assessing sequence comparison methods with the average precision criterion.

MOTIVATION: Comprehensive performance assessment is important for improving sequence database search methods. Sensitivity, selectivity and speed are three major yet usually conflicting evaluation criteria. The average precision (AP) measure aims to combine the sensitivity and selectivity features of a search algorithm. It can be easily visualized and extended to analyze results from a set of queries. Finally, the time-AP plot can clearly show the overall performance of different search methods. RESULTS: Experiments are performed based on the SCOP database. Popular sequence comparison algorithms, namely Smith-Waterman (SSEARCH), FASTA, BLAST and PSI-BLAST are evaluated. We find that (1) the low-complexity segment filtration procedure in BLAST actually harms its overall search quality; (2) AP scores of different search methods are approximately in proportion of the logarithm of search time; and (3) homologs in protein families with many members tend to be more obscure than those in small families. This measure may be helpful for developing new search algorithms and can guide researchers in selecting most suitable search methods. AVAILABILITY: Test sets and source code of this evaluation tool are available upon request.

Algorithms↗

Evaluation of a statistically derived decision tree for the cytodiagnosis of fine needle aspirates of the breast (FNAB).

A decision tree for the diagnosis of FNAB was derived from defined human observations using a rule induction method, C4.5 (a derivative of the ID3 algorithm). This algorithm is an implementation of the top-down induction method where the tree is determined iteratively by adding those nodes and branches which maximize the information gain at each step. The tree was derived from a training set of 200 FNAB with known outcome using 10 defined features (from one observer) and patient age. The tree contained a total of seven nodes (six observable features and patient age) with eight endpoints (four benign, four malignant). The tree was applied to a test set of 400 further FNAB with observations from the training observer and produced a sensitivity of 95%, specificity of 93% and a positive predictive value (PPV) of a malignant result of 89%. Four trainee pathologists were given a training session on the observable features and then used the tree to determine outcome in a further 50 FNAB. The observers were blind to clinical details apart from age and the endpoints were coded with letters and not labelled benign or malignant. The results from these observers produced ranges of sensitivity 80-96%, specificity 64-92%, PPV 73-92% and kappa statistics (with known outcome) 0.6-0.8. Reported difficulties in using the tree included estimation of nuclear size. These results were worse than the performance of the observers on a further 50 cases without using the decision tree (sensitivity 80-100%, specificity 72-100%, PPV 78-100%, kappa 0.72-0.92). The original 50 case test set was rerandomized and the four trainee observers made all 10 defined observations on each specimen without using the decision tree; these observations were then used to derive decisions from the tree. The performance from this method was similar to that using selected features from the tree, suggesting that observation of all features together does not improve the reliability of each specific observation. The poor performance of this tree suggests that this methodology may be unsuitable for producing decision support aids for diagnostic or training purposes in this domain.

Biopsy, Needle↗

Association of prognosis in surgically treated lung cancer patients with cytometric, histometric and ligand histochemical properties: with an emphasis on structural entropy.

OBJECTIVE: To explore new tumor features for refined category formation that permits the tailoring of individualized treatment schemes in lung cancer. STUDY DESIGN: Survival data on patients from six independent studies on cases with surgically treated lung cancer, primary lung carcinoids or metastasizing breast carcinoma (including data on primary breast carcinoma) were analyzed by nonhierarchic multivariant discriminant analysis with respect to a set of cytometric/histometric and immunohistochemical/ligand histochemical parameters. The number of stem lines, S-phase-related tumor cell fraction and the extent of structural entropy and its current were measured. In addition, the expression of binding capacities for histo-blood group trisacharides, galectins, the alpha/beta-interferon antagonist sarcolectin, the lymphokine macrophage migration inhibitory factor and a monoclonal antibody to the Le(y) epitope was monitored for insight into aspects of immunologic and biologic behavior. RESULTS: In all studies, a correlation between tumor parameters, according to TNM stage and survival, was seen. In order to refine this category formation, at least certain selected features should provide an even more stringent association than TNM stages. Indeed, statistical correlation of the cytometric and histometric parameters as well as the expression of receptors for the two histo-blood group trisaccharides, ligands for the galectins (CL-16, CL-14) and macrophage migration inhibitory factor was stronger than that of TNM stage. A large amount of the current of structural entropy was especially highly significantly associated with poor survival. This observation could be verified in each of the different studies. CONCLUSION: The obtained data strongly support the notion that thermodynamic evaluation of tumor growth focusing on the "entropy distance" of the tumor from its environment is a promising perspective warranting extended studies. Additionally, glycohistochemical features, including binding capacities for histo-blood group trisaccharides, have the potential to aid in establishment of a biologic marker set for tumor staging.

Biomarkers, Tumor↗

Transfer Learning across Material Properties Using Center-Environment Features: From Energetics to Mechanical Properties in Multicomponent Mo Alloys.

Transfer learning (TL) provides a viable approach to mitigate data scarcity in materials informatics. While conventional TL focuses on predicting identical properties across different systems, this work demonstrates a cross-property extension of TL from energy to mechanical properties via end-to-end model weight pre-training and fine-tuning: knowledge learned from predicting substitution energies is transferred to predict distinctly different mechanical properties, substantially improving computational efficiency given the typically higher cost of acquiring target-domain data. To accelerate computational alloy design, machine learning models using center-environment (CE) features were first developed to predict substitution energies of alloying elements in molybdenum (Mo)-based alloys. The Random Forest models achieved the optimal performance and transferability-R2 = 0.97, 〈MAE〉 = 0.11 eV, and 〈RMSE〉 = 0.16 eV-against the density functional theory (DFT) benchmark. The model dependency of feature selection and importance analysis was discussed. The transferability of the energy models was validated on unknown systems with new elements. Subsequently, the energy models were fine-tuned using limited mechanical property data to construct energy-to-property (E2P) TL models capable of predicting elastic properties, including bulk modulus, Young's modulus, shear modulus, and elastic constants, achieving an improved accuracy over the non-transferred ML by ∼10-30%, with its transferability verified by additional DFT calculations. This cross-property E2P transfer learning framework opens a new avenue for accelerating computational materials discovery and may be extended to other multiproperty predictions governed by similar physical principles.

center-environment feature↗

Differential distribution of Fos-like immunoreactivity in the spinal trigeminal nucleus after noxious and innocuous thermal and chemical stimulation of rat cornea.

Corneal afferent nerves project to two spatially distinct sites within the spinal trigeminal nucleus: the subnucleus interpolaris/caudalis transition and the subnucleus caudalis/upper cervical spinal cord transition. The role of these two regions in processing corneal input is uncertain. To determine if neurons in these regions encode different features of an applied corneal stimulus, immunoreactivity for the immediate early gene protein product, Fos, was quantified in barbiturate-anesthetized rats. Intensity was varied across thermal (thermal probe 5, 35, 42, 52 degrees C; radiant heat of approximately 45 degrees C) stimuli and compared with that seen after mustard oil (5 microliters, 20%) or mineral oil application. All stimuli increased the number of Fos-positive neurons located at the ventrolateral pole of the subnucleus interpolaris/caudalis transition compared with unstimulated controls. By contrast, only 52 degrees C thermal probe and mustard oil produced an additional peak of Fos-positive neurons within the superficial laminae at the subnucleus caudalis/cervical cord transition. Further, the magnitudes of the bimodal peaks of Fos produced by 52 degrees C thermal probe and mustard oil stimuli were different quantitatively. Mustard oil caused a greater Fos response at the subnucleus interpolaris/caudalis transition than 52 degrees C thermal probe stimulation, whereas the opposite was true at the subnucleus caudalis/cervical cord transition. Double-labeling revealed that Fos immunoreactive neurons within the spinal trigeminal nucleus were restricted to regions densely labeled for calcitonin gene-related peptide. These results indicate that select features of corneal stimuli such as modality are encoded differently by neurons in the trigeminal subnucleus interpolaris/caudalis transition compared with those located in the subnucleus caudalis/cervical cord transition. It is likely that neurons in these two brainstem regions subserve different aspects of corneal sensation.

Animals↗

Detecting genomic features under weak selective pressure: the example of codon usage in animals and plants.

Large scale experiments of gene inactivation in yeast have shown that 50% of genes have no detectable impact on the phenotype, and similar observations have been made in other model organisms. This apparent paradox is probably due to the fact that many genes only have a marginal contribution to the fitness of organisms. Because of the size of populations and the number of generations that can be studied in laboratories, experimental approaches only permit to detect functional elements that have a strong phenotypic impact. Comparative sequence analysis can help to solve this problem: the analysis of sequences evolution permits to detect the action of selection, and hence to reveal functional features of genomes. This approach will be illustrated by the study of synonymous codon usage in animals and plants.

Animals↗

Three-dimensional reconstruction of temporal bone from computed tomographic scans on a personal computer.

The advantages of computer reconstruction of anatomical structures from computed tomographic scans are common knowledge by now. Unfortunately, to date most reconstructions have required the use of large computers and/or have entailed tedious manual contour tracing. The system described here allows largely automatic detection of surfaces in computed tomographic scans plus the usual display capabilities including feature selection, magnification, rotation, shading, and slicing as well as measurement of lengths and angles. It runs on a normal International Business Machines AT-compatible computer with a medium-resolution video card.

Child↗

Discontinuing therapy in childhood acute lymphocytic leukemia. A multicentric survey in Italy.

The results of discontinuing therapy in children with acute lymphocytic leukemia observed at four associated institutions are presented. Of the 247 patients who achieved complete remission, 122 (49.3%) reached the point of discontinuing therapy after 2-4 years of continuous remission. The median period off therapy was 13 months with a range of 1-69 months. Of the 122 children removed from therapy, 27 (22.1%) relapsed, mainly in the bone marrow; relapses occurred 1-32 months after cessation of therapy (median ten months) with only two relapses occurring later than two years. By actuarial analysis, 57% of the patients are projected in continuous remission after five years from cessation of therapy. Neither selected features at diagnosis nor single modalities of treatment were found to predict whether relapse would occur after discontinuing therapy. Long-term remission and possibly cure can be expected in over one-third of newly diagnosed children with ALL after 2-4 years of antileukemic treatment.

Adolescent↗