Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

In vitro evolution used to define a protein recognition site within a large RNA domain.

A minimum of 460 nucleotides of 16S ribosomal RNA are needed to fold the target site for E. coli ribosomal protein S4, although a much smaller region within this large domain is protected from chemical reagents by the protein. Starting with a 531-nucleotide tRNA fragment, cycles of mutagenesis, selection with S4, and amplification ('in vitro evolution') were used to obtain a pool of 30 RNA sequences selected for S4 recognition but approximately 30% different from wild type. Numerous compensatory base pair changes have largely preserved the same secondary structure among these RNAs as found in wild-type sequences. A 20-base deletion and a single nucleotide insertion are among several unusual features found in most of the selected sequences and also prevalent among other prokaryotic rRNAs. Most of the compensatory base changes and selected features are located outside of the region protected by S4 from chemical reagents. It was unexpected that S4 would select for RNA structures throughout such a large domain; the selected features are probably contributing indirectly to S4 recognition by promoting correct tertiary folding of the region actually contacted by S4. The role of S4 may be to stabilize this domain (nearly one-third of the 16S rRNA) in its proper conformation for ribosome function.

Base Sequence↗

Synthesis and evaluation of (pyridylmethylene)tetrahydronaphthalenes/-indanes and structurally modified derivatives: potent and selective inhibitors of aldosterone synthase.

Elevated aldosterone levels are key effectors for the development and progression of congestive heart failure and myocardial fibrosis. Recently, we proposed inhibition of aldosterone synthase (CYP11B2) as an innovative strategy for the treatment of these diseases. In this study, the synthesis and biological evaluation of E- and Z-(pyridylmethylene)tetrahydronaphthalenes and -indanes (1a,b-38a) is described. The activity of the compounds was determined using human CYP11B2, and the selectivity was evaluated toward the human steroidogenic enzymes CYP11B1, CYP19, and CYP17. The biological results revealed a few rather selective inhibitors of CYP11B1, some compounds inhibiting both CYP11B1 and CYP11B2, and a large number of highly selective inhibitors of CYP11B2. The most active inhibitor was the 3-pyridyl compound 5a (IC(50) = 7 nM). The pyrimidyl-substituted derivative 28a was found to be the most selective CYP11B2 inhibitor (IC(50) = 27 nM) in this series, showing a 120-fold selectivity for CYP11B1 (IC(50) = 3179 nM). Molecular modeling, i.e., examination of the electronic and steric features of selected compounds and homology modeling and docking, was used to understand the structure-activity/-selectivity relationships.

Adrenal Cortex Hormones↗

Approximations of selected standard English sentences by speakers of black English.

The purpose of this study was to investigate the changes in six selected linguistic features as a function of age among children from the North Texas area who utilized features of black English in their oral language. A total of 48 children were tested, with eight selected from each chronological age group from four through nine years. Each subject's sole task was to imitate a series of Standard English sentences. The results indicated that, while a gradual decrease in black English characteristics occurred as a function of age, there was a tendency for the most pronounced dialectal decline in the language of the subjects to occur between six and seven years of age for most of the linguistic features. However, it was observed that black English caracteristics did not completely disappear from the surface structure of even the oldest subjects in the study.

Black or African American↗

A completely automated CAD system for mass detection in a large mammographic database.

Mass localization plays a crucial role in computer-aided detection (CAD) systems for the classification of suspicious regions in mammograms. In this article we present a completely automated classification system for the detection of masses in digitized mammographic images. The tool system we discuss consists in three processing levels: (a) Image segmentation for the localization of regions of interest (ROIs). This step relies on an iterative dynamical threshold algorithm able to select iso-intensity closed contours around gray level maxima of the mammogram. (b) ROI characterization by means of textural features computed from the gray tone spatial dependence matrix (GTSDM), containing second-order spatial statistics information on the pixel gray level intensity. As the images under study were recorded in different centers and with different machine settings, eight GTSDM features were selected so as to be invariant under monotonic transformation. In this way, the images do not need to be normalized, as the adopted features depend on the texture only, rather than on the gray tone levels, too. (c) ROI classification by means of a neural network, with supervision provided by the radiologist's diagnosis. The CAD system was evaluated on a large database of 3369 mammographic images [2307 negative, 1062 pathological (or positive), containing at least one confirmed mass, as diagnosed by an expert radiologist]. To assess the performance of the system, receiver operating characteristic (ROC) and free-response ROC analysis were employed. The area under the ROC curve was found to be Az = 0.783 +/- 0.008 for the ROI-based classification. When evaluating the accuracy of the CAD against the radiologist-drawn boundaries, 4.23 false positives per image are found at 80% of mass sensitivity.

Algorithms↗

Spike morphology, location, and frequency in benign epilepsy with centrotemporal spikes.

The literature on benign epilepsy with centrotemporal spikes reports a constellation of neurophysiologic features in selected populations with heterogeneous methodologies. The aim of this study was to determine the specific electroencephalographic (EEG) features (spike morphology, location, and frequency and associated background slowing) in a broad population-based cohort identified through EEG laboratories. The mean spike frequency in the awake state was 9.3 per minute (95% confidence interval 6.5-12.0), in drowsiness, 21.2 per minute (16.7-25.6); and in sleep, 45.6 per minute (38.3-52.8), where 60% of patients had > 40 discharges per minute. In five patients, spike train rates occupied > 80% of the sleep record, and in nine patients, they occupied 61% to 80%. An ambulatory overnight record did not add new information comparing early-onset sleep with a mean spike frequency of 37.1 per minute (27.3-46.9) with slow-wave sleep, 36.0 per minute (27.3-44.7). Patients with benign epilepsy with centrotemporal spikes have a high spike burden, which can impact on cognitive function.

Adolescent↗

Neural models of motion integration and segmentation.

A neural model is developed of how motion integration and segmentation processes compute global motion percepts. Figure-ground properties, such as occlusion, influence which motion signals determine the percept. For visible apertures, a line's extrinsic terminators do not specify true line motion. For invisible apertures, a line's intrinsic terminators create veridical feature tracking signals, which are amplified before they propagate across space and are integrated with ambiguous motion signals within line interiors. This integration process is the result of several processing stages: directional transient cells respond to image transients and input to a directional short-range filter that selectively boosts feature tracking signals. Competitive interactions further boost feature tracking signals and create speed-selective receptive fields. A long-range filter gives rise to true directional cells by pooling signals over multiple orientations and opposite contrast polarities. A distributed population code of speed tuning realizes a size-speed correlation, whereby activations of multiple spatially short-range filters of different sizes are transformed into speed-tuned cell responses. These mechanisms use transient cell responses, output thresholds that covary with filter size, and competition. The model reproduces empirically derived speed discrimination curves and simulates data showing how visual speed perception and discrimination are affected by stimulus contrast.

Models, Neurological↗

Causal circuit tracing reveals distinct computational architectures in single-cell foundation models: inhibitory dominance, biological coherence, and cross-model convergence.

MOTIVATION: Sparse autoencoders (SAEs) decompose foundation-model activations into interpretable features, but the model-internal causal interactions between those features (i.e. what ablating one feature does to the others, as distinct from the biological causal structure of the underlying cells)-and how those model-internal relationships relate to biological structure-are uncharacterized in single-cell foundation models. RESULTS: We introduce model-internal causal circuit tracing-zeroing one SAE feature at a source layer and measuring the resulting change in all downstream SAE features, for each of 120 source features-and apply it to Geneformer V2-316M and scGPT whole-human across four conditions (96&#xa0;892 ablation-derived edges, 80&#xa0;191 forward passes). On annotation-selected source features, edges share GO/KEGG/Reactome/STRING/TRRUST ontology terms at 50.9%-68.5%, a 2.9-6.2&#xd7; enrichment over a configuration-preserving permutation null (P<.002); on 20 randomly sampled source features this attenuates to 21.5%-26.3%-still 2.5-3.1&#xd7; above null-quantifying the annotation-selection contribution. Inhibitory dominance (fraction of ablation edges with d<0, i.e. source activation supports downstream target) is 65.5%-89.4%. scGPT produces larger raw per-edge effects (mean |d|=1.40 versus 1.05); after feature-share normalization, Geneformer is stronger (paired gene-pair ratio 0.64 on 33&#xa0;301 shared pairs). Cross-model consensus yields 1142 architecture-invariant domain pairs (ordered pairs of GO biological-process categories "A&#x2192;B" each connected by at least one ablation edge in both models; 10.6&#xd7; enrichment over permutation null; P<.001). Circuit edge magnitude explains <1% of the variance in marginal driver-gene coexpression on the same cells (R2=0.010, n=31&#xa0;176): the graph encodes structure beyond bivariate correlation. Against a matched-cell-type ENCODE ChIP-seq prior, circuit-predicted transcription factor (TF)&#x2192;target pairs are enriched 2.06&#xd7; (Fisher OR 5.84), markedly higher than 1.12&#xd7; against TRRUST; direct ChIP-seq-supported target pairs show 10-30&#xd7; larger CRISPRi sign-bias-corrected excess than indirect pairs. Gene-level CRISPRi validation on Replogle K562 and the noncancer RPE1 arm (and a true primary-T-cell control from Shifrut E, Carnevale J, Tobin V et&#xa0;al. Genome-wide CRISPR screens in primary human T cells reveal key regulators of immune function. Cell 2018; 175: 1958-71.e15) after sign-bias correction shows excess over baseline of +0.03 and +0.35 percentage points on K562 and RPE1, respectively (baseline already 52%-56% from sign marginals); effect-magnitude Spearman correlations &#x3c1;&#x2248;0. Bootstrap and per-cell-type stability (N&#x2208;{50,100,200}; B cell, CD4&#xa0;+ T, macrophage) give Pearson r&#x2265;0.97 on shared edges with 100% sign agreement; edge Jaccard grows monotonically with sample size. The circuit graph is therefore highly reproducible as an effect-size map, cell type specific in edge identity, consistent with coexpression encoding, and weakly but detectably enriched for ChIP-seq-supported direct regulatory edges. AVAILABILITY AND IMPLEMENTATION: https://github.com/Biodyn-AI/bio-sae-circuits (Python). Archival DOI: 10.5281/zenodo.19,633,166 (Zenodo).

Humans↗

Complex social behaviour can select for variability in visual features: a case study in Polistes wasps.

The ability to recognize individuals is common in animals; however, we know little about why the phenotypic variability necessary for individual recognition has evolved in some animals but not others. One possibility is that natural selection favours variability in some social contexts but not in others. Polistes fuscatus wasps have variable facial and abdominal markings used for individual recognition within their complex societies. Here, I explore whether social behaviour can select for variability by examining the relationship between social behaviour and variability in visual features (marking variability) across social wasp taxa. Analysis using a concentrated changes test demonstrates that marking variability is significantly associated with nesting strategy. Species with flexible nest-founding strategies have highly variable markings, whereas species without flexible nest-founding strategies have low marking variability. These results suggest that: (i) individual recognition may be widespread in the social wasps, and (ii) natural selection may play a role in the origin and maintenance of the variable distinctive markings. Theoretical and empirical evidence suggests that species with flexible nesting strategies have reproductive transactions, a type of complex social behaviour predicted to require individual recognition. Therefore, the reproductive transactions of flexible species may select for highly variable individuals who are easy to identify as individuals. Further, selection for distinctiveness may provide an alternative explanation for the evolution of phenotypic diversity.

Animals↗

Metaplastic breast tumors with a dominant fibromatosis-like phenotype have a high risk of local recurrence.

BACKGROUND: In the current study the authors describe the clinicopathologic characteristics of a low grade variant of spindle cell metaplastic tumors of the breast. Previously these tumors have been considered within a larger group recognized as metaplastic carcinoma, including cases with higher grade features. METHODS: Breast tumors comprised predominantly of low grade spindle cells, with sparse low grade epithelial elements, were selected. Clinical features as well as macroscopic, microscopic, and immunohistochemical findings were reviewed with emphasis on the biologic behavior and the differential diagnosis from other spindle cell lesions. RESULTS: Of 30 tumors fulfilling strict criteria, 20 contained squamous or glandular elements associated with the spindle cells. Ten tumors were comprised entirely of low grade spindle cells with limited clustered epithelioid cells. At the periphery, all tumors showed a proliferation of bland spindle cells infiltrating the adjacent parenchyma and mimicking fibromatosis. The epithelioid cells and some spindle cells expressed both vimentin and one or more cytokeratins. Seven of eight patients treated by excisional biopsy developed local recurrence, whereas only one of ten patients treated with wide excisional biopsy developed a local recurrence. No distant or regional metastases occurred. CONCLUSIONS: The presence of limited clusters of epithelioid cells along with a dominant fibromatosis-like pattern may be unique in the breast. The biologic potential of the fibromatosis-like, spindle cell, metaplastic breast tumors most likely is defined by their major histologic phenotype; they are capable of local recurrence with no demonstrated distant spread or regional metastases, as in pure fibromatosis of the breast.

Adult↗

Unsupervised feature dimension reduction for classification of MR spectra.

We present an unsupervised feature dimension reduction method for the classification of magnetic resonance spectra. The technique preserves spectral information, important for disease profiling. We propose to use this technique as a preprocessing step for computationally demanding wrapper-based feature subset selection. We show that the classification accuracy on an independent test set can be sustained while achieving considerable feature reduction. Our method is applicable to other classification techniques, such as neural networks, support vector machines, etc.

Candida↗

Pulmonary CT image classification with evolutionary programming.

RATIONALE AND OBJECTIVES: It is often difficult to classify information in medical images from derived features. The purpose of this research was to investigate the use of evolutionary programming as a tool for selecting important features and generating algorithms to classify computed tomographic (CT) images of the lung. MATERIALS AND METHODS: Training and test sets consisting of 11 features derived from multiple lung CT images were generated, along with an indicator of the target area from which features originated. The images included five parameters based on histogram analysis, 11 parameters based on run length and co-occurrence matrix measures, and the fractal dimension. Two classification experiments were performed. In the first, the classification task was to distinguish between the subtle but known differences between anterior and posterior portions of transverse lung CT sections. The second classification task was to distinguish normal lung CT images from emphysematous images. The performance of the evolutionary programming approach was compared with that of three statistical classifiers that used the same training and test sets. RESULTS: Evolutionary programming produced solutions that compared favorably with those of the statistical classifiers. In separating the anterior from the posterior lung sections, the evolutionary programming results were better than two of the three statistical approaches. The evolutionary programming approach correctly identified all the normal and abnormal lung images and accomplished this by using less features than the best statistical method. CONCLUSION: The results of this study demonstrate the utility of evolutionary programming as a tool for developing classification algorithms.

Algorithms↗

An optimized index of human cardiovascular adaptation to simulated weightlessness.

Prolonged exposure to weightlessness is known to produce a variety of cardiovascular changes, some of which may influence the astronaut's performance during a mission. In order to find a reliable indicator of cardiovascular adaptation to weightlessness, we analyzed data from nine male subjects after a 24-hour period of normal activity and after a period of simulated weightlessness produced by two hours in a launch position followed by 20 hours of 6 degrees head-down tilt plus pharmacologically induced diuresis (furosemide). Heart rate, arterial pressure, thoracic fluid index, and radial flow were analyzed. Autoregressive spectral estimation and decomposition were used to obtain the spectral components of each variable from the subjects in the supine position during pre- and post-simulated weightlessness. We found a significant decrease in heart rate power and an increase in thoracic fluid index power in the high frequency region (0.2-0.45 Hz) and significant increases in radial flow and arterial pressure powers in the low frequency region (<0.2 Hz) in response to simulated weightlessness. However, due to the variability among subjects, any single variable appeared limited as a dependable index of cardiovascular adaptation to weightlessness. The backward elimination algorithm was then used to select the best discriminatory features from these spectral components. Fisher's linear discriminant and Bayes' quadratic discriminant were used to combine the selected features to obtain an optimal index of adaptation to simulated weightlessness. Results showed that both techniques provided improved discriminant performance over any single variable and thus have the potential for use as an index to track adaptation and prescribe countermeasures to the effects of weightlessness.

Adaptation, Physiological↗

The syndrome of Gogi (word meaning) aphasia. Selective impairment of kanji processing.

In 1943, Imura described an aphasic syndrome shown by Japanese patients and designated it as Gogi ("word-meaning") aphasia. Salient features are selective impairment of processing kanji or Chinese characters and difficulty in finding access to the lexicon in both production and reception, with preservation of processing kana or phonetic signs, and fluent oral repetition. A patient with this syndrome is presented, with emphasis on the nature of his kanji impairment. Cases of Gogi aphasia in the literature are reviewed and contrasted to cases of Broca's aphasia with selective impairment of kana processing. The implications for a neurolinguistic model of language processing in aphasia are discussed.

Adult↗

Automated location of dysplastic fields in colorectal histology using image texture analysis.

Automation in histopathology is an attractive concept and recent advances in the application of computerized expert systems and machine vision have made automated image analysis of histological images possible. Systems capable of complete automation not only require the ability to segment tissue features and grade histological abnormalities, but, must also be capable of locating diagnostically useful areas from within complex histological scenes. This is the first stage of the diagnostic process. The object of this study was to develop criteria for the automatic identification of focal areas of colorectal dysplasia from a background of histologically normal tissue. Fields of view representing normal colorectal mucosa (n = 120) and dysplastic mucosa (n = 120) were digitally captured and subjected to image texture analysis. Two features were selected as being the most important in the discrimination of normal and adenomatous colorectal mucosa. The first was a feature of the co-occurrence matrix and the second was the number of low optical density pixels in the image. A linear classification rule defined using these two features was capable of correctly classifying 86 per cent of a series of training images into their correct groups. In addition, large histological scenes were digitally captured, split into their component images, analysed according to texture, and classified as normal or abnormal using the previously defined classification rule. Maps of the histological scenes were constructed and in most cases, dysplastic colorectal mucosa was correctly identified on the basis of image texture: 83 per cent of test images were correctly classified. This study demonstrates that abnormalities in low-power tissue morphology can be identified using quantitative image analysis. The identification of diagnostically useful fields advances the potential of automated systems in histopathology: these regions could than be scrutinized at high power using knowledge-guided image segmentation for disease grading. Systems of this kind have the potential to provide objectivity, unbiased sampling, and valuable diagnostic decision support.

Adenomatous Polyps↗

The importance of internal loops within RNA substrates of ADAR1.

Adenosine deaminases that act on RNA (ADARs) are a family of RNA editing enzymes that convert adenosines to inosines within double-stranded RNA (dsRNA). Although ADARs deaminate perfectly base-paired dsRNA promiscuously, deamination is limited to a few, selected adenosines within dsRNA containing mismatches, bulges and internal loops. As a first step in understanding how RNA structural features promote selectivity, we investigated the role of internal loops within ADAR substrates. We observed that a dsRNA helix is deaminated at the same sites whether it exists as a free molecule or is flanked by internal loops. Thus, internal loops delineate helix ends for ADAR1. Since ADAR1 deaminates short RNAs at fewer adenosines than long RNAs, loops decrease the number of deaminations within an RNA by dividing a long RNA into shorter substrates. For a series of symmetric internal loops related in sequence, larger loops (>/=six nucleotides) acted as helix ends, whereas smaller loops (</=four nucleotides) did not. Our work provides the first information about how secondary structure within ADAR substrates dictates selectivity, and suggests a rational approach for delineating minimal substrates for RNAs deaminated by ADARs in vivo.

Adenosine Deaminase↗

Improving quantitative trait loci mapping resolution in experimental crosses by the use of genotypically selected samples.

One of the key factors contributing to the success of a quantitative trait locus (QTL) mapping experiment is the precision with which QTL positions can be estimated. We show, using simulations, that QTL mapping precision for an experimental cross can be increased by the use of a genotypically selected sample of individuals rather than an unselected sample of the same size. Selection is performed using a previously described method that optimizes the complementarity of the crossover sites within the sample. Although the increase in precision is accompanied by a decrease in QTL detection power at markers distant from QTL, only a modest increase in marker density is needed to obtain equivalent power over the whole map. Selected samples also show a slight reduction in the number of false-positive QTL. We find that two features of selected samples independently contribute to these effects: an increase in the number of crossover sites and increased evenness in crossover spacing. We provide an empirical formula for crossover enrichment in selected samples that is useful in experimental design and data analysis. For QTL studies in which the phenotyping is more of a limiting factor than the generation of individuals and the scoring of genotypes, selective sampling is an attractive strategy for increasing genome-wide QTL map resolution.

Chromosome Mapping↗

Multiresolution segmentation of three-dimensional medical images using mathematical morphology techniques.

A semi-automatic method for three-dimensional segmentation of medical images is proposed. A multiresolution representation is achieved through the application of morphological filters, which assures causality for image extrema. This allows for a compact scale space representation, in which each extremum is assigned a scale value. Interactive selection of the interesting extrema of the image is carried out, aided by this scale information and other relevant features. Extrema selected are then used as markers in three-dimensional watersheds calculation. The system has been developed and tested under low cost platforms, and can be the base for totally automatic, knowledge based segmentation systems.

Artificial Intelligence↗

Automatic selection of uncontaminated electromyogram as applied to respiratory muscle fatigue.

An automatic procedure for detecting artifacts in the electromyogram (EMG) has been developed and applied to a study of respiratory muscle fatigue. Signal segments are characterized by a set of features, the normal variations of which have been estimated in a training session. From the features are calculated a classification variable, which expresses the degree of deviation from normal conditions. A deviation larger than a certain threshold value designates a segment as disturbed. The study deals with the choice of features, the selection of a suitable segment length, and the determination of an optimal classification threshold. The four features chosen include measures of amplitude symmetry, extreme excursions in the signal tracing, the signal-to-noise ratio, and the shape of the EMG power spectrum. Recordings from three subjects were used for the evaluation of the method. The results indicate that a segment length of 250 ms is appropriate. Accepting a 10% rate of false detections, the average rate of missed detections was 2.2%.

Diaphragm↗