Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Cloning of a partial cDNA for rat interleukin-12 (IL-12) and analysis of IL-12 expression in vivo.

Experimental models of autoimmunity in the rat may feature selective activation of either the Th1 or Th2 subset of helper T cells. Interleukin-12 (IL-12) is a key cytokine in the development of Th1 responses. In order to study IL-12 in the rat we used polymerase chain reaction (PCR) primers based on murine IL-12 to amplify a partial cDNA from rat tissue. The product was cloned and sequenced: it shows 94% nucleotide identity with the murine gene and 94% identity of predicted amino acid sequence. Primers based on the rat IL-12 sequence were used to analyse IL-12 expression in vivo using semi-quantitative PCR. We studied RNA from lymphoid tissues of two rat strains which differ in their response to mercuric chloride (HgCl2): Brown Norway (BN) rats develop autoimmunity with a predominant Th2 response; Lewis rats are resistant. Interleukin-12 expression was higher in Lewis than BN, and higher in spleen than lymph node. After HgCl2, IL-12 expression increased in BN towards the time when the autoimmune response autoregulates. Variation in baseline levels of IL-12 expression may account for the Th2 predisposition of BN rats compared to Lewis rats; IL-12 may play a role in the autoregulation of the Th2 response induced by HgCl2.

Amino Acid Sequence↗

Nodal staging of colorectal carcinomas from quantitative and qualitative aspects. Can lymphatic mapping help staging?

Retrospective data analysis was performed to determine the minimum number of lymph nodes required for the staging of colorectal carcinomas, and a prospective feasibility study was carried out to identify sentinel nodes in order to clarify whether these may predict the nodal status. From among 240 colorectal carcinoma specimens investigated between 1996 and 1998, 224 tumors were analyzed for their nodal status. Lymphatic mapping with vital patent blue dye injection into the peritumoral sub-serosal layer was performed in 25 patients. Blue nodes were identified by the pathologist in the unfixed specimen immediately after the resection of the bowel and were assessed separately. Of the 123 node-positive carcinomas, 40 had more than 3 nodes involved. The nodal positivity increased substantially when more than 6 nodes were assessed. The cumulative percentage analysis demonstrated that ideally 16 and 13 nodes should be obtained for the identification of any nodal involvement or the involvement of more than 3 nodes, respectively. Lymphatic mapping was successful in 24 patients (96%). Blue nodes were predictive of the nodal status in 19 cases (79%), and were the only sites of metastasis in 2 patients (15% of the node-positive cases). Lymphatic mapping with the vital blue dye technique does not seem to facilitate the staging of colorectal cancers, at least in our patient population with relatively large and deeply infiltrating tumors, and unless the technique is improved or other selective features of lymph nodes are found, all lymph nodes should be assessed. A minimum of 6 nodes, and an optimum of 16 nodes or more, are suggested from these series.

Adult↗

Predicting interpretability of metabolome models based on behavior, putative identity, and biological relevance of explanatory signals.

Powerful algorithms are required to deal with the dimensionality of metabolomics data. Although many achieve high classification accuracy, the models they generate have limited value unless it can be demonstrated that they are reproducible and statistically relevant to the biological problem under investigation. Random forest (RF) generates models, without any requirement for dimensionality reduction or feature selection, in which individual variables are ranked for significance and displayed in an explicit manner. In metabolome fingerprinting by mass spectrometry, each metabolite can be represented by signals at several m/z. Exploiting a prior understanding of expected biochemical differences between sample classes, we aimed to develop meaningful metrics relevant to the significance both of the overall RF model and individual, potentially explanatory, signals. Pair-wise comparison of related plant genotypes with strong phenotypic differences demonstrated that robust models are not only reproducible but also logically structured, highlighting correlated m/z derived from just a small number of explanatory metabolites reflecting the biological differences between sample classes. RF models were also generated by using groupings of samples known to be increasingly phenotypically similar. Although classification accuracy was often reasonable, we demonstrated reproducibly in both Arabidopsis and potato a performance threshold based on margin statistics beyond which such models showed little structure indicative of either generalizability or further biological interpretability. In a multiclass problem using 25 Arabidopsis genotypes, despite the complicating effects of ecotype background and secondary metabolome perturbations common to several mutations, the ranking of metabolome signals by RF provided scope for deeper interpretability.

Arabidopsis↗

Opposing effects of estradiol and progesterone on oxytocin receptors in rabbit uterus.

Estradiol-17beta administration to young (10- to 12-week-old) rabbits to produce the "estrogen-dominated" uterus increased the uterine contractile response to both oxytocin and methacholine in vitro. In "progesterone-dominated" uteri, obtained from rabbits that received progesterone for 4 days after estrogen pretreatment, the contractile response to oxytocin in vitro was selectively abolished; the response to methacholine was unaffected. Parallel changes were observed in the concentration (but not affinity) of specific sites in uterine microsomal membranes that bind [(3)H]oxytocin with selectivity features expected for oxytocin receptors. Thus, estrogen-dominated uteri have an increased number of specific [(3)H]oxytocin binding sites per mg of membrane protein relative to untreated controls, whereas specific oxytocin binding sites are reduced to barely detectable levels in the progesterone-dominated uterus. Similar results are obtained when binding sites are measured in membranes from the myometrium of estrogen- or progesterone-dominated uteri. Short-term (24-hr) progesterone administration to estrogen-pretreated rabbits decreased, but did not abolish, specific [(3)H]oxytocin binding; the concentration of specific [(3)H]oxytocin binding sites was reduced without influence on the affinity of these sites. A sublethal dose of actinomycin D, administered over a 24-hr period to rabbits pretreated with estradiol for 4 days, likewise reduced specific oxytocin binding; additive effects were not observed when progesterone and actinomycin D were administered together. These results suggest that the regulatory effects of estrogens and progesterone upon the rabbit uterine contractile response to oxytocin are achieved, at least in part, by the opposing actions of these steroids in regulating the number of oxytocin receptors in smooth muscle cells. Estradiol increased the concentration of uterine oxytocin receptors; the maintenance of high receptor levels appears to depend upon the continuous de novo synthesis of oxytocin receptors. In contrast, progesterone, like actinomycin D, appears to act at the nuclear locus to repress synthesis of oxytocin receptors.

Animals↗

Cross-talk between G-protein and protein kinase C modulation of N-type calcium channels is dependent on the G-protein beta subunit isoform.

The modulation of N-type calcium current by protein kinases and G-proteins is a factor in the fine tuning of neurotransmitter release. We have previously shown that phosphorylation of threonine 422 in the alpha(1B) calcium channel domain I-II linker region resulted in a dramatic reduction in somatostatin receptor-mediated G-protein inhibition of the channels and that the I-II linker consequently serves as an integration center for cross-talk between protein kinase C (PKC) and G-proteins (Hamid, J., Nelson, D., Spaetgens, R., Dubel, S. J., Snutch, T. P., and Zamponi, G. W. (1999) J. Biol. Chem. 274, 6195-6202). Here we show that opioid receptor-mediated inhibition of N-type channels is affected to a lesser extent compared with that seen with somatostatin receptors, hinting at the possibility that PKC/G-protein cross-talk might be dependent on the G-protein subtype. To address this issue, we have examined the effects of four different types of G-protein beta subunits on both wild type and mutant alpha(1B) calcium channels in which residue 422 has been replaced by glutamate to mimic PKC-dependent phosphorylation and on channels that have been directly phosphorylated by protein kinase C. Our data show that phosphorylation or mutation of residue 422 antagonizes the effect of Gbeta(1) on channel activity, whereas Gbeta(2), Gbeta(3), and Gbeta(4) are not affected. Our data therefore suggest that the observed cross-talk between G-proteins and protein kinase C modulation of N-type channels is a selective feature of the Gbeta(1) subunit.

Calcium Channels, N-Type↗

Aspartate residues of the Glu-Glu-Asp-Asp (EEDD) pore locus control selectivity and permeation of the T-type Ca(2+) channel alpha(1G).

The structural determinant of the permeation and selectivity properties of high voltage-activated (HVA) Ca(2+) channels is a locus formed by four glutamate residues (EEEE), one in each P-region of the domains I-IV of the alpha(1) subunit. We tested whether the divergent aspartate residues of the EEDD locus of low voltage-activated (LVA or T-type) Ca(2+) channels account for the distinctive permeation and selectivity features of these channels. Using the whole-cell patch-clamp technique in the HEK293 expression system, we studied the properties of the alpha(1G) T-type, the alpha(1C) L-type Ca(2+) channel subunits, and alpha(1G) pore mutants, containing aspartate-to-glutamate conversions in domain III, domain IV, or both. Three characteristic features of HVA Ca(2+) channel permeation, i.e. (a) Ba(2+) over Ca(2+) permeability, (b) Ca(2+)/Ba(2+) anomalous mole fraction effect (AMFE), and (c) high Cd(2+) sensitivity, were conferred on the domain III mutant (EEED) of alpha(1G). In contrast, the relative Ca(2+)/Ba(2+) permeability and the lack of AMFE of the alpha(1G) wild type channel were retained in the domain IV mutant (EEDE). The double mutant (EEEE) displayed AMFE and a Cd(2+) sensitivity similar to that of alpha(1C), but currents were larger in Ca(2+)- than in Ba(2+)-containing solutions. The mutation in domain III, but not that in domain IV, consistently displayed outward fluxes of monovalent cations. H(+) blocked Ca(2+) currents in all mutants more efficiently than in alpha(1G). In addition, activation curves of all mutants were displaced to more positive voltages and had a larger slope factor than in alpha(1G) wild type. We conclude that the aspartate residues of the EEDD locus of the alpha(1G) Ca(2+) channel subunit not only control its permeation properties, but also affect its activation curve. The mutation of both divergent aspartates only partially confers HVA channel permeation properties to the alpha(1G) Ca(2+) channel subunit.

Aspartic Acid↗

The conformation of a signal peptide bound by Escherichia coli preprotein translocase SecA.

To understand the structural nature of signal sequence recognition by the preprotein translocase SecA, we have characterized the interactions of a signal peptide corresponding to a LamB signal sequence (modified to enhance aqueous solubility) with SecA by NMR methods. One-dimensional NMR studies showed that the signal peptide binds SecA with a moderately fast exchange rate (Kd approximately 10(-5) m). The line-broadening effects observed from one-dimensional and two-dimensional NMR spectra indicated that the binding mode does not equally immobilize all segments of this peptide. The positively charged arginine residues of the n-region and the hydrophobic residues of the h-region had less mobility than the polar residues of the c-region in the SecA-bound state, suggesting that this peptide has both electrostatic and hydrophobic interactions with the binding pocket of SecA. Transferred nuclear Overhauser experiments revealed that the h-region and part of the c-region of the signal peptide form an alpha-helical conformation upon binding to SecA. One side of the hydrophobic core of the helical h-region appeared to be more strongly bound in the binding pocket, whereas the extreme C terminus of the peptide was not intimately involved. These results argue that the positive charges at the n-region and the hydrophobic helical h-region are the selective features for recognition of signal sequences by SecA and that the signal peptide-binding site on SecA is not fully buried within its structure.

Adenosine Triphosphatases↗

Assessing individual genetic susceptibility to metabolic syndrome: interpretable machine learning method.

BACKGROUND: Genome-wide association studies have provided profound insights into the genetic aetiology of metabolic syndrome (MetS). However, there is a lack of machine-learning (ML)-based predictive models to assess individual genetic susceptibility to MetS. This study utilized single-nucleotide polymorphisms (SNPs) as variables and employed ML-based genetic risk score (GRS) models to predict the occurrence of MetS, bringing it closer to clinical application. METHODS: Feature selection was performed using Least Absolute Shrinkage and Selection Operator. Six ML algorithms were employed to construct GRS models. A fivefold cross-validation was utilized to aid in the internal validation of models. The receiver operating characteristic (ROC) curve was used to select the better-performing GRS model. The SHapley Additive exPlanations (SHAP) was then applied to interpret the model. After extracting GRS, stratified analysis of BMI, age and gender was performed. Finally, these conventional risk factors and GRS were integrated through multivariate logistic regression to establish a combined model. RESULTS: A total of 17 SNPs were selected for analysis. Among the GRS models, the extreme gradient boosting (XGBoost) model demonstrated superior discriminative performance (AUC = 0.837). The XGBoost's optimal robustness was also validated through five-fold cross-validation (mean ROC-AUC = 0.706). The XGBoost-based SHAP algorithm not only elucidated the global effects of 17 SNPs across all samples, but also described the interaction between SNPs, providing a visual representation of how SNPs impact the prediction of MetS in an individual. There was a strong correlation between GRS and MetS risk, particularly observed among young individuals, males and overweight individuals. Furthermore, the model combining conventional risk factors and GRS exhibited excellent discriminative performance (AUC = 0.962) and outstanding robustness (mean ROC-AUC = 0.959). CONCLUSION: This study established a reliable XGBoost-based GRS model and a GRS prediction platform (https://metabolicsyndromeapps.shinyapps.io/geneticriskscore/) to assess individual genetic susceptibility to MetS. This model has high interpretability and can provide personalized reference for determining the necessity of primary prevention measures for MetS. Additionally, there may be interactions between traditional risk factors and GRS, and the integration of both in a comprehensive model is useful in the prediction of MetS occurrence.

Humans↗

Stage-specific ROMO1 in rheumatoid arthritis: predictive immune insights into the MIF pathway and HLA-DR/IL2RA axis via integrated GWAS, transcriptomic, single-cell, and spatial profiling.

Emerging evidence links reactive oxygen species modulator 1 (ROMO1), a key mitochondrial ROS regulator, to rheumatoid arthritis (RA) pathogenesis. However, its exact mechanism remains elusive given the conflicting evidence about its specific function. We used a four-level integrative framework combining multi-omics data and literature‑supported mechanistic inference. At the genetic level, Mendelian randomization (MR) was performed to explore potential causal relationships between ROMO1, IL2RA, HLA-DR, MIF, and RA risk, followed by differential expression analysis and machine learning-based feature selection to identify key mROS genes. The temporal expression dynamics of ROMO1 were assessed in RA progression. At the cellular and tissue levels, we integrated single-cell RNA sequencing and spatial transcriptomics to map cell-type-specific expression and synovial localization of ROMO1-related immune cells and pathways. Finally, our multi-omics findings were contextualized with literature-supported mechanistic inference. (1) MR results were consistent with a potential protective effect of ROMO1 on RA (OR = 0.52) and its potential regulation of risk factors IL2RA (OR = 0.46) and HLA-DR (OR = 0.40). Conversely, IL2RA (OR = 1.42), HLA-DR (OR = 1.88), and MIF (OR = 1.17) were positively associated with RA risk. Additionally, ROMO1 was identified as a top candidate diagnostic predictor with stage-specific dynamics: downregulated in the early but upregulated in the late/remission stages. (2) Single-cell RNA sequencing showed ROMO1's cell-specific expression in CD14+ HLA-DR+ CD74+ monocytes and CD4+ IL2RA+ T cells. Cell communication analysis further suggested that these cells may participate in MIF pathway regulation. Spatial transcriptomics subsequently identified that ROMO1-related cells localized to synovial pathological regions, with MIF pathway changes correlated with RA progression. (3) Finally, literature-supported mechanistic inference suggests that ROMO1 may modulate mROS levels to promote anti-inflammatory M2 macrophage polarization, which could theoretically contribute to reduced systemic inflammation and the alleviation of multi-organ decline in RA. This integrated multi-omics investigation, supported by literature-based mechanistic inference, suggests ROMO1 as a stage-dependent biomarker candidate and potential immune regulator in RA.

Humans↗

Data mining and knowledge discovery in predictive toxicology.

This article describes the knowledge discovery process in predictive toxicology. This process consists of five major steps (i) feature calculation, (ii) feature selection, (iii) model induction, (iv) model validation and (v) interpretation of predictions and models. Data mining is a part of the knowledge discovery process and consists of the application of data analysis and discovery algorithms, which can be useful in all of the above steps. A brief review of suitable algorithms and their advantages and disadvantages is given for each knowledge discovery step, followed by a more detailed description of a problem-specific implementation of the lazar prediction system.

Algorithms↗

Prediction of mammalian toxicity of organophosphorus pesticides from QSTR modeling.

Quantitative structure-toxicity relationship (QSTR) models were derived for estimating the acute oral toxicity of organophosphorus pesticides to male and female rats. The 51 chemicals of the training set and the nine compounds of the external testing set were described by means of autocorrelation vectors encoding lipophilicity, molar refractivity, H-bonding acceptor ability (HBA) and H-bonding donor ability (HBD) of the molecules. A feature selection was employed for selecting the most relevant autocorrelation descriptors. A PLS regression analysis and an artificial neural network (ANN) were used for deriving models accounting for the sex of the organisms in the estimation of the toxicity of pesticides. The best results were obtained with an 8/4/1 ANN model trained with theback-propagation and conjugate gradient descent algorithms. The root mean square residual (RMSR) values for the training set and the external testing set equaled 0.29 and 0.26, respectively.

Animals↗

Prediction of acute mammalian toxicity of organophosphorus pesticide compounds from molecular structure.

A quantitative structure-activity relationship (QSAR) investigation was done for the acute oral mammalian toxicity (LD50) of a set of 54 organophosphorus pesticide compounds. The compounds were represented with calculated molecular structure descriptors, which encoded their topological, electronic, and geometrical features. Feature selection was done with a genetic algorithm to find subsets of descriptors that would support a high quality computational neural network (CNN) model to link the structural descriptors to the -log(mmol/kg) values for the compounds. The best seven-descriptor non-linear CNN model found had an rms error of 0.22 log units for the training set compounds and 0.25 log units for the prediction set compounds.

Animals↗

Naïve observers' perceptions of family drawings by 7-year-olds with disorganized attachment histories.

Previous research has succeeded in distinguishing among drawings made by children with histories of organized attachment relationships (secure, avoidant, and resistant); however, drawings of children with histories of disorganized attachment have yet to be systematically investigated. The purpose of this study was to determine whether naïve observers would respond differentially to family drawings of 7-year-olds who were classified in infancy as disorganized vs. organized. Seventy-three undergraduate students from one university and 78 from a second viewed 50 family drawings of 7-year-olds (25 by children with organized infant attachment and 25 by children with disorganized infant attachment). Participants were asked to (1) circle the emotion that best described their reaction to the drawings and (2) rate the drawings on 6 bipolar scales. Drawings from children classified as disorganized in infancy evoked positive emotion labels less often and negative emotion labels more often than those children classified as organized. Furthermore, drawings from children classified as disorganized in infancy received higher ratings on scales for disorganization, carelessness, family chaos, bizarreness, uneasiness, and dysfunction. These data indicate that naive observers are relatively successful in distinguishing selected features of drawings by children with histories of disorganized vs. organized attachment.

Affect↗

Classification of surface EMG signals using harmonic wavelet packet transform.

In this paper, an efficient method based on the discrete harmonic wavelet packet transform (DHWPT) is presented to classify surface electromyographic (SEMG) signals. After the relative energy of SEMG signals in each frequency band had been extracted by the DHWPT, a genetic algorithm was utilized to select appropriate features in order to reduce the feature dimensionality. Then, the selected features were used as the input vectors to a neural network classifier to discriminate four types of prosthesis movements. Compared with other classification methods, the proposed method provided high classification accuracy in experimental research. In addition, this method could also save a lot of computational time because the DHWPT has a fast algorithm based on the fast Fourier transform for numerical implementation.

Animals↗

Towards adaptive classification for BCI.

Non-stationarities are ubiquitous in EEG signals. They are especially apparent in the use of EEG-based brain-computer interfaces (BCIs): (a) in the differences between the initial calibration measurement and the online operation of a BCI, or (b) caused by changes in the subject's brain processes during an experiment (e.g. due to fatigue, change of task involvement, etc). In this paper, we quantify for the first time such systematic evidence of statistical differences in data recorded during offline and online sessions. Furthermore, we propose novel techniques of investigating and visualizing data distributions, which are particularly useful for the analysis of (non-)stationarities. Our study shows that the brain signals used for control can change substantially from the offline calibration sessions to online control, and also within a single session. In addition to this general characterization of the signals, we propose several adaptive classification schemes and study their performance on data recorded during online experiments. An encouraging result of our study is that surprisingly simple adaptive methods in combination with an offline feature selection scheme can significantly increase BCI performance.

Adaptation, Physiological↗

Status of electronic reporting of notifiable conditions in the United States and Europe.

OBJECTIVE: To improve the computer connectivity and network strategies to connect U.S. county health departments (CHDs), state health departments (SHDs), and the Centers for Disease Control (CDC) for reporting notifiable conditions. METHODS: HSPNET-L mailing list discussions and individual Internet communications were used to compare selected features of notifiable conditions networking in the United States, France, and the United Kingdom. RESULTS: In the US, the CHD is the agency that first responds to an infectious disease outbreak on receiving notifications from physicians. Prompt recognition by the SHD that a widespread outbreak has occurred depends on the way in which county data are received, the "age" of the data, and the time taken to analyze them. Similarly, the recognition of the national scale of the outbreak depends on the promptness with which SHDs report to the CDC and the age of the data. An analysis of the French Communicable Disease Network suggests that an expansion of electronic links between US CHDs and SHDs will improve timeliness. Electronic data exchange allows CHDs to set up a local database and reduces transcription errors, mailing costs, and telephone costs. CONCLUSION: A fuller use of e-mail or other electronic communication by US CHDs will allow them to use a local database as a tool for managing local disease outbreaks more effectively and independently. Federal and state agency access to the CHD databases will enable early reporting of epidemic outbreaks. Periodic posting of public health information on Internet servers is recommended for immediate access to the public health data by Internet users worldwide.

Centers for Disease Control and Prevention, U.S.↗

Estimating population structure using epigenome-wide methylation data.

Population stratification is one of the source of inflation in epigenome-wide association studies (EWAS) when not properly accounted for. To address this, we developed methylation population scores (MPSs) to predict genetic principal components (GPCs) using a feature selection approach. We used multi-ethnic DNA methylation data from Illumina EPIC arrays across five cohorts, including MESA (n&#xa0;=&#xa0;929), CARDIA (n&#xa0;=&#xa0;1123), JHS (n&#xa0;=&#xa0;1365), ARIC (n&#xa0;=&#xa0;2338), and HCHS/SOL (n&#xa0;=&#xa0;1475), randomly splitting participants into training (85%) and test (15%) sets. Within each cohort, associations between GPCs and CpG sites were estimated using linear regression adjusting for age, sex, smoking and alcohol use, race/ethnicity, body mass index, and cell type proportions, followed by meta-analysis and selection of CpGs with FDR <0.05. We then applied a two-stage weighted least squares Lasso regression to construct MPSs, adjusting for the aforementioned covariates. In the test dataset, MPSs showed strong correlation with GPCs, with R&#xb2; ranging from 0.27 (MPS7 vs. GPC7) to 0.98 (MPS1 vs. GPC1). Visualization demonstrated that MPSs recapitulated the pattern shown by GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups and outperformed methylation-based principal components constructed using alternative published methods. Additionally, MPSs showed comparable performance to GPCs in reducing inflation in EWAS. Overall, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations, and provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent.

Humans↗

Tclass: tumor classification system based on gene expression profile.

A method that incorporates feature selection into Fisher's linear discriminant analysis for gene expression based tumor classification and a corresponding program Tclass were developed. The proposed method was applied to a public gene expression data set for colon cancer that consists of 22 normal and 40 tumor colon tissue samples to evaluate its performance for classification. Preliminary results demonstrated that using only a subset of genes ranging from 3 to 10 can achieve high classification accuracy.

Colonic Neoplasms↗