Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Quantitation of the immunocytochemical assay for estrogen receptor protein (ER-ICA) in human breast cancer by television imaging.

A Quantimet 720D Image Analysis System has been programmed for light microscopic evaluation of the nuclear estrogen receptor distribution in frozen sections of human breast cancer stained by the peroxidase-antiperoxidase method using monoclonal antibodies to estrogen receptor protein (ER). This method provides precise criteria for distinguishing ER-positive and -negative cells and a sensitive and reproducible means for densitometric quantification of the staining patterns. Although imaging sequence and graphic analysis are automated by computer programs, light pen interaction provides supervision of feature selection. Imaging of the immunocytochemical assay (ER-ICA) in 50 patients revealed marked heterogeneity of nuclear estrogen receptor concentration varying over a nearly 100 fold concentration range. Various ER concentration patterns were evident: (I) distributions with a single peak (CV = 5%) present at various concentration levels; (II) bimodal distributions, revealing co-existent ER-positive and ER-negative subpopulations; (III) multimodal distributions with a number of resolvable concentration peaks; and (IV) highly skewed distributions with or without discernible peaks, frequently extending over the entire concentration range. Statistical methods of de-convolution were applied to determine the frequency and ER concentration characteristics of component subpopulations in the mosaic cases and for resolving the proportion of ER-positive and -negative cells. An approach for evaluating nuclear ER content in conjunction with ER concentration patterns in individual patients revealed whether spread in the ER concentration distribution resulted from differences in nuclear ER content or from variability in nuclear volume distribution.

Antibodies, Monoclonal↗

Clustering of the SOM easily reveals distinct gene expression patterns: results of a reanalysis of lymphoma study.

BACKGROUND: A method to evaluate and analyze the massive data generated by series of microarray experiments is of utmost importance to reveal the hidden patterns of gene expression. Because of the complexity and the high dimensionality of microarray gene expression profiles, the dimensional reduction of raw expression data and the feature selections necessary for, for example, classification of disease samples remains a challenge. To solve the problem we propose a two-level analysis. First self-organizing map (SOM) is used. SOM is a vector quantization method that simplifies and reduces the dimensionality of original measurements and visualizes individual tumor sample in a SOM component plane. Next, hierarchical clustering and K-means clustering is used to identify patterns of gene expression useful for classification of samples. RESULTS: We tested the two-level analysis on public data from diffuse large B-cell lymphomas. The analysis easily distinguished major gene expression patterns without the need for supervision: a germinal center-related, a proliferation, an inflammatory and a plasma cell differentiation-related gene expression pattern. The first three patterns matched the patterns described in the original publication using supervised clustering analysis, whereas the fourth one was novel. CONCLUSIONS: Our study shows that by using SOM as an intermediate step to analyze genome-wide gene expression data, the gene expression patterns can more easily be revealed. The "expression display" by the SOM component plane summarises the complicated data in a way that allows the clinician to evaluate the classification options rather than giving a fixed diagnosis.

Cluster Analysis↗

Computational protein biomarker prediction: a case study for prostate cancer.

BACKGROUND: Recent technological advances in mass spectrometry pose challenges in computational mathematics and statistics to process the mass spectral data into predictive models with clinical and biological significance. We discuss several classification-based approaches to finding protein biomarker candidates using protein profiles obtained via mass spectrometry, and we assess their statistical significance. Our overall goal is to implicate peaks that have a high likelihood of being biologically linked to a given disease state, and thus to narrow the search for biomarker candidates. RESULTS: Thorough cross-validation studies and randomization tests are performed on a prostate cancer dataset with over 300 patients, obtained at the Eastern Virginia Medical School using SELDI-TOF mass spectrometry. We obtain average classification accuracies of 87% on a four-group classification problem using a two-stage linear SVM-based procedure and just 13 peaks, with other methods performing comparably. CONCLUSIONS: Modern feature selection and classification methods are powerful techniques for both the identification of biomarker candidates and the related problem of building predictive models from protein mass spectrometric profiles. Cross-validation and randomization are essential tools that must be performed carefully in order not to bias the results unfairly. However, only a biological validation and identification of the underlying proteins will ultimately confirm the actual value and power of any computational predictions.

Biomarkers, Tumor↗

Gene/protein name recognition based on support vector machine using dictionary as features.

BACKGROUND: Automated information extraction from biomedical literature is important because a vast amount of biomedical literature has been published. Recognition of the biomedical named entities is the first step in information extraction. We developed an automated recognition system based on the SVM algorithm and evaluated it in Task 1.A of BioCreAtIvE, a competition for automated gene/protein name recognition. RESULTS: In the work presented here, our recognition system uses the feature set of the word, the part-of-speech (POS), the orthography, the prefix, the suffix, and the preceding class. We call these features "internal resource features", i.e., features that can be found in the training data. Additionally, we consider the features of matching against dictionaries to be external resource features. We investigated and evaluated the effect of these features as well as the effect of tuning the parameters of the SVM algorithm. We found that the dictionary matching features contributed slightly to the improvement in the performance of the f-score. We attribute this to the possibility that the dictionary matching features might overlap with other features in the current multiple feature setting. CONCLUSION: During SVM learning, each feature alone had a marginally positive effect on system performance. This supports the fact that the SVM algorithm is robust on the high dimensionality of the feature vector space and means that feature selection is not required.

Algorithms↗

miTarget: microRNA target gene prediction using a support vector machine.

BACKGROUND: MicroRNAs (miRNAs) are small noncoding RNAs, which play significant roles as posttranscriptional regulators. The functions of animal miRNAs are generally based on complementarity for their 5' components. Although several computational miRNA target-gene prediction methods have been proposed, they still have limitations in revealing actual target genes. RESULTS: We implemented miTarget, a support vector machine (SVM) classifier for miRNA target gene prediction. It uses a radial basis function kernel as a similarity measure for SVM features, categorized by structural, thermodynamic, and position-based features. The latter features are introduced in this study for the first time and reflect the mechanism of miRNA binding. The SVM classifier produces high performance with a biologically relevant data set obtained from the literature, compared with previous tools. We predicted significant functions for human miR-1, miR-124a, and miR-373 using Gene Ontology (GO) analysis and revealed the importance of pairing at positions 4, 5, and 6 in the 5' region of a miRNA from a feature selection experiment. We also provide a web interface for the program. CONCLUSION: miTarget is a reliable miRNA target gene prediction tool and is a successful application of an SVM classifier. Compared with previous tools, its predictions are meaningful by GO analysis and its performance can be improved given more training examples.

Algorithms↗

Human fetal neuroblast and neuroblastoma transcriptome analysis confirms neuroblast origin and highlights neuroblastoma candidate genes.

BACKGROUND: Neuroblastoma tumor cells are assumed to originate from primitive neuroblasts giving rise to the sympathetic nervous system. Because these precursor cells are not detectable in postnatal life, their transcription profile has remained inaccessible for comparative data mining strategies in neuroblastoma. This study provides the first genome-wide mRNA expression profile of these human fetal sympathetic neuroblasts. To this purpose, small islets of normal neuroblasts were isolated by laser microdissection from human fetal adrenal glands. RESULTS: Expression of catecholamine metabolism genes, and neuronal and neuroendocrine markers in the neuroblasts indicated that the proper cells were microdissected. The similarities in expression profile between normal neuroblasts and malignant neuroblastomas provided strong evidence for the neuroblast origin hypothesis of neuroblastoma. Next, supervised feature selection was used to identify the genes that are differentially expressed in normal neuroblasts versus neuroblastoma tumors. This approach efficiently sifted out genes previously reported in neuroblastoma expression profiling studies; most importantly, it also highlighted a series of genes and pathways previously not mentioned in neuroblastoma biology but that were assumed to be involved in neuroblastoma pathogenesis. CONCLUSION: This unique dataset adds power to ongoing and future gene expression studies in neuroblastoma and will facilitate the identification of molecular targets for novel therapies. In addition, this neuroblast transcriptome resource could prove useful for the further study of human sympathoadrenal biogenesis.

Databases, Genetic↗

Machine learning-guided risk stratification in elderly AML based on genomic, immunophenotypic and therapeutic profiles.

BACKGROUND: Elderly patients with acute myeloid leukemia (AML) exhibit considerable biological and clinical heterogeneity, hindering precise prognosis. Existing prognostic systems inadequately capture the complexity of elderly AML due to their reliance on data from younger cohorts and omission of key factors like immunophenotypic markers and therapeutic profiles. This study aimed to develop and internally validate a machine learning-based prognostic model specifically tailored to elderly AML patients. METHODS: A total of 156 patients were analyzed using a two-stage modeling strategy. Clinical and genomic variables were modeled first, followed by independent analysis of immunophenotypic features. Feature selection was performed using multilayer perceptron (MLP) and random forest (RF), while multivariate Cox regression was used for final model construction. Internal validation was conducted using 1000 bootstrap iterations to assess model stability and performance. RESULTS: The model demonstrated strong predictive performance, with a concordance index (C-index) of 0.702. Time-dependent area under the curve (AUC) and calibration plots confirmed accurate prediction of 1-, 3-, and 5-year overall survival. Decision curve analysis indicated favorable net benefit across a range of threshold probabilities. Key independent prognostic factors identified included TP53 mutations, high CD13 expression, and IDH2 mutations. CONCLUSION: This model provides a robust and interpretable tool for individualized risk stratification in elderly AML. By integrating genomic, immunophenotypic, and therapeutic variables, it may help optimize treatment decisions and improve outcomes for this vulnerable population. Future efforts should focus on external validation and integration of dynamic biomarkers.

Humans↗

Construction of a prognostic model for gastric cancer based on immune infiltration and microenvironment, and exploration of MEF2C gene function.

BACKGROUND: Advanced gastric cancer (GC) exhibits a high recurrence rate and a dismal prognosis. Myocyte enhancer factor 2c (MEF2C) was found to contribute to the development of various types of cancer. Therefore, our aim is to develop a prognostic model that predicts the prognosis of GC patients and initially explore the role of MEF2C in immunotherapy for GC. METHODS: Transcriptome sequence data of GC was obtained from The Cancer Genome Atlas (TCGA), the Gene Expression Omnibus (GEO) and PRJEB25780 cohort for subsequent immune infiltration analysis, immune microenvironment analysis, consensus clustering analysis and feature selection for definition and classification of gene M and N. Principal component analysis (PCA) modeling was performed based on gene M and N for the calculation of immune checkpoint inhibitor (ICI) Score. Then, a Nomogram was constructed and evaluated for predicting the prognosis of GC patients, based on univariate and multivariate Cox regression. Functional enrichment analysis was performed to initially investigate the potential biological mechanisms. Through Genomics of Drug Sensitivity in Cancer (GDSC) dataset, the estimated IC50 values of several chemotherapeutic drugs were calculated. Tumor-related transcription factors (TFs) were retrieved from the Cistrome Cancer database and utilized our model to screen these TFs, and weighted correlation network analysis (WGCNA) was performed to identify transcription factors strongly associated with immunotherapy in GC. Finally, 10 patients with advanced GC were enrolled from Sun Yat-sen University Cancer Center, including paired tumor tissues, paracancerous tissues and peritoneal metastases, for preparing sequencing library, in order to perform external validation. RESULTS: Lower ICI Score was correlated with improved prognosis in both the training and validation cohorts. First, lower mutant-allele tumor heterogeneity (MATH) was associated with lower ICI Score, and those GC patients with lower MATH and lower ICI Score had the best prognosis. Second, regardless of the T or N staging, the low ICI Score group had significantly higher overall survival (OS) compared to the high ICI Score group. For its mechanisms, consistently, for Camptothecin, Doxorubicin, Mitomycin, Docetaxel, Cisplatin, Vinblastine, Sorafenib and Paclitaxel, all of the IC50 values were significantly lower in the low ICI Score group compared to the high ICI Score group. As a result, based on univariate and multivariate Cox regression, ICI Score was considered to be an independent prognostic factor for GC. And our Nomogram showed good agreement between predicted and actual probabilities. Based on CIBERSORT deconvolution analysis, there was difference of immune cell composition found between high and low ICI Score groups, probably affecting the efficacy of immunotherapy. Then, MEF2C, a tumor-related transcription factor, was screened out by WGCNA analysis. Higher MEF2C expression is significantly correlated with a worse OS. Moreover, its higher expression is also negatively correlated with tumor mutation burden (TMB) and microsatellite instability (MSI), but positively correlated with several immunosuppressive molecules, indicating MEF2C may exert its influence on tumor development by upregulating immunosuppressive molecules. Finally, based on transcriptome sequencing data on 10 paired tumor tissues from Sun Yat-sen University Cancer Center, MEF2C expression was significantly lower in paracancerous tissues compared to tumor tissues and peritoneal metastases, and it was also lower in tumor tissues compared to peritoneal metastases, indicating a potential positive association between MEF2C expression and tumor invasiveness. CONCLUSIONS: Our prognostic model can effectively predict outcomes and facilitate stratification GC patients, offering valuable insights for clinical decision-making. The identified transcription factor MEF2C can serve as a biomarker for assessing the efficacy of immunotherapy for GC.

Humans↗

Distinct immune-metabolic phenotypes underlie poor coronary collateral circulation.

BACKGROUND: Coronary collateral circulation (CCC) significantly impacts myocardial perfusion and clinical outcomes in coronary artery disease patients, yet the underlying molecular heterogeneity remains inadequately characterized. OBJECTIVE: To identify distinct molecular phenotypes in patients with poor CCC, validate these phenotypes using clinical parameters, and evaluate their prognostic implications. METHODS: This study enrolled 149 patients (80 with good CCC and 69 with poor CCC) for high-throughput proteomic profiling. Unsupervised consensus clustering identified molecular subtypes within poor CCC patients, followed by differential expression analysis and KEGG pathway enrichment. Boruta feature selection was implemented, and multiple machine learning algorithms were tested on clinical data, with XGBoost optimization (accuracy 80.0%, F1-score 80.31%) and SHAP value interpretation. External validation was performed using the MIMIC database. Kaplan-Meier analysis and Cox regression models assessed major adverse cardiovascular events (MACE). RESULTS: Two distinct phenotypes emerged among poor CCC patients: Cluster 1 (n&#x2009;=&#x2009;39, Complement-Driven Vascular Remodeling [CDVR]) and Cluster 2 (n&#x2009;=&#x2009;30, Immuno-Thrombotic Myocardial Dysfunction [ITMD]). An XGBoost model incorporating fasting glucose, eosinophil percentage, and HbA1c achieved excellent discrimination (AUC&#x2009;>&#x2009;0.91). External validation confirmed the phenotype-specific clinical patterns. Notably, Cluster 2 demonstrated significantly higher MACE incidence compared to Cluster 1 (Log-rank p&#x2009;<&#x2009;0.05), with KEGG analysis revealing significant upregulation of platelet activation, diabetic cardiomyopathy, and metabolic pathways in the ITMD phenotype. CONCLUSION: Poor CCC encompasses distinct immune-metabolic phenotypes that can be accurately classified using integrated proteomic-clinical modeling. This classification enables more precise risk stratification and may guide personalized therapeutic strategies for coronary artery disease patients with inadequate collateralization.

Humans↗

Radiogenomic MRI biomarkers for noninvasive prediction of GPC3 expression and tumor microenvironment in hepatocellular carcinoma.

BACKGROUND: Glypican-3 (GPC3) is frequently overexpressed in hepatocellular carcinoma (HCC) and plays a key role in immune and metabolic remodeling of the tumor microenvironment. Reliable noninvasive biomarkers for predicting GPC3 status could improve patient stratification and support precision immunotherapy. METHODS: This multicenter retrospective study included 274 patients with pathologically confirmed hepatocellular carcinoma from three institutions, 34 external cases with MRI from The Cancer Imaging Archive, and 363 transcriptomic profiles from The Cancer Genome Atlas. Contrast-enhanced T1-weighted imaging and diffusion-weighted imaging were analyzed. Tumor and peritumoral regions were segmented manually and radiomic features extracted using PyRadiomics. Feature selection was performed with correlation filtering and least absolute shrinkage and selection operator regression. Machine learning classifiers including logistic regression, random forest, support vector machine, k-nearest neighbor, and decision tree were trained with 10-fold cross-validation and tested on independent external cohorts. A radiomics score was calculated for each patient. Radiogenomic analysis correlated radiomics scores with transcriptomic data using weighted gene co-expression network analysis. Hub genes and enriched pathways were identified, and immune infiltration and predicted immunotherapy response were assessed using computational methods. RESULTS: The random forest model using contrast-enhanced T1-weighted imaging achieved an area under the curve of 0.966 in training and 0.935 in internal validation. The integrated contrast-enhanced T1-weighted imaging plus diffusion-weighted imaging model reached an internal validation area under the curve of 0.979. In external testing, the best performance was obtained with a support vector machine model (area under the curve 0.756). Radiomics scores were significantly correlated with GPC3 expression (R&#x2009;=&#x2009;0.78, p&#x2009;<&#x2009;0.05). Transcriptomic analysis identified a 10-gene signature enriched in hypoxia and lipid metabolism pathways that stratified patients into prognostic subgroups (concordance index 0.720, hazard ratio 4.07, p&#x2009;<&#x2009;0.0001). High-risk patients had greater immune infiltration and a lower predicted immune evasion score, suggesting a potential benefit from immunotherapy. CONCLUSIONS: MRI-based radiomics models can noninvasively predict GPC3 expression in hepatocellular carcinoma. Radiomics scores reflect underlying hypoxia and lipid metabolism pathways and stratify patients by prognosis and predicted immunotherapy response. These findings support radiogenomics as a translational approach to imaging-guided precision treatment in hepatocellular carcinoma.

Humans↗

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve&#x2013;based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC&#x2009;=&#x2009;0.844). The NODM cohort was stratified into high- (n&#x2009;=&#x2009;2,362) and low-risk (n&#x2009;=&#x2009;5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans↗

Accelerated Barnes maze test in mice for assessment of stress effects on memory.

Repeated restraint stress in rodents impairs spatial memory in a Y-maze test and induces hippocampal neuronal changes that last up to 5 d after the stressor ends. Our goal was to implement a Barnes maze spatial memory test in mice that could be used to validate our findings of social stress induced Y-maze impairment. We measured performance of mice in 5- and 9-day test paradigms previously used in rats and mice, respectively. Selecting features from each paradigm, we implemented a 5-d test (pre-training, training (4 trials/d/3 d) and probe testing for assessment of spatial memory in mice. Stress consisted of placing each test mouse in a stainless steel perforated box (25.5 cm x 21.5 cm x 16.5 cm) within an aggressor's home cage for 6 h/d for 21 d; direct agonistic encounters occurred randomly throughout stress periods. Barnes maze pre-training (habituation) was on day 21 of the stress exposures. In a preliminary experiment, mice that habituated following their last stressor performed poorly relative to unstressed and to those not habituated prior to the last stressor, as demonstrated by a greater latency to escape and more errors. We conclude that acute stress in a chronic stress paradigm may impair spatial memory acquisition.

Animals↗

Impact of age on cortisol secretory dynamics basally and as driven by nutrient-withdrawal stress.

The present study tests the clinical hypothesis that aging impairs homeostatic adaptations of cortisol secretion to stress. To this end, we implemented a short-term 3.5-day fast as an ethically acceptable metabolic stressor in eight young (ages 18-35 yr) and eight older (ages 60-72 yr) healthy men. Volunteers were studied in randomly ordered fed vs. fasting sessions. To capture the more complex dynamics of cortisol's feedback control, blood was sampled every 10 min for 24 h for later RIA of serum cortisol concentrations and quantitation of the pulsatile, entropic, and 24-h rhythmic modes of cortisol release using deconvolution analysis, the approximate entropy statistic, and cosine regression, respectively. The stress of fasting elevated the mean (24-h) serum cortisol concentration equivalently in the two age cohorts [i.e. from 7.2 +/- 0.35 to 11.6 +/- 0.71 microg/dL in young men and from 7.7 +/- 0.39 to 12.6 +/- 0.59 microg/dL in older individuals (P < 10(-7))]. The rise in integrated cortisol output was driven mechanistically by selective augmentation of cortisol secretory burst mass (P = 0.002). The resultant daily (pulsatile) cortisol secretion rate increased significantly but equally in young (from 94 +/- 6.3 to 151 +/- 15 microg/dL x day) and older (from 85 +/- 5.4 to 145 +/- 7.3 microg/dL x day) volunteers (P < 10(-4)). Nutrient restriction also prompted a marked reduction in the quantifiable regularity of (univariate) cortisol release patterns in both cohorts (P < 10(-4)). However, older men showed loss of joint synchrony of cortisol and LH secretion even in the fed state, which failed to change with metabolic stress (P < 10(-6)). In addition, older individuals maintained a premature (early-day) cortisol elevation in the fed state and unexpectedly evolved an anomalous further cortisol phase advance of 99 +/- 16 min during fasting (P < 10(-5)). Caloric deprivation in aging men also disproportionately elevated the mesor of 24-h rhythmic cortisol release (P = 10(-7)) and elicited a greater increment in the mean day-night variation in cortisol secretory-burst mass (P < 0.01 vs. young controls). Lastly, short-term caloric depletion in older subjects paradoxically normalized their age-associated suppression of the 24-h rhythm in cortisol interburst intervals. In summary, acute metabolic stress in healthy aging men (compared with young individuals) unmasks distinct, albeit complex, disruption of cortisol homeostasis. These dynamic anomalies impact the feedback-dependent and time-sensitive coupling of pulsatile and 24-h rhythmic cortisol secretion. Nutrient-withdrawal stress in the older male heightens the cortisol phase disparity already evident in fed elderly individuals. Conversely, the stress of fasting in young men paradoxically reproduces selected features of the aging unstressed (fed) cortisol axis; viz., abrogation of joint cortisol-LH synchrony and suppression of the normal diurnal variation in cortisol burst frequency. Whether fasting would unveil analogous disruption of feedback-dependent control of the corticotropic axis in healthy aging women is not yet known.

Adolescent↗

Pharmacotherapy for aphasia.

Selected features of aphasia may reflect disruption of specific neurotransmitter systems. Pharmacotherapy focused on these aphasic symptoms may improve language performance following stroke. We attempted to restore speech fluency in a patient with long-standing transcortical motor aphasia by treating his symptoms of hesitancy and impaired initiation of speech with bromocriptine. During therapy his language performance improved substantially, due to reduced latency of response, decreased paraphasias, and increased naming ability. After cessation of drug therapy his language returned to baseline.

Aphasia↗

Some applications of categorical data analysis to epidemiological studies.

Several examples of categorized data from epidemiological studies are analyzed to illustrate that more informative analysis than tests of independence can be performed by fitting models. All of the analyses fit into a unified conceptual framework that can be performed by weighted least squares. The methods presented show how to calculate point estimate of parameters, asymptotic variances, and asymptotically valid chi 2 tests. The examples presented are analysis of relative risks estimated from several 2 x 2 tables, analysis of selected features of life tables, construction of synthetic life tables from cross-sectional studies, and analysis of dose-response curves.

Actuarial Analysis↗

Sampling irregularity perturbs visual reconstruction.

We explored the human observer's ability to detect and discriminate sine-wave and square-wave gratings that were sampled at intervals varying from 4.7 to 9.4 arcmin. To study the effect of sampling irregularity on visual performance, we varied the position of each line sample on the basis of a Gaussian probability distribution, the standard deviation of which varied from 0 (regular sampling) to 4.7 arcmin (highly irregular sampling). The results indicate that irregular sampling has no systematic effect on the observer's ability merely to detect the presence of a sine- or square-wave grating. In contrast, sampling irregularity strongly impairs the subject's ability to discriminate between these waveforms. A model based on the convolution of difference-of-Gaussians-type weighting profiles predicted that sampling irregularity should have little to no effect on the output of a channel tuned to the third harmonic of the square-wave grating. The findings thus suggest the existence of a sampling scheme in the visual system. This scheme is based on local feature-selective mechanisms, probably edge detectors, that are highly sensitive to the relative position of the sample points in the space domain.

Discrimination, Psychological↗

AWGE-ESPCA: An edge sparse PCA model based on adaptive noise elimination regularization and weighted gene network for Hermetia illucens genomic data analysis.

Hermetia illucens is an important insect resource. Studies have shown that exploring the effects of Cu2+-stressed on the growth and development of the Hermetia illucens genome holds significant scientific importance. There are three major challenges in the current studies of Hermetia illucens genomic data analysis: firstly, the lack of available genomic data which limits researchers in Hermetia illucens genomic data analysis. Secondly, to the best of our knowledge, there are no Artificial Intelligence (AI) feature selection models designed specifically for Hermetia illucens genome. Unlike human genomic data, noise in Hermetia illucens data is a more serious problem. Third, how to choose those genes located in the pathway enrichment region. Existing models assume that each gene probe has the same priori weight. However, researchers usually pay more attention to gene probes which are in the pathway enrichment region. Based on the above challenges, we initially construct experiments and establish a new Cu2+-stressed Hermetia illucens growth genome dataset. Subsequently, we propose AWGE-ESPCA: an edge Sparse PCA model based on adaptive noise elimination regularization and weighted gene network. The AWGE-ESPCA model innovatively proposes an adaptive noise elimination regularization method, effectively addressing the noise challenge in Hermetia illucens genomic data. We also integrate the known gene-pathway quantitative information into the Sparse PCA(SPCA) framework as a priori knowledge, which allows the model to filter out the gene probes in pathway-rich regions as much as possible. Ultimately, this study conducts five independent experiments and compared four latest Sparse PCA models as well as representative supervised and unsupervised baseline models to validate the model performance. The experimental results demonstrate the superior pathway and gene selection capabilities of the AWGE-ESPCA model. Ablation experiments validate the role of the adaptive regularizer and network weighting module. To summarize, this paper presents an innovative unsupervised model for Hermetia illucens genome analysis, which can effectively help researchers identify potential biomarkers. In addition, we also provide a working AWGE - ESPCA model code in the address: https://github.com/yhyresearcher/AWGE_ESPCA.

Animals↗

A framework for block-wise missing data in multi-omics.

High-throughput technologies have generated vast amounts of omic data. It is a consensus that the integration of diverse omics sources improves predictive models and biomarker discovery. However, managing multiple omics data poses challenges such as data heterogeneity, noise, high-dimensionality and missing data, especially in block-wise patterns. This study addresses the challenges of high dimensionality and block-wise missing data through a regularization and constrained-based approach. The methodology is implemented in the R package bwm for binary and continuous response variables, and applied to breast cancer and exposome multi-omics datasets, achieving strong performance even in scenarios with missing data present in all omics. In binary classification task, our proposed model achieves accuracy in the range of 86% to 92%, and F1 in the range of 68% to 79%. And, in regression task the correlation between true and predicted responses is in the range of 72% to 76%. However, there is a slight decline in performance metrics as the percentage of missing data increases. In scenarios where block-wise missing data affects multiple omics, the model performance actually surpasses that of scenarios where missing data is present in only one omics. One possible explanation for this might be that the other scenarios introduce a greater diversity of observation profiles, leading to a more robust model. Depending on the specific omics being studied, there is greater consistency in feature selection when comparing block-wise missing data scenarios.

Humans↗