Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Collective posterior inference from highly variable empirical replicates.

High-throughput experimental platforms now routinely generate data from dozens or hundreds of independent observations. Simulation-based inference (SBI) offers a powerful framework for estimating model parameters from such complex datasets, but standard methods struggle to scale to the noisy multiple-replicates regime without incurring prohibitive computational costs or careful hyperparameter tuning. Here, we introduce a new method for fast and robust collective posterior inference from multiple independent replicates using a robust product-of-experts aggregation scheme that automatically mitigates the influence of outliers. Evaluating it on synthetic and empirical evolutionary datasets, we find it achieves state-of-the-art estimation accuracy and computational efficiency, including inference from noisy observations. Our method is compatible with any SBI framework, providing a scalable, plug-and-play solution for inference from noisy multiple-replicate datasets.

Computational Biology↗

NLCD: A method to discover nonlinear causal relations among genes.

Distinguishing correlation from causation is a fundamental challenge in many scientific fields, including biology, especially when interventions like randomized controlled trials are infeasible and only observational data are available. Methods based on statistical tests of conditional independence within the Mendelian Randomization framework can detect causality between two observed variables that are each associated with a third instrumental variable. However, these methods for detecting causal relationships between traits (e.g., two gene expression or clinical traits associated with a genetic variant, all observed in the same population) often assume a linear relationship, thereby hindering the discovery of causal gene networks from genomics data. We have developed NLCD, a method for NonLinear Causal Discovery from genomics data based on nonlinear regression modeling and conditional feature importance scoring. NLCD uses these techniques to extend the statistical tests in an existing linear causal discovery method called the Causal Inference Test (CIT). We benchmarked NLCD against current state-of-the-art methods: CIT, Findr, and MRPC. On simulated datasets, NLCD performs comparably to most methods in detecting linear relations (Average AUPRC (Area Under the Precision-Recall Curve) of NLCD = 0.94, CIT = 0.94, Findr = 0.94, and MRPC = 0.99), and outperforms them in detecting nonlinear (sine and sawtooth type) relations between two genes (Average AUPRC of NLCD = 0.76, CIT = 0.60, Findr = 0.56, and MRPC = 0.73). When tested on a nonlinear subset of a yeast genomic dataset to recover known causal relations involving transcription factors, NLCD and CIT performed comparable to each other and slightly better than Findr and MRPC (Average AUPRC of NLCD = 0.82, CIT = 0.81, Findr = 0.71, and MRPC = 0.54). On application to a human genomic dataset, NLCD revealed active causal gene pairs (IRF1 → PSME1 and HLA-C → HLA-T) in the muscle tissue, and clarified the promises and challenges in discovering causal gene networks in tissues under in vivo human settings.

Humans↗

Early history of mammals is elucidated with the ENCODE multiple species sequencing data.

Understanding the early evolution of placental mammals is one of the most challenging issues in mammalian phylogeny. Here, we addressed this question by using the sequence data of the ENCODE consortium, which include 1% of mammalian genomes in 18 species belonging to all main mammalian lineages. Phylogenetic reconstructions based on an unprecedented amount of coding sequences taken from 218 genes resulted in a highly supported tree placing the root of Placentalia between Afrotheria and Exafroplacentalia (Afrotheria hypothesis). This topology was validated by the phylogenetic analysis of a new class of genomic phylogenetic markers, the conserved noncoding sequences. Applying the tests of alternative topologies on the coding sequence dataset resulted in the rejection of the Atlantogenata hypothesis (Xenarthra grouping with Afrotheria), while this test rejected the second alternative scenario, the Epitheria hypothesis (Xenarthra at the base), when using the noncoding sequence dataset. Thus, the two datasets support the Afrotheria hypothesis; however, none can reject both of the remaining topological alternatives.

Animals↗

Creating and validating an algorithm to measure AIDS mortality in the adult population using verbal autopsy.

BACKGROUND: Vital registration and cause of death reporting is incomplete in the countries in which the HIV epidemic is most severe. A reliable tool that is independent of HIV status is needed for measuring the frequency of AIDS deaths and ultimately the impact of antiretroviral therapy on mortality. METHODS AND FINDINGS: A verbal autopsy questionnaire was administered to caregivers of 381 adults of known HIV status who died between 1998 and 2003 in Manicaland, eastern Zimbabwe. Individuals who were HIV positive and did not die in an accident or during childbirth (74%; n = 282) were considered to have died of AIDS in the gold standard. Verbal autopsies were randomly allocated to a training dataset (n = 279) to generate classification criteria or a test dataset (n = 102) to verify criteria. A rule-based algorithm created to minimise false positives had a specificity of 66% and a sensitivity of 76%. Eight predictors (weight loss, wasting, jaundice, herpes zoster, presence of abscesses or sores, oral candidiasis, acute respiratory tract infections, and vaginal tumours) were included in the algorithm. In the test dataset of verbal autopsies, 69% of deaths were correctly classified as AIDS/non-AIDS, and it was not necessary to invoke a differential diagnosis of tuberculosis. Presence of any one of these criteria gave a post-test probability of AIDS death of 0.84. CONCLUSIONS: Analysis of verbal autopsy data in this rural Zimbabwean population revealed a distinct pattern of signs and symptoms associated with AIDS mortality. Using these signs and symptoms, demographic surveillance data on AIDS deaths may allow for the estimation of AIDS mortality and even HIV prevalence.

Adolescent↗

A gene expression signature predicts survival of patients with stage I non-small cell lung cancer.

BACKGROUND: Lung cancer is the leading cause of cancer-related death in the United States. Nearly 50% of patients with stages I and II non-small cell lung cancer (NSCLC) will die from recurrent disease despite surgical resection. No reliable clinical or molecular predictors are currently available for identifying those at high risk for developing recurrent disease. As a consequence, it is not possible to select those high-risk patients for more aggressive therapies and assign less aggressive treatments to patients at low risk for recurrence. METHODS AND FINDINGS: In this study, we applied a meta-analysis of datasets from seven different microarray studies on NSCLC for differentially expressed genes related to survival time (under 2 y and over 5 y). A consensus set of 4,905 genes from these studies was selected, and systematic bias adjustment in the datasets was performed by distance-weighted discrimination (DWD). We identified a gene expression signature consisting of 64 genes that is highly predictive of which stage I lung cancer patients may benefit from more aggressive therapy. Kaplan-Meier analysis of the overall survival of stage I NSCLC patients with the 64-gene expression signature demonstrated that the high- and low-risk groups are significantly different in their overall survival. Of the 64 genes, 11 are related to cancer metastasis (APC, CDH8, IL8RB, LY6D, PCDHGA12, DSP, NID, ENPP2, CCR2, CASP8, and CASP10) and eight are involved in apoptosis (CASP8, CASP10, PIK3R1, BCL2, SON, INHA, PSEN1, and BIK). CONCLUSIONS: Our results indicate that gene expression signatures from several datasets can be reconciled. The resulting signature is useful in predicting survival of stage I NSCLC and might be useful in informing treatment decisions.

Algorithms↗

Assessing the emergence time of SARS-CoV-2 zoonotic spillover.

Understanding the evolution of Severe Acute Respiratory Syndrome Coronavirus (SARS-CoV-2) and its relationship to other coronaviruses in the wild is crucial for preventing future virus outbreaks. While the origin of the SARS-CoV-2 pandemic remains uncertain, mounting evidence suggests the direct involvement of the bat and pangolin coronaviruses in the evolution of the SARS-CoV-2 genome. To unravel the early days of a probable zoonotic spillover event, we analyzed genomic data from various coronavirus strains from both human and wild hosts. Bayesian phylogenetic analysis was performed using multiple datasets, using strict and relaxed clock evolutionary models to estimate the occurrence times of key speciation, gene transfer, and recombination events affecting the evolution of SARS-CoV-2 and its closest relatives. We found strong evidence supporting the presence of temporal structure in datasets containing SARS-CoV-2 variants, enabling us to estimate the time of SARS-CoV-2 zoonotic spillover between August and early October 2019. In contrast, datasets without SARS-CoV-2 variants provided mixed results in terms of temporal structure. However, they allowed us to establish that the presence of a statistically robust clade in the phylogenies of gene S and its receptor-binding (RBD) domain, including two bat (BANAL) and two Guangdong pangolin coronaviruses (CoVs), is due to the horizontal gene transfer of this gene from the bat CoV to the pangolin CoV that occurred in the middle of 2018. Importantly, this clade is closely located to SARS-CoV-2 in both phylogenies. This phylogenetic proximity had been explained by an RBD gene transfer from the Guangdong pangolin CoV to a very recent ancestor of SARS-CoV-2 in some earlier works in the field before the BANAL coronaviruses were discovered. Overall, our study provides valuable insights into the timeline and evolutionary dynamics of the SARS-CoV-2 pandemic.

Animals↗

WilsonGenAI a deep learning approach to classify pathogenic variants in Wilson Disease.

BACKGROUND: Advances in Next Generation Sequencing have made rapid variant discovery and detection widely accessible. To facilitate a better understanding of the nature of these variants, American College of Medical Genetics and Genomics and the Association of Molecular Pathologists (ACMG-AMP) have issued a set of guidelines for variant classification. However, given the vast number of variants associated with any disorder, it is impossible to manually apply these guidelines to all known variants. Machine learning methodologies offer a rapid way to classify large numbers of variants, as well as variants of uncertain significance as either pathogenic or benign. Here we classify ATP7B genetic variants by employing ML and AI algorithms trained on our well-annotated WilsonGen dataset. METHODS: We have trained and validated two algorithms: TabNet and XGBoost on a high-confidence dataset of manually annotated, ACMG & AMP classified variants of the ATP7B gene associated with Wilson's Disease. RESULTS: Using an independent validation dataset of ACMG & AMP classified variants, as well as a patient set of functionally validated variants, we showed how both algorithms perform and can be used to classify large numbers of variants in clinical as well as research settings. CONCLUSION: We have created a ready to deploy tool, that can classify variants linked with Wilson's disease as pathogenic or benign, which can be utilized by both clinicians and researchers to better understand the disease through the nature of genetic variants associated with it.

Hepatolenticular Degeneration↗

Data clustering in life sciences.

Clustering has a wide range of applications in life sciences and over the years has been used in many areas ranging from the analysis of clinical information, phylogeny, genomics, and proteomics. The primary goal of this article is to provide an overview of the various issues involved in clustering large biological datasets, describe the merits and underlying assumptions of some of the commonly used clustering approaches, and provide insights on how to cluster datasets arising in various areas within life sciences. We also provide a brief introduction to CLUTO, a general purpose toolkit for clustering various datasets, with an emphasis on its applications to problems and analysis requirements within life sciences.

Algorithms↗

Review of image-guided radiation therapy.

Image-guided radiation therapy represents a new paradigm in the field of high-precision radiation medicine. A synthesis of recent technological advances in medical imaging and conformal radiation therapy, image-guided radiation therapy represents a further expansion in the recent push for maximizing targeting capabilities with high-intensity radiation dose deposition limited to the true target structures, while minimizing radiation dose deposited in collateral normal tissues. By improving this targeting discrimination, the therapeutic ratio may be enhanced significantly. The principle behind image-guided radiation therapy relies heavily on the acquisition of serial image datasets using a variety of medical imaging platforms, including computed tomography, ultrasound and magnetic resonance imaging. These anatomic and volumetric image datasets are now being augmented through the addition of functional imaging. The current interest in positron-emitted tomography represents a good example of this sort of functional information now being correlated with anatomic localization. As the sophistication of imaging datasets grows, the precise 3D and 4D positions of the target and normal structures become of great relevance, leading to a recent exploration of real- or near-real-time positional replanning of the radiation treatment localization coordinates. This 'adaptive' radiotherapy explicitly recognizes that both tumors and normal tissues change position in time and space during a multiweek course of treatment, and even within a single treatment fraction. As targets and normal tissues change, the attenuation of radiation beams passing through these structures will also change, thus adding an additional level of imprecision in targeting unless these changes are taken into account. All in all, image-guided radiation therapy can be seen as further progress in the development of minimally invasive highly targeted cytotoxic therapies with the goal of substituting remote technologies for direct contact on the part of an operator or surgeon. Although data demonstrating clear-cut superiority of this new high-tech paradigm compared with more conventional radiation treatment approaches are scant, the emergence of preliminary data from several early studies shows that interest in this field is broad based and robust. As outcomes data accumulate, it is very likely that this field will continue to expand greatly. Although at present most of the work is being performed at major academic centers, the enthusiastic adoption of many of the devices and approaches being developed for this field suggest a rapid penetration into the community and the use of the technology by teams of specialists in the fields of radiation medicine, radiation physics and various branches of surgery. A recent survey of practitioners predicted very widespread adoption within the next 10 years.

Humans↗

Reduced VEPH1 expression is associated with an invasive phenotype and poor prognosis in clear cell renal cell carcinoma.

BACKGROUND: Clear cell renal cell carcinoma (ccRCC) remains a clinically heterogeneous urologic malignancy, and improved biomarkers are needed to refine prognostic stratification. VEPH1 has been implicated in cancer biology, but its role in ccRCC is incompletely defined. This study aimed to investigate the expression, prognostic relevance, and functional effects of VEPH1 in ccRCC. METHODS: VEPH1 transcript expression and prognostic relevance were evaluated using The Cancer Genome Atlas Kidney Renal Clear Cell Carcinoma (TCGA-KIRC) dataset and the University of Alabama at Birmingham Cancer Data Analysis Portal (UALCAN) and validated in paired ccRCC and adjacent normal renal tissues. The ability of VEPH1 transcript expression to distinguish tumor from normal tissues within the TCGA-KIRC dataset was assessed by receiver operating characteristic analysis. Gain- and loss-of-function experiments were performed in 786-O and 769-P ccRCC cells to determine the effects of VEPH1 on epithelial-mesenchymal transition (EMT)-related markers, migration, and invasion. AKT and ERK phosphorylation was evaluated by western blotting. RESULTS: VEPH1 transcript expression was significantly lower in ccRCC tissues than in normal renal tissues and distinguished tumor from normal samples within the TCGA-KIRC dataset. Low VEPH1 transcript expression was associated with poorer overall survival. Validation in 11 paired clinical specimens confirmed reduced VEPH1 messenger RNA (mRNA) and VEPH1 protein expression in tumor tissues. Functionally, VEPH1 overexpression increased E-cadherin, decreased N-cadherin, and suppressed migration and invasion, whereas partial VEPH1 knockdown produced the opposite changes. In exploratory signaling analyses, VEPH1 overexpression was associated with reduced AKT and ERK phosphorylation without altering total AKT or ERK levels. CONCLUSIONS: Reduced VEPH1 transcript expression was associated with poorer overall survival, whereas experimental VEPH1 depletion was associated with invasive and EMT-related features in ccRCC cells. VEPH1 may represent a candidate prognostic indicator in ccRCC; however, its relationship with AKT and ERK signaling and its clinical relevance require further mechanistic and independent-cohort validation.

Clear cell renal cell carcinoma (ccRCC)↗

Peeling off the hidden genetic heterogeneities of cancers based on disease-relevant functional modules.

Discovering molecular heterogeneities in phenotypically defined disease is of critical importance both for understanding pathogenic mechanisms of complex diseases and for finding efficient treatments. Recently, it has been recognized that cellular phenotypes are determined by the concerted actions of many functionally related genes in modular fashions. The underlying modular mechanisms should help the understanding of hidden genetic heterogeneities of complex diseases. We defined a putative disease module to be the functional gene groups in terms of both biological process and cellular localization, which are significantly enriched with genes highly variably expressed across the disease samples. As a validation, we used two large cancer datasets to evaluate the ability of the modules for correctly partitioning samples. Then, we sought the subtypes of complex diffuse large B-cell lymphoma (DLBCL) using a public dataset. Finally, the clinical significance of the identified subtypes was verified by survival analysis. In two validation datasets, we achieved highly accurate partitions that best fit the clinical cancer phenotypes. Then, for the notoriously heterogeneous DLBCL, we demonstrated that two partitioned subtypes using an identified module ("cellular response to stress") had very different 5-year overall rates (65% vs. 14%) and were highly significantly (P < 0.007) correlated with the clinical survival rate. Finally, we built a multivariate Cox proportional-hazard prediction model that included 4 genes as risk predictors for survival over DLBCL. The proposed modular approach is a promising computational strategy for peeling off genetic heterogeneities and understanding the modular mechanisms of human diseases such as cancers.

Databases, Genetic↗

Promoter classifier: software package for promoter database analysis.

Promoter Classifier is a package of seven stand-alone Windows-based C++ programs allowing the following basic manipulations with a set of promoter sequences: (i) calculation of positional distributions of nucleotides averaged over all promoters of the dataset; (ii) calculation of the averaged occurrence frequencies of the transcription factor binding sites and their combinations; (iii) division of the dataset into subsets of sequences containing or lacking certain promoter elements or combinations; (iv) extraction of the promoter subsets containing or lacking CpG islands around the transcription start site; and (v) calculation of spatial distributions of the promoter DNA stacking energy and bending stiffness. All programs have a user-friendly interface and provide the results in a convenient graphical form. The Promoter Classifier package is an effective tool for various basic manipulations with eukaryotic promoter sequences that usually are necessary for analysis of large promoter datasets. The program Promoter Divider is described in more detail as a representative component of the package.

Algorithms↗

Contrasting patterns of transcript abundance in tumour tissue and cancer cell lines.

Comparison of data on transcript abundance in ovarian, prostate and colon tumours with the corresponding cancer cell lines was used to assess the similarities of expression profiles. Although transcript abundances in tumours and cell lines were positively correlated, there were substantial differences with respect to the overall expression pattern. Compared with tumours, cancer cell lines showed more variable patterns of transcript abundance among tissue types. In the ovary and colon, cancer cell lines showed greater overall transcript abundance than normal tissue; this increase was much more marked in the case of the colon. However, in the prostate, cancer cell lines showed overall reduced transcript abundance when compared with normal tissue. Principal component analyses, applied separately to each tissue type, showed that approximately 80% of the variance was explained by overall expression level differences, which were maintained across normal tissue, tumour tissue and cancer cell lines. The remaining variance ( approximately 20%) could be attributed to contrasts in expression pattern among normal tissue, tumour tissue and cancer cell lines. In each dataset and in a combined dataset of transcripts shared among the three datasets, principal components revealed both contrasts in expression pattern between tumour tissue and cancer cell lines, and common features in the expression pattern of cancer cell lines that were distinct from those of tumour tissue and were shared across the different tissue types. These results imply that data on gene expression in cancer cell lines should be used with caution in inferring gene expression of in vivo tumours.

Biomarkers, Tumor↗

ARL6IP1 Inhibits Breast Cancer Tumor Progression by Targeting OLFM4 to Regulate Glycolysis.

INTRODUCTION: ARL6IP1 has been linked to cancer progression, but its precise role in BC, particularly in metabolism and its interaction with an OLFM4, remains unclear. AIMS: This study aimed to investigate the role of ADP-ribosylation factor-like 6 interacting protein 1 (ARL6IP1) in breast cancer (BC) cell behavior and metabolism and explore its interaction with an olfactomedin-4 (OLFM4) as a potential therapeutic target. OBJECTIVE: The objective of this study was to determine the effects of ARL6IP1 knockdown on BC cell proliferation, invasion, migration, apoptosis, oxidative stress, and glycolysis. Additionally, this study also explored the interaction between ARL6IP1 and OLFM4 and their combined role in BC progression and metabolism. METHODS: Key gene modules in the GSE73540 dataset were identified through weighted gene co-expression network analysis (WGCNA). Three BC-related datasets (GSE73540, GSE22820, and GSE36295) and The Cancer Genome Atlas (TCGA) were applied for additional examination of differentially expressed genes (DEGs). Intersection analysis selected ARL6IP1 as a hub gene for prognostic analysis. In vitro experiments investigated how ARL6IP1 knockdown influences BC cell proliferation, invasion, migration, apoptosis, epithelial-mesenchymal transition (EMT), oxidative stress, and glycolysis. The connection between ARL6IP1 and an OLFM4 was confirmed using Co-immunoprecipitation (Co-IP), and their roles in BC tumor progression and glycolysis were evaluated. RESULTS: ARL6IP1 was elevated in BC datasets and linked with poor BC prognosis. Experiments demonstrated that knockdown of ARL6IP1 significantly reduced BC cell growth while promoting apoptosis and oxidative stress. Besides, ARL6IP1 knockdown reduced glycolysis, as manifested by decreased extracellular acidification rate (ECAR), glucose consumption, adenosine triphosphate (ATP) levels, and lactate production while increasing mitochondrial respiration (OCR). Co-IP validated the connection between ARL6IP1 and OLFM4, and OLFM4 overexpression partially counteracted the suppression of glycolysis and cell behavior resulting from ARL6IP1 knockdown. CONCLUSION: ARL6IP1 is a critical regulator of BC progression, influencing glycolysis, mitochondrial function, and key cellular behaviors. Targeting the ARL6IP1-OLFM4 axis offers a promising therapeutic strategy for managing BC.

Humans↗

Elucidating the Mechanism of Xiaoqinglong Decoction in Chronic Urticaria Treatment: An Integrated Approach of Network Pharmacology, Bioinformatics Analysis, Molecular Docking, and Molecular Dynamics Simulations.

INTRODUCTION: Xiaoqinglong Decoction (XQLD) is a traditional Chinese medicinal formula commonly used to treat chronic urticaria (CU). However, its underlying therapeutic mechanisms remain incompletely characterized. This study employed an integrated approach combining network pharmacology, bioinformatics, molecular docking, and molecular dynamics simulations to identify the active components, potential targets, and related signaling pathways involved in XQLD's therapeutic action against CU, thereby providing a mechanistic foundation for its clinical application. METHODS: The active components of XQLD and their corresponding targets were identified using the Traditional Chinese Medicine Systems Pharmacology (TCMSP) database. CU-related targets were retrieved from the OMIM and GeneCards databases. Subsequently, core components and targets were determined via protein-protein interaction (PPI) network analysis and component-target-pathway network construction. Topological analyses were performed using Cytoscape software to prioritize core nodes within these networks. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were conducted via the DAVID database to identify enriched biological processes and signaling pathways. Molecular docking was performed to evaluate binding interactions between key components and core targets, while molecular dynamics (MD) simulations were employed to assess the stability of the component-target complexes with the lowest binding energy. Finally, CU-related targets of XQLD were validated using datasets from the Gene Expression Omnibus (GEO) database. RESULTS: A total of 135 active components and 249 potential targets of XQLD were identified, alongside 1,711 CU-related targets. Core components, such as quercetin, kaempferol, beta-sitosterol, naringenin, stigmasterol, and luteolin, exhibited high degree values in the constructed networks. The core targets identified included AKT1, TNF, IL6, TP53, PTGS2, CASP3, BCL2, ESR1, PPARG, and MAPK3. GO and KEGG pathway enrichment analyses revealed the PI3K-Akt signaling pathway as a central regulatory mechanism. Molecular docking studies demonstrated strong binding affinities between active components and core targets, with the stigmasterol-AKT1 complex exhibiting the lowest binding energy (-11.4 kcal/mol) and high stability in MD simulations. Validation using GEO datasets identified 12 core genes shared between CU-related targets and XQLD-associated targets, including PTGS2 and IL6, which were also prioritized as core targets in the network pharmacology analyses. DISCUSSION: This study comprehensively integrates multidisciplinary approaches to clarify the potential molecular mechanisms of XQLD in treating CU, highlighting its multitarget and multipathway synergistic effects. Molecular docking and dynamics simulations confirm the stable interaction between stigmasterol and the core target AKT1. Additionally, GEO dataset analysis verifies the pathogenic relevance of targets such as PTGS2 and IL6, significantly enhancing the credibility of our findings. These results provide a modern scientific basis for the traditional therapeutic effects of XQLD on CU and have important implications for developing multitarget treatments for this condition. However, this study mainly relies on database mining and computational simulations. Further in vitro and in vivo experimental validations are needed to confirm the predicted component-target-pathway interactions. CONCLUSION: This study identifies the active components, potential targets, and pathways through which XQLD exerts therapeutic effects on CU. These findings provide a theoretical foundation for further mechanistic studies and support their clinical application in the treatment of CU.

Molecular Docking Simulation↗

Genome-wide Association Studies of the Pathogenic Sphingosine-1-Phosphate Gene in Ulcerative Colitis.

BACKGROUND: Ulcerative colitis (UC) is a chronic inflammatory bowel disease that can lead to malignancies over time. Sphingosine-1-phosphate (S1P) receptor signaling affects lymphocyte trafficking and vascular integrity, influencing intestinal inflammation. This study aimed to identify S1P-related key genes in UC. METHODS: Differentially expressed genes (DEGs) between the UC and control groups were analyzed in the GSE87473 (training) dataset. Genes overlapping between the DEGs and S1P-related genes were considered candidate genes. These genes were incorporated into machine learning algorithms and subjected to expression analysis to identify key genes. Gene functions were determined through a gene&#x2013;gene interaction network, enrichment analysis, and immune cell infiltration analysis. In addition, transcription factor&#x2013;mRNA and mRNA&#x2013;miRNA&#x2013;lncRNA networks were constructed. Finally, reverse transcription&#x2013;quantitative polymerase chain reaction (RT-qPCR) was performed to evaluate the expression of key candidate genes in UC and control tissues. RESULTS: This study identified two key genes (SPHK2 and SPNS2) associated with UC. Notably, SPHK2 expression was lower and SPNS2 expression was higher in the UC group in both training and validation datasets and in clinical UC tissues (RT-qPCR). The area under the curve values of SPHK2 and SPNS2 exceeded 0.7 in both datasets, indicating that the genes had good diagnostic efficacy for UC. Consistently, the nomogram showed that the two genes had promising diagnostic value in UC. SPHK2 and SPNS2 were found to be localized to the plasma membrane. The correlations of the two genes with different immune cells showed significantly opposite trends. In particular, SPHK2 had the strongest positive correlation with M2 macrophages (r = 0.6) and the strongest negative correlation with neutrophils. Moreover, mRNA&#x2013;miRNA&#x2013;lncRNA and transcription factor&#x2013; mRNA networks of the key genes were constructed. CONCLUSION: This study suggests that SPHK2 and SPNS2 are key genes associated with UC, highlighting their potential as effective diagnostic biomarkers.

Humans↗

Weighted analysis of paired microarray experiments.

In microarray experiments quality often varies, for example between samples and between arrays. The need for quality control is therefore strong. A statistical model and a corresponding analysis method is suggested for experiments with pairing, including designs with individuals observed before and after treatment and many experiments with two-colour spotted arrays. The model is of mixed type with some parameters estimated by an empirical Bayes method. Differences in quality are modelled by individual variances and correlations between repetitions. The method is applied to three real and several simulated datasets. Two of the real datasets are of Affymetrix type with patients profiled before and after treatment, and the third dataset is of two-colour spotted cDNA type. In all cases, the patients or arrays had different estimated variances, leading to distinctly unequal weights in the analysis. We suggest also plots which illustrate the variances and correlations that affect the weights computed by our analysis method. For simulated data the improvement relative to previously published methods without weighting is shown to be substantial.

Journal Article↗

Linear data mining the Wichita clinical matrix suggests sleep and allostatic load involvement in chronic fatigue syndrome.

OBJECTIVES: To provide a mathematical introduction to the Wichita (KS, USA) clinical dataset, which is all of the nongenetic data (no microarray or single nucleotide polymorphism data) from the 2-day clinical evaluation, and show the preliminary findings and limitations, of popular, matrix algebra-based data mining techniques. METHODS: An initial matrix of 440 variables by 227 human subjects was reduced to 183 variables by 164 subjects. Variables were excluded that strongly correlated with chronic fatigue syndrome (CFS) case classification by design (for example, the multidimensional fatigue inventory [MFI] data), that were otherwise self reporting in nature and also tended to correlate strongly with CFS classification, or were sparse or nonvarying between case and control. Subjects were excluded if they did not clearly fall into well-defined CFS classifications, had comorbid depression with melancholic features, or other medical or psychiatric exclusions. The popular data mining techniques, principle components analysis (PCA) and linear discriminant analysis (LDA), were used to determine how well the data separated into groups. Two different feature selection methods helped identify the most discriminating parameters. RESULTS: Although purely biological features (variables) were found to separate CFS cases from controls, including many allostatic load and sleep-related variables, most parameters were not statistically significant individually. However, biological correlates of CFS, such as heart rate and heart rate variability, require further investigation. CONCLUSIONS: Feature selection of a limited number of variables from the purely biological dataset produced better separation between groups than a PCA of the entire dataset. Feature selection highlighted the importance of many of the allostatic load variables studied in more detail by Maloney and colleagues in this issue [1] , as well as some sleep-related variables. Nonetheless, matrix linear algebra-based data mining approaches appeared to be of limited utility when compared with more sophisticated nonlinear analyses on richer data types, such as those found in Maloney and colleagues [1] and Goertzel and colleagues [2] in this issue.

Adult↗