Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “GEO database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

A hill-climbing approach for automatic gridding of cDNA microarray images.

Image and statistical analysis are two important stages of cDNA microarrays. Of these, gridding is necessary to accurately identify the location of each spot while extracting spot intensities from the microarray images and automating this procedure permits high-throughput analysis. Due to the deficiencies of the equipment used to print the arrays, rotations, misalignments, high contamination with noise and artifacts, and the enormous amount of data generated, solving the gridding problem by means of an automatic system is not trivial. Existing techniques to solve the automatic grid segmentation problem cover only limited aspects of this challenging problem and require the user to specify the size of the spots, the number of rows and columns in the grid, and boundary conditions. In this paper, a hill-climbing automatic gridding and spot quantification technique is proposed which takes a microarray image (or a subgrid) as input and makes no assumptions about the size of the spots, rows, and columns in the grid. The proposed method is based on a hill-climbing approach that utilizes different objective functions. The method has been found to effectively detect the grids on microarray images drawn from databases from GEO and the Stanford genomic laboratories.

Algorithms↗

The lacrimal gland transcriptome is an unusually rich source of rare and poorly characterized gene transcripts.

PURPOSE: To sequence and comprehensively analyze human and mouse lacrimal gland transcriptomes as part of the NEIBank project. METHODS: cDNA libraries generated from normal human and mouse lacrimal glands were sequenced and analyzed by PHRED, RepeatMasker, BLAST, and GRIST. Human "lacrimal-preferred genes" and putative gene regulatory elements were respectively identified in UniGene and ConSite, and gene clustering was analyzed by chromosomal mapping. "Hypothetical proteins," identified by keyword search, were verified by genomic alignment and queried in the Conserved Domain database and GEO Profiles. RESULTS: The top six transcripts in human and mouse differed, revealing a previously unappreciated molecular divergence. The human transcriptome is enriched with transcripts from 29 lacrimal-preferred genes and a content of poorly characterized hypothetical proteins, proportionally greater than in all other tissues. Only 45% of lacrimal preferred, but 71% of hypotheticals, have mouse orthologs. Many of the latter display apparently altered cancer expression in the CGAP SAGE library collection-often in keeping with predicted WD40, protein kinase, Src homology 2 and 3, RhoGEF, and pleckstrin homology domains involved in cell signaling. At the genomic level, lacrimal-expressed genes show some evidence of clustering, particularly on human chromosomes 9 and 12. Binding sites for TFAP2A, FOXC1, and other transcription factors are predicted. CONCLUSIONS: Interspecies divergence cautions against use of mouse models of human dry eye syndromes. Lacrimal preferred and hypothetical proteins, gene clustering, and putative gene regulatory elements together provide new clues for a molecular understanding of lacrimal gland function and mechanisms of coordinated tissue-specific transcriptional regulation.

Aged↗

Hereditary profiles of disorderly transcription?

BACKGROUND: Microscopic examination of living cells often reveals that cells from some cell strains appear to be in a permanent state of disarray without obvious reason. In all probability such a disorderly state affects cell functioning. The aim of this study was to establish whether a disorderly state could occur that adversely affects gene expression profiles and whether such a state might have biomedical consequences. To this end, the expression profiles of the 14 genes of the proteasome derived from the GEO SAGE database were utilized as a model system. RESULTS: By adopting the overall expression profile as the standard for normal expression, deviation in transcription was frequently observed. Each deviating tissue exhibited its own characteristic profile of over-expressed and under-expressed genes. Moreover such a specific deviating profile appeared to be epigenetic in origin and could be stably transmitted to a clonal derivative e.g. from a precancerous normal tissue to its tumor. A significantly greater degree of deviation was observed in the expression profiles from the tumor tissues. The changes in the expression of different genes display a network of interdependencies. Therefore our hypothesis is that deviating profiles reflect disorder in the localization of genes within the nucleus. The underlying cause(s) for these disorderly states remain obscure; it could be noise and/or deterministic chaos. Presence of mutational damage does not appear to be predominantly involved. CONCLUSION: As disturbances in expression profiles frequently occur and have biomedical consequences, its determination could prove of value in several fields of biomedical research.

Journal Article↗

Genome-wide identification and characterization of 1-amino-cyclopropane-1- carboxylate synthase (ACS) gene family in Carica papaya and expression insights in response to hormone stress.

ACC-synthase (1-aminocyclopropane-1-carboxylate synthase), also known as the ACS gene, plays a pivotal role in ethylene production, which is of great importance in the fruit ripening process for producing saleable yield (marketable fruit). The ACS gene family presumably controls stress responses, plant growth and development, and particularly fruit ripening. Computational biology was used as an essential tool to identify seven ACS genes in Carica papaya (red hermaphrodite) using an RNA-seq database (NCBI GEO). Further, the phylogenetic relationships of ACS genes determined gene family resemblance in the genomes of Hordeum vulgare, Musa acuminata, C. papaya, and Arabidopsis thaliana; therefore, the identified gene families were further classified into four distinct clades (Type-I, Type-II, Type-III, and Type-IV) in alignment with the well-established Arabidopsis classification. Moreover, encompassing gene structure, domain motifs, cis-element phylogenetic profiling, synteny, and transcriptomic profiling unveiled latent structural and functional attributes within CpACS genes. Through segmental duplication of CpACS, insights into evolutionary duplication events were predicted. The paralogous behavior of ACS genes in C. papaya and a comprehensive transcriptomic analysis demonstrated both up- and down-regulation patterns in response to ethylene treatment at different time points during the fruit ripening process, using the papaya manual handbook V2 (2021). Gene expression showed upregulation of two essential CpACS genes, CpACS5 and CpACS6. RT-qPCR validates the expression of these important genes during fruit ripening. However, one gene, CpACS7, is expressed in the later stages of fruit development. Our results demonstrated novel avenues for understanding the expression pathways of the ACS gene family in red hermaphrodite papaya, and most of these genes were linked to regulating various abiotic stresses, plant growth, and fruit development.

Carica↗

Mitochondria related gene signature serves as prognosis prediction and risk stratification of cholangiocarcinoma.

BACKGROUND: Cholangiocarcinoma (CHOL) is a highly aggressive biliary malignancy with poor clinical outcomes and limited effective prognostic biomarkers. Mitochondrial dysfunction participates in multiple oncological processes of CHOL, yet the prognostic roles of mitochondria‑related genes (MRGs) remain poorly understood. This study aimed to characterize MRGs expression in CHOL and develop a molecular prognostic model for predicting patient survival and guiding clinical management. METHODS: RNA sequencing (RNA-seq) and clinical data of CHOL were obtained from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) (GSE89748) databases. Differentially expressed MRGs were identified, and 10 machine learning algorithms were used to construct prognostic models. The optimal model (highest average C-index) was selected to establish a mitochondria-related risk score (MRRS), which was validated internally and externally. A nomogram integrating clinical factors and MRRS was developed, and biological mechanisms were explored via functional and immune analyses. RESULTS: A 3-MRG (MAP3K1, MRPL18, PYGB) prognostic signature was constructed, stratifying patients into high- and low-risk groups with significantly different overall survival. The model showed high predictive accuracy, with an area under the curve (AUC) up to 0.845, and MRRS was an independent prognostic factor. The signature was associated with mitochondrial pathways, and the high-risk group had distinct immune infiltration and mutation profiles. CONCLUSIONS: A validated MRG prognostic model effectively stratifies CHOL patients and has potential clinical value for prognosis prediction. Further validation in larger cohorts is needed to confirm its applicability.

Cholangiocarcinoma (CHOL)↗

CDC20B Dysregulation: Links to Tumor Prognosis and Immunity.

OBJECTIVE: This study aimed to clarify the pan-cancer expression pattern, upstream regulatory mechanisms, prognostic relevance, and immune associations of CDC20B. METHOD: Using public databases (GTEx, GEO, and TCGA), we examined CDC20B expression and its associations with prognosis and tumor immunity across multiple cancers. Immunohistochemistry (IHC) on an independent clinical cohort was performed to validate CDC20B upregulation in tumor tissues. Promoter methylation, genetic alterations, and immune infiltration were analyzed using bioinformatics tools (cBioPortal, UALCAN, TIMER2.0, ESTIMATE). Functional enrichment was assessed by GSEA and single-cell state analysis (CancerSEA). RESULTS: CDC20B was markedly upregulated in most tumor types (p < 0.001), with strong diagnostic efficiency (AUC > 0.7 in 15 cancers) and potential regulation by promoter hypomethylation. IHC confirmed its overexpression in clinical tumor tissues. However, the prognostic impact of CDC20B was cancer-type-specific: high expression correlated with poor overall survival in UCS, LGG, KIRC, and OV, but with favorable survival in BRCA, LUAD, and PAAD. CDC20B expression was associated with immune infiltration patterns, showing negative correlations with ImmuneScore in most cancers but positive correlations with CD8+ T cells in PAAD. Functional analyses indicated involvement in EMT, KRAS/NF-&#x3ba;B signaling, and DNA damage response pathways. DISCUSSION: The dual prognostic role of CDC20B suggests context-dependent functions, likely influenced by tumor microenvironment composition and underlying oncogenic programs. Promoter hypomethylation emerges as a potential epigenetic driver of overexpression. The associations with immune modulation and genomic instability suggest that CDC20B is a candidate biomarker, though causal relationships require experimental validation. CONCLUSION: CDC20B may contribute to tumor progression in a context-dependent manner, with its prognostic impact varying across cancer types. Its role in tumor immunity and oncogenic pathways warrants further investigation, particularly in stratified patient populations.

CDC20B↗

Contribution of Geographic Information Systems and location models to planning of wastewater systems.

This paper presents the contributions of Geographic Information Systems (GIS) and location models towards planning regional wastewater systems (sewers and wastewater treatment plants) serving small agglomerations, i.e. agglomerations with less than 2,000 inhabitants. The main goal was to develop a decision support tool for tracing and locating regional wastewater systems. The main results of the model are expressed in terms of number, capacity and location of Wastewater Treatment Plants (WWTP) and the length of main sewers. The decision process concerning the location and capacity of wastewater systems has a number of parameters that can be optimized. These parameters include the total sewer length and number, capacity and location of WWTP. The optimization of parameters should lead to the minimization of construction and operation costs of the integrated system. Location models have been considered as tools for decision support, mainly when a geo-referenced database can be used. In these cases, the GIS may represent an important role for the analysis of data and results especially in the preliminary stage of planning and design. After selecting the spatial location model and the heuristics, two greedy algorithms were implemented in Visual Basic for Applications on the ArcGIS software environment. To illustrate the application of these algorithms a case study was developed, in a rural area located in the central part of Portugal.

Algorithms↗

ITTACA: a new database for integrated tumor transcriptome array and clinical data analysis.

Transcriptome microarrays have become one of the tools of choice for investigating the genes involved in tumorigenesis and tumor progression, as well as finding new biomarkers and gene expression signatures for the diagnosis and prognosis of cancer. Here, we describe a new database for Integrated Tumor Transcriptome Array and Clinical data Analysis (ITTACA). ITTACA centralizes public datasets containing both gene expression and clinical data. ITTACA currently focuses on the types of cancer that are of particular interest to research teams at Institut Curie: breast carcinoma, bladder carcinoma and uveal melanoma. A web interface allows users to carry out different class comparison analyses, including the comparison of expression distribution profiles, tests for differential expression and patient survival analyses. ITTACA is complementary to other databases, such as GEO and SMD, because it offers a better integration of clinical data and different functionalities. It also offers more options for class comparison analyses when compared with similar projects such as Oncomine. For example, users can define their own patient groups according to clinical data or gene expression levels. This added flexibility and the user-friendly web interface makes ITTACA especially useful for comparing personal results with the results in the existing literature. ITTACA is accessible online at http://bioinfo.curie.fr/ittaca.

Breast Neoplasms↗

Expression profilings of 39 genes selected by ANOVA could separate precursors of murine dendritic cells and macrophages.

Dendritic cells (DCs) and macrophages share some stages in the development and function of antigen presentation. But it is difficult to separate them from their precursors. We used one-way ANOVA (analysis of variances) on murine expression profilings of several hematopoietic cells associated with DCs and macrophages to find the genes with great differences across the cell groups. These groups were the DCs from spleen, cultivated DCs, DC precursors, DC progenitors, DC progenitor cell lines, hematopoietic stem cell (HSC), and bone marrow-derived macrophages. The data of expression profilings were all downloaded from GEO and ArrayExpress database. After the normalization of 11 housekeeping genes across 42 arrays, we got 39 genes (44 probesets) by analysis of one-way ANOVA (Bonferroni step-down) with p values cutoff of 0.05. These genes (probesets) could separate the hematopoietic cells well by the methods of unsupervised hierarchical clustering and principal component analysis (PCA). The class prediction also indicated that these genes could separate the precursors of DC and macrophages with 20 arrays composed of 5 cell types with the same normalization. The accuracy rate of class prediction was 90% (18/20). The genes selected by one-way ANOVA included those of MHC (major histocompatibility complex) and defense of immunity, cell adhesion, chemokine or its receptors, and transcription factors. The results indicated that these 39 genes could separate precursors of DC and macrophages very clearly. It was suggested that these genes might represent some important molecules that related with the precursors of DCs and macrophages, and were worthy for further study.

Animals↗

The effect of soil type and climate on hookworm (Necator americanus) distribution in KwaZulu-Natal, South Africa.

We investigated environmental factors influencing the distribution of hookworm infection in KwaZulu-Natal, South Africa. Prevalence data were sourced from previous studies and additional surveys carried out to supplement the database. When geo-referenced the data revealed that higher prevalences are limited to areas below 150 m above sea level, and low prevalences to areas above this altitude. Using univariate analysis we investigated the differences in environmental factors in the two areas. The relationship between hookworm prevalence, altitude and climate-derived variables was assessed using Pearson correlation coefficient, and that of soil type using the t-test. Multivariate analysis was used to determine environmental factors that combine best to provide favourable conditions for hookworm distribution. The results revealed that areas 150 m above sea level, i.e. inland, supported low mean hookworm prevalences (x = 6, n = 21), and were characterized by soils with a clay content of more than 45%, variable temperatures and moderate rainfall. Hookworm prevalence also decreased southwards as temperatures became slightly cooler, rainfall remained more-or-less constant and the coastal plain narrowed. In the multivariate model prevalence was most significantly correlated with the mean daily minimum temperature for January followed by the mean number of rainy days for January. This indicates the importance of summer conditions in the transmission of hookworm infection in KwaZulu-Natal and suggests that transmission may be seasonal.

Adolescent↗

Sex-related effect on gene expression in the mouse meibomian gland.

PURPOSE: Sex-related differences have been identified in the anatomy and physiology of the meibomian gland. We hypothesize that these differences are due, at least in part, to variations in gene expression. This study's objective was to determine whether sex-related differences do exist in meibomian gland gene expression. We also sought to elucidate whether such differences, if any, might be (a) analogous to those known to occur in the lacrimal gland and (b) due to the effect of sex steroids. METHODS: Meibomian glands were obtained from young adult male and female BALB/c mice (n=7 to 15 mice per sex per experiment), pooled according to sex and processed for the isolation of RNA. Samples were evaluated for differentially expressed mRNAs by using CodeLink Bioarrays and GEM 1 and 2 gene chips. Bioarray data were analyzed with GeneSifter. Net software and also compared with microarray data in GEO and GeneSifter databases. RESULTS: Our results demonstrate that sex has a significant influence on the expression of 164 genes in the mouse meibomian gland. These genes are involved in a broad spectrum of biological processes, molecular functions, and cellular components, including such activities as metabolism, catalysis, cell growth and maintenance, membrane architecture, nucleic acid binding, transcription, and signal transduction. In addition, the nature of the sex-related variations in meibomian gland gene expression is quite different from those in the lacrimal gland and appear to be mediated in part by the action of androgens, but not estrogens or progestins. CONCLUSIONS: These findings support our hypothesis that sex-related differences exist in gene expression of the meibomian gland.

Animals↗

The PowerAtlas: a power and sample size atlas for microarray experimental design and research.

BACKGROUND: Microarrays permit biologists to simultaneously measure the mRNA abundance of thousands of genes. An important issue facing investigators planning microarray experiments is how to estimate the sample size required for good statistical power. What is the projected sample size or number of replicate chips needed to address the multiple hypotheses with acceptable accuracy? Statistical methods exist for calculating power based upon a single hypothesis, using estimates of the variability in data from pilot studies. There is, however, a need for methods to estimate power and/or required sample sizes in situations where multiple hypotheses are being tested, such as in microarray experiments. In addition, investigators frequently do not have pilot data to estimate the sample sizes required for microarray studies. RESULTS: To address this challenge, we have developed a Microrarray PowerAtlas. The atlas enables estimation of statistical power by allowing investigators to appropriately plan studies by building upon previous studies that have similar experimental characteristics. Currently, there are sample sizes and power estimates based on 632 experiments from Gene Expression Omnibus (GEO). The PowerAtlas also permits investigators to upload their own pilot data and derive power and sample size estimates from these data. This resource will be updated regularly with new datasets from GEO and other databases such as The Nottingham Arabidopsis Stock Center (NASC). CONCLUSION: This resource provides a valuable tool for investigators who are planning efficient microarray studies and estimating required sample sizes.

Algorithms↗

Bioinformatics identification and validation of pyroptosis-related gene for ischemic stroke.

BACKGROUND: Ischemic stroke (IS) is one of the common and frequent diseases with extremely high lethality and disability in the world, and there is no effective treatment at present. This study aimed to screen hub genes involved in cerebral ischemia/reperfusion injury (CIRI) and pyroptosis, and explore promising intervention targets. METHODS: CIRI-related genes (GSE202659 and GSE131193) and pyroptosis-related genes (PRGs) in mice were obtained from the Gene Expression Omnibus (GEO) and GeneCards database. We screened for LASSO regression to construct a prognostic model of GSE131193 and PRGs and examined by GSE137482. The functional enrichment analysis of Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), Gene Set Enrichment Analysis (GSEA) and Gene Set Variation Analysis (GSVA) were performed on pyroptosis-related differentially expressed genes (PRDEGs) of GSE202659.The key modules for CIRI and pyroptosis were identified by Weight Gene Co-expression Network Analysis (WGCNA). Subsequently, Protein-protein Interaction (PPI) network and the Cytoscape was constructed to screen out hub genes. Used the starBase to predict miRNA interacting with hub genes and constructed mRNA-miRNA-lncRNA interaction networks. CIRI-related Molecular Subtypes were constructed for hub genes. The relationship between immune cells and hub genes was verified via CIBERSORT. Finally, we selected C57BL/6 mice to construct models to confirm hub genes by enzyme linked immunosorbent assay (ELISA), reverse transcription-polymerase chain reaction (RT-PCR), western blot, and Immunofluorescence. RESULTS: A total of 272 PRGs and 35 PRDEGs were screened. An eight-gene risk prediction models were established (AUC&#x2009;=&#x2009;0.868). GO, KEGG, GSEA and GSVA analyses revealed that PRDEGs were mainly involved in positive regulation of cytokine production, and NOD-like receptor signaling pathway. And then, seven hub genes (Irf1, Icam1, Tlr2, Tnf, Cebpb, Il1rn, and Casp8) were identified by PPI. Icam1, Tnf, Cebpb, Il1rn, and Casp8 had high expression profiles in Cluster2 by hierarchical clustering. The immune infiltration analysis results showed that among the hub genes, Cebpb, Il1rn, and Casp8, showed a significant positive correlation with the degree of NK.Actived, and Icam1 showed a significant negative correlation with B.Cells.Memory. The results of animal experiments significantly demonstrated an upregulation of Irf1, Icam1, Tlr2, Cebpb, and Il1rn. CONCLUSION: Our finding indicated that Irf1, Icam1, Tlr2, Cebpb, and Il1rn are hub genes associated with pyroptosis, and these genes are all associated with different immune cells, so as to provide new targets for the prevention and treatment of IS from the perspective of pyroptosis.

Pyroptosis↗

Transcription profiling of renal cell carcinoma.

AIMS: Our aim was to prepare a comprehensive catalogue of the changes in gene expression accompanying the development and progression of renal cell carcinoma, and to correlate these with histo-pathological, cytogenetic and clinical findings. METHODS: mRNA samples from paired neoplastic and non-cancerous human kidney tissue were labeled and hybridized in duplicate against high-density cDNA arrays. Two array technologies were used: 31,500-element transcriptome-wide nylon arrays for hybridization with 37 radioactively labelled sample pairs, and 4200-element kidney- and cancer-specific glass microarrays for hybridization with 19 fluorescently labelled sample pairs. RESULTS: We identified more than 1700 cDNA clones that show differential transcription levels in kidney tumor tissue compared to normal kidney tissue. The functional classification of 389 annotated genes provided views of the changes in the activities of specific biological processes in renal cancer. Among the biological processes with a large proportion of up-regulated genes we found cell adhesion, signal transduction, and nucleotide metabolism. Down-regulated processes included small molecule transport, ion homeostasis, and oxygen and radical metabolism. Furthermore, we explored the feasibility of molecular diagnosis for renal cell tumors using cDNA microarrays on glass slides, investigating the association of transcription levels with tumor type, progression, and a putative prognostic variable. The experimental data is available from the GEO gene expression database (http://www.ncbi.nlm.nih.gov/geo; accession no. GSE3), and a comprehensive presentation of the results is available in the web supplement (http://www.dkfz-heidelberg.de/abt0840/whuber/rcc). CONCLUSION: Transcription profiling using high-density cDNA arrays is a powerful method with the potential to improve cancer diagnosis and prognosis. The identification and classification of differentially transcribed genes, as described in our study, is the beginning of a more complete understanding of kidney cancer.

Carcinoma, Renal Cell↗

Database resources of the National Center for Biotechnology Information.

In addition to maintaining the GenBank nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides analysis and retrieval resources for the data in GenBank and other biological data made available through NCBI's Web site. NCBI resources include Entrez, the Entrez Programming Utilities, My NCBI, PubMed, PubMed Central, Entrez Gene, the NCBI Taxonomy Browser, BLAST, BLAST Link(BLink), Electronic PCR, OrfFinder, Spidey, Splign, RefSeq, UniGene, HomoloGene, ProtEST, dbMHC, dbSNP, Cancer Chromosomes, Entrez Genome, Genome Project and related tools, the Trace and Assembly Archives, the Map Viewer, Model Maker, Evidence Viewer, Clusters of Orthologous Groups (COGs), Viral Genotyping Tools, Influenza Viral Resources, HIV-1/Human Protein Interaction Database, Gene Expression Omnibus (GEO), Entrez Probe, GENSAT, Online Mendelian Inheritance in Man (OMIM), Online Mendelian Inheritance in Animals (OMIA), the Molecular Modeling Database (MMDB), the Conserved Domain Database (CDD), the Conserved Domain Architecture Retrieval Tool (CDART) and the PubChem suite of small molecule databases. Augmenting many of the Web applications are custom implementations of the BLAST program optimized to search specialized data sets. These resources can be accessed through the NCBI home page at www.ncbi.nlm.nih.gov.

Animals↗

Database resources of the National Center for Biotechnology Information.

In addition to maintaining the GenBank nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides data retrieval systems and computational resources for the analysis of data in GenBank and other biological data made available through NCBI's website. NCBI resources include Entrez, Entrez Programming Utilities, PubMed, PubMed Central, Entrez Gene, the NCBI Taxonomy Browser, BLAST, BLAST Link (BLink), Electronic PCR, OrfFinder, Spidey, RefSeq, UniGene, HomoloGene, ProtEST, dbMHC, dbSNP, Cancer Chromosomes, Entrez Genomes and related tools, the Map Viewer, Model Maker, Evidence Viewer, Clusters of Orthologous Groups (COGs), Retroviral Genotyping Tools, HIV-1/Human Protein Interaction Database, SAGEmap, Gene Expression Omnibus (GEO), Online Mendelian Inheritance in Man (OMIM), the Molecular Modeling Database (MMDB), the Conserved Domain Database (CDD) and the Conserved Domain Architecture Retrieval Tool (CDART). Augmenting many of the Web applications are custom implementations of the BLAST program optimized to search specialized datasets. All of the resources can be accessed through the NCBI home page at http://www.ncbi.nlm.nih.gov.

Amino Acid Sequence↗

Large scale data mining approach for gene-specific standardization of microarray gene expression data.

MOTIVATION: The identification of the change of gene expression in multifactorial diseases, such as breast cancer is a major goal of DNA microarray experiments. Here we present a new data mining strategy to better analyze the marginal difference in gene expression between microarray samples. The idea is based on the notion that the consideration of gene's behavior in a wide variety of experiments can improve the statistical reliability on identifying genes with moderate changes between samples. RESULTS: The availability of a large collection of array samples sharing the same platform in public databases, such as NCBI GEO, enabled us to re-standardize the expression intensity of a gene using its mean and variation in the wide variety of experimental conditions. This approach was evaluated via the re-identification of breast cancer-specific gene expression. It successfully prioritized several genes associated with breast tumor, for which the expression difference between normal and breast cancer cells was marginal and thus would have been difficult to recognize using conventional analysis methods. Maximizing the utility of microarray data in the public database, it provides a valuable tool particularly for the identification of previously unrecognized disease-related genes. AVAILABILITY: A user friendly web-interface (http://compbio.sookmyung.ac.kr/~lage/) was constructed to provide the present large-scale approach for the analysis of GEO microarray data (GS-LAGE server).

Algorithms↗

Anoikis classification of lung squamous cell carcinoma reveals correlation with clinical prognosis and immune characteristics.

BACKGROUND: Anoikis is a new mode of cell death that has been shown to correlate significantly with tumors. However, the clinical prognostic significance of anoikis in lung squamous cell carcinoma (LUSC) remains poorly studied. METHODS: The differentially expressed ARGs and candidate genes were selected by the differential analysis to construct a predictive model. Independent prognostic gene was determined by Cox and LASSO analysis and we used the HCC95 and NCI H520 cell line to verify the gene function. We used the data from TCGA, GEO, GeneCards, and Harmonizome databases to analyze the immune microenvironment, functional enrichment, and drug sensitivity analysis. RESULTS: We identified 717 differentially expressed and selected 3 ARGs (FADD, SNAI1, and BAG4) to construct a predictive model. We found that SNAI1 is an independent prognostic gene and confirmed that knocking out the SNAI1 inhibited the HCC95/NCI H520 cell proliferation. We used single-sample gene-set enrichment analysis (ssGSEA) to evaluate the immune infiltration based on the 3 ARG expression levels. We constructed a risk score and provided a visual representation of the prophetic implications of the ARGs-based signature through a nomogram. We found 15 susceptible drugs in the high-risk group and 15 sensitive drugs in the low-risk group by the drug sensitivity analysis. CONCLUSION: We used ARGs to construct a prognosis model for LUSC that can accurately predict the prognosis of LUSC patients. ARGs, especially SNAI1, play an essential role in developing LUSC. These findings could provide individualized treatment plans and new research ideas for LUSC patients.

Humans↗