Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Comprehensive genomic profiling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Functional genomics by integrated analysis of metabolome and transcriptome of Arabidopsis plants over-expressing an MYB transcription factor.

The integration of metabolomics and transcriptomics can provide precise information on gene-to-metabolite networks for identifying the function of unknown genes unless there has been a post-transcriptional modification. Here, we report a comprehensive analysis of the metabolome and transcriptome of Arabidopsis thaliana over-expressing the PAP1 gene encoding an MYB transcription factor, for the identification of novel gene functions involved in flavonoid biosynthesis. For metabolome analysis, we performed flavonoid-targeted analysis by high-performance liquid chromatography-mass spectrometry and non-targeted analysis by Fourier-transform ion-cyclotron mass spectrometry with an ultrahigh-resolution capacity. This combined analysis revealed the specific accumulation of cyanidin and quercetin derivatives, and identified eight novel anthocyanins from an array of putative 1800 metabolites in PAP1 over-expressing plants. The transcriptome analysis of 22,810 genes on a DNA microarray revealed the induction of 38 genes by ectopic PAP1 over-expression. In addition to well-known genes involved in anthocyanin production, several genes with unidentified functions or annotated with putative functions, encoding putative glycosyltransferase, acyltransferase, glutathione S-transferase, sugar transporters and transcription factors, were induced by PAP1. Two putative glycosyltransferase genes (At5g17050 and At4g14090) induced by PAP1 expression were confirmed to encode flavonoid 3-O-glucosyltransferase and anthocyanin 5-O-glucosyltransferase, respectively, from the enzymatic activity of their recombinant proteins in vitro and results of the analysis of anthocyanins in the respective T-DNA-inserted mutants. The functional genomics approach through the integration of metabolomics and transcriptomics presented here provides an innovative means of identifying novel gene functions involved in plant metabolism.

Arabidopsis↗

Genome-Wide Identification and Expression Pattern of the ANK Gene Family in Sorghum bicolor Under Salt Stress.

The Ankyrin-repeat proteins (ANKs) play a key role in plant development and in response to abiotic stress. This research identified family members of the ANK genes in Sorghum bicolor at the whole-genome level, analyzed their sequence characteristics, evolutionary relationships, and expression patterns, and provided a scientific basis for elucidating the functionality of SbANK genes and for salt-tolerant breeding. Using bioinformatics methods, this study conducted a comprehensive identification of the SbANK gene family, analyzing its physicochemical properties, domain composition, chromosomal distribution, colinearity relationships, promoter cis-acting elements, and conserved protein motifs. Transcriptomic data and qRT-PCR were used to detect changes in their expression under salt stress. A total of 186 ANK family members were identified in the Sorghum bicolor genome, classified into 13 subfamilies and unevenly distributed across 10 chromosomes. Intra-species colinearity analysis revealed 7 pairs of duplicated genes, while inter-species colinearity analysis showed that S. bicolor and Oryza sativa share 88 pairs of orthologs, far exceeding the number found in Arabidopsis thaliana (11 pairs). Promoter analysis indicated that SbANK genes are enriched with cis-acting elements associated with hormone responses (particularly MeJA elements, accounting for 51.7%) and stress responses (particularly anaerobic-inducible elements, accounting for 60.9%). Transcriptomic expression analysis revealed that SbANK genes exhibit distinct tissue specificity, with the ANK-IQ subfamily highly expressed in leaves and the ANK-M subfamily showing the most widespread response under salt stress. Expression levels of the 10 candidate genes showing the most significant responses to salt stress were analyzed using qRT-PCR. The results indicated that SbANK91, SbANK135, and SbANK136 were significantly upregulated under 200 mmol/L NaCl treatment. The SbANK family is distinguished by a large number of member genes and structural diversity, with the ANK-M subfamily being the primary group responding to salt stress. SbANK91, SbANK135, and SbANK136 are identified as putative candidate genes for salt stress responses.

Sorghum↗

Role of histone-like proteins H-NS and StpA in expression of virulence determinants of uropathogenic Escherichia coli.

The histone-like protein H-NS is a global regulator in Escherichia coli that has been intensively studied in nonpathogenic strains. However, no comprehensive study on the role of H-NS and its paralogue, StpA, in gene expression in pathogenic E. coli has been carried out so far. Here, we monitored the global effects of H-NS and StpA in a uropathogenic E. coli isolate by using DNA arrays. Expression profiling revealed that more than 500 genes were affected by an hns mutation, whereas no effect of StpA alone was observed. An hns stpA double mutant showed a distinct gene expression pattern that differed in large part from that of the hns single mutant. This suggests a direct interaction between the two paralogues and the existence of distinct regulons of H-NS and an H-NS/StpA heteromeric complex. hns mutation resulted in increased expression of alpha-hemolysin, fimbriae, and iron uptake systems as well as genes involved in stress adaptation. Furthermore, several other putative virulence genes were found to be part of the H-NS regulon. Although the lack of H-NS, either alone or in combination with StpA, has a huge impact on gene expression in pathogenic E. coli strains, its effect on virulence is ambiguous. At a high infection dose, hns mutants trigger more sudden lethality due to their increased acute toxicity in murine urinary tract infection and sepsis models. At a lower infectious dose, however, mutants lacking H-NS are attenuated through their impaired growth rate, which can only partially be compensated for by the higher expression of numerous virulence factors.

Animals↗

Combining gene expression data from different generations of oligonucleotide arrays.

BACKGROUND: One of the important challenges in microarray analysis is to take full advantage of previously accumulated data, both from one's own laboratory and from public repositories. Through a comparative analysis on a variety of datasets, a more comprehensive view of the underlying mechanism or structure can be obtained. However, as we discover in this work, continual changes in genomic sequence annotations and probe design criteria make it difficult to compare gene expression data even from different generations of the same microarray platform. RESULTS: We first describe the extent of discordance between the results derived from two generations of Affymetrix oligonucleotide arrays, as revealed in cluster analysis and in identification of differentially expressed genes. We then propose a method for increasing comparability. The dataset we use consists of a set of 14 human muscle biopsy samples from patients with inflammatory myopathies that were hybridized on both HG-U95Av2 and HG-U133A human arrays. We find that the use of the probe set matching table for comparative analysis provided by Affymetrix produces better results than matching by UniGene or LocusLink identifiers but still remains inadequate. Rescaling of expression values for each gene across samples and data filtering by expression values enhance comparability but only for few specific analyses. As a generic method for improving comparability, we select a subset of probes with overlapping sequence segments in the two array types and recalculate expression values based only on the selected probes. We show that this filtering of probes significantly improves the comparability while retaining a sufficient number of probe sets for further analysis. CONCLUSIONS: Compatibility between high-density oligonucleotide arrays is significantly affected by probe-level sequence information. With a careful filtering of the probes based on their sequence overlaps, data from different generations of microarrays can be combined more effectively.

Biopsy↗

Transcriptomic and proteomic characterization of the Fur modulon in the metal-reducing bacterium Shewanella oneidensis.

The availability of the complete genome sequence for Shewanella oneidensis MR-1 has permitted a comprehensive characterization of the ferric uptake regulator (Fur) modulon in this dissimilatory metal-reducing bacterium. We have employed targeted gene mutagenesis, DNA microarrays, proteomic analysis using liquid chromatography-mass spectrometry, and computational motif discovery tools to define the S. oneidensis Fur regulon. Using this integrated approach, we identified nine probable operons (containing 24 genes) and 15 individual open reading frames (ORFs), either with unknown functions or encoding products annotated as transport or binding proteins, that are predicted to be direct targets of Fur-mediated repression. This study suggested, for the first time, possible roles for four operons and eight ORFs with unknown functions in iron metabolism or iron transport-related functions. Proteomic analysis clearly identified a number of transporters, binding proteins, and receptors related to iron uptake that were up-regulated in response to a fur deletion and verified the expression of nine genes originally annotated as pseudogenes. Comparison of the transcriptome and proteome data revealed strong correlation for genes shown to be undergoing large changes at the transcript level. A number of genes encoding components of the electron transport system were also differentially expressed in a fur deletion mutant. The gene omcA (SO1779), which encodes a decaheme cytochrome c, exhibited significant decreases in both mRNA and protein abundance in the fur mutant and possessed a strong candidate Fur-binding site in its upstream region, thus suggesting that omcA may be a direct target of Fur activation.

Bacterial Proteins↗

Systematic learning of gene functional classes from DNA array expression data by using multilayer perceptrons.

Recent advances in microarray technology have opened new ways for functional annotation of previously uncharacterised genes on a genomic scale. This has been demonstrated by unsupervised clustering of co-expressed genes and, more importantly, by supervised learning algorithms. Using prior knowledge, these algorithms can assign functional annotations based on more complex expression signatures found in existing functional classes. Previously, support vector machines (SVMs) and other machine-learning methods have been applied to a limited number of functional classes for this purpose. Here we present, for the first time, the comprehensive application of supervised neural networks (SNNs) for functional annotation. Our study is novel in that we report systematic results for ~100 classes in the Munich Information Center for Protein Sequences (MIPS) functional catalog. We found that only ~10% of these are learnable (based on the rate of false negatives). A closer analysis reveals that false positives (and negatives) in a machine-learning context are not necessarily "false" in a biological sense. We show that the high degree of interconnections among functional classes confounds the signatures that ought to be learned for a unique class. We term this the "Borges effect" and introduce two new numerical indices for its quantification. Our analysis indicates that classification systems with a lower Borges effect are better suitable for machine learning. Furthermore, we introduce a learning procedure for combining false positives with the original class. We show that in a few iterations this process converges to a gene set that is learnable with considerably low rates of false positives and negatives and contains genes that are biologically related to the original class, allowing for a coarse reconstruction of the interactions between associated biological pathways. We exemplify this methodology using the well-studied tricarboxylic acid cycle.

Algorithms↗

Pathway analysis of coronary atherosclerosis.

Large-scale gene expression studies provide significant insight into genes differentially regulated in disease processes such as cancer. However, these investigations offer limited understanding of multisystem, multicellular diseases such as atherosclerosis. A systems biology approach that accounts for gene interactions, incorporates nontranscriptionally regulated genes, and integrates prior knowledge offers many advantages. We performed a comprehensive gene level assessment of coronary atherosclerosis using 51 coronary artery segments isolated from the explanted hearts of 22 cardiac transplant patients. After histological grading of vascular segments according to American Heart Association guidelines, isolated RNA was hybridized onto a customized 22-K oligonucleotide microarray, and significance analysis of microarrays and gene ontology analyses were performed to identify significant gene expression profiles. Our studies revealed that loss of differentiated smooth muscle cell gene expression is the primary expression signature of disease progression in atherosclerosis. Furthermore, we provide insight into the severe form of coronary artery disease associated with diabetes, reporting an overabundance of immune and inflammatory signals in diabetics. We present a novel approach to pathway development based on connectivity, determined by language parsing of the published literature, and ranking, determined by the significance of differentially regulated genes in the network. In doing this, we identify highly connected "nexus" genes that are attractive candidates for therapeutic targeting and followup studies. Our use of pathway techniques to study atherosclerosis as an integrated network of gene interactions expands on traditional microarray analysis methods and emphasizes the significant advantages of a systems-based approach to analyzing complex disease.

Adult↗

Reduction of hematopoietic cell-specific tyrosine phosphatase SHP-1 gene expression in natural killer cell lymphoma and various types of lymphomas/leukemias : combination analysis with cDNA expression array and tissue microarray.

To investigate the lymphomagenesis of NK/T lymphoma, we comprehensively and systematically analyzed the expression pattern of the human NK/T cell line (NK-YS) genome by cDNA expression array and tissue microarray. We detected significant changes in the gene expression of NK-YS cell line: an increase in 18 and a decrease in 20 genes compared to normal NK cells or peripheral blood mononuclear cells. Among these genes, we found a strong decrease in hematopoietic cell specific protein-tyrosine-phosphatase SH-PTP1 (SHP1) mRNA by cDNA expression array and reverse transcriptase-polymerase chain reaction. Further analysis with standard immunohistochemistry and tissue microarray, which used 207 paraffin-embedded specimens of various kinds of malignant lymphomas, showed that 100% of NK/T lymphoma specimens and more than 95% of various types of malignant lymphoma were negative for SHP1 protein expression. On the other hand, SHP1 protein was strongly expressed in the mantle zone and interfollicular zone lymphocytes in reactive lymphoid hyperplasia specimens. In addition, various kinds of hematopoietic cell lines, particularly the highly aggressive lymphoma/leukemia lines, lacked SHP1 expression in vitro, suggesting that loss of SHP1 expression may be related to not only malignant transformation, but also tumor cell aggressiveness. SHP1 expression could not be induced in either of two NK/T cell lines by phorbol ester, suggesting that genetic impairment or modification with methylation of SHP1 DNA could be one of the critical events in the pathogenesis of NK/T lymphoma. This evidence strongly suggests that loss of SHP1 gene expression plays an important role in multistep tumorigenesis, possibly as an anti-oncogene in the wide range of lymphomas/leukemias as well as NK/T lymphomas.

DNA, Complementary↗

Early Transcriptional Changes in Neutrophil-Mediated Processes Following Recanalization After Ischemic Stroke.

BACKGROUND: Ischemic stroke is a leading cause of death and long-term disability worldwide. Recanalization therapies, including thrombolysis and mechanical thrombectomy, restore blood flow, yet many patients experience poor outcomes, a phenomenon known as futile recanalization. Given the short therapeutic window for ischemic stroke, identifying early biomarkers to guide targeted interventions and improve outcomes is critical. METHODS: Using a murine middle cerebral occlusion model that mimics a large vessel occlusion with recanalization, a comprehensive microarray analysis from blood samples collected immediately and 3 hours after recanalization (N=44) was performed. Differentially expressed genes, enrichment pathways, immune cell proportions, enriched cell markers, predicted micro-RNAs, and transcription factors were identified using RStudio. Findings in mice were validated with rat middle cerebral artery occlusion (GSE21136) and patients with stroke (GSE16561) data sets to confirm transcriptional changes in peripheral blood postrecanalization. RESULTS: Il1r2, Cd55, Mmp8, Cd14, and Cd69 were early biomarkers poststroke and postrecanalization. Cross-validation revealed Vcan as a differentially expressed gene conserved across species, making it a novel ischemic marker detected as early as 3 hours postrecanalization (4 hours after middle cerebral artery occlusion) in mice, 24 hours after recanalization in rats (middle cerebral artery occlusion-thrombectomy), and within 24 hours from onset in humans receiving recombinant tissue plasminogen activator-thrombolysis. CIBERSORTx and ImmuCellAI-mouse deconvolution showed neutrophil elevation postrecanalization. Leukocyte and neutrophil activation pathways were enriched early after stroke in mice and humans, with stronger upregulation in the female sex. Several regulatory micro-RNAs were identified, and Nuclear Factor Erythroid 4 (NFE4) and Metal Regulatory Transcription Factor 1 (MTF1) emerged as key transcription factors. A coregulatory network underlying neutrophil activity was constructed, highlighting its central role in early responses to ischemia and recanalization, which was enriched in the female sex. CONCLUSIONS: We identified novel early genomic markers for ischemia and recanalization, including the conserved marker Vcan, and highlighted age- and sex-specific immune responses. Mapping a neutrophil-centered coregulatory network provides mechanistic insight into futile recanalization and supports the development of targeted therapies to improve clinical outcomes.

Animals↗

An atlas of gene expression from seed to seed through barley development.

Assaying relative and absolute levels of gene expression in a diverse series of tissues is a central step in the process of characterizing gene function and a necessary component of almost all publications describing individual genes or gene family members. However, throughout the literature, such studies lack consistency in genotype, tissues analyzed, and growth conditions applied, and, as a result, the body of information that is currently assembled is fragmented and difficult to compare between different studies. The development of a comprehensive platform for assaying gene expression that is available to the entire research community provides a major opportunity to assess whole biological systems in a single experiment. It also integrates detailed knowledge and information on individual genes into a unified framework that provides both context and resource to explore their contributions in a broader biological system. We have established a data set that describes the expression of 21,439 barley genes in 15 tissues sampled throughout the development of the barley cv. Morex grown under highly controlled conditions. Rather than attempting to address a specific biological question, our experiment was designed to provide a reference gene expression data set for barley researchers; a gene expression atlas and a comparative data set for those investigating genes or regulatory networks in other plant species. In this paper we describe the tissues sampled and their transcriptomes, and provide summary information on genes that are either specifically expressed in certain tissues or show correlated expression patterns across all 15 tissue samples. Using specific examples and an online tutorial, we describe how the data set can be interrogated for patterns and levels of barley gene expression and how the resulting information can be used to generate and/or test specific biological hypotheses.

Databases, Genetic↗

Comprehensive bioinformatics analysis identifies candidate ciliogenesis-related genes preferentially associated with N0-stage lung squamous cell carcinoma.

PURPOSE: There is few research on which genes play an important role in tumors without lymph metastasis. This study aimed to identify candidate molecular alterations preferentially associated with N0-stage LUSC. METHODS: we conducted a comprehensive bioinformatics analysis using publicly available The Cancer Genome Atlas (TCGA) data. Differentially expressed genes (DEGs) were identified separately by comparing N0 tumors and N+ tumors with normal lung tissues. Genes dysregulated in both N0 and N+ tumors were excluded to identify candidate N0-associated genes PPI networks were constructed using STRING and Cytoscape, with module analysis performed via MCODE. Hub genes were identified using multiple Cytohubba algorithms. Functional enrichment analyses were conducted using GO, and KEGG pathways using DAVID. Gene interaction networks were further explored using GeneMANIA. Immune cell infiltration was evaluated with TIMER. Associations with pathological stage and patient survival were assessed using GEPIA and other relevant tools. RESULTS: A total of 1103 candidate N0-associated DEGs were identified, including 748 upregulated and 355 downregulated genes. The PPI network contained five major MCODE clusters. One cluster (MCODE 4) included TTC30A, TTC30B, BBS7, and KIF3B genes implicated in ciliogenesis. TTC30B showed significant differential expression across pathological stages in the overall LUSC cohort. Seven consensus hub genes (ERBB2, CHUK, CASP8, NOTCH1, HNF4A, CREBBP, and IRS1) were identified based on their consistent ranking across multiple CytoHubba algorithms. Upregulated candidate N0-associated genes were primarily enriched in immune-related processes, including B-cell-mediated immunity and humoral responses, whereas downregulated genes were enriched in lysosomal and trans-Golgi network-related pathways. Exploratory immune infiltration analyses identified associations between the four ciliogenesis-related genes and several immune cell populations. CONCLUSIONS: This study identified candidate molecular signatures preferentially associated with N0-stage LUSC, including ciliogenesis-related genes and consensus hub genes. These findings provide hypotheses regarding molecular features of N0-stage LUSC and warrant further validation in independent cohorts and experimental studies.

Humans↗

Construction of expression-ready cDNA clones for KIAA genes: manual curation of 330 KIAA cDNA clones.

We have accumulated information on protein-coding sequences of uncharacterized human genes, which are known as KIAA genes, through cDNA sequencing. For comprehensive functional analysis of the KIAA genes, it is necessary to prepare a set of cDNA clones which direct the synthesis of functional KIAA gene products. However, since the KIAA cDNAs were derived from long mRNAs (> 4 kb), it was not expected that all of them were full-length. Thus, as the first step toward preparing these clones, we evaluated the integrity of protein-coding sequences of KIAA cDNA clones through comparison with homologous protein entries in the public database. As a result, 1141 KIAA cDNAs had at least one homologous entry in the database, and 619 of them (54%) were found to be truncated at the 5' and/or 3' ends. In this study, 290 KIAA cDNA clones were tailored to be full-length or have considerably longer sequences than the original clones by isolating additional cDNA clones and/or connected parts of additional cDNAs or PCR products of the missing portion to the original cDNA clone. Consequently, 265, 8, and 17 predicted CDSs of KIAA cDNA clones were increased in the amino-, carboxy-, and both terminal sequences, respectively. In addition, 40 cDNA clones were modified to remove spurious interruption of protein-coding sequences. The total length of the resultant extensions at amino- and carboxy-terminals of KIAA gene products reached 97,000 and 7,216 amino acid residues, respectively, and various protein domains were found in these extended portions.

Cloning, Molecular↗

Gene discovery in oral squamous cell carcinoma through the Head and Neck Cancer Genome Anatomy Project: confirmation by microarray analysis.

The near completion of the human genome project and the recent development of novel, highly sensitive high-throughput techniques have now afforded the unique opportunity to perform a comprehensive molecular characterization of normal, precancerous, and malignant cells, including those derived from squamous carcinomas of the head and neck (HNSCC). As part of these efforts, representative cDNA libraries from patient sets, comprising of normal and malignant squamous epithelium, were generated and contributed to the Head and Neck Cancer Genome Anatomy Project (HN-CGAP). Initial analysis of the sequence information indicated the existence of many novel genes in these libraries [Oral Oncol 36 (2000) 474]. In this study, we surveyed the available sequence information using bioinformatic tools and identified a number of known genes that were differentially expressed in normal and malignant epithelium. Furthermore, this effort resulted in the identification of 168 novel genes. Comparison of these clones to the human genome identified clusters in loci that were not previously recognized as being altered in HNSCC. To begin addressing which of these novel genes are frequently expressed in HNSCC, their DNA was used to construct an oral-cancer-specific microarray, which was used to hybridize alpha-(33)P dCTP labeled cDNA derived from five HNSCC patient sets. Initial assessment demonstrated 10 clones to be highly expressed (>2-fold) in the normal squamous epithelium, while 14 were highly represented in the malignant counterpart, in three of the five patient sets, thus suggesting that a subset of these newly discovered transcripts might be highly expressed in this tumor type. These efforts, together with other multi-institutional genomic and proteomic initiatives are expected to contribute to the complete understanding of the molecular pathogenesis of HNSCCs, thus helping to identify new markers for the early detection of preneoplastic lesions and novel targets for pharmacological intervention in this disease.

Aged↗

Comprehensive analysis of gene expression patterns of hedgehog-related genes.

BACKGROUND: The Caenorhabditis elegans genome encodes ten proteins that share sequence similarity with the Hedgehog signaling molecule through their C-terminal autoprocessing Hint/Hog domain. These proteins contain novel N-terminal domains, and C. elegans encodes dozens of additional proteins containing only these N-terminal domains. These gene families are called warthog, groundhog, ground-like and quahog, collectively called hedgehog (hh)-related genes. Previously, the expression pattern of seventeen genes was examined, which showed that they are primarily expressed in the ectoderm. RESULTS: With the completion of the C. elegans genome sequence in November 2002, we reexamined and identified 61 hh-related ORFs. Further, we identified 49 hh-related ORFs in C. briggsae. ORF analysis revealed that 30% of the genes still had errors in their predictions and we improved these predictions here. We performed a comprehensive expression analysis using GFP fusions of the putative intergenic regulatory sequence with one or two transgenic lines for most genes. The hh-related genes are expressed in one or a few of the following tissues: hypodermis, seam cells, excretory duct and pore cells, vulval epithelial cells, rectal epithelial cells, pharyngeal muscle or marginal cells, arcade cells, support cells of sensory organs, and neuronal cells. Using time-lapse recordings, we discovered that some hh-related genes are expressed in a cyclical fashion in phase with molting during larval development. We also generated several translational GFP fusions, but they did not show any subcellular localization. In addition, we also studied the expression patterns of two genes with similarity to Drosophila frizzled, T23D8.1 and F27E11.3A, and the ortholog of the Drosophila gene dally-like, gpn-1, which is a heparan sulfate proteoglycan. The two frizzled homologs are expressed in a few neurons in the head, and gpn-1 is expressed in the pharynx. Finally, we compare the efficacy of our GFP expression effort with EST, OST and SAGE data. CONCLUSION: No bona-fide Hh signaling pathway is present in C. elegans. Given that the hh-related gene products have a predicted signal peptide for secretion, it is possible that they constitute components of the extracellular matrix (ECM). They might be associated with the cuticle or be present in soluble form in the body cavity. They might interact with the Patched or the Patched-related proteins in a manner similar to the interaction of Hedgehog with its receptor Patched.

Animals↗

SPINE: an integrated tracking database and data mining approach for identifying feasible targets in high-throughput structural proteomics.

High-throughput structural proteomics is expected to generate considerable amounts of data on the progress of structure determination for many proteins. For each protein this includes information about cloning, expression, purification, biophysical characterization and structure determination via NMR spectroscopy or X-ray crystallography. It will be essential to develop specifications and ontologies for standardizing this information to make it amenable to retrospective analysis. To this end we created the SPINE database and analysis system for the Northeast Structural Genomics Consortium. SPINE, which is available at bioinfo.mbb.yale.edu/nesg or nesg.org, is specifically designed to enable distributed scientific collaboration via the Internet. It was designed not just as an information repository but as an active vehicle to standardize proteomics data in a form that would enable systematic data mining. The system features an intuitive user interface for interactive retrieval and modification of expression construct data, query forms designed to track global project progress and external links to many other resources. Currently the database contains experimental data on 985 constructs, of which 740 are drawn from Methanobacterium thermoautotrophicum, 123 from Saccharomyces cerevisiae, 93 from Caenorhabditis elegans and the remainder from other organisms. We developed a comprehensive set of data mining features for each protein, including several related to experimental progress (e.g. expression level, solubility and crystallization) and 42 based on the underlying protein sequence (e.g. amino acid composition, secondary structure and occurrence of low complexity regions). We demonstrate in detail the application of a particular machine learning approach, decision trees, to the tasks of predicting a protein's solubility and propensity to crystallize based on sequence features. We are able to extract a number of key rules from our trees, in particular that soluble proteins tend to have significantly more acidic residues and fewer hydrophobic stretches than insoluble ones. One of the characteristics of proteomics data sets, currently and in the foreseeable future, is their intermediate size ( approximately 500-5000 data points). This creates a number of issues in relation to error estimation. Initially we estimate the overall error in our trees based on standard cross-validation. However, this leaves out a significant fraction of the data in model construction and does not give error estimates on individual rules. Therefore, we present alternative methods to estimate the error in particular rules.

Animals↗

TmaDB: a repository for tissue microarray data.

BACKGROUND: Tissue microarray (TMA) technology has been developed to facilitate large, genome-scale molecular pathology studies. This technique provides a high-throughput method for analyzing a large cohort of clinical specimens in a single experiment thereby permitting the parallel analysis of molecular alterations (at the DNA, RNA, or protein level) in thousands of tissue specimens. As a vast quantity of data can be generated in a single TMA experiment a systematic approach is required for the storage and analysis of such data. DESCRIPTION: To analyse TMA output a relational database (known as TmaDB) has been developed to collate all aspects of information relating to TMAs. These data include the TMA construction protocol, experimental protocol and results from the various immunocytological and histochemical staining experiments including the scanned images for each of the TMA cores. Furthermore the database contains pathological information associated with each of the specimens on the TMA slide, the location of the various TMAs and the individual specimen blocks (from which cores were taken) in the laboratory and their current status i.e. if they can be sectioned into further slides or if they are exhausted. TmaDB has been designed to incorporate and extend many of the published common data elements and the XML format for TMA experiments and is therefore compatible with the TMA data exchange specifications developed by the Association for Pathology Informatics community. Finally the design of the database is made flexible such that TMA experiments from several types of cancer can be stored in a single database, which incorporates the national minimum data set required for pathology reports supported by the Royal College of Pathologists (UK). CONCLUSION: TmaDB will provide a comprehensive repository for TMA data such that a large number of results from the numerous immunostaining experiments can be efficiently compared for each of the TMA cores. This will allow a systematic, large-scale comparison of tumour samples to facilitate the identification of gene products of clinical importance such as therapeutic or prognostic markers. In addition this work will contribute to the establishment of a standard for reporting TMA data analogous to MIAME in the description of microarray data.

Data Interpretation, Statistical↗

Genome-wide atlas of gene expression in the adult mouse brain.

Molecular approaches to understanding the functional circuitry of the nervous system promise new insights into the relationship between genes, brain and behaviour. The cellular diversity of the brain necessitates a cellular resolution approach towards understanding the functional genomics of the nervous system. We describe here an anatomically comprehensive digital atlas containing the expression patterns of approximately 20,000 genes in the adult mouse brain. Data were generated using automated high-throughput procedures for in situ hybridization and data acquisition, and are publicly accessible online. Newly developed image-based informatics tools allow global genome-scale structural analysis and cross-correlation, as well as identification of regionally enriched genes. Unbiased fine-resolution analysis has identified highly specific cellular markers as well as extensive evidence of cellular heterogeneity not evident in classical neuroanatomical atlases. This highly standardized atlas provides an open, primary data resource for a wide variety of further studies concerning brain organization and function.

Animals↗

Comprehensive In Silico Analysis Identifies MSTO1 and LIG1 as Candidate Biomarkers With Diagnostic and Prognostic Relevance in Hepatocellular Carcinoma.

BACKGROUND: Hepatocellular carcinoma (HCC) is the most common primary liver malignancy and remains a major cause of cancer-related mortality worldwide. Its poor clinical outcomes are largely attributed to late-stage diagnosis and the limited accuracy of currently available diagnostic and prognostic biomarkers. Therefore, identifying novel molecular markers with improved sensitivity, specificity, and therapeutic relevance is essential for enhancing early detection and guiding personalized treatment strategies. AIMS: To identify and prioritize novel candidate HCC biomarkers with diagnostic and prognostic value and potential therapeutic vulnerability using integrated multi-omics, survival, functional dependency, and tumor microenvironment analyses. METHODS AND RESULTS: We examined the mRNA and protein expression levels of 8 DEGs in HCC tissues in the TCGA and CPTAC datasets using UALCAN, which showed that MSTO1 and LIG1 were overexpressed consistently in HCC relative to normal liver tissues. Moreover, elevated expression levels of these genes were significantly associated with higher tumor grade and advanced stage. Kaplan-Meier plotter survival data confirmed that increased expression of MSTO1 and LIG1 was associated with poorer overall survival. The DepMap CRISPR knockout data confirmed a functional dependency of both genes in HCC cell lines. CBioPortal analyses provided characterization of genomic alterations and enabled enrichment analysis of co-expressed genes, and the TCGA-UALCAN pan-cancer analyses supported the assessment of tissue specificity across tumor types. TIMER3 analyses linked candidate gene expression with immune cell infiltration patterns. Diagnostic performance by ROC analysis showed excellent discrimination for MSTO1 (AUC = 0.987) and good discrimination for LIG1 (AUC = 0.897). Multivariate Cox regression with Benjamini-Hochberg FDR correction across the eight genes supported MSTO1 as a candidate independent prognostic factor after adjustment for tumor stage, grade, etiology, age, and sex (HR = 1.29, p = 0.035), whilst LIG1 showed no independent prognostic value. Promoter methylation of MSTO1 and ADH4, assessed via UALCAN, showed that both genes were significantly differentially methylated in the promoter region of primary HCC tissues compared with normal liver tissues. Our study also confirmed the biological and clinical relevance of established HCC biomarkers: TERT, IRAK1, and ADH4. CONCLUSION: MSTO1 and LIG1 emerged as candidate diagnostic biomarkers in HCC. Additionally, MSTO1 showed a candidate prognostic association with overall survival that remained significant after adjusting for tumor stage, grade, and etiology, as well as patients' age, but not after further adjustment for AFP status. Functional data also highlighted MSTO1 as a candidate therapeutic dependency. On the other hand, LIG1 showed no independent prognostic association in either multivariate model. Their differential expression and functional essentiality in HCC cell lines highlighted their value for further experimental and independent-cohort validation before potential integration into biomarker development pipelines aimed at improving early detection and targeted therapy in HCC.

Humans↗