Search PubMed⌕ Search

Biomedical subjects

RIKEN GER Group

Publications and source records attributed to RIKEN GER Group.

At least 19 recordsLinked to original sources

Identification of putative noncoding RNAs among the RIKEN mouse full-length cDNA collection.

With the sequencing and annotation of genomes and transcriptomes of several eukaryotes, the importance of noncoding RNA (ncRNA)-RNA molecules that are not translated to protein products-has become more evident. A subclass of ncRNA transcripts are encoded by highly regulated, multi-exon, transcriptional units, are processed like typical protein-coding mRNAs and are increasingly implicated in regulation of many cellular functions in eukaryotes. This study describes the identification of candidate functional ncRNAs from among the RIKEN mouse full-length cDNA collection, which contains 60,770 sequences, by using a systematic computational filtering approach. We initially searched for previously reported ncRNAs and found nine murine ncRNAs and homologs of several previously described nonmouse ncRNAs. Through our computational approach to filter artifact-free clones that lack protein coding potential, we extracted 4280 transcripts as the largest-candidate set. Many clones in the set had EST hits, potential CpG islands surrounding the transcription start sites, and homologies with the human genome. This implies that many candidates are indeed transcribed in a regulated manner. Our results demonstrate that ncRNAs are a major functional subclass of processed transcripts in mammals.

Animals↗

Exploration of the cell-cycle genes found within the RIKEN FANTOM2 data set.

The cell cycle is one of the most fundamental processes within a cell. Phase-dependent expression and cell-cycle checkpoints require a high level of control. A large number of genes with varying functions and modes of action are responsible for this biology. In a targeted exploration of the FANTOM2-Variable Protein Set, a number of mouse homologs to known cell-cycle regulators as well as novel members of cell-cycle families were identified. Focusing on two prototype cell-cycle families, the cyclins and the NIMA-related kinases (NEKs), we believe we have identified all of the mouse members of these families, 24 cyclins and 10 NEKs, and mapped them to ENSEMBL transcripts. To attempt to globally identify all potential cell cycle-related genes within mouse, the MGI (Mouse Genome Database) assignments for the RIKEN Representative Set (RPS) and the results from two homology-based queries were merged. We identified 1415 genes with possible cell-cycle roles, and 1758 potential paralogs. We comment on the genes identified in this screen and evaluate the merits of each approach.

Amino Acid Sequence↗

Identification and analysis of chromodomain-containing proteins encoded in the mouse transcriptome.

The chromodomain is 40-50 amino acids in length and is conserved in a wide range of chromatic and regulatory proteins involved in chromatin remodeling. Chromodomain-containing proteins can be classified into families based on their broader characteristics, in particular the presence of other types of domains, and which correlate with different subclasses of the chromodomains themselves. Hidden Markov model (HMM)-generated profiles of different subclasses of chromodomains were used here to identify sequences encoding chromodomain-containing proteins in the mouse transcriptome and genome. A total of 36 different loci encoding proteins containing chromodomains, including 17 novel loci, were identified. Six of these loci (including three apparent pseudogenes, a novel HP1 ortholog, and two novel Msl-3 transcription factor-like proteins) are not present in the human genome, whereas the human genome contains four loci (two CDY orthologs and two apparent CDY pseudogenes) that are not present in mouse. A number of these loci exhibit alternative splicing to produce different isoforms, including 43 novel variants, some of which lack the chromodomain. The likely functions of these proteins are discussed in relation to the known functions of other chromodomain-containing proteins within the same family.

Acetyltransferases↗

Cytokine-related genes identified from the RIKEN full-length mouse cDNA data set.

To identify novel cytokine-related genes, we searched the set of 60,770 annotated RIKEN mouse cDNA clones (FANTOM2 clones), using keywords such as cytokine itself or cytokine names (such as interferon, interleukin, epidermal growth factor, fibroblast growth factor, and transforming growth factor). This search produced 108 known cytokines and cytokine-related products such as cytokine receptors, cytokine-associated genes, or their products (enhancers, accessory proteins, cytokine-induced genes). We found 15 clusters of FANTOM2 clones that are candidates for novel cytokine-related genes. These encoded products with strong sequence similarity to guanylate-binding protein (GBP-5), interleukin-1 receptor-associated kinase 2 (IRAK-2), interleukin 20 receptor alpha isoform 3, a member of the interferon-inducible proteins of the Ifi 200 cluster, four members of the membrane-associated family 1-8 of interferon-inducible proteins, one p27-like protein, and a hypothetical protein containing a Toll/Interleukin receptor domain. All four clones representing novel candidates of gene products from the family contain a novel highly conserved cross-species domain. Clones similar to growth factor-related products included transforming growth factor beta-inducible early growth response protein 2 (TIEG-2), TGFbeta-induced factor 2, integrin beta-like 1, latent TGF-binding protein 4S, and FGF receptor 4B. We performed a detailed sequence analysis of the candidate novel genes to elucidate their likely functional properties.

Amino Acid Sequence↗

Impact of alternative initiation, splicing, and termination on the diversity of the mRNA transcripts encoded by the mouse transcriptome.

We analyzed the FANTOM2 clone set of 60,770 RIKEN full-length mouse cDNA sequences and 44,122 public mRNA sequences. We developed a new computational procedure to identify and classify the forms of splice variation evident in this data set and organized the results into a publicly accessible database that can be used for future expression array construction, structural genomics, and analyses of the mechanism and regulation of alternative splicing. Statistical analysis shows that at least 41% and possibly as much as 60% of multiexon genes in mouse have multiple splice forms. Of the transcription units with multiple splice forms, 49% contain transcripts in which the apparent use of an alternative transcription start (stop) is accompanied by alternative splicing of the initial (terminal) exon. This implies that alternative transcription may frequently induce alternative splicing. The fact that 73% of all exons with splice variation fall within the annotated coding region indicates that most splice variation is likely to affect the protein form. Finally, we compared the set of constitutive (present in all transcripts) exons with the set of cryptic (present only in some transcripts) exons and found statistically significant differences in their length distributions, the nucleotide distributions around their splice junctions, and the frequencies of occurrence of several short sequence motifs.

Alternative Splicing↗

Comparative analysis of apoptosis and inflammation genes of mice and humans.

Apoptosis (programmed cell death) plays important roles in many facets of normal mammalian physiology. Host-pathogen interactions have provided evolutionary pressure for apoptosis as a defense mechanism against viruses and microbes, sometimes linking apoptosis mechanisms with inflammatory responses through NFkappaB induction. Proteins involved in apoptosis and NFkappaB induction commonly contain evolutionarily conserved domains that can serve as signatures for identification by bioinformatics methods. Using a combination of public (NCBI) and private (RIKEN) databases, we compared the repertoire of apoptosis and NFkappaB-inducing genes in humans and mice from cDNA/EST/genomic data, focusing on the following domain families: (1) Caspase proteases; (2) Caspase recruitment domains (CARD); (3) Death Domains (DD); (4) Death Effector Domains (DED); (5) BIR domains of Inhibitor of Apoptosis Proteins (IAPs); (6) Bcl-2 homology (BH) domains of Bcl-2 family proteins; (7) Tumor Necrosis Factor (TNF)-family ligands; (8) TNF receptors (TNFR); (9) TIR domains; (10) PAAD (PYRIN; PYD, DAPIN); (11) nucleotide-binding NACHT domains; (12) TRAFs; (13) Hsp70-binding BAG domains; (14) endonuclease-associated CIDE domains; and (15) miscellaneous additional proteins. After excluding redundancy due to alternative splice forms, sequencing errors, and other considerations, we identified cDNAs derived from a total of 227 human genes among these domain families. Orthologous murine genes were found for 219 (96%); in addition, several unique murine genes were found, which appear not to have human orthologs. This mismatch may be due to the still fragmentary information about the mouse genome or genuine differences between mouse and human repertoires of apoptotic genes. With this caveat, we discuss similarities and differences in human and murine genes from these domain families.

Animals↗

Discovery of imprinted transcripts in the mouse transcriptome using large-scale expression profiling.

Candidate imprinted transcriptional units in the mouse genome were identified systematically from 27,663 FANTOM2 full-length mouse cDNA clones by expression profiling. Large-scale cDNA microarrays were used to detect differential expression dependent upon chromosomal parent of origin by comparing the mRNA levels in the total tissue of 9.5 dpc parthenogenote and androgenote mouse embryos. Of the FANTOM2 transcripts, 2114 were identified as candidates on the basis of the array data. Of these, 39 mapped to known imprinted regions of the mouse genome, 56 were considered as nonprotein-coding RNAs, and 159 were natural antisense transcripts. The imprinted expression of two transcripts located in the mouse chromosomal region syntenic to the human Prader-Willi syndrome region was confirmed experimentally. We further mapped all candidate imprinted transcripts to the mouse and human genome and were shown in correlation with the imprinting disease loci. These data provide a major resource for understanding the role of imprinting in mammalian inherited traits.

Animals↗

Continued discovery of transcriptional units expressed in cells of the mouse mononuclear phagocyte lineage.

The current RIKEN transcript set represents a significant proportion of the mouse transcriptome but transcripts expressed in the innate and acquired immune systems are poorly represented. In the present study we have assessed the complexity of the transcriptome expressed in mouse macrophages before and after treatment with lipopolysaccharide, a global regulator of macrophage gene expression, using existing RIKEN 19K arrays. By comparison to array profiles of other cells and tissues, we identify a large set of macrophage-enriched genes, many of which have obvious functions in endocytosis and phagocytosis. In addition, a significant number of LPS-inducible genes were identified. The data suggest that macrophages are a complex source of mRNA for transcriptome studies. To assess complexity and identify additional macrophage expressed genes, cDNA libraries were created from purified populations of macrophage and dendritic cells, a functionally related cell type. Sequence analysis revealed a high incidence of novel mRNAs within these cDNA libraries. These studies provide insights into the depths of transcriptional complexity still untapped amongst products of inducible genes, and identify macrophage and dendritic cell populations as a starting point for sampling the inducible mammalian transcriptome.

Animals↗

Systematic expression profiling of the mouse transcriptome using RIKEN cDNA microarrays.

The number of known mRNA transcripts in the mouse has been greatly expanded by the RIKEN Mouse Gene Encyclopedia project. Validation of their reproducible expression in a tissue is an important contribution to the study of functional genomics. In this report, we determine the expression profile of 57,931 clones on 20 mouse tissues using cDNA microarrays. Of these 57,931 clones, 22,928 clones correspond to the FANTOM2 clone set. The set represents 20,234 transcriptional units (TUs) out of 33,409 TUs in the FANTOM2 set. We identified 7206 separate clones that satisfied stringent criteria for tissue-specific expression. Gene Ontology terms were assigned for these 7206 clones, and the proportion of 'molecular function' ontology for each tissue-specific clone was examined. These data will provide insights into the function of each tissue. Tissue-specific gene expression profiles obtained using our cDNA microarrays were also compared with the data extracted from the GNF Expression Atlas based on Affymetrix microarrays. One major outcome of the RIKEN transcriptome analysis is the identification of numerous nonprotein-coding mRNAs. The expression profile was also used to obtain evidence of expression for putative noncoding RNAs. In addition, 1926 clones (70%) of 2768 clones that were categorized as "unknown EST," and 1969 (58%) clones of 3388 clones that were categorized as "unclassifiable" were also shown to be reproducibly expressed.

Animals↗

G protein-coupled receptor genes in the FANTOM2 database.

G protein-coupled receptors (GPCRs) comprise the largest family of receptor proteins in mammals and play important roles in many physiological and pathological processes. Gene expression of GPCRs is temporally and spatially regulated, and many splicing variants are also described. In many instances, different expression profiles of GPCR gene are accountable for the changes of its biological function. Therefore, it is intriguing to assess the complexity of the transcriptome of GPCRs in various mammalian organs. In this study, we took advantage of the FANTOM2 (Functional Annotation Meeting of Mouse cDNA 2) project, which aimed to collect full-length cDNAs inclusively from mouse tissues, and found 410 candidate GPCR cDNAs. Clustering of these clones into transcriptional units (TUs) reduced this number to 213. Out of these, 165 genes were represented within the known 308 GPCRs in the Mouse Genome Informatics (MGI) resource. The remaining 48 genes were new to mouse, and 14 of them had no clear mammalian ortholog. To dissect the detailed characteristics of each transcript, tissue distribution pattern and alternative splicing were also ascertained. We found many splicing variants of GPCRs that may have a relevance to disease occurrence. In addition, the difficulty in cloning tissue-specific and infrequently transcribed GPCRs is discussed further.

Alternative Splicing↗

Analysis of the mouse transcriptome for genes involved in the function of the nervous system.

We analyzed the mouse Representative Transcript and Protein Set for molecules involved in brain function. We found full-length cDNAs of many known brain genes and discovered new members of known brain gene families, including Family 3 G-protein coupled receptors, voltage-gated channels, and connexins. We also identified previously unknown candidates for secreted neuroactive molecules. The existence of a large number of unique brain ESTs suggests an additional molecular complexity that remains to be explored.A list of genes containing CAG stretches in the coding region represents a first step in the potential identification of candidates for hereditary neurological disorders.

Adenine↗

Systematic characterization of the zinc-finger-containing proteins in the mouse transcriptome.

Zinc-finger-containing proteins can be classified into evolutionary and functionally divergent protein families that share one or more domains in which a zinc ion is tetrahedrally coordinated by cysteines and histidines. The zinc finger domain defines one of the largest protein superfamilies in mammalian genomes;46 different conserved zinc finger domains are listed in InterPro (http://www.ebi.ac.uk/InterPro). Zinc finger proteins can bind to DNA, RNA, other proteins, or lipids as a modular domain in combination with other conserved structures. Owing to this combinatorial diversity, different members of zinc finger superfamilies contribute to many distinct cellular processes, including transcriptional regulation, mRNA stability and processing, and protein turnover. Accordingly, mutations of zinc finger genes lead to aberrations in a broad spectrum of biological processes such as development, differentiation, apoptosis, and immunological responses. This study provides the first comprehensive classification of zinc finger proteins in a mammalian transcriptome. Specific detailed analysis of the SP/Krüppel-like factors and the E3 ubiquitin-ligase RING-H2 families illustrates the importance of such an analysis for a more comprehensive functional classification of large protein families. We describe the characterization of a new family of C2H2 zinc-finger-containing proteins and a new conserved domain characteristic of this family, the identification and characterization of Sp8, a new member of the Sp family of transcriptional regulators, and the identification of five new RING-H2 proteins.

Alternative Splicing↗

Phosphoregulators: protein kinases and protein phosphatases of mouse.

With the completion of the human and mouse genome sequences, the task now turns to identifying their encoded transcripts and assigning gene function. In this study, we have undertaken a computational approach to identify and classify all of the protein kinases and phosphatases present in the mouse gene complement. A nonredundant set of these sequences was produced by mining Ensembl gene predictions and publicly available cDNA sequences with a panel of InterPro domains. This approach identified 561 candidate protein kinases and 162 candidate protein phosphatases. This cohort was then analyzed using TribeMCL protein sequence similarity clustering followed by CLUSTALV alignment and hierarchical tree generation. This approach allowed us to (1) distinguish between true members of the protein kinase and phosphatase families and enzymes of related biochemistry, (2) determine the structure of the families, and (3) suggest functions for previously uncharacterized members. The classifications obtained by this approach were in good agreement with previous schemes and allowed us to demonstrate domain associations with a number of clusters. Finally, we comment on the complementary nature of cDNA and genome-based gene detection and the impact of the FANTOM2 transcriptome project.

Animals↗

A comprehensive transcript map of the mouse Gnas imprinted complex.

The recent publication of the FANTOM mouse transcriptome has provided a unique opportunity to study the diversity of transcripts arising from a single gene locus. We have focused on the Gnas complex, as imprinting loci themselves provide unique insights into transcriptional regulation. Thirteen full-length cDNAs from the FANTOM2 set were mapped to the Gnas locus. These represented one previously described transcript and 12 putative new transcripts. Of these, eight were found to be differentially expressed from either the maternal or paternal allele. Two clones extended Nespas in the 3' direction, providing evidence of antisense transcription spanning a 30-kb genomic region from a single allele. The transcripts were summarized into six transcriptional units, Nespas, Nesp, Gnasxl, F7, exon 1A, and Gnas. The resolution of the Gnas transcript map by the FANTOM2 clones revealed a pattern of alternate splicing. In addition to the transcripts described previously as splicing onto exon 2 of Gnas, each new sense transcript had an alternate short 3'UTR independent of Gnas. Both spliced and unspliced variants of the new imprinted sense transcripts were found. Whereas the functional significance of these alternate transcripts is not known, the availability of the FANTOM clones has provided remarkable insights into the repertoire of transcripts in the Gnas complex locus.

Animals↗

Comprehensive analysis of the mouse metabolome based on the transcriptome.

The complete set of cDNAs encoding the enzymes of known metabolic pathways has not previously been available for any mammal. Here, transcripts encoding the metabolic pathways of the mouse (mouse metabolome) were reconstructed by making use of the KEGG metabolic pathway database and gene ontology (GO) assignment to the mouse representative transcript and protein set (RTPS), which contains all available mouse transcript sequences including the FANTOM set of RIKEN mouse cDNA clones. By assigning EC numbers extracted from the molecular function ontology in GO, the known mouse transcriptome was predicted to encode enzymes with 726 unique EC numbers. Of these, 648 EC numbers were newly assigned based on the FANTOM set. The mouse metabolome confirmed by cDNA analysis includes almost all of the enzymes of well known pathways such as the tricarboxylic acid cycle and urea cycle. On the other hand, analysis of enzymes required for the tryptophan metabolism pathway revealed a lack of connectivity, indicating that cDNAs/genes encoding several key enzymes remain to be identified. The information derived from coexpression from the cDNA microarray analysis of enzymes of known function may lead to identification of the missing components of the metabolome, and will add new insights into the connectivity of the mammalian metabolic pathways.

Amino Acids↗

Mouse proteome analysis.

A general overview of the protein sequence set for the mouse transcriptome produced during the FANTOM2 sequencing project is presented here. We applied different algorithms to characterize protein sequences derived from a nonredundant representative protein set (RPS) and a variant protein set (VPS) of the mouse transcriptome. The functional characterization and assignment of Gene Ontology terms was done by analysis of the proteome using InterPro. The Superfamily database analyses gave a detailed structural classification according to SCOP and provide additional evidence for the functional characterization of the proteome data. The MDS database analysis revealed new domains which are not presented in existing protein domain databases. Thus the transcriptome gives us a unique source of data for the detection of new functional groups. The data obtained for the RPS and VPS sets facilitated the comparison of different patterns of protein expression. A comparison of other existing mouse and human protein sequence sets (e.g., the International Protein Index) demonstrates the common patterns in mammalian proteomes. The analysis of the membrane organization within the transcriptome of multiple eukaryotes provides valuable statistics about the distribution of secretory and transmembrane proteins

Animals↗

Human disease genes and their cloned mouse orthologs: exploration of the FANTOM2 cDNA sequence data set.

The FANTOM2 cDNA sequence data set is an excellent model to demonstrate the power of large-scale cDNA sequencing, with the goal of providing a full-length transcript sequence for each mouse gene. This data set enhances the use of the mouse as a model for human disease. Here we identify mouse cDNA sequences in the FANTOM2 data set for a set of 67 human disease genes that as of May 2002 had no corresponding mouse cDNA annotated in the Mouse Genome Informatics (MGI) database. These 67 human disease genes include genes related to neurological and eye disorders and cancer. We also present a list of the human disease genes and their cloned mouse orthologs found in two public databases, LocusLink and MGI. Allelic variant and gene functional information available in MGI provides additional information relative to these mouse models, whereas computed sequence-based connections at NCBI support facile navigation through multiple genomes.

Alleles↗

The comparative proteomics of ubiquitination in mouse.

Ubiquitination is a common posttranslational modification in eukaryotic cells, influencing many fundamental cellular processes. Defects in ubiquitination and the processes it mediates are involved in many human disease states. The ubiquitination of a substrate involves four classes of enzymes:a ubiquitin-activating enzyme (E1), a ubiquitin-conjugating enzyme (E2), a ubiquitin protein ligase (E3), and a de-ubiquitinating enzyme (DUB). A substantial number of E1s (four), E2s (13), E3s (97), and DUBs (six) that were previously unknown in the mouse are included in the FANTOM2 Representative Transcript and Protein Set (RTPS). Many of the genes encoding these proteins will constitute promising candidates for involvement in disease. In addition, the RTPS provides the basis for the most comprehensive survey of ubiquitination-associated proteins across eukaryotes undertaken to date. Comparisons of these proteins across human and other organisms suggest that eukaryotic evolution has been associated with an increase in the number and diversity of E3s (possessing either zinc-finger RING, F-box, or HECT domains) and DUBs (containing the ubiquitin thiolesterase family 2 domain). These increases in numbers are too large to be accounted for by the presence of fragmentary proteins in the data sets examined. Much of this innovation appears to have been associated with the emergence of multicellular organisms, and subsequently of vertebrates, increasing the opportunity for complex regulation of ubiquitination-mediated cellular and developmental processes.

Animals↗