Search PubMed⌕ Search

Biomedical subjects

Takahiro Arakawa

Publications and source records attributed to Takahiro Arakawa.

At least 19 recordsLinked to original sources

Microfluidics assisted synthesis of well-defined spherical polymeric microcapsules and their utilization as potential encapsulants.

In this article, the development of a novel technique to fabricate spherical polymeric microcapsules by utilizing microfluidic technology is presented. Atom transfer radical polymerization (ATRP) was employed to synthesize well-defined amphiphilic block copolymers. An organic polymer solution was constrained to adopt the spherical droplets in a continuous water phase at a T-junction microchannel, and the generation of the droplets was studied quantitatively. The flow conditions of two immiscible solutions were adjusted for the successful generation of the polymer droplets. The morphology of the microcapsules was examined. The efficiency of these polymer microcapsules as containers for the storage and controlled release of loaded molecules was evaluated by encapsulating the microcapsules with Congo-red dye and investigating the release performance using temperature controlled UV-VIS spectroscopy.

Capsules↗

Protein-protein interactions of the hyperthermophilic archaeon Pyrococcus horikoshii OT3.

BACKGROUND: Although 2,061 proteins of Pyrococcus horikoshii OT3, a hyperthermophilic archaeon, have been predicted from the recently completed genome sequence, the majority of proteins show no similarity to those from other organisms and are thus hypothetical proteins of unknown function. Because most proteins operate as parts of complexes to regulate biological processes, we systematically analyzed protein-protein interactions in Pyrococcus using the mammalian two-hybrid system to determine the function of the hypothetical proteins. RESULTS: We examined 960 soluble proteins from Pyrococcus and selected 107 interactions based on luciferase reporter activity, which was then evaluated using a computational approach to assess the reliability of the interactions. We also analyzed the expression of the assay samples by western blot, and a few interactions by in vitro pull-down assays. We identified 11 hetero-interactions that we considered to be located at the same operon, as observed in Helicobacter pylori. We annotated and classified proteins in the selected interactions according to their orthologous proteins. Many enzyme proteins showed self-interactions, similar to those seen in other organisms. CONCLUSION: We found 13 unannotated proteins that interacted with annotated proteins; this information is useful for predicting the functions of the hypothetical Pyrococcus proteins from the annotations of their interacting partners. Among the heterogeneous interactions, proteins were more likely to interact with proteins within the same ortholog class than with proteins of different classes. The analysis described here can provide global insights into the biological features of the protein-protein interactions in P. horikoshii.

Genes, Archaeal↗

Libraries enriched for alternatively spliced exons reveal splicing patterns in melanocytes and melanomas.

It is becoming increasingly clear that alternative splicing enables the complex development and homeostasis of higher organisms. To gain a better understanding of how splicing contributes to regulatory pathways, we have developed an alternative splicing library approach for the identification of alternatively spliced exons and their flanking regions by alternative splicing sequence enriched tags sequencing. Here, we have applied our approach to mouse melan-c melanocyte and B16-F10Y melanoma cell lines, in which 5,401 genes were found to be alternatively spliced. These genes include those encoding important regulatory factors such as cyclin D2, Ilk, MAPK12, MAPK14, RAB4, melastatin 1 and previously unidentified splicing events for 436 genes. Real-time PCR further identified cell line-specific exons for Tmc6, Abi1, Sorbs1, Ndel1 and Snx16. Thus, the ASL approach proved effective in identifying splicing events, which suggest that alternative splicing is important in melanoma development.

Alternative Splicing↗

Absolute expression values for mouse transcripts: re-annotation of the READ expression database by the use of CAGE and EST sequence tags.

The RIKEN expression array database (READ) provides comprehensive gene expression data for the mouse, which were obtained as relative values from microarray double-staining experiments with E17.5 mRNA as common reference. To assign absolute expression values for mouse transcripts within READ, we applied the E17.5 reference sample to CAGE (cap analysis of gene expression) and expressed sequence tag (EST) high-throughput tag sequencing. Newly assigned values within the READ database were validated by comparison to expression data from serial analysis of gene expression, CAGE and EST experiments. These experiments confirmed the great significance of the absolute expression values within the improved READ database. The new Absolute READ database on absolute expression data is available under.

Animals↗

Solution structure of the SEA domain from the murine homologue of ovarian cancer antigen CA125 (MUC16).

Human CA125, encoded by the MUC16 gene, is an ovarian cancer antigen widely used for a serum assay. Its extracellular region consists of tandem repeats of SEA domains. In this study we determined the three-dimensional structure of the SEA domain from the murine MUC16 homologue using multidimensional NMR spectroscopy. The domain forms a unique alpha/beta sandwich fold composed of two alpha helices and four antiparallel beta strands and has a characteristic turn named the TY-turn between alpha1 and alpha2. The internal mobility of the main chain is low throughout the domain. The residues that form the hydrophobic core and the TY-turn are fully conserved in all SEA domain sequences, indicating that the fold is common in the family. Interestingly, no other residues are conserved throughout the family. Thus, the sequence alignment of the SEA domain family was refined on the basis of the three-dimensional structure, which allowed us to classify the SEA domains into several subfamilies. The residues on the surface differ between these subfamilies, suggesting that each subfamily has a different function. In the MUC16 SEA domains, the conserved surface residues, Asn-10, Thr-12, Arg-63, Asp-75, Asp-112, Ser-115, and Phe-117, are clustered on the beta sheet surface, which may be functionally important. The putative epitope (residues 58-77) for anti-MUC16 antibodies is located around the beta2 and beta3 strands. On the other hand the tissue tumor marker MUC1 has a SEA domain belonging to another subfamily, and its GSVVV motif for proteolytic cleavage is located in the short loop connecting beta2 and beta3.

Amino Acid Sequence↗

Solution structure of a BolA-like protein from Mus musculus.

The BolA-like proteins are widely conserved from prokaryotes to eukaryotes. The BolA-like proteins seem to be involved in cell proliferation or cell-cycle regulation, but the molecular function is still unknown. Here we determined the structure of a mouse BolA-like protein. The overall topology is alphabetabetaalphaalphabetaalpha, in which beta(1) and beta(2) are antiparallel, and beta(3) is parallel to beta(2). This fold is similar to the class II KH fold, except for the absence of the GXXG loop, which is well conserved in the KH fold. The conserved residues in the BolA-like proteins are assembled on the one side of the protein.

Amino Acid Sequence↗

EICO (Expression-based Imprint Candidate Organizer): finding disease-related imprinted genes.

We have developed an integrated database that is specialized for the study of imprinted disease genes. The database contains novel candidate imprinted genes identified by the RIKEN full-length mouse cDNA microarray study, information on validated single nucleotide polymorphisms (SNPs) to confirm imprinting using reciprocal mouse crosses and the predicted physical position of imprinting-related disease loci in the mouse and human genomes. It has two user-friendly search interfaces: the SNP-central view (MuSCAT: MoUse SNP CATalog) and the candidate gene-central view (CITE: Candidate Imprinted Transcripts by Expression). The database, EICO (Expression-based Imprint Candidate Organizer), can be accessed via the World Wide Web (http://fantom2.gsc.riken.jp/EICODB/) and the DAS client software. These data and interfaces facilitate understanding of the mechanism of imprinting in mammalian inherited traits.

Animals↗

FREP: a database of functional repeats in mouse cDNAs.

The FREP database (http://facts.gsc.riken.go.jp/FREP/) contains 31 396 RepeatMasker-identified non-redundant variant repeat sequences derived from 16,527 mouse cDNAs with protein-coding potential. The repeats were computationally associated with potential effects on transcriptional variation, translation, protein function or involvement in disease to identify Functional REPeats (FREPs). FREPs are defined by the (i) occurrence of exon-exon boundaries in repeats, (ii) presence of polyadenylation sites in 3'UTR-located repeats, (iii) effect on translation, (iv) position in the protein- coding region or protein domains or (v) conditional association with disease MeSH terms. Currently the database contains 9261 (29.5%) inferred FREPs derived from 6861 (41.5%) mouse cDNAs. Integrated evidence of the functional assignments and dynamically generated sequence similarity search results support the exploration and annotation of functional, ancestral or taxon-specific repeats. Keyword and pre-selected feature searches (e.g. coding sequence-repeat or splice site-repeat relations) support intuitive database querying as well as the retrieval of repeat sequences. Integrated sequence search and alignment tools allow the analysis of known or identification of new functional repeat candidates. FREP is a unique resource for illuminating the role of transposons and repetitive sequences in shaping the coding part of the mouse transcriptome and for selecting the appropriate experimental model to study diseases with suspected repeat etiology contributions.

Animals↗

Identification of unique transcripts from a mouse full-length, subtracted inner ear cDNA library.

A small-scale full-length library construction approach was developed to facilitate production of a mouse full-length cDNA encyclopedia representing approximately 250 enriched, normalized, and/or subtracted cDNA libraries. One library produced using this approach was a subtracted adult mouse inner ear cDNA library (sIEa). The average size of the inserts was approximately 2.5 kb, with the majority ranging from 0.5 to 7.0 kb. From this library 22,574 sequence reads were obtained from 15,958 independent clones. Sequencing and chromosomal localization established 5240 clusters, with 1302 clusters being unique and 359 representing new ESTs. Our sIEa library contributed 56.1% of the 7773 nonredundant Unigene clusters associated with the four mouse inner ear libraries in the NCBI dbEST. Based on homologous chromosomal regions between human and mouse, we identified 1018 UniGene clusters associated with the deafness locus critical regions. Of these, 59 clusters were found only in our sIEa library and represented approximately 50% of the identified critical regions.

Amino Acid Sequence↗

Solution structure of the RWD domain of the mouse GCN2 protein.

GCN2 is the alpha-subunit of the only translation initiation factor (eIF2alpha) kinase that appears in all eukaryotes. Its function requires an interaction with GCN1 via the domain at its N-terminus, which is termed the RWD domain after three major RWD-containing proteins: RING finger-containing proteins, WD-repeat-containing proteins, and yeast DEAD (DEXD)-like helicases. In this study, we determined the solution structure of the mouse GCN2 RWD domain using NMR spectroscopy. The structure forms an alpha + beta sandwich fold consisting of two layers: a four-stranded antiparallel beta-sheet, and three side-by-side alpha-helices, with an alphabetabetabetabetaalphaalpha topology. A characteristic YPXXXP motif, which always occurs in RWD domains, forms a stable loop including three consecutive beta-turns that overlap with each other by two residues (triple beta-turn). As putative binding sites with GCN1, a structure-based alignment allowed the identification of several surface residues in alpha-helix 3 that are characteristic of the GCN2 RWD domains. Despite the apparent absence of sequence similarity, the RWD structure significantly resembles that of ubiquitin-conjugating enzymes (E2s), with most of the structural differences in the region connecting beta-strand 4 and alpha-helix 3. The structural architecture, including the triple beta-turn, is fundamentally common among various RWD domains and E2s, but most of the surface residues on the structure vary. Thus, it appears that the RWD domain is a novel structural domain for protein-binding that plays specific roles in individual RWD-containing proteins.

Amino Acid Sequence↗

Large-scale collection and characterization of promoters of human and mouse genes.

We report the generation and initial characterization of a large-scale collection of sequences of putative promoter regions (PPRs) of human and mouse genes. Based on our unique collection of 400,225 and 580,209 human and mouse full-length cDNAs, we determined exact transcriptional start sites (TSSs). Using positional information of the TSSs, we could retrieve adjacent sequences as PPRs for 8,793 and 6,875 human and mouse genes, respectively. The positions of the PPRs were 4 kb upstream to previously reported 5'-ends of cDNAs on average, demonstrating that full-length cDNA information is indispensable for this purpose. Among those PPRs supported by experimentally validated TSSs, 3,324 could be paired as mutually homologous genes between human and mouse and were used for the comprehensive comparative studies. The sequence identities in the proximal regions of the TSSs were 45% on average, and 22,794 putative transcription factor binding sites that are conserved between human and mouse were identified. The data resource created in the present work and the results of the sequences' initial characterization should lay the firm foundation for deciphering the transcriptional modulations of human genes. All the data were deposited and made available through a database for comparative studies, DBTSS.

Animals↗

Cap analysis gene expression for high-throughput analysis of transcriptional starting point and identification of promoter usage.

We introduce cap analysis gene expression (CAGE), which is based on preparation and sequencing of concatamers of DNA tags deriving from the initial 20 nucleotides from 5' end mRNAs. CAGE allows high-throughout gene expression analysis and the profiling of transcriptional start points (TSP), including promoter usage analysis. By analyzing four libraries (brain, cortex, hippocampus, and cerebellum), we redefined more accurately the TSPs of 11-27% of the analyzed transcriptional units that were hit. The frequency of CAGE tags correlates well with results from other analyses, such as serial analysis of gene expression, and furthermore maps the TSPs more accurately, including in tissue-specific cases. The high-throughput nature of this technology paves the way for understanding gene networks via correlation of promoter usage and gene transcriptional factor expression.

Animals↗

Empirical analysis of transcriptional activity in the Arabidopsis genome.

Functional analysis of a genome requires accurate gene structure information and a complete gene inventory. A dual experimental strategy was used to verify and correct the initial genome sequence annotation of the reference plant Arabidopsis. Sequencing full-length cDNAs and hybridizations using RNA populations from various tissues to a set of high-density oligonucleotide arrays spanning the entire genome allowed the accurate annotation of thousands of gene structures. We identified 5817 novel transcription units, including a substantial amount of antisense gene transcription, and 40 genes within the genetically defined centromeres. This approach resulted in completion of approximately 30% of the Arabidopsis ORFeome as a resource for global functional experimentation of the plant proteome.

Arabidopsis↗

Collection, mapping, and annotation of over 28,000 cDNA clones from japonica rice.

We collected and completely sequenced 28,469 full-length complementary DNA clones from Oryza sativa L. ssp. japonica cv. Nipponbare. Through homology searches of publicly available sequence data, we assigned tentative protein functions to 21,596 clones (75.86%). Mapping of the cDNA clones to genomic DNA revealed that there are 19,000 to 20,500 transcription units in the rice genome. Protein informatics analysis against the InterPro database revealed the existence of proteins presented in rice but not in Arabidopsis. Sixty-four percent of our cDNAs are homologous to Arabidopsis proteins.

Alternative Splicing↗

Targeting a complex transcriptome: the construction of the mouse full-length cDNA encyclopedia.

We report the construction of the mouse full-length cDNA encyclopedia,the most extensive view of a complex transcriptome,on the basis of preparing and sequencing 246 libraries. Before cloning,cDNAs were enriched in full-length by Cap-Trapper,and in most cases,aggressively subtracted/normalized. We have produced 1,442,236 successful 3'-end sequences clustered into 171,144 groups, from which 60,770 clones were fully sequenced cDNAs annotated in the FANTOM-2 annotation. We have also produced 547,149 5' end reads,which clustered into 124,258 groups. Altogether, these cDNAs were further grouped in 70,000 transcriptional units (TU),which represent the best coverage of a transcriptome so far. By monitoring the extent of normalization/subtraction, we define the tentative equivalent coverage (TEC),which was estimated to be equivalent to >12,000,000 ESTs derived from standard libraries. High coverage explains discrepancies between the very large numbers of clusters (and TUs) of this project,which also include non-protein-coding RNAs,and the lower gene number estimation of genome annotations. Altogether,5'-end clusters identify regions that are potential promoters for 8637 known genes and 5'-end clusters suggest the presence of almost 63,000 transcriptional starting points. An estimate of the frequency of polyadenylation signals suggests that at least half of the singletons in the EST set represent real mRNAs. Clones accounting for about half of the predicted TUs await further sequencing. The continued high-discovery rate suggests that the task of transcriptome discovery is not yet complete.

Animals↗

Subtraction of cap-trapped full-length cDNA libraries to select rare transcripts.

The normalization and subtraction of highly expressed cDNAs from relatively large tissues before cloning dramatically enhanced the gene discovery by sequencing for the mouse full-length cDNA encyclopedia, but these methods have not been suitable for limited RNA materials. To normalize and subtract full-length cDNA libraries derived from limited quantities of total RNA, here we report a method to subtract plasmid libraries excised from size-unbiased amplified lambda phage cDNA libraries that avoids heavily biasing steps such as PCR and plasmid library amplification. The proportion of full-length cDNAs and the gene discovery rate are high, and library diversity can be validated by in silico randomization.

Gene Expression Profiling↗

Functional annotation of a full-length Arabidopsis cDNA collection.

Full-length complementary DNAs (cDNAs) are essential for the correct annotation of genomic sequences and for the functional analysis of genes and their products. We isolated 155,144 RIKEN Arabidopsis full-length (RAFL) cDNA clones. The 3'-end expressed sequence tags (ESTs) of 155,144 RAFL cDNAs were clustered into 14,668 nonredundant cDNA groups, about 60% of predicted genes. We also obtained 5' ESTs from 14,034 nonredundant cDNA groups and constructed a promoter database. The sequence database of the RAFL cDNAs is useful for promoter analysis and correct annotation of predicted transcription units and gene products. Furthermore, the full-length cDNAs are useful resources for analyses of the expression profiles, functions, and structures of plant proteins.

Arabidopsis↗