Search PubMed⌕ Search

Biomedical subjects

Sumio Sugano

Publications and source records attributed to Sumio Sugano.

At least 55 records · Page 3Linked to original sources

DBTSS, DataBase of Transcriptional Start Sites: progress report 2004.

DBTSS (http://dbtss.hgc.jp) was originally constructed based on a collection of experimentally determined TSSs of human genes. Since its first release in 2002, it has been updated several times. First, the amount of stored data has increased significantly: e.g. the number of clones that match both the RefSeq mRNA set and the genome sequence has increased from 111,382 to 190,964, now covering 1,234 genes. Second, the positions of SNPs in dbSNP were displayed on the upstream regions of contained human genes. Third, DBTSS now covers other species such as mouse and the human malaria parasite. It will become a central database containing data for many more species with oligo-capping and related methods. Lastly, the database now serves for comparative promoter analyses: in the current version, comparative views of potentially orthologous promoters from human and mouse are presented with an additional function of searching potential transcription-factor binding sites, which are either conserved or diverged between species.

Animals↗

Full-malaria 2004: an enlarged database for comparative studies of full-length cDNAs of malaria parasites, Plasmodium species.

Full-malaria (http://fullmal.ims.u-tokyo.ac.jp), a database for full-length cDNAs from the human malaria parasite, Plasmodium falciparum has been updated in at least three points. (i) We added 8934 sequences generated from the addition of new libraries, so that our collection of 11,424 full-length cDNAs covers 1375 (25%) of the estimated number of the entire 5409 parasite genes. (ii) All of our full-length cDNAs and GenBank EST sequences were mapped to genomic sequences together with publicly available annotated genes and other predictions. This precisely determined the gene structures and positions of the transcriptional start sites, which are indispensable for the identification of the promoter regions. (iii) A total of 4257 cDNA sequences were newly generated from murine malaria parasites, Plasmodium yoelii yoelii. The genome/cDNA sequences were compared at both nucleotide and amino acid levels, with those of P.falciparum, and the sequence alignment for each gene is presented graphically. This part of the database serves as a versatile platform to elucidate the function(s) of malaria genes by a comparative genomic approach. It should also be noted that all of the cDNAs represented in this database are supported by physical cDNA clones, which are publicly and freely available, and should serve as indispensable resources to explore functional analyses of malaria genomes.

Animals↗

Promoter prediction analysis on the whole human genome.

Promoter prediction programs (PPPs) are important for in silico gene discovery without support from expressed sequence tag (EST)/cDNA/mRNA sequences, in the analysis of gene regulation and in genome annotation. Contrary to previous expectations, a comprehensive analysis of PPPs reveals that no program simultaneously achieves sensitivity and a positive predictive value >65%. PPP performances deduced from a limited number of chromosomes or smaller data sets do not hold when evaluated at the level of the whole genome, with serious inaccuracy of predictions for non-CpG-island-related promoters. Some PPPs even perform worse than, or close to, pure random guessing.

Algorithms↗

Analysis of small human proteins reveals the translation of upstream open reading frames of mRNAs.

To find novel short coding sequences from accumulated full-length cDNA sequences, proteomic analysis of small proteins expressed in human leukemia K562 cells was performed using high-resolution nanoflow liquid chromatography coupled with electrospray ionization tandem mass spectrometry. Our analysis led to the identification of 54 proteins not more than 100 amino acids in length, including four novel ones. These novel short coding sequences were all located upstream of the longest open reading frame (ORF) of the corresponding cDNA. Our findings indicate that the translation of short ORFs occurs in vivo whether or not there exists a longer coding region in the downstream of the mRNA. This investigation provides the first direct evidence of translation of upstream ORFs in human cells, which could greatly change the current outline of the human proteome.

Amino Acid Sequence↗

Sequence comparison of human and mouse genes reveals a homologous block structure in the promoter regions.

Comparative sequence analysis was carried out for the regions adjacent to experimentally validated transcriptional start sites (TSSs), using 3324 pairs of human and mouse genes. We aligned the upstream putative promoter sequences over the 1-kb proximal regions and found that the sequence conservation could not be further extended at, on average, 510 bp upstream positions of the TSSs. This discontinuous manner of the sequence conservation revealed a "block" structure in about one-third of the putative promoter regions. Consistently, we also observed that G+C content and CpG frequency were significantly different inside and outside the blocks. Within the blocks, the sequence identity was uniformly 65% regardless of their length. About 90% of the previously characterized transcription factor binding sites were located within those blocks. In 46% of the blocks, the 5' ends were bounded by interspersed repetitive elements, some of which may have nucleated the genomic rearrangements. The length of the blocks was shortest in the promoters of genes encoding transcription factors and of genes whose expression patterns are brain specific, which suggests that the evolutional diversifications in the transcriptional modulations should be the most marked in these populations of genes.

Animals↗

The status, quality, and expansion of the NIH full-length cDNA project: the Mammalian Gene Collection (MGC).

The National Institutes of Health's Mammalian Gene Collection (MGC) project was designed to generate and sequence a publicly accessible cDNA resource containing a complete open reading frame (ORF) for every human and mouse gene. The project initially used a random strategy to select clones from a large number of cDNA libraries from diverse tissues. Candidate clones were chosen based on 5'-EST sequences, and then fully sequenced to high accuracy and analyzed by algorithms developed for this project. Currently, more than 11,000 human and 10,000 mouse genes are represented in MGC by at least one clone with a full ORF. The random selection approach is now reaching a saturation point, and a transition to protocols targeted at the missing transcripts is now required to complete the mouse and human collections. Comparison of the sequence of the MGC clones to reference genome sequences reveals that most cDNA clones are of very high sequence quality, although it is likely that some cDNAs may carry missense variants as a consequence of experimental artifact, such as PCR, cloning, or reverse transcriptase errors. Recently, a rat cDNA component was added to the project, and ongoing frog (Xenopus) and zebrafish (Danio) cDNA projects were expanded to take advantage of the high-throughput MGC pipeline.

Animals↗

Large-scale collection and characterization of promoters of human and mouse genes.

We report the generation and initial characterization of a large-scale collection of sequences of putative promoter regions (PPRs) of human and mouse genes. Based on our unique collection of 400,225 and 580,209 human and mouse full-length cDNAs, we determined exact transcriptional start sites (TSSs). Using positional information of the TSSs, we could retrieve adjacent sequences as PPRs for 8,793 and 6,875 human and mouse genes, respectively. The positions of the PPRs were 4 kb upstream to previously reported 5'-ends of cDNAs on average, demonstrating that full-length cDNA information is indispensable for this purpose. Among those PPRs supported by experimentally validated TSSs, 3,324 could be paired as mutually homologous genes between human and mouse and were used for the comprehensive comparative studies. The sequence identities in the proximal regions of the TSSs were 45% on average, and 22,794 putative transcription factor binding sites that are conserved between human and mouse were identified. The data resource created in the present work and the results of the sequences' initial characterization should lay the firm foundation for deciphering the transcriptional modulations of human genes. All the data were deposited and made available through a database for comparative studies, DBTSS.

Animals↗

Fugu ESTs: new resources for transcription analysis and genome annotation.

The draft Fugu rubripes genome was released in 2002, at which time relatively few cDNAs were available to aid in the annotation of genes. The data presented here describe the sequencing and analysis of 24,398 expressed sequence tags (ESTs) generated from 15 different adult and juvenile Fugu tissues, 74% of which matched protein database entries. Analysis of the EST data compared with the Fugu genome data predicts that approximately 10,116 gene tags have been generated, covering almost one-third of Fugu predicted genes. This represents a remarkable economy of effort. Comparison with the Washington University zebrafish EST assemblies indicates strong conservation within fish species, but significant differences remain. This potentially represents divergence of sequence in the 5' terminal exons and UTRs between these two fish species, although clearly, complete EST data sets are not available for either species. This project provides new Fugu resources, and the analysis adds significant weight to the argument that EST programs remain an essential resource for genome exploitation and annotation. This is particularly timely with the increasing availability of draft genome sequence from different organisms and the mounting emphasis on gene function and regulation.

Animals↗

Inhibition of angiogenesis and vascular leakiness by angiopoietin-related protein 4.

Angiopoietins and angiopoietin-related proteins (ARPs) have been shown to regulate angiogenesis, a process essential for various neovascular diseases including tumors. Here, we identify ARP4/fasting-induced adipose factor/peroxisome proliferator-activated receptor gamma angiopoietin-related as a novel antiangiogenic modulatory factor. We hypothesized that ARP4 may regulate angiogenesis. In vitro experiments using purified recombinant ARP4 protein revealed that ARP4 markedly inhibited the proliferation, chemotaxis, and tubule formation of endothelial cells. Moreover, using corneal neovascularization and Miles permeability assays, we found that both vascular endothelial growth factor-induced in vivo angiogenesis and vascular leakiness were significantly inhibited by the addition of ARP4. Finally, we found remarkable suppression of tumor growth within the dermal layer associated with decreased numbers of invading blood vessels in transgenic mice that express ARP4 in the skin driven by the keratinocyte promoter. These findings demonstrate that ARP4 functions as a novel antiangiogenic modulatory factor and indicate a potential therapeutic effect of ARP4 in neoplastic diseases.

Angiogenesis Inhibitors↗

Expression profiling and characterization of 4200 genes cloned from primary neuroblastomas: identification of 305 genes differentially expressed between favorable and unfavorable subsets.

Neuroblastoma (NBL), one of the most common childhood solid tumors, has a distinct nature in different prognostic subgroups: NBL in patients under 1 year of age usually regresses spontaneously, whereas that in patients over 1 year of age often grows aggressively and eventually kills the patient. To understand the molecular mechanism of biology and tumorigenesis of NBL, we decided to perform a comprehensive approach to unveil the gene expression profiles among the NBL subsets. We constructed the subset-specific oligo-capping cDNA libraries from the primary NBL tissues with favorable (F: stage 1, high expression of TrkA and a single copy of MYCN) and unfavorable (UF: stage 3 or 4, decreased expression of TrkA and MYCN amplification) characteristics and randomly cloned 4654 cDNAs. Among 4243 cDNAs sequenced successfully, 1799 (42.4%) were the genes with unknown function. Excluding the housekeeping genes, an expression profile of each subset was extremely different. To determine the genes expressed differentially between F and UF subsets, we performed semiquantitative reverse transcriptase (RT)-PCR for each of the 1842 independent genes using RNA obtained from 16 F and 16 UF NBLs as template. This revealed that 278 genes were highly expressed in the F subset as compared to the UF one, while, surprisingly, only 27 genes were expressed at higher levels in the UF rather than the F subset. These differentially expressed genes included 194 genes with unknown function. Many of the genes expressed at high levels in the F subset were related to catecholamine biosynthesis, small GTPases, synapse formation, synaptic vesicle transport, and transcription factors regulating differentiation of the neural crest-derived cells. On the other hand, the genes expressed at high levels in the UF subset included transcription factors and/or receptors that might regulate neuronal growth and differentiation. The chromosomal mapping of those genes showed some clusters. Thus, our mass-identification and characterization of the differentially expressed genes between the subsets may become a powerful tool for finding the important genes of NBL as well as developing new diagnostic and therapeutic strategies against aggressive NBL.

Chromosome Mapping↗

Neuroblastoma oligo-capping cDNA project: toward the understanding of the genesis and biology of neuroblastoma.

Neuroblastoma (NBL) is a common pediatric cancer originated from the neuronal precursor cells of sympathoadrenal lineage. NBLs show a variety of clinical phenotypes from spontaneous regression to malignant progression with acquirement of resistance to therapy. To understand the molecular mechanism of the genesis, progression, and regression of NBL, we need to identify key molecules determining the neuronal development of sympathoadrenal lineage. To this end, we have performed the NBL cDNA project. It includes (1) mass-cloning of the expressed genes from oligo-capping cDNA libraries derived from primary NBLs with different clinical and biological features; (2) mass-identification of differentially expressed genes between favorable and unfavorable subsets; and (3) molecular and functional analyses of the novel genes, which could be useful prognostic indicators. To date, 10,000 cDNA clones in total, approximately 40% of which contained novel sequences, were randomly picked up and DNA sequenced. We have identified approximately 500 differentially expressed genes between favorable and unfavorable subsets of NBL, among which more than 250 were the genes with unknown function.

DNA, Complementary↗

Collection, mapping, and annotation of over 28,000 cDNA clones from japonica rice.

We collected and completely sequenced 28,469 full-length complementary DNA clones from Oryza sativa L. ssp. japonica cv. Nipponbare. Through homology searches of publicly available sequence data, we assigned tentative protein functions to 21,596 clones (75.86%). Mapping of the cDNA clones to genomic DNA revealed that there are 19,000 to 20,500 transcription units in the rice genome. Protein informatics analysis against the InterPro database revealed the existence of proteins presented in rice but not in Arabidopsis. Sixty-four percent of our cDNAs are homologous to Arabidopsis proteins.

Alternative Splicing↗

Isolation and functional analysis of the melanoma specific promoter region of human GD3 synthase gene.

Human GD3 synthase gene consisted of five exons and span about 135 kilobases. The 5'-flanking region lacked canonical TATA and CAAT boxes, but contained SP1 binding site(s) as in rat and mouse. The promoter activity in the 5'-flanking region (-2262 approximately +1) became definite when SV40 enhancer was added to the reporter plasmid. Luciferase assay with deletion mutants suggested the existence of a silencer region between -2262 and -978 nt similarly with those in mouse and rat. They also commonly contained a GT/CG repeat sequence at upstream of -1200 approximately -1300 nt, suggesting that they form Z-type DNA, and are involved in the gene regulation.

5' Flanking Region↗

The 5' terminal oligopyrimidine tract of human elongation factor 1A-1 gene functions as a transcriptional initiator and produces a variable number of Us at the transcriptional level.

The human elongation factor 1A-1 (eEF1A-1) gene is a member of the 5' terminal oligopyrimidine tract (5' TOP) gene family, and the number of thymidines (Ts) at the 5' TOP of cDNAs corresponding to this gene is known to show variation. Here we determined the 5'-end sequences of 125 eEF1A-1 clones and the complete sequences of 19 eEF1A-1 clones from an oligo-capped cDNA library and showed that variation in the number of Ts is generated by an in vivo process, not by an in vitro artifact during the construction of the cDNA library. Moreover, using green fluorescent protein transgenic mice, we demonstrated that the variation in T number is probably generated during or after transcription. We also introduced various mutations in the mRNA start site of this gene, particularly in the T stretch at the 5' TOP, and examined the effects on the promoter activity. The results showed that at least three Ts must exist at the 5' TOP for the high transcriptional activity of the eEF1A-1 gene promoter. Many other housekeeping genes, including ribosomal protein genes, are also members of the 5' TOP gene family, and the 5' TOP sequence may be an important core-promoter element of these genes.

5' Untranslated Regions↗

Gene expression profiling reveals the mechanism and pathophysiology of mouse liver regeneration.

Comprehensive analysis of the changes in gene expression during liver regeneration was carried out by using an in-house microarray composed of 2,304 distinct mouse liver cDNA clones. Mice were subjected to partial two-thirds hepatectomy, and changes in mRNA levels were monitored up to 48 h. Of the 2,304 genes analyzed, 496 genes showed expression levels measurable at all time points after the partial hepatectomy. 317 genes were up- or down-regulated 2-fold or more at least at one time point during liver regeneration and were classified into eight clusters based on their expression patterns. With a more stringent cut-off value of +/-2 S.D., 68 genes were listed and were classified into five clusters. In these two analyses with different clustering criteria, functionally categorized genes showed similar cluster distributions. Genes involved in protein synthesis and posttranslational processing were significantly enriched in the cluster characterized by rapid gene activation and subsequent persistence. This suggests the importance of modulating the efficiency of protein supply and/or altering the composition of protein population from the early phase of hepatocyte proliferation. Genes for two major liver functions, i.e. plasma protein secretion and intermediate metabolism were enriched in distinct clusters exhibiting the features of gradual gene activation and sustained repression, respectively. Therefore, these genes are differentially regulated during the regeneration, possibly leading to changes in the flow of amino acids and energy from enzyme proteins to plasma proteins in their synthesis. Thus, clustering analysis of expression patterns of functionally classified genes gave insights into mechanism and pathophysiology of liver regeneration.

Animals↗

Large-scale identification and characterization of human genes that activate NF-kappaB and MAPK signaling pathways.

We have carried out a large-scale identification and characterization of human genes that activate the NF-kappaB and MARK signaling pathways. We constructed full-length cDNA libraries using the oligo-capping method and prepared an arrayed cDNA pool consisting of 150 000 cDNAs randomly isolated from the libraries. For analysis of the NF-kappaB signaling pathway, we introduced each of the cDNAs into human embryonic kidney 293 cells and examined whether it activated the transcription of a luciferase reporter gene driven by a promoter containing the consensus NF-kappaB binding sites. In total, we identified 299 cDNAs that activate the NF-kappaB pathway, and we classified them into 83 genes, including 30 characterized activator genes of the NF-kappaB pathway, 28 genes whose involvement in the NF-kappaB pathways have not been characterized and 25 novel genes. We then carried out a similar analysis for the identification of genes that activate the MARK pathway, utilizing the same cDNA resource. We assayed 145 000 cDNAs and identified 57 genes that activate the MARK pathway. Interestingly, 27 genes were overlapping between the NF-kappaB and the MAPK pathways, which may indicate that these genes play cross-talking roles between these two pathways.

Amino Acid Sequence↗