Search PubMed⌕ Search

Biomedical subjects

Susumu Goto

Publications and source records attributed to Susumu Goto.

At least 19 recordsLinked to original sources

EGassembler: online bioinformatics service for large-scale processing, clustering and assembling ESTs and genomic DNA fragments.

Expressed sequence tag (EST) sequencing has proven to be an economically feasible alternative for gene discovery in species lacking a draft genome sequence. Ongoing large-scale EST sequencing projects feel the need for bioinformatics tools to facilitate uniform EST handling. This brings about a renewed importance for a universal tool for processing and functional annotation of large sets of ESTs. EGassembler (http://egassembler.hgc.jp/) is a web server, which provides an automated as well as a user-customized analysis tool for cleaning, repeat masking, vector trimming, organelle masking, clustering and assembling of ESTs and genomic fragments. The web server is publicly available and provides the community a unique all-in-one online application web service for large-scale ESTs and genomic DNA clustering and assembling. Running on a Sun Fire 15K supercomputer, a significantly large volume of data can be processed in a short period of time. The results can be used to functionally annotate genes, to facilitate splice alignment analysis, to link the transcripts to genetic and physical maps, design microarray chips, to perform transcriptome analysis and to map to KEGG metabolic pathways. The service provides an excellent bioinformatics tool to research groups in wet-lab as well as an all-in-one-tool for sequence handling to bioinformatics researchers.

Computational Biology↗

Extraction of phylogenetic network modules from the metabolic network.

BACKGROUND: In bio-systems, genes, proteins and compounds are related to each other, thus forming complex networks. Although each organism has its individual network, some organisms contain common sub-networks based on function. Given a certain sub-network, the distribution of organisms common to it represents the diversity of its function. RESULTS: We extracted such "common" sub-networks, defined as "phylogenetic network modules," using phylogenetic profiles and cluster analysis. The enzymes in the same "phylogenetic network module" have similar phylogenetic profiles and related functions. These modules are shown to be phylogenetic building blocks. Furthermore, the network of the modules illustrated hierarchical feature as well as the network of enzymes involved in the metabolism. CONCLUSION: We conclude that phylogenetic network modules are evolutionary conserved functional units in the metabolic network. We claim that our concept of phylogenetic modules provides a more accurate understanding of the evolution of biological networks.

Animals↗

Genome sequence of the cat pathogen, Chlamydophila felis.

Chlamydophila felis (Chlamydia psittaci feline pneumonitis agent) is a worldwide spread pathogen for pneumonia and conjunctivitis in cats. Herein, we determined the entire genomic DNA sequence of the Japanese C. felis strain Fe/C-56 to understand the mechanism of diseases caused by this pathogen. The C. felis genome is composed of a circular 1,166,239 bp chromosome encoding 1005 protein-coding genes and a 7552 bp circular plasmid. Comparison of C. felis gene contents with other Chlamydia species shows that 795 genes are common in the family Chlamydiaceae species and 47 genes are specific to C. felis. Phylogenetic analysis of the common genes reveals that most of the orthologue sets exhibit a similar divergent pattern but 14 C. felis genes accumulate more mutations, implicating that these genes may be involved in the evolutional adaptation to the C. felis-specific niche. Gene distribution and orthologue analyses reveal that two distinctive regions, i.e. the plasticity zone and frequently gene-translocated regions (FGRs), may play important but different roles for chlamydial genome evolution. The genomic DNA sequence of C. felis provides information for comprehension of diseases and elucidation of the chlamydial evolution.

Adaptation, Biological↗

ODB: a database of operons accumulating known operons across multiple genomes.

Operon structures play an important role in co-regulation in prokaryotes. Although over 200 complete genome sequences are now available, databases providing genome-wide operon information have been limited to certain specific genomes. Thus, we have developed an ODB (Operon DataBase), which provides a data retrieval system of known operons among the many complete genomes. Additionally, putative operons that are conserved in terms of known operons are also provided. The current version of our database contains about 2000 known operon information in more than 50 genomes and about 13 000 putative operons in more than 200 genomes. This system integrates four types of associations: genome context, gene co-expression obtained from microarray data, functional links in biological pathways and the conservation of gene order across the genomes. These associations are indicators of the genes that organize an operon, and the combination of these indicators allows us to predict more reliable operons. Furthermore, our system validates these predictions using known operon information obtained from the literature. This database integrates known literature-based information and genomic data. In addition, it provides an operon prediction tool, which make the system useful for both bioinformatics researchers and experimental biologists. Our database is accessible at http://odb.kuicr.kyoto-u.ac.jp/.

Databases, Nucleic Acid↗

From genomics to chemical genomics: new developments in KEGG.

The increasing amount of genomic and molecular information is the basis for understanding higher-order biological systems, such as the cell and the organism, and their interactions with the environment, as well as for medical, industrial and other practical applications. The KEGG resource (http://www.genome.jp/kegg/) provides a reference knowledge base for linking genomes to biological systems, categorized as building blocks in the genomic space (KEGG GENES) and the chemical space (KEGG LIGAND), and wiring diagrams of interaction networks and reaction networks (KEGG PATHWAY). A fourth component, KEGG BRITE, has been formally added to the KEGG suite of databases. This reflects our attempt to computerize functional interpretations as part of the pathway reconstruction process based on the hierarchically structured knowledge about the genomic, chemical and network spaces. In accordance with the new chemical genomics initiatives, the scope of KEGG LIGAND has been significantly expanded to cover both endogenous and exogenous molecules. Specifically, RPAIR contains curated chemical structure transformation patterns extracted from known enzymatic reactions, which would enable analysis of genome-environment interactions, such as the prediction of new reactions and new enzyme genes that would degrade new environmental compounds. Additionally, drug information is now stored separately and linked to new KEGG DRUG structure maps.

Biotransformation↗

Effects of post-electrophoretic analysis on variance in gel-based proteomics.

2D electrophoresis (2DE) is a prominent separation method for complex proteomes. Although recent advances have increased the utility of this method in quantitative proteomics studies, many sources of variance still exist. This review discusses the post-electrophoretic sources of variance in current 2DE analysis. The essential improvements in protein visualization and software algorithms that have made 2DE a leading quantitative proteomics method are briefly reviewed. A number of shortcomings in the post-electrophoretic analysis of 2DE data that require further attention are highlighted. Topics discussed include protein visualization and image acquisition, internal standards and normalization methods, background subtraction algorithms, normality of distribution, and the need for standardized tests for the evaluation of 2DE analysis software packages.

Electrophoresis, Gel, Two-Dimensional↗

Extraction of leukemia specific glycan motifs in humans by computational glycomics.

There have been almost no standard methods for conducting computational analyses on glycan structures in comparison to DNA and proteins. In this paper, we present a novel method for extracting functional motifs from glycan structures using the KEGG/GLYCAN database. First, we developed a new similarity measure for comparing glycan structures taking into account the characteristic mechanisms of glycan biosynthesis, and we tested its ability to classify glycans of different blood components in the framework of support vector machines (SVMs). The results show that our method can successfully classify glycans from four types of human blood components: leukemic cells, erythrocyte, serum, and plasma. Next, we extracted characteristic functional motifs of glycans considered to be specific to each blood component. We predicted the substructure alpha-D-Neup5Ac-(2-->3)-beta-D-Galp-(1-->4)-D-GlcpNAc as a leukemia specific glycan motif. Based on the fact that the Agrocybe cylindracea galectin (ACG) specifically binds to the same substructure, we conducted an experiment using cell agglutination assay and confirmed that this fungal lectin specifically recognized human leukemic cells.

Biomarkers, Tumor↗

MRP1 mutated in the L0 region transports SN-38 but not leukotriene C4 or estradiol-17 (beta-D-glucuronate).

Multidrug resistance protein 1 (MRP1) is an ATP-binding cassette transporter that confers multidrug resistance on tumor cells. Much convincing evidence has accumulated that MRP1 transports most substances in a GSH-dependent manner. On the other hand, several reports have revealed that MRP1 can transport some substrates independently of GSH; however, the importance of GSH-independent transport activity is not well established and the mechanistic differences between GSH-dependent and -independent transport by MRP1 are unclear. We previously demonstrated that the amino acids W261 and K267 in the L0 region of MRP1 were important for leukotriene C4 (LTC4) transport activity of MRP1 and for GSH-dependent photolabeling of MRP1 with azidophenyl agosterol-A (azidoAG-A). In this paper, we further tested the effect of W222L, W223L and R230A mutations in MRP1, designated dmL0MRP1, on MRP1 transport activity. SN-38 is an active metabolic form of CPT-11 that is one of the most promising anti-cancer drugs. Membrane vesicles prepared from cells expressing dmL0MRP1 could transport SN-38, but not LTC4 or estradiol-17 (beta-D-glucuronate), and could not be photolabeled with azidoAG-A. These data suggested that SN-38 was transported by a different mechanism than that of GSH-dependent transport. Understanding the GSH-independent transport mechanism of MRP1, and identification of drugs that are transported by this mechanism, will be critical for combating MRP1-mediated drug resistance. We performed a pairwise comparison of compounds that are transported by MRP1 in a GSH-dependent or -independent manner. These data indicated that it may be possible to predict compounds that are transported by MRP1 in a GSH-independent manner.

Base Sequence↗

Prediction of glycan structures from gene expression data based on glycosyltransferase reactions.

MOTIVATION: Glycan chains are synthesized by a combination of several kinds of glycosyltransferases (GTs). Thus, once we know the repertoire of GTs in the genome, in the transcriptome or in the proteome, it should in principle be possible to predict the repertoire of possible glycan structures in an organism or at a specific stage of the cell. Here, we show that a repertoire of glycan structures can be predicted from the set of GTs in the transcriptome. That is, using knowledge about glycan structure characteristics, we can predict glycan structures from incomplete or noisy data such as DNA microarray data. RESULTS: First, we constructed a reaction pattern library consisting of bond-formation patterns of GT reactions and investigated the co-occurrence frequencies of all reaction patterns in the glycan database. This was followed by the prediction of glycan structures using this library and a co-occurrence score. A penalty score was also implemented in the prediction method. Then we examined the performance of prediction by the leave-one-out cross validation method using individual reaction pattern profiles in the KEGG GLYCAN database as virtual expression profiles. The accuracy of prediction was 81%. Finally, we applied the prediction method to real expression data. Using expression profiles from the human carcinoma cell, glycan structures with sialic acid and sialyl Lewis X epitope were predicted, which corresponded well with experimental results.

Algorithms↗

KEGG as a glycome informatics resource.

Bioinformatics approaches to carbohydrate research have recently begun using large amounts of protein and carbohydrate data. In this field called glycome informatics, the foremost necessity is a comprehensive resource for genome-scale bioinformatics analysis of glycan data. Although the accumulation of experimental data may be useful as a reference of biological and biochemical information on carbohydrates, this is insufficient for bioinformatics analysis. Thus, we have developed a glycome informatics resource (http://www.genome.jp/kegg/glycan/) in KEGG (Kyoto Encyclopedia of Genes and Genomes), an integrated knowledge base of protein networks, genomic information, and chemical information. This review describes three noteworthy features: (1) GLYCAN, a database of carbohydrate structures; (2) glycan-related pathways; and (3) Composite Structure Map (CSM), a map illustrating all possible variations of carbohydrate structures within organisms. GLYCAN includes two useful tools: an intuitive drawing tool called KegDraw, and an efficient glycan search and alignment tool called KEGG Carbohydrate Matcher (KCaM). KEGG's glycan biosynthesis and metabolism pathways, integrating carbohydrate structures, proteins, and reactions, are also a pivotal resource. CSM is constructed as a bridge between carbohydrate functions and structures. CSM is able to display, for example, expression data of glycosyltransferases in a compact manner. In all the KEGG resources, various objects including KEGG pathways, chemical compounds, as well as carbohydrate structures are commonly represented as graphs, which are widely studied and utilized in the computer science field.

Carbohydrates↗

Conservation of gene co-regulation between two prokaryotes: Bacillus subtilis and Escherichia coli.

We measured conservation of gene co-regulation between two distantly related prokaryotes, B. subtilis and E. coli. The co-regulation between genes was extracted from knowledge of regulation of genes stored in databases. For B. subtilis operons, we obtained the data set from ODB which we have developed and, for the regulons, we used DBTBS. For E. coli data set, we used known regulons derived from RegulonDB. We obtained a reliable data set of co-regulated genes in B. subtilis and E. coli. About 60-80 % of gene pairs conserved co-regulation relationships, so co-regulation between genes are highly conserved even between distantly related species. To measure the functional relationship between these conserved genes, we used KEGG PATHWAY and COG. When two co-regulated genes are in the same biological pathway in KEGG or share the same functional category in COG, we assume that they have the same function. As a result, we also found that many conserved co-regulated gene pairs share the same functions. These observations would help to predict gene co-regulation and protein functions.

Bacillus subtilis↗

Comprehensive analysis and prediction of synthetic lethality using subcellular locations.

The lethality of a gene is a fundamental and representative measure for understanding the function of a gene and its associated bio-systems. Recently, many research groups have started focusing on the concept of synthetic lethality. The synthetic lethality between genes is defined by the combination of mutations in two genes causing cell death. Here, we confirm that synthetic lethality and cellular location have close relationships among the Saccharomyces cerevisiae genes. Furthermore, we attempt the prediction of candidate gene pairs with synthetic lethality. The prediction is based on the hierarchical aspect model (HAM) which learns from a data set of cellular location to estimate a likelihood value indicating the synthetic lethality between genes.

Cell Death↗

A global representation of the carbohydrate structures: a tool for the analysis of glycan.

Glycan resources have been developed of late, such as carbohydrate databases, analysis tools, and algorithms for analysis of carbohydrate features. With this background, bioinformatics approaches to carbohydrate research have recently begun using a large amount of protein and carbohydrate data. This paper introduces one of these projects that elucidates the range of carbohydrate structures. In this study, the variety of carbohydrate structures have been enumerated in a global tree structure called variation trees, using the KEGG GLYCAN database, which is a public-domain glycan resource for bioinformatics analysis. Additionally, a glycosyltransferase mapping list of glycosyltransferases and their catalyzing glycosidic linkages was constructed. From this, we present the composite structure map (CSM), which is a structural variation map integrating its variation trees and glycosyltransferase map list. CSM is able to display, for example, expression data of glycosyltransferases in a compact manner, illustrating its versatility as a new bioinformatics resource and tool capable of analyzing carbohydrate structures on a global scale. These resources are available at http://www.genome.jp/kegg/glycan/.

Carbohydrate Conformation↗

Computational assignment of the EC numbers for genomic-scale analysis of enzymatic reactions.

The EC (Enzyme Commission) numbers represent a hierarchical classification of enzymatic reactions, but they are also commonly utilized as identifiers of enzymes or enzyme genes in the analysis of complete genomes. This duality of the EC numbers makes it possible to link the genomic repertoire of enzyme genes to the chemical repertoire of metabolic pathways, the process called metabolic reconstruction. Unfortunately, there are numerous reactions known to be present in various pathways, but they will never get EC numbers because the EC number assignment requires published articles on full characterization of enzymes. Here we report a computerized method to automatically assign the EC numbers up to the sub-subclasses, i.e., without the fourth serial number for substrate specificity, given pairs of substrates and products. The method is based on a new classification scheme of enzymatic reactions, named the RC (reaction classification) number. Each reaction in the current dataset of the EC numbers is first decomposed into reactant pairs. Each pair is then structurally aligned to identify the reaction center, the matched region, and the difference region. The RC number represents the conversion patterns of atom types in these three regions. We examined the correspondence between computationally assigned RC numbers and manually assigned EC numbers by the jackknife cross-validation test and found that the EC sub-subclasses could be assigned with the accuracy of about 90%. Furthermore, we examined the correlation with genomic information as represented by the KEGG ortholog clusters (OC) and confirmed that the RC numbers are correlated not only with elementary reaction mechanisms but also with protein families.

Computers↗

Fast and accurate database homology search using upper bounds of local alignment scores.

MOTIVATION: It is widely recognized that homology search and ortholog clustering are very useful for analyzing biological sequences. However, recent growth of sequence database size makes homolog detection difficult, and rapid and accurate methods are required. RESULTS: We present a novel method for fast and accurate homology detection, assuming that the Smith-Waterman (SW) scores between all similar sequence pairs in a target database are computed and stored. In this method, SW alignment is computed only if the upper bound, which is derived from our novel inequality, is higher than the given threshold. In contrast to other methods such as FASTA and BLAST, this method is guaranteed to find all sequences whose scores against the query are higher than the specified threshold. Results of computational experiments suggest that the method is dozens of times faster than SSEARCH if genome sequence data of closely related species are available.

Algorithms↗

KCaM (KEGG Carbohydrate Matcher): a software tool for analyzing the structures of carbohydrate sugar chains.

KCaM (KEGG Carbohydrate Matcher) is a tool for the analysis of carbohydrate sugar chains, or glycans. It consists of a web-based graphical user interface that allows users to enter glycans easily with the mouse. The glycan structure is then transformed into our KCF (KEGG Chemical Function) file format and sent to our program which implements an efficient tree-structure alignment algorithm, similar to sequence alignment algorithms but for branched tree structures. Users can also retrieve glycan tree structures in KCF format from their local computers for visualization over the web. The tree-matching algorithm provides several options for performing different types of tree-matching procedures on glycans. These options consist of whether to incorporate gaps in a match, whether to take the linkage information into consideration and local versus global alignment. The results of this program are returned as a list of glycan structures in order of similarity based on these options. The actual alignment can be viewed graphically, and the annotation information can also be viewed easily since all this information is linked with KEGG's comprehensive suite of genomic data. Analogously to BLAST, users are thus able to compare glycan structures of interest with glycans from different glycan databases using a variety of tree-alignment options. KCaM is currently available at http://glycan.genome.ad.jp.

Algorithms↗

The KEGG resource for deciphering the genome.

A grand challenge in the post-genomic era is a complete computer representation of the cell and the organism, which will enable computational prediction of higher-level complexity of cellular processes and organism behavior from genomic information. Toward this end we have been developing a knowledge-based approach for network prediction, which is to predict, given a complete set of genes in the genome, the protein interaction networks that are responsible for various cellular processes. KEGG at http://www.genome.ad.jp/kegg/ is the reference knowledge base that integrates current knowledge on molecular interaction networks such as pathways and complexes (PATHWAY database), information about genes and proteins generated by genome projects (GENES/SSDB/KO databases) and information about biochemical compounds and reactions (COMPOUND/GLYCAN/REACTION databases). These three types of database actually represent three graph objects, called the protein network, the gene universe and the chemical universe. New efforts are being made to abstract knowledge, both computationally and manually, about ortholog clusters in the KO (KEGG Orthology) database, and to collect and analyze carbohydrate structures in the GLYCAN database.

Animals↗

Extraction of phylogenetic network modules from prokayrote metabolic pathways.

In the post-genomic era, it is important to analyze interaction networks that include genes, proteins, enzymes and compounds such as a metabolic pathway. Every organism has such networks individually. However, several parts of them are conserved in different organisms. The purpose of this analysis is to extract sub-networks composed of these common elements through the phylogenetic analysis. We extracted network modules from metabolic pathways using phylogenetic profile and cluster analysis. The enzymes of these modules are related by evolutionary and functional correlation. Our results give a valuable insight into the evolution of metabolic pathways.

Bacteria↗