Search PubMed⌕ Search

Biomedical subjects

Minoru Kanehisa

Publications and source records attributed to Minoru Kanehisa.

18 recordsLinked to original sources

The evolutionary repertoires of the eukaryotic-type ABC transporters in terms of the phylogeny of ATP-binding domains in eukaryotes and prokaryotes.

ABC (ATP-binding cassette) transporters play an important role in the communication of various substrates across cell membranes. They are ubiquitous in prokaryotes and eukaryotes, and eukaryotic types (EK-types) are distinguished from prokaryotic types (PK-types) in terms of their genes and domain organizations. The EK-types and PK-types mainly consist of exporters and importers, respectively. Prokaryotes have both the EK-types and the PK-types. The EK-types in prokaryotes are usually called "bacterial multidrug ABC transporters," but they are not well characterized in comparison with the multidrug ABC transporters in eukaryotes. Thus, an exhaustive search of the EK-types among diverse organisms and detailed sequence classification and analysis would elucidate the evolutionary history of EK-types. It would also help shed some light on the fundamental repertoires of the wide variety of substrates through which multidrug ABC transporters in eukaryotes communicate. In this work, we have identified the EK-type ABC transporters in 126 prokaryotes using the profiles of the ATP-binding domain (NBD) of the EK-type ABC transporters from 12 eukaryotes. As a result, 11 clusters were identified from 1,046 EK-types ABC transporters. In particular, two large novel clusters emerged, corresponding to the bacterial multidrug ABC transporters related to the ABCB and ABCC families in eukaryotes, respectively. In the genomic context, most of these genes are located alone or adjacent to genes from the same clusters. Additionally, to detect functional divergences in the NBDs, the Kullback-Leibler divergence was measured among these bacterial multidrug transporters. As a result, several putative functional regions were identified, some corresponding to the predicted secondary structures. We also analyzed a phylogeny of the EK-type ABC transporters in both prokaryotes and eukaryotes, which revealed that the EK-type ABC transporters in prokaryotes have certain repertoires corresponding to the conventional ABC protein groups in eukaryotes. On the basis of these findings, we propose an updated evolutionary hypothesis in which the EK-type ABC transporters in both eukaryotes and prokaryotes consisted of several kinds of ABC transporters in putative ancestor cells before the divergence of eukaryotic and prokaryotic cells.

ATP-Binding Cassette Transporters↗

Application of a new probabilistic model for recognizing complex patterns in glycans.

MOTIVATION: The study of carbohydrate sugar chains, or glycans, has been one of slow progress mainly due to the difficulty in establishing standard methods for analyzing their structures and biosynthesis. Glycans are generally tree structures that are more complex than linear DNA or protein sequences, and evidence shows that patterns in glycans may be present that spread across siblings and into further regions that are not limited by the edges in the actual tree structure itself. Current models were not able to capture such patterns. RESULTS: We have applied a new probabilistic model, called probabilistic sibling-dependent tree Markov model (PSTMM), which is able to inherently capture such complex patterns of glycans. Not only is the ability to capture such patterns important in itself, but this also implies that PSTMM is capable of performing multiple tree structure alignments efficiently. We prove through experimentation on actual glycan data that this new model is extremely useful for gaining insight into the hidden, complex patterns of glycans, which are so crucial for the development and functioning of higher level organisms. Furthermore, we also show that this model can be additionally utilized as an innovative approach to multiple tree alignment, which has not been applied to glycan chains before. This extension on the usage of PSTMM may be a major step forward for not only the structural analysis of glycans, but it may consequently prove useful for discovering clues into their function.

Algorithms↗

KCaM (KEGG Carbohydrate Matcher): a software tool for analyzing the structures of carbohydrate sugar chains.

KCaM (KEGG Carbohydrate Matcher) is a tool for the analysis of carbohydrate sugar chains, or glycans. It consists of a web-based graphical user interface that allows users to enter glycans easily with the mouse. The glycan structure is then transformed into our KCF (KEGG Chemical Function) file format and sent to our program which implements an efficient tree-structure alignment algorithm, similar to sequence alignment algorithms but for branched tree structures. Users can also retrieve glycan tree structures in KCF format from their local computers for visualization over the web. The tree-matching algorithm provides several options for performing different types of tree-matching procedures on glycans. These options consist of whether to incorporate gaps in a match, whether to take the linkage information into consideration and local versus global alignment. The results of this program are returned as a list of glycan structures in order of similarity based on these options. The actual alignment can be viewed graphically, and the annotation information can also be viewed easily since all this information is linked with KEGG's comprehensive suite of genomic data. Analogously to BLAST, users are thus able to compare glycan structures of interest with glycans from different glycan databases using a variety of tree-alignment options. KCaM is currently available at http://glycan.genome.ad.jp.

Algorithms↗

The KEGG resource for deciphering the genome.

A grand challenge in the post-genomic era is a complete computer representation of the cell and the organism, which will enable computational prediction of higher-level complexity of cellular processes and organism behavior from genomic information. Toward this end we have been developing a knowledge-based approach for network prediction, which is to predict, given a complete set of genes in the genome, the protein interaction networks that are responsible for various cellular processes. KEGG at http://www.genome.ad.jp/kegg/ is the reference knowledge base that integrates current knowledge on molecular interaction networks such as pathways and complexes (PATHWAY database), information about genes and proteins generated by genome projects (GENES/SSDB/KO databases) and information about biochemical compounds and reactions (COMPOUND/GLYCAN/REACTION databases). These three types of database actually represent three graph objects, called the protein network, the gene universe and the chemical universe. New efforts are being made to abstract knowledge, both computationally and manually, about ortholog clusters in the KO (KEGG Orthology) database, and to collect and analyze carbohydrate structures in the GLYCAN database.

Animals↗

Response to oxidative stress involves a novel peroxiredoxin gene in the unicellular cyanobacterium Synechocystis sp. PCC 6803.

Exposure to methyl viologen in the presence of light facilitates the production of superoxide that gives severe damage on photosynthetic apparatus as well as many cellular processes in cyanobacteria and plants. The effects of methyl viologen on global gene expression of a unicellular cyanobacterium Synechocystis sp. strain PCC 6803 were determined by DNA microarray. The ORFs sll1621, slr1738, slr0074, slr0075, and slr0589 were significantly induced by treatment of methyl viologen for 15 min commonly under conditions of normal and high light. One of these genes, slr1738, which encodes a ferric uptake repressor (Fur)-type transcriptional regulator, is located divergently next to another induced gene, sll1621, in the genome. We deleted slr1738, and compared the global gene expression patterns of this mutant to that of wild type under non-stressed conditions. It was found that sll1621 was derepressed to the greatest extent, while many other genes including slr0589 but not slr0074 or slr0075 were derepressed to lesser extent in the mutant. Genetic disruption of sll1621, which encodes a putative type 2 peroxiredoxin, indicates that it is essential for aerobic phototrophic growth in both liquid and solid media in high light and on solid medium even in low light. Slr1738 was prepared as a His-tagged recombinant protein and shown to specifically bind to the intergenic region between sll1621 and slr1738. The binding was enhanced by dithiothreitol and abolished by hydrogen peroxide. We concluded that the Fur homolog, Slr1738, plays a regulatory role in the induction of a potent antioxidant gene, sll1621, in response to oxidative stress.

Amino Acid Sequence↗

Development of a chemical structure comparison method for integrated analysis of chemical and genomic information in the metabolic pathways.

Cellular functions result from intricate networks of molecular interactions, which involve not only proteins and nucleic acids but also small chemical compounds. Here we present an efficient algorithm for comparing two chemical structures of compounds, where the chemical structure is treated as a graph consisting of atoms as nodes and covalent bonds as edges. On the basis of the concept of functional groups, 68 atom types (node types) are defined for carbon, nitrogen, oxygen, and other atomic species with different environments, which has enabled detection of biochemically meaningful features. Maximal common subgraphs of two graphs can be found by searching for maximal cliques in the association graph, and we have introduced heuristics to accelerate the clique finding and to detect optimal local matches (simply connected common subgraphs). Our procedure was applied to the comparison and clustering of 9383 compounds, mostly metabolic compounds, in the KEGG/LIGAND database. The largest clusters of similar compounds were related to carbohydrates, and the clusters corresponded well to the categorization of pathways as represented by the KEGG pathway map numbers. When each pathway map was examined in more detail, finer clusters could be identified corresponding to subpathways or pathway modules containing continuous sets of reaction steps. Furthermore, it was found that the pathway modules identified by similar compound structures sometimes overlap with the pathway modules identified by genomic contexts, namely, by operon structures of enzyme genes.

Algorithms↗

Prediction of protein subcellular locations by support vector machines using compositions of amino acids and amino acid pairs.

MOTIVATION: The subcellular location of a protein is closely correlated to its function. Thus, computational prediction of subcellular locations from the amino acid sequence information would help annotation and functional prediction of protein coding genes in complete genomes. We have developed a method based on support vector machines (SVMs). RESULTS: We considered 12 subcellular locations in eukaryotic cells: chloroplast, cytoplasm, cytoskeleton, endoplasmic reticulum, extracellular medium, Golgi apparatus, lysosome, mitochondrion, nucleus, peroxisome, plasma membrane, and vacuole. We constructed a data set of proteins with known locations from the SWISS-PROT database. A set of SVMs was trained to predict the subcellular location of a given protein based on its amino acid, amino acid pair, and gapped amino acid pair compositions. The predictors based on these different compositions were then combined using a voting scheme. Results obtained through 5-fold cross-validation tests showed an improvement in prediction accuracy over the algorithm based on the amino acid composition only. This prediction method is available via the Internet.

Algorithms↗

Identification of a new cryptochrome class. Structure, function, and evolution.

Cryptochrome flavoproteins, which share sequence homology with light-dependent DNA repair photolyases, function as photoreceptors in plants and circadian clock components in animals. Here, we coupled sequencing of an Arabidopsis cryptochrome gene with phylogenetic, structural, and functional analyses to identify a new cryptochrome class (cryptochrome DASH) in bacteria and plants, suggesting that cryptochromes evolved before the divergence of eukaryotes and prokaryotes. The cryptochrome crystallographic structure, reported here for Synechocystis cryptochrome DASH, reveals commonalities with photolyases in DNA binding and redox-dependent function, despite distinct active-site and interaction surface features. Whole genome transcriptional profiling together with experimental confirmation of DNA binding indicated that Synechocystis cryptochrome DASH functions as a transcriptional repressor.

Amino Acid Sequence↗

Bioinformatics in the post-sequence era.

In the past decade, bioinformatics has become an integral part of research and development in the biomedical sciences. Bioinformatics now has an essential role both in deciphering genomic, transcriptomic and proteomic data generated by high-throughput experimental technologies and in organizing information gathered from traditional biology. Sequence-based methods of analyzing individual genes or proteins have been elaborated and expanded, and methods have been developed for analyzing large numbers of genes or proteins simultaneously, such as in the identification of clusters of related genes and networks of interacting proteins. With the complete genome sequences for an increasing number of organisms at hand, bioinformatics is beginning to provide both conceptual bases and practical methods for detecting systemic functional behaviors of the cell and the organism.

Computational Biology↗

Extracting active pathways from gene expression data.

MOTIVATION: A promising way to make sense out of gene expression profiles is to relate them to the activity of metabolic and signalling pathways. Each pathway usually involves many genes, such as enzymes, which can themselves participate in many pathways. The set of all known pathways can therefore be represented by a complex network of genes. Searching for regularities in the set of gene expression profiles with respect to the topology of this gene network is a way to automatically extract active pathways and their associated patterns of activity. METHOD: We present a method to perform this task, which consists in encoding both the gene network and the set of profiles into two kernel functions, and performing a regularized form of canonical correlation analysis between the two kernels. RESULTS: When applied to publicly available expression data the method is able to extract biologically relevant expression patterns, as well as pathways with related activity.

Algorithms↗

DNA microarray analysis of redox-responsive genes in the genome of the cyanobacterium Synechocystis sp. strain PCC 6803.

Whole-genome DNA microarrays were used to evaluate the effect of the redox state of the photosynthetic electron transport chain on gene expression in Synechocystis sp. strain PCC 6803. Two specific inhibitors of electron transport, 3-(3,4-dichlorophenyl)-1,1-dimethylurea (DCMU) and 2,5-dibromo-3-methyl-6-isopropyl-p-benzoquinone (DBMIB), were added to the cultures, and changes in accumulation of transcripts were examined. About 140 genes were highlighted as reproducibly affected by the change in the redox state of the photosynthetic electron transport chain. It was shown that some stress-responsive genes but not photosynthetic genes were under the control of the redox state of the plastoquinone pool in Synechocystis sp. strain PCC 6803.

Bacterial Proteins↗

Update of MAGEST: Maboya Gene Expression patterns and Sequence Tags.

MAGEST is a database for maternal gene expression information for an ascidian, Halocynthia roretzi. The ascidian has become an animal model in developmental biological research because it shows a simple developmental process, and belongs to one of the chordate groups. Various data are deposited into the MAGEST database, e.g. the 3'- and 5'-tag sequences from the fertilized egg cDNA library, the results of similarity searches against GenBank and the expression data from whole mount in situ hybridization. Over the last 2 years, the data retrieval systems have been improved in several aspects, and the tag sequence entries have increased to over 20 000 clones. Additionally, we constructed a database, translated MAGEST, for the amino acid fragment sequences predicted from the EST data sets. Using this information comprehensively, we should obtain new information on gene functions. The MAGEST database is accessible via the Internet at http://www.genome.ad.jp/magest/.

Amino Acid Sequence↗

LIGAND: database of chemical compounds and reactions in biological pathways.

LIGAND is a composite database comprising three sections: COMPOUND for the information about metabolites and other chemical compounds, REACTION for the collection of substrate-product relations representing metabolic and other reactions, and ENZYME for the information about enzyme molecules. The current release (as of September 7, 2001) includes 7298 compounds, 5166 reactions and 3829 enzymes. In addition to the keyword search provided by the DBGET/LinkDB system, a substructure search to the COMPOUND and REACTION sections is now available through the World Wide Web (http://www.genome.ad.jp/ligand/). LIGAND may be also downloaded by anonymous FTP (ftp://ftp.genome.ad.jp/pub/kegg/ligand/).

Animals↗

The KEGG databases at GenomeNet.

The Kyoto Encyclopedia of Genes and Genomes (KEGG) is the primary database resource of the Japanese GenomeNet service (http://www.genome.ad.jp/) for understanding higher order functional meanings and utilities of the cell or the organism from its genome information. KEGG consists of the PATHWAY database for the computerized knowledge on molecular interaction networks such as pathways and complexes, the GENES database for the information about genes and proteins generated by genome sequencing projects, and the LIGAND database for the information about chemical compounds and chemical reactions that are relevant to cellular processes. In addition to these three main databases, limited amounts of experimental data for microarray gene expression profiles and yeast two-hybrid systems are stored in the EXPRESSION and BRITE databases, respectively. Furthermore, a new database, named SSDB, is available for exploring the universe of all protein coding genes in the complete genomes and for identifying functional links and ortholog groups. The data objects in the KEGG databases are all represented as graphs and various computational methods are developed to detect graph features that can be related to biological functions. For example, the correlated clusters are graph similarities which can be used to predict a set of genes coding for a pathway or a complex, as summarized in the ortholog group tables, and the cliques in the SSDB graph are used to annotate genes. The KEGG databases are updated daily and made freely available (http://www.genome.ad.jp/kegg/).

Animals↗

Screening for the target gene of cyanobacterial cAMP receptor protein SYCRP1.

The target genes for SYCRP1, a cyanobacterial cAMP receptor protein, were surveyed using a DNA microarray method. Total RNAs were extracted from a wild-type strain and a sycrp1 disruptant of Synechocystis sp. PCC 6803, and the respective gene expression levels were compared. The expression levels of six genes (slr1667, slr1168, slr2015, slr2016, slr2017 and slr2018) were clearly decreased by the disruption of the sycrp1 gene. The data suggest that slr1667 and slr1668 constitute one operon and the other four genes constitute another operon. Transcription start points for the first genes of these putative operons, which are slr1667 and slr2015, were determined by primer extension experiments. Gel mobility shift assays and DNase 1 footprint analyses were carried out to explore the binding of SYCRP1 to the putative promoter regions of slr1667 and slr2015. SYCRP1 bound to the specific site in the 5' upstream region of slr1667 from positions -170 to -155 relative to the transcription start point, while it did not bind to the 5' upstream region of slr2015. It was concluded that SYCRP1 regulates the expression of the slr1667 gene directly by binding to a specific site in its promoter.

Bacterial Proteins↗

A two-component Mn2+-sensing system negatively regulates expression of the mntCAB operon in Synechocystis.

Mn is an essential component of the oxygen-evolving machinery of photosynthesis and is an essential cofactor of several important enzymes, such as Mn-superoxide dismutase and Mn-catalase. The availability of Mn in the environment varies, and little is known about the mechanisms for maintaining cytoplasmic Mn(2+) ion homeostasis. Using a DNA microarray, we screened knockout libraries of His kinases and response regulators of Synechocystis sp PCC 6803 to identify possible participants in this process. We identified a His kinase, ManS, which might sense the extracellular concentration of Mn(2+) ions, and a response regulator, ManR, which might regulate the expression of the mntCAB operon for the ABC-type transporter of Mn(2+) ions. Furthermore, analysis with the DNA microarray and by reverse transcription PCR suggested that ManS produces a signal that activates ManR, which represses the expression of the mntCAB operon. At low concentrations of Mn(2+) ions, ManS does not generate a signal, with resulting inactivation of ManR and subsequent expression of the mntCAB operon.

ATP-Binding Cassette Transporters↗

The KEGG database.

KEGG (http://www.genome.ad.jp/kegg/) is a suite of databases and associated software for understanding and simulating higher-order functional behaviours of the cell or the organism from its genome information. First, KEGG computerizes data and knowledge on protein interaction networks (PATHWAY database) and chemical reactions (LIGAND database) that are responsible for various cellular processes. Second, KEGG attempts to reconstruct protein interaction networks for all organisms whose genomes are completely sequenced (GENES and SSDB databases). Third, KEGG can be utilized as reference knowledge for functional genomics (EXPRESSION database) and proteomics (BRITE database) experiments. I will review the current status of KEGG and report on new developments in graph representation and graph computations.

Amino Acid Sequence↗

Extraction of organism groups from phylogenetic profiles using independent component analysis.

In recent years, the analysis of orthologous genes based on phylogenetic profiles has received popularity in bioinfomatics. We propose a new method to extract organism groups and their hierarchy from phylogenetic profiles using the independent component analysis (ICA). The method involves first finding independent axes in the projected space from the multivariate data matrix representing phylogenetic profiles for a number of orthologous genes. Then the extracted axes are correlated with major organism groups, according to the extent of affiliation of axes scores for all the genes to specific organisms. The ICA was applied to the phylogenetic profiles created for 2,875 orthologs in 77 organisms by using the KEGG/GENES database. The 9 extracted components out of 18 predefined components well represented the organism groups as categorized in KEGG. Furthermore, we performed the cluster analysis and obtained the hierarchy of organism groups.

Animals↗