Search PubMed⌕ Search

Biomedical subjects

Hiroyuki Toh

Publications and source records attributed to Hiroyuki Toh.

At least 19 recordsLinked to original sources

Partial correlation coefficient between distance matrices as a new indicator of protein-protein interactions.

MOTIVATION: The computational prediction of protein-protein interactions is currently a major issue in bioinformatics. Recently, a variety of co-evolution-based methods have been investigated toward this goal. In this study, we introduced a partial correlation coefficient as a new measure for the degree of co-evolution between proteins, and proposed its use to predict protein-protein interactions. RESULTS: The accuracy of the prediction by the proposed method was compared with those of the original mirror tree method and the projection method previously developed by our group. We found that the partial correlation coefficient effectively reduces the number of false positives, as compared with other methods, although the number of false negatives increased in the prediction by the partial correlation coefficient. AVAILABILITY: The R script for the prediction of protein-protein interactions reported in this manuscript is available at http://timpani.genome.ad.jp/~parco/

Amino Acid Sequence↗

Identification of rDNA-specific non-LTR retrotransposons in Cnidaria.

Ribosomal RNA genes are abundant repetitive sequences in most eukaryotes. Ribosomal DNA (rDNA) contains many insertions derived from mobile elements including non-long terminal repeat (non-LTR) retrotransposons. R2 is the well-characterized 28S rDNA-specific non-LTR retrotransposon family that is distributed over at least 4 bilaterian phyla. R2 is a large family sharing the same insertion specificity and classified into 4 clades (R2-A, -B, -C, and -D) based on the N-terminal domain structure and the phylogeny. There is no observation of horizontal transfer of R2; therefore, the origin of R2 dates back to before the split between protostomes and deuterostomes. Here, we in silico identified 1 R2 element from the sea anemone Nematostella vectensis and 2 R2-like retrotransposons from the hydrozoan Hydra magnipapillata. R2 from N. vectensis was inserted into the 28S rDNA like other R2, but the R2-like elements from H. magnipapillata were inserted into the specific sequence in the highly conserved region of the 18S rDNA. We designated the Hydra R2-like elements R8. R8 is inserted at 37 bp upstream from R7, another 18S rDNA-specific retrotransposon family. There is no obvious sequence similarity between targets of R2 and R8, probably because they recognize long DNA sequences. Domain structure and phylogeny indicate that R2 from N. vectensis is the member of the R2-D clade, and R8 from H. magnipapillata belongs to the R2-A clade despite its different sequence specificity. These results suggest that R2 had been generated before the split between cnidarians and bilaterians and that R8 is a retrotransposon family that changed its target from the 28S rDNA to the 18S rDNA.

Amino Acid Sequence↗

Positive selection in the ComC-ComD system of Streptococcal Species.

Competence-stimulating peptide (CSP) and ComD of the streptococcal species are a pheromone and its receptor, respectively, involved in the regulation of competence for natural genetic transformation. We show here that these molecules have undergone positive selection. This study is the first report of positive selection due to competition among bacterial populations.

Amino Acid Sequence↗

GASH: an improved algorithm for maximizing the number of equivalent residues between two protein structures.

BACKGROUND: We introduce GASH, a new, publicly accessible program for structural alignment and superposition. Alignments are scored by the Number of Equivalent Residues (NER), a quantitative measure of structural similarity that can be applied to any structural alignment method. Multiple alignments are optimized by conjugate gradient maximization of the NER score within the genetic algorithm framework. Initial alignments are generated by the program Local ASH, and can be supplemented by alignments from any other program. RESULTS: We compare GASH to DaliLite, CE, and to our earlier program Global ASH on a difficult test set consisting of 3,102 structure pairs, as well as a smaller set derived from the Fischer-Eisenberg set. The extent of alignment crossover, as well as the completeness of the initial set of alignments are examined. The quality of the superpositions is evaluated both by NER and by the number of aligned residues under three different RMSD cutoffs (2,4, and 6A). In addition to the numerical assessment, the alignments for several biologically related structural pairs are discussed in detail. CONCLUSION: Regardless of which criteria is used to judge the superposition accuracy, GASH achieves the best overall performance, followed by DaliLite, Global ASH, and CE. In terms of CPU usage, DaliLite CE and GASH perform similarly for query proteins under 500 residues, but for larger proteins DaliLite is faster than GASH or CE. Both an http interface and a simple object application protocol (SOAP) interface to the GASH program are available at http://www.pdbj.org/GASH/.

Algorithms↗

ASIAN: a web server for inferring a regulatory network framework from gene expression profiles.

The standard workflow in gene expression profile analysis to identify gene function is the clustering by various metrics and techniques, and the following analyses, such as sequence analyses of upstream regions. A further challenging analysis is the inference of a gene regulatory network, and some computational methods have been intensively developed to deduce the gene regulatory network. Here, we describe our web server for inferring a framework of regulatory networks from a large number of gene expression profiles, based on graphical Gaussian modeling (GGM) in combination with hierarchical clustering (http://eureka.ims.u-tokyo.ac.jp/asian). GGM is based on a simple mathematical structure, which is the calculation of the inverse of the correlation coefficient matrix between variables, and therefore, our server can analyze a wide variety of data within a reasonable computational time. The server allows users to input the expression profiles, and it outputs the dendrogram of genes by several hierarchical clustering techniques, the cluster number estimated by a stopping rule for hierarchical clustering and the network between the clusters by GGM, with the respective graphical presentations. Thus, the ASIAN (Automatic System for Inferring A Network) web server provides an initial basis for inferring regulatory relationships, in that the clustering serves as the first step toward identifying the gene function.

Cluster Analysis↗

The inference of protein-protein interactions by co-evolutionary analysis is improved by excluding the information about the phylogenetic relationships.

MOTIVATION: The prediction of protein-protein interactions is currently an important issue in bioinformatics. The mirror tree method uses evolutionary information to predict protein-protein interactions. However, it has been recognized that predictions by the mirror tree method lead to many false positives. The incentive of our study was to solve this problem by improving the method of extracting the co-evolutionary information regarding the protein pairs. RESULTS: We developed a novel method to predict protein-protein interactions from co-evolutionary information in the framework of the mirror tree method. The originality is the use of the projection operator to exclude the information about the phylogenetic relationships among the source organisms from the distance matrix. Each distance matrix was transformed into a vector for the operation. The vector is referred to as a 'phylogenetic vector'. We have proposed three ways to extract the phylogenetic information: (1) using the 16S rRNA from the same source organisms as the proteins under consideration, (2) averaging the phylogenetic vectors and (3) analyzing the principal components of the phylogenetic vectors. We examined the performance of the proposed methods to predict interacting protein pairs from Escherichia coli, using experimentally verified data. Our method was successful, and it drastically reduced the number of false positives in the prediction. AVAILABILITY: The R script for the prediction of protein-protein interactions reported in this manuscript is available at http://timpani.genome.ad.jp/~proj/ CONTACT: sato@kuicr.kyoto-u.ac.jp SUPPLEMENTARY INFORMATION: The information is also available at the same site as the R script.

Algorithms↗

Prediction of interfaces for oligomerizations of G-protein coupled receptors.

Several lines of biochemical and pharmacological evidence have suggested that some G-protein-coupled receptors (GPCRs) form homo oligomers, hetero oligomers or both. The GPCRs oligomerizations are considered to be related to signal transduction and some diseases. Therefore, an accurate prediction of the residues that interact upon oligomerization interface would further our understanding of signal transduction and the diseases in which GPCRs are involved. One of the complications for such a prediction is that the interfaces differ with the subtypes, even within the same GPCR family. Focusing on the distribution of residues conserved on the molecular surface in a particular subtype, we developed a new method to predict the interface for the GPCR oligomers, and applied it to several subtypes of known GPCRs to check the sensitivity. Subsequently, we found that predicted interfaces of rhodopsin, D(2) dopamine receptor and beta(2) adrenergic receptor agreed with the experimentally suggested interfaces, despite difference in the interface region among the three subtypes. Moreover, a highly conserved residue detected from the D(2) dopamine receptor corresponded to a residue involved in a missense change found in the large family of myoclonus dystonia. Our observation suggests the possibility that the disease is caused by the disorder of the oligomerization, although the molecular mechanism of the disease has not been revealed yet. The benefits and the pitfalls of the new method will be discussed, based on the results of the applications.

Algorithms↗

MAFFT version 5: improvement in accuracy of multiple sequence alignment.

The accuracy of multiple sequence alignment program MAFFT has been improved. The new version (5.3) of MAFFT offers new iterative refinement options, H-INS-i, F-INS-i and G-INS-i, in which pairwise alignment information are incorporated into objective function. These new options of MAFFT showed higher accuracy than currently available methods including TCoffee version 2 and CLUSTAL W in benchmark tests consisting of alignments of >50 sequences. Like the previously available options, the new options of MAFFT can handle hundreds of sequences on a standard desktop computer. We also examined the effect of the number of homologues included in an alignment. For a multiple alignment consisting of approximately 8 sequences with low similarity, the accuracy was improved (2-10 percentage points) when the sequences were aligned together with dozens of their close homologues (E-value < 10(-5)-10(-20)) collected from a database. Such improvement was generally observed for most methods, but remarkably large for the new options of MAFFT proposed here. Thus, we made a Ruby script, mafftE.rb, which aligns the input sequences together with their close homologues collected from SwissProt using NCBI-BLAST.

Reproducibility of Results↗

A novel representation of protein sequences for prediction of subcellular location using support vector machines.

As the number of complete genomes rapidly increases, accurate methods to automatically predict the subcellular location of proteins are increasingly useful to help their functional annotation. In order to improve the predictive accuracy of the many prediction methods developed to date, a novel representation of protein sequences is proposed. This representation involves local compositions of amino acids and twin amino acids, and local frequencies of distance between successive (basic, hydrophobic, and other) amino acids. For calculating the local features, each sequence is split into three parts: N-terminal, middle, and C-terminal. The N-terminal part is further divided into four regions to consider ambiguity in the length and position of signal sequences. We tested this representation with support vector machines on two data sets extracted from the SWISS-PROT database. Through fivefold cross-validation tests, overall accuracies of more than 87% and 91% were obtained for eukaryotic and prokaryotic proteins, respectively. It is concluded that considering the respective features in the N-terminal, middle, and C-terminal parts is helpful to predict the subcellular location.

Amino Acids, Basic↗

A study of archaeal enzymes involved in polar lipid synthesis linking amino acid sequence information, genomic contexts and lipid composition.

Cellular membrane lipids, of which phospholipids are the major constituents, form one of the characteristic features that distinguish Archaea from other organisms. In this study, we focused on the steps in archaeal phospholipid synthetic pathways that generate polar lipids such as archaetidylserine, archaetidylglycerol, and archaetidylinositol. Only archaetidylserine synthase (ASS), from Methanothermobacter thermautotrophicus, has been experimentally identified. Other enzymes have not been fully examined. Through database searching, we detected many archaeal hypothetical proteins that show sequence similarity to members of the CDP alcohol phosphatidyltransferase family, such as phosphatidylserine synthase (PSS), phosphatidylglycerol synthase (PGS) and phosphatidylinositol synthase (PIS) derived from Bacteria and Eukarya. The archaeal hypothetical proteins were classified into two groups, based on the sequence similarity. Members of the first group, including ASS from M. thermautotrophicus, were closely related to PSS. The rough agreement between PSS homologue distribution within Archaea and the experimentally identified distribution of archaetidylserine suggested that the hypothetical proteins are ASSs. We found that an open reading frame (ORF) tends to be adjacent to that of ASS in the genome, and that the order of the two ORFs is conserved. The sequence similarity of phosphatidylserine decarboxylase to the product of the ORF next to the ASS gene, together with the genomic context conservation, suggests that the ORF encodes archaetidylserine decarboxylase, which may transform archaetidylserine to archaetidylethanolamine. The second group of archaeal hypothetical proteins was related to PGS and PIS. The members of this group were subjected to molecular phylogenetic analysis, together with PGSs and PISs and it was found that they formed two distinct clusters in the molecular phylogenetic tree. The distribution of members of each cluster within Archaea roughly corresponded to the experimentally identified distribution of archaetidylglycerol or archaetidylinositol. The molecular phylogenetic tree patterns and the correspondence to the membrane compositions suggest that the two clusters in this group correspond to archaetidylglycerol synthases and archaetidylinositol synthases. No archaeal hypothetical protein with sequence similarity to known phosphatidylcholine synthases was detected in this study.

Archaea↗

[Bioinformatics].

Explore the source record for details and available documents.

Amino Acid Sequence↗

Improvement in the accuracy of multiple sequence alignment program MAFFT.

In 2002, we developed and released a rapid multiple sequence alignment program MAFFT that was designed to handle a huge (up to approximately 5,000 sequences) and long data (approximately 2,000 aa or approximately 5,000 nt) in a reasonable time on a standard desktop PC. As for the accuracy, however, the previous versions (v.4 and lower) of MAFFT were outperformed by ProbCons and TCoffee v.2, both of which were released in 2004, in several benchmark tests. Here we report a recent extension of MAFFT that aims to improve the accuracy with as little cost of calculation time as possible. The extended version of MAFFT (v.5) has new iterative refinement options, G-INS-i and L-INS-i (collectively denoted as [GL]-INS-i in this report). These options use a new objective function combining the weighted sum-of-pairs (WSP) score and a score similar to COFFEE derived from all pairwise alignments. We discuss the improvement in accuracy brought by this extension, mainly using two benchmark tests released very recently, BAliBASE v.3 (for protein alignments) and BRAliBASE (for RNA alignments). According to BAliBASE v.3, the overall average accuracy of L-INS-i was higher than those of other methods successively released in 2004, although the difference among the most accurate methods (ProbCons, TCoffee v.2 and new options of MAFFT) was small. The advantage in accuracy of [GL]-INS-i became greater for the alignments consisting of approximately 50-100 sequences. By utilizing this feature of MAFFT, we also examined another possible approach to improve the accuracy by incorporating homolog information collected from database. The [GL]-INS-i options are applicable to aligning up to approximately 200 sequences, although not applicable to thousands of sequences because of time and space complexities.

Amino Acid Sequence↗

Detecting local structural similarity in proteins by maximizing number of equivalent residues.

A new algorithm for superimposing protein structures based on maximizing the number of spatially equivalent residues is introduced. The algorithm works in three distinct steps. First, the optimal residue map is calculated by structural alignment. By default, the double dynamic programming algorithm, as implemented in the program ASH, was used for the structure alignment step, but we also present results based on alignments imported from three other programs (Dali, CE, and VAST).Second, the structures are spatially superimposed such that the effective number of equivalent residues (NER)--aligned residue pairs that can be spatially overlapped--is maximized. The NER score is an analytic, differentiable similarity function that rewards spatially equivalent residues but ignores non-equivalent ones. Maximization of the NER score results in accurate superpositions in cases where root mean square deviation (RMSD) minimization fails. Third, the NER function is used in conjunction with traditional dynamic programming to realign the structures based on the proximity of residues in the superposition. Results are presented for a wide range of superposition problems and compared to results from Dali, CE, and VAST. In addition, several structure-structure pairs that show only partial similarity are discussed, and results are compared to those from the LGA, SARF2, and ThreeCa programs.

Algorithms↗

ASIAN: a website for network inference.

UNLABELLED: We constructed a website for inferring a network by applying the graphical Gaussian model, from a large amount of data, including redundant information. The available tools on the website are based on a system, named ASIAN (Automatic System for Inferring A Network), in combination with the two methods in our previous papers, which were designed to analyze gene expression profiles on a genomic scale. One of the remarkable features of the website is its ability to infer a network, concomitant with hierarchical clustering and the following estimation of cluster boundaries. AVAILABILITY: http://eureka.ims.u-tokyo.ac.jp/asian

Computer Simulation↗

How does a topological inversion change the evolutionary constraints on membrane proteins?

The members of the aquaporin family and those of the ClC chloride ion channel family consist of two-fold tandem repeats. The orientation of the N-terminal domain against membrane is opposite to that of the C-terminal domain. Several lines of evidence suggest that the extracellular and the cytoplasmic environments impose different evolutionary constraints on proteins (e.g. positive-inside rule). Therefore, the different constraints would affect the corresponding regions of the two domains, which are exposed to the different environments. To examine this hypothesis, the N- and the C-terminal domains were aligned and the difference in residue composition or conservation pattern between the two domains was calculated at each alignment site by several methods. Then, the residues corresponding to the sites exhibiting significant difference were mapped onto the tertiary structure. In spite of the difference in the methods, the mapped residues clustered on the pore surface of the channel; in contrast, the number of the residues mapped on the extracellular or cytoplasmic sides of the proteins was small. A minor modification of the methods improved the sensitivity to detect sites related to the positive-inside rule. The results support our hypothesis about the relationship between the topological inversion and the different constraints.

Amino Acid Sequence↗

Identification of cryptochrome DASH from vertebrates.

A new type of cryptochrome, CRY-DASH, has been recently identified. The CRY-DASH proteins constitute the fifth subfamily of the photolyase/cryptochrome family. CRY-DASHs have been identified from Synechocystis sp. PCC 6803, Vibrio cholerae, and Arabidopsis thaliana. The Synechocystis CRY-DASH was the first cryptochrome identified from bacteria, and its biochemical features and tertiary structure have been extensively investigated. To determine how broadly the subfamily is distributed within living organisms, we searched for new CRY-DASH candidates within several databases. We found five sequences as new CRY-DASH candidates, which are derived from four marine bacteria and Neurospora crassa. We also found many CRY-DASH candidates from the EST databases, which included sequences from fish and amphibians. We cloned and sequenced the cDNAs of the zebrafish and Xenopus laevis candidates, based on the EST sequences. The proteins encoded by the two genes were purified and characterized. Both proteins contained folate and flavin cofactors, and have a weak DNA photolyase activity. A phylogenetic analysis revealed that the seven candidates actually belong to the new type of cryptochrome subfamily. This is the first report of the CRY-DASH members from vertebrates and fungi.

Amino Acid Sequence↗

Archaeal-type rhodopsins in Chlamydomonas: model structure and intracellular localization.

Phototaxis in the unicellular green alga Chlamydomonas reinhardtii is mediated by rhodopsin-type photoreceptor(s). Recent expressed sequence tag database from the Kazusa DNA Research Institute has provided the basis for unequivocal identification of two archaeal-type rhodopsins in it. Here we demonstrate that one is located near the eyespot, wherein the photoreceptor(s) has long been thought to be enriched, along with the results of bioinformatic analyses. Secondary structure prediction showed that the second putative transmembrane helices (helix B) of these rhodopsins are rich in glutamate residues, and homology modeling suggested that some additional intra- or intermolecular interactions are necessary for opsin-like folding of the N-terminal ca. 300-aa membrane spanning domains of 712 and 737-aa polypeptides. These results complement physiological and electrophysiological experiments combined with the manipulation of their expression [O.A. Sineshchekov, K.H. Jung, J.H. Spudich, Proc. Natl. Sci. USA 99 (2002) 8689; G. Nagel, D. Olig, M. Fuhrmann, S. Kateriya, A.M. Musti, E. Bamberg, P. Hegemann, Science 296 (2002) 2395].

Algal Proteins↗