Search PubMed⌕ Search

Biomedical subjects

Masaru Tomita

Publications and source records attributed to Masaru Tomita.

At least 19 recordsLinked to original sources

Archaeal Pyrococcus furiosus thymidylate synthase 1 is an RNA-binding protein.

Using a stem-loop RNA oligonucleotide (19-mer) containing an AUG sequence in the loop region as a probe, we screened the protein library from a hyperthermophilic archaeon, Pyrococcus furiosus, and found that a flavin-dependent thymidylate synthase, Pf-Thy1 (Pyrococcus furiosus thymidylate synthase 1), possessed RNA-binding activity. Recombinant Pf-Thy1 was able to bind to the stem-loop structure at a high temperature (75 degrees C) with an apparent dissociation constant of 0.6 microM. A similar stem-loop RNA structure was located around the translation start AUG codon of Pf-Thy1 RNA, and gel-shift analysis revealed that Pf-Thy1 could also bind to this stem-loop structure. In vitro translation analysis using chimaeric constructs containing the stem-loop sequence in their Pf-Thy1 RNA and a luciferase reporter gene indicated that the stem-loop structure acted as an inhibitory regulator of translation by preventing the binding of its Shine-Dalgarno-like sequence by positioning it in the stem region. Addition of Pf-Thy1 into the in vitro translation system also inhibited translation. These results suggested that this class of thymidylate synthases may autoregulate their own translation in a manner analogous to that of the well characterized thymidylate synthase A proteins, although there is no significant amino acid sequence similarity between them.

Amino Acid Sequence↗

Hybrid dynamic/static method for large-scale simulation of metabolism.

BACKGROUND: Many computer studies have employed either dynamic simulation or metabolic flux analysis (MFA) to predict the behaviour of biochemical pathways. Dynamic simulation determines the time evolution of pathway properties in response to environmental changes, whereas MFA provides only a snapshot of pathway properties within a particular set of environmental conditions. However, owing to the large amount of kinetic data required for dynamic simulation, MFA, which requires less information, has been used to manipulate large-scale pathways to determine metabolic outcomes. RESULTS: Here we describe a simulation method based on cooperation between kinetics-based dynamic models and MFA-based static models. This hybrid method enables quasi-dynamic simulations of large-scale metabolic pathways, while drastically reducing the number of kinetics assays needed for dynamic simulations. The dynamic behaviour of metabolic pathways predicted by our method is almost identical to that determined by dynamic kinetic simulation. CONCLUSION: The discrepancies between the dynamic and the hybrid models were sufficiently small to prove that an MFA-based static module is capable of performing dynamic simulations as accurately as kinetic models. Our hybrid method reduces the number of biochemical experiments required for dynamic models of large-scale metabolic pathways by replacing suitable enzyme reactions with a static module.

Animals↗

Dynamic simulation of red blood cell metabolism and its application to the analysis of a pathological condition.

BACKGROUND: Cell simulation, which aims to predict the complex and dynamic behavior of living cells, is becoming a valuable tool. In silico models of human red blood cell (RBC) metabolism have been developed by several laboratories. An RBC model using the E-Cell simulation system has been developed. This prototype model consists of three major metabolic pathways, namely, the glycolytic pathway, the pentose phosphate pathway and the nucleotide metabolic pathway. Like the previous model by Joshi and Palsson, it also models physical effects such as osmotic balance. This model was used here to reconstruct the pathology arising from hereditary glucose-6-phosphate dehydrogenase (G6PD) deficiency, which is the most common deficiency in human RBC. RESULTS: Since the prototype model could not reproduce the state of G6PD deficiency, the model was modified to include a pathway for de novo glutathione synthesis and a glutathione disulfide (GSSG) export system. The de novo glutathione (GSH) synthesis pathway was found to compensate partially for the lowered GSH concentrations resulting from G6PD deficiency, with the result that GSSG could be maintained at a very low concentration due to the active export system. CONCLUSION: The results of the simulation were consistent with the estimated situation of real G6PD-deficient cells. These results suggest that the de novo glutathione synthesis pathway and the GSSG export system play an important role in alleviating the consequences of G6PD deficiency.

Computer Simulation↗

Space in systems biology of signaling pathways--towards intracellular molecular crowding in silico.

How cells utilize intracellular spatial features to optimize their signaling characteristics is still not clearly understood. The physical distance between the cell-surface receptor and the gene expression machinery, fast reactions, and slow protein diffusion coefficients are some of the properties that contribute to their intricacy. This article reviews computational frameworks that can help biologists to elucidate the implications of space in signaling pathways. We argue that intracellular macromolecular crowding is an important modeling issue, and describe how recent simulation methods can reproduce this phenomenon in either implicit, semi-explicit or fully explicit representation.

Algorithms↗

GC-compositional strand bias around transcription start sites in plants and fungi.

BACKGROUND: A GC-compositional strand bias or GC-skew (=(C-G)/(C+G)), where C and G denote the numbers of cytosine and guanine residues, was recently reported near the transcription start sites (TSS) of Arabidopsis genes. However, it is unclear whether other eukaryotic species have equally prominent GC-skews, and the biological meaning of this trait remains unknown. RESULTS: Our study confirmed a significant GC-skew (C > G) in the TSS of Oryza sativa (rice) genes. The full-length cDNAs and genomic sequences from Arabidopsis and rice were compared using statistical analyses. Despite marked differences in the G+C content around the TSS in the two plants, the degrees of bias were almost identical. Although slight GC-skew peaks, including opposite skews (C < G), were detected around the TSS of genes in human and Drosophila, they were qualitatively and quantitatively different from those identified in plants. However, plant-like GC-skew in regions upstream of the translation initiation sites (TIS) in some fungi was identified following analyses of the expressed sequence tags and/or genomic sequences from other species. On the basis of our dataset, we estimated that > 70 and 68% of Arabidopsis and rice genes, respectively, had a strong GC-skew (> 0.33) in a 100-bp window (that is, the number of C residues was more than double the number of G residues in a +/-100-bp window around the TSS). The mean GC-skew value in the TSS of highly-expressed genes in Arabidopsis was significantly greater than that of genes with low expression levels. Many of the GC-skew peaks were preferentially located near the TSS, so we examined the potential value of GC-skew as an index for TSS identification. Our results confirm that the GC-skew can be used to assist the TSS prediction in plant genomes. CONCLUSION: The GC-skew (C > G) around the TSS is strictly conserved between monocot and eudicot plants (ie. angiosperms in general), and a similar skew has been observed in some fungi. Highly-expressed Arabidopsis genes had overall a more marked GC-skew in the TSS compared to genes with low expression levels. We therefore propose that the GC-skew around the TSS in some plants and fungi is related to transcription. It might be caused by mutations during transcription initiation or the frequent use of transcription factor-biding sites having a strand preference. In addition, GC-skew is a good candidate index for TSS prediction in plant genomes, where there is a lack of correlation among CpG islands and genes.

Arabidopsis↗

Computational analysis suggests that alternative first exons are involved in tissue-specific transcription in rice (Oryza sativa).

MOTIVATION: Transcription start site selection and alternative splicing greatly contribute to diversifying gene expression. Recent studies have revealed the existence of alternative first exons, but most have involved mammalian genes, and as yet the regulation of usage of alternative first exons has not been clarified, especially in plants. RESULTS: We systematically identified putative alternative first exon transcripts in rice, verified the candidates using RT-PCR, and searched for the promoter elements that might regulate the alternative first exons. As a result, we detected a number of unreported alternative first exons, some of which are regulated in a tissue-specific manner. SUPPLEMENTARY INFORMATION: http://www.bioinfo.sfc.keio.ac.jp/research/intron.

Alternative Splicing↗

In silico diagnosis of inherently inhibited gene expression focusing on initial codon combinations.

The translation start site, immediately downstream from the start codon, is a dominant factor for gene expression in Escherichia coli. At present, no method exists to improve the expression level of cloned genes, since it remains difficult to find the best codon combination within the region. We determined the expression parameters that correspond to all sense codons within the first four codons using GFPuv which encodes a derivative of green fluorescent protein. Using a genetic algorithm (GA)-based computer program, these parameters were incorporated in a simple, static model for the prediction of translation efficiency, and optimized to the expression level for 137 randomly isolated GFPuv genes. The calculated initial translation index (ITI), also proven for the DsRed2 gene that encodes a red fluorescent protein, should provide a solution to overcome the gene expression problem in cloned genes whose expression is often inherently blocked at the translation process. The proposed method facilitates heterologous protein production in E. coli, the most commonly used host in biological and industrial fields.

Cloning, Molecular↗

Large-scale prediction of cationic metabolite identity and migration time in capillary electrophoresis mass spectrometry using artificial neural networks.

We developed a computational technique to assist in the large-scale identification of charged metabolites. The electrophoretic mobility of metabolites in capillary electrophoresis-mass spectrometry (CE-MS) was predicted from their structure, using an ensemble of artificial neural networks (ANNs). Comparison between relative migration times of 241 various cations measured by CE-MS and predicted by a trained ANN ensemble produced a correlation coefficient of 0.931. When we used our technique to characterize all metabolites listed in the KEGG ligand database, the correct compounds among the top three candidates were predicted in 78.0% of cases. We suggest that this approach can be used for the prediction of the migration time of any cation and that it represents a powerful method for the identification of uncharacterized CE-MS peaks in metabolome analysis.

Cations↗

All systems go: launching cell simulation fueled by integrated experimental biology data.

Biological simulation serves to unify the basic elements of systems biology, namely, model selection, experimentation and model refinement. To select biochemical models for simulation, metabolome analysis can be performed using capillary electrophoresis or liquid chromatography coupled with mass spectrometry. In this manner, selected models can be elaborated with temporal/spatial gene and protein expression data obtained from model organisms such as Escherichia coli. The E. coli single gene deletion mutant library (KO collection) and His-tag/GFP-fusion single open reading frame clone expression library (ASKA) are powerful resources for this task. The integration of parallel experimental datasets into dynamic simulation tools forms the remaining challenge for the systematic analysis and elucidation of biological networks and holds promise for biotechnological applications.

Cell Physiological Phenomena↗

The SCO2299 gene from Streptomyces coelicolor A3(2) encodes a bifunctional enzyme consisting of an RNase H domain and an acid phosphatase domain.

The SCO2299 gene from Streptomyces coelicolor encodes a single peptide consisting of 497 amino acid residues. Its N-terminal region shows high amino acid sequence similarity to RNase HI, whereas its C-terminal region bears similarity to the CobC protein, which is involved in the synthesis of cobalamin. The SCO2299 gene suppressed a temperature-sensitive growth defect of an Escherichia coli RNase H-deficient strain, and the recombinant SCO2299 protein cleaved an RNA strand of RNA.DNA hybrid in vitro. The N-terminal domain of the SCO2299 protein, when overproduced independently, exhibited RNase H activity at a similar level to the full length protein. On the other hand, the C-terminal domain showed no CobC-like activity but an acid phosphatase activity. The full length protein also exhibited acid phosphatase activity at almost the same level as the C-terminal domain alone. These results indicate that RNase H and acid phosphatase activities of the full length SCO2299 protein depend on its N-terminal and C-terminal domains, respectively. The physiological functions of the SCO2299 gene and the relation between RNase H and acid phosphatase remain to be determined. However, the bifunctional enzyme examined here is a novel style in the Type 1 RNase H family. Additionally, S. coelicolor is the first example of an organism whose genome contains three active RNase H genes.

Acid Phosphatase↗

Reverse engineering of biochemical equations from time-course data by means of genetic programming.

Increased research aimed at simulating biological systems requires sophisticated parameter estimation methods. All current approaches, including genetic algorithms, need pre-existing equations to be functional. A generalized approach to predict not only parameters but also biochemical equations from only observable time-course information must be developed and a computational method to generate arbitrary equations without knowledge of biochemical reaction mechanisms must be developed. We present a technique to predict an equation using genetic programming. Our technique can search topology and numerical parameters of mathematical expression simultaneously. To improve the search ability of numeric constants, we added numeric mutation to the conventional procedure. As case studies, we predicted two equations of enzyme-catalyzed reactions regarding adenylate kinase and phosphofructokinase. Our numerical experimental results showed that our approach could obtain correct topology and parameters that were close to the originals. The mean errors between given and simulation-predicted time-courses were 1.6 x 10(-5)% and 2.0 x 10(-3)%, respectively. Our equation prediction approach can be applied to identify metabolic reactions from observable time-courses.

Algorithms↗

Cleavage of double-stranded RNA by RNase HI from a thermoacidophilic archaeon, Sulfolobus tokodaii 7.

ST0753, the orthologous gene of Type 1 RNase H found in a thermoacidophilic archaeon, Sulfolobus tokodaii, was analyzed. The recombinant ST0753 protein exhibited RNase H activity in both in vivo and in vitro assays. The protein expressed in an RNase H-deficient mutant Escherichia coli strain functioned to suppress the temperature-sensitive phenotype associated with the lack of RNase H. The in vitro characteristics of the gene's RNase H activity were similar to those of Halobacterium RNase HI, the first archaeal Type 1 RNase H to be characterized. Surprisingly, the S.tokodaii RNase HI cleaved not only the RNA strand of an RNA/DNA hybrid but also an RNA strand of an RNA/RNA duplex in the presence of Mn2+ or Co2+. The result of gel filtration column chromatography showed this double-stranded RNA-dependent RNase (dsRNase) activity was coincident with S.tokodaii RNase HI. A site-directed mutagenesis study of essential amino acids for RNase H activity indicated that this activity also affected dsRNase activity. A single amino acid replacement of Asp-125 by Asn resulted in loss of dsRNase activity but not RNase H activity, suggesting that amino acid residues required for dsRNase activity seemed slightly different from those of RNase H activity. Some reverse transcriptases from retroelements can cleave double-stranded RNA, and this activity requires the RNase H domain. Similarities in primary structure and biochemical characteristics between S.tokodaii RNase HI and reverse transcriptases imply that the S.tokodaii enzyme might be derived from the RNase H domain of reverse transcriptase.

Amino Acid Sequence↗

Toward large-scale modeling of the microbial cell for computer simulation.

In the post-genomic era, the large-scale, systematic, and functional analysis of all cellular components using transcriptomics, proteomics, and metabolomics, together with bioinformatics for the analysis of the massive amount of data generated by these "omics" methods are the focus of intensive research activities. As a consequence of these developments, systems biology, whose goal is to comprehend the organism as a complex system arising from interactions between its multiple elements, becomes a more tangible objective. Mathematical modeling of microorganisms and subsequent computer simulations are effective tools for systems biology, which will lead to a better understanding of the microbial cell and will have immense ramifications for biological, medical, environmental sciences, and the pharmaceutical industry. In this review, we describe various types of mathematical models (structured, unstructured, static, dynamic, etc.), of microorganisms that have been in use for a while, and others that are emerging. Several biochemical/cellular simulation platforms to manipulate such models are summarized and the E-Cell system developed in our laboratory is introduced. Finally, our strategy for building a "whole cell metabolism model", including the experimental approach, is presented.

Biotechnology↗

Identification of the first archaeal Type 1 RNase H gene from Halobacterium sp. NRC-1: archaeal RNase HI can cleave an RNA-DNA junction.

All the archaeal genomes sequenced to date contain a single Type 2 RNase H gene. We found that the genome of a halophilic archaeon, Halobacterium sp. NRC-1, contains an open reading frame with similarity to Type 1 RNase H. The protein encoded by the Vng0255c gene, possessed amino acid sequence identities of 33% with Escherichia coli RNase HI and 34% with a Bacillus subtilis RNase HI homologue. The B. subtilis RNase HI homologue, however, lacks amino acid sequences corresponding to a basic protrusion region of the E. coli RNase HI, and the Vng0255c has the similar deletion. As this deletion apparently conferred a complete loss of RNase H activity on the B. subtilis RNase HI homologue protein, the Vng0255c product was expected to exhibit no RNase H activity. However, the purified recombinant Vng0255c protein specifically cleaved an RNA strand of the RNA/DNA hybrid in vitro, and when the Vng0255c gene was expressed in an E. coli strain MIC2067 it could suppress the temperature-sensitive growth defect associated with the loss of RNase H enzymes of this strain. These results in vitro and in vivo strongly indicate that the Halobacterium Vng0255c is the first archaeal Type 1 RNase H. This enzyme, unlike other Type 1 RNases H, was able to cleave an Okazaki fragment-like substrate at the junction between the 3'-side of ribonucleotide and 5'-side of deoxyribonucleotide. It is likely that the archaeal Type 1 RNase H plays a role in the removal of the last ribonucleotide of the RNA primer from the Okazaki fragment during DNA replication.

Amino Acid Sequence↗

The 'weighted sum of relative entropy': a new index for synonymous codon usage bias.

Shannon entropy from information theory has been applied to estimate the degree of deviation from equal usage of synonymous codons; however, previous attempts have failed to take into account all three aspects of amino acid usage, i.e. (i) the number of distinct amino acids, (ii) their relative frequencies, and (iii) their degree of codon degeneracy. A new index taking into account all of these aspects is proposed. The index, designated as the 'weighted sum of relative entropy' (E(w)), is defined as the sum of the relative entropy of each amino acid weighted by its relative frequency in the sequence. In this paper, we demonstrate that E(w) allows us to avoid some amino acid usage biases and can yield results contradictory to those obtained by previous methods.

Algorithms↗

A new role for expressed pseudogenes as ncRNA: regulation of mRNA stability of its homologous coding gene.

We have earlier generated a mutant mouse in a course of making a transgenic line that exhibited interesting heterozygote phenotypes, which exhibited failure to thrive, severe bone deformities, and polycystic kidneys. This mutant mouse provided a clue to uncover a unique role of expressed pseudogenes. In this mutant the transgene was integrated into the vicinity of the expressing pseudogene of Makorin1 called Makorin1-p1. This insertion reduced transcription of the Makorin1-p1, resulting in destabilization of the Makorin1 mRNA in trans via a cis-acting RNA decay element within the 5' region of Makorin1 that is homologous between Makorin1 and Makorin1-p1. These findings demonstrate a novel and specific regulatory role of an expressed pseudogene as well as functional significance for noncoding RNAs. Next, we developed an original algorithm to determine how many pseudogenes are expressed. Based on our examination 2-3% of human processed pseudogenes are expressed using the most strict criteria. Interestingly, the mouse has a much smaller proportion of expressed pseudogenes (0.5-1%). Pseudogenes are functionally less constrained, and have accumulated more mutations than translated genes. If they have some functions in gene regulation, this property would allow more rapid functional diversification than protein-coding genes. In addition, some genetic phenomena that exhibit incomplete penetrance might be attributed to "mutation" or "variation" of pseudogenes.

Animals↗

A general computational model of mitochondrial metabolism in a whole organelle scale.

UNLABELLED: A computational tool for mitochondrial systems biology has been developed as a simulation model of E-Cell2, a publicly available simulation system. The general model consists of 58 enzymatic reactions and 117 metabolites, representing the respiratory chain, the TCA cycle, the fatty acid beta-oxidation and the inner-membrane transport system. It is based on previously published enzyme kinetics studies in the literature; we have successfully integrated and packaged them into a single large model. The model can be easily extended and modified so that mitochondrial biologists/physiologists can integrate their own models and evaluate them in the context of the whole organelle metabolism. AVAILABILITY: The mitochondrial model is bundled up with E-Cell2 simulation system, which can be downloaded from http://www.e-cell.org. CD-ROMs are also available and are distributed at major conferences. SUPPLEMENTARY INFORMATION: All the kinetic data are available via http://www.e-cell.org

Animals↗

A multi-algorithm, multi-timescale method for cell simulation.

MOTIVATION: Many important problems in cell biology require the dense nonlinear interactions between functional modules to be considered. The importance of computer simulation in understanding cellular processes is now widely accepted, and a variety of simulation algorithms useful for studying certain subsystems have been designed. Many of these are already widely used, and a large number of models constructed on these existing formalisms are available. A significant computational challenge is how we can integrate such sub-cellular models running on different types of algorithms to construct higher order models. RESULTS: A modular, object-oriented simulation meta-algorithm based on a discrete-event scheduler and Hermite polynomial interpolation has been developed and implemented. It is shown that this new method can efficiently handle many components driven by different algorithms and different timescales. The utility of this simulation framework is demonstrated further with a 'composite' heat-shock response model that combines the Gillespie-Gibson stochastic algorithm and deterministic differential equations. Dramatic improvements in performance were obtained without significant accuracy drawbacks. A multi-timescale demonstration of coupled harmonic oscillators is also shown.

Algorithms↗