Search PubMed⌕ Search

Biomedical subjects

Mitiko Go

Publications and source records attributed to Mitiko Go.

13 recordsLinked to original sources

Coverage of whole proteome by structural genomics observed through protein homology modeling database.

We have been developing FAMSBASE, a protein homology-modeling database of whole ORFs predicted from genome sequences. The latest update of FAMSBASE ( http://daisy.nagahama-i-bio.ac.jp/Famsbase/ ), which is based on the protein three-dimensional (3D) structures released by November 2003, contains modeled 3D structures for 368,724 open reading frames (ORFs) derived from genomes of 276 species, namely 17 archaebacterial, 130 eubacterial, 18 eukaryotic and 111 phage genomes. Those 276 genomes are predicted to have 734,193 ORFs in total and the current FAMSBASE contains protein 3D structure of approximately 50% of the ORF products. However, cases that a modeled 3D structure covers the whole part of an ORF product are rare. When portion of an ORF with 3D structure is compared in three kingdoms of life, in archaebacteria and eubacteria, approximately 60% of the ORFs have modeled 3D structures covering almost the entire amino acid sequences, however, the percentage falls to about 30% in eukaryotes. When annual differences in the number of ORFs with modeled 3D structure are calculated, the fraction of modeled 3D structures of soluble protein for archaebacteria is increased by 5%, and that for eubacteria by 7% in the last 3 years. Assuming that this rate would be maintained and that determination of 3D structures for predicted disordered regions is unattainable, whole soluble protein model structures of prokaryotes without the putative disordered regions will be in hand within 15 years. For eukaryotic proteins, they will be in hand within 25 years. The 3D structures we will have at those times are not the 3D structure of the entire proteins encoded in single ORFs, but the 3D structures of separate structural domains. Measuring or predicting spatial arrangements of structural domains in an ORF will then be a coming issue of structural genomics.

Amino Acid Sequence↗

Large-scale identification and characterization of alternative splicing variants of human gene transcripts using 56,419 completely sequenced and manually annotated full-length cDNAs.

We report the first genome-wide identification and characterization of alternative splicing in human gene transcripts based on analysis of the full-length cDNAs. Applying both manual and computational analyses for 56,419 completely sequenced and precisely annotated full-length cDNAs selected for the H-Invitational human transcriptome annotation meetings, we identified 6877 alternative splicing genes with 18 297 different alternative splicing variants. A total of 37,670 exons were involved in these alternative splicing events. The encoded protein sequences were affected in 6005 of the 6877 genes. Notably, alternative splicing affected protein motifs in 3015 genes, subcellular localizations in 2982 genes and transmembrane domains in 1348 genes. We also identified interesting patterns of alternative splicing, in which two distinct genes seemed to be bridged, nested or having overlapping protein coding sequences (CDSs) of different reading frames (multiple CDS). In these cases, completely unrelated proteins are encoded by a single locus. Genome-wide annotations of alternative splicing, relying on full-length cDNAs, should lay firm groundwork for exploring in detail the diversification of protein function, which is mediated by the fast expanding universe of alternative splicing variants.

Alternative Splicing↗

Alternative splicing in human transcriptome: functional and structural influence on proteins.

Alternative splicing is a molecular mechanism that produces multiple proteins from a single gene, and is thought to produce variety in proteins translated from a limited number of genes. Here we analyzed how alternative splicing produced variety in protein structure and function, by using human full-length cDNAs on the assumption that all of the alternatively spliced mRNAs were translated to proteins. We found that the length of alternatively spliced amino acid sequences, in most cases, fell into a size shorter than that of average protein domain. We evaluated comprehensively the presumptive three-dimensional structures of the alternatively spliced products to assess the impact of alternative splicing on gene function. We found that more than half of the products encoded proteins which were involved in signal transduction, transcription and translation, and more than half of alternatively spliced regions comprised interaction sites between proteins and their binding partners, including substrates, DNA/RNA, and other proteins. Intriguingly, 67% of the alternatively spliced isoforms showed significant alterations to regions of the protein structural core, which likely resulted in large conformational change. Based on those findings, we speculate that there are a large number of cases that alternative splicing modulates protein networks through significant alteration in protein conformation.

Alternative Splicing↗

Survey of conserved alternative splicing events of mRNAs encoding SR proteins in land plants.

The serine/arginine-rich (SR) protein family plays an important role in constitutive and alternative splicing (AS). These proteins regulate AS in a tissue-specific and stress-responsive manner. Pre-mRNAs encoding SR proteins are often alternatively spliced, and these AS events may be important for the regulation of AS events of other pre-mRNAs. In this study, we analyzed AS events of SR proteins in Arabidopsis thaliana and Oryza sativa (rice). We found three sets of AS events conserved between Arabidopsis and rice. These conserved AS events were found in the plant-novel-SR protein, SC35-like (SCL), and two-Zn-knuckles-type 9G8 subfamilies. Each member of these subfamilies has at least one RNA recognition motif (RRM) and at least one intron in the RRM-encoded region. We found that the conserved AS events occurred in these introns and, in each case, the conserved AS events resulted in mature mRNAs encoding proteins with incomplete RRMs. To search for the evolutionary origin of these AS events, we analyzed SR proteins in Physcomitrella patens (moss) in addition to those in Arabidopsis and rice. We found moss homologues of the plant-novel-SR protein, SCL, and the two-Zn-knuckles-type 9G8 subfamilies in silico, and these homologues have long introns at the same location of the conserved AS sites in Arabidopsis and rice. Such long introns are quite specific for alternatively spliced introns concerning the Arabidopsis SR protein genes. The long introns found in the moss SR protein genes strongly suggested that conserved AS events in moss SR protein genes might be similar to those in Arabidopsis and rice. We traced the evolutionary origin of the conserved AS events to 400 MYA, when plants first invaded land. These events are likely important in the regulation of whole AS events and likely contribute to the complicated transcriptome described by AS. The complicated transcriptome created by regulated AS events might have provided plants tolerance against droughts or temperature shifts and given them the ability to live on land.

Alternative Splicing↗

An empirical approach for detecting nucleotide-binding sites on proteins.

Protein structure data in the PDB (Protein Data Bank) were used to construct empirical scores of nucleotide-protein interactions. A simple strategy to evaluate the spatial distribution of protein atoms around the base moieties of nucleotides was applied to categorize adenine, guanine, nicotinamide and flavin nucleotide-binding sites. In addition to the known nucleotide-binding motifs, the empirical scores detected several other features that were shared among proteins with different folds. The empirical scores were also used to predict the binding sites on protein molecules and a comprehensive test of the prediction system was performed. As a result, adenine, guanine, nicotinamide and flavin sites were detected with efficiencies of 31, 29, 32 and 40%, respectively. The predictions were judged to be successful if the predicted base with the best score was located within a 3.0 A r.m.s.d. from the known ligand positions.

Amino Acid Motifs↗

Disulfide linkages and a three-dimensional structure model of the extracellular ligand-binding domain of guanylyl cyclase C.

Guanylyl cyclase C (GC-C) is a single-transmembrane receptor that is specifically activated by endogenous ligands, including guanylin, and the exogenous ligand, heat-stable enterotoxin. Using combined HPLC separation and MS analysis techniques the positions of the disulfide linkages in the extracellular ligand-binding domain (ECD) of GC-C were determined to be between Cys7-Cys94, Cys72-Cys77, Cys101-Cys128 and Cys179-Cys226. Furthermore, a three-dimensional structural model of the ECD was constructed by homology modeling, using the structure of the ECD of GC-A as a template (van den Akker et al., 2000, Nature, 406: 101-104) and the information of the disulfide linkages. Although the GC-C model was similar to the known structure of GC-A, importantly its ligand-binding site appears to be located on the quite different region from that in GC-A.

Amino Acid Sequence↗

Role of KaiC phosphorylation in the circadian clock system of Synechococcus elongatus PCC 7942.

In the cyanobacterium Synechococcus elongatus PCC 7942, KaiA, KaiB, and KaiC are essential proteins for the generation of a circadian rhythm. KaiC is proposed as a negative regulator of the circadian expression of all genes in the genome, and its phosphorylation is regulated positively by KaiA and negatively by KaiB and shows a circadian rhythm in vivo. To study the functions of KaiC phosphorylation in the circadian clock system, we identified two autophosphorylation sites, Ser-431 and Thr-432, by using mass spectrometry (MS). We generated Synechococcus mutants in which these residues were substituted for alanine by using site-directed mutagenesis. Phosphorylation of KaiC was reduced in the single mutants and was completely abolished in the double mutant, indicating that KaiC is also phosphorylated at these sites in vivo. These mutants lost circadian rhythm, indicating that phosphorylation at each of the two sites is essential for the control of the circadian oscillation. Although the nonphosphorylatable mutant KaiC was able to form a hexamer in vitro, it failed to form a clock protein complex with KaiA, KaiB, and SasA in the Synechococcus cells. When nonphosphorylatable KaiC was overexpressed, the kaiBC promoter activity was only transiently repressed. These results suggest that KaiC phosphorylation regulates its transcriptional repression activity by controlling its binding affinity for other clock proteins.

Bacterial Proteins↗

Structure and function of a family 10 beta-xylanase chimera of Streptomyces olivaceoviridis E-86 FXYN and Cellulomonas fimi Cex.

The catalytic domain of xylanases belonging to glycoside hydrolase family 10 (GH10) can be divided into 22 modules (M1 to M22; Sato, Y., Niimura, Y., Yura, K., and Go, M. (1999) Gene (Amst.) 238, 93-101). Inspection of the crystal structure of a GH10 xylanase from Streptomyces olivaceoviridis E-86 (SoXyn10A) revealed that the catalytic domain of GH10 xylanases can be dissected into two parts, an N-terminal larger region and C-terminal smaller region, by the substrate binding cleft, corresponding to the module border between M14 and M15. It has been suggested that the topology of the substrate binding clefts of GH10 xylanases are not conserved (Charnock, S. J., Spurway, T. D., Xie, H., Beylot, M. H., Virden, R., Warren, R. A. J., Hazlewood, G. P., and Gilbert, H. J. (1998) J. Biol. Chem. 273, 32187-32199). To facilitate a greater understanding of the structure-function relationship of the substrate binding cleft of GH10 xylanases, a chimeric xylanase between SoXyn10A and Xyn10A from Cellulomonas fimi (CfXyn10A) was constructed, and the topology of the hybrid substrate binding cleft established. At the three-dimensional level, SoXyn10A and CfXyn10A appear to possess 5 subsites, with the amino acid residues comprising subsites -3 to +1 being well conserved, although the +2 subsites are quite different. Biochemical analyses of the chimeric enzyme along with SoXyn10A and CfXyn10A indicated that differences in the structure of subsite +2 influence bond cleavage frequencies and the catalytic efficiency of xylooligosaccharide hydrolysis. The hybrid enzyme constructed in this study displays fascinating biochemistry, with an interesting combination of properties from the parent enzymes, resulting in a low production of xylose.

Catalytic Domain↗

Het-PDB Navi.: a database for protein-small molecule interactions.

The genomes of more than 100 species have been sequenced, and the biological functions of encoded proteins are now actively being researched. Protein function is based on interactions between proteins and other molecules. One approach to assuming protein function based on genomic sequence is to predict interactions between an encoded protein and other molecules. As a data source for such predictions, knowledge regarding known protein-small molecule interactions needs to be compiled. We have, therefore, surveyed interactions between proteins and other molecules in Protein Data Bank (PDB), the protein three-dimensional (3D) structure database. Among 20,685 entries in PDB (April, 2003), 4,189 types of small molecules were found to interact with proteins. Biologically relevant small molecules most often found in PDB were metal ions, such as calcium, zinc, and magnesium. Sugars and nucleotides were the next most common. These molecules are known to act as cofactors for enzymes and/or stabilizers of proteins. In each case of interactions between a protein and small molecule, we found preferred amino acid residues at the interaction sites. These preferences can be the basis for predicting protein function from genomic sequence and protein 3D structures. The data pertaining to these small molecules were collected in a database named Het-PDB Navi., which is freely available at http://daisy.nagahama-i-bio.ac.jp/golab/hetpdbnavi.html and linked to the official PDB home page.

Adenosine Triphosphate↗

Distinct interaction of versican/PG-M with hyaluronan and link protein.

The proteoglycan aggregate is the major structural component of the cartilage matrix, comprising hyaluronan (HA), link protein (LP), and a large chondroitin sulfate (CS) proteoglycan, aggrecan. Here, we found that another member of aggrecan family, versican, biochemically binds to both HA and LP. Functional analyses of recombinant looped domains (subdomains) A, B, and B' of the N-terminal G1 domain revealed that the B-B' segment of versican is adequate for binding to HA and LP, whereas A and B-B' of aggrecan bound to LP and HA, respectively. BIAcore trade mark analyses showed that the A subdomain of versican G1 enhances HA binding but has a negligible effect on LP binding. Overlay sensorgrams demonstrated that versican G1 or its B-B' segment forms a complex with both HA and LP. We generated a molecular model of the B-B' segment, in which a deletion and an insertion of B' and B are critical for stable structure and HA binding. These results provide important insights into the mechanisms of formation of the proteoglycan aggregate and HA binding of molecules containing the link module.

Amino Acid Sequence↗

Chondroitin sulfate synthase-2. Molecular cloning and characterization of a novel human glycosyltransferase homologous to chondroitin sulfate glucuronyltransferase, which has dual enzymatic activities.

Chondroitin sulfate is found in a variety of tissues as proteoglycans and consists of repeating disaccharide units of N-acetylgalactosamine and glucuronic acid residues with sulfate residues at various places. We found a novel human gene (GenBank accession number AB086063) that possesses a sequence homologous with the human chondroitin sulfate glucuronyltransferase gene which we recently cloned and characterized. The full-length open reading frame encodes a typical type II membrane protein comprising 775 amino acids. The protein had a domain containing beta 3-glycosyltransferase motif but lacked a typical beta 4-glycosyltransferase motif, which is the same as chondroitin sulfate glucuronyltransferase, whereas chondroitin synthase had both domains. The putative catalytic domain was expressed in COS-7 cells as a soluble enzyme. Surprisingly, both glucuronyltransferase and N-acetylgalactosaminyltransferase activities were observed when chondroitin, chondroitin sulfate, and their oligosaccharides were used as the acceptor substrates. The reaction products were identified to have the linkage of GlcUA beta 1-3GalNAc and GalNAc beta 1-4GlcUA at the non-reducing terminus of chondroitin for glucuronyltransferase activity and N-acetylgalactosaminyltransferase activity, respectively. Quantitative real time PCR analysis revealed that the transcripts were ubiquitously expressed in various human tissues but highly expressed in the pancreas, ovary, placenta, small intestine, and stomach. These results indicate that this enzyme could synthesize chondroitin sulfate chains as a chondroitin sulfate synthase that has both glucuronyltransferase and N-acetylgalactosaminyltransferase activities. Sequence analysis based on three-dimensional structure revealed the presence of not typical but significant beta 4-glycosyltransferase architecture.

Amino Acid Motifs↗

Enlarged FAMSBASE: protein 3D structure models of genome sequences for 41 species.

Enlarged FAMSBASE is a relational database of comparative protein structure models for the whole genome of 41 species, presented in the GTOP database. The models are calculated by Full Automatic Modeling System (FAMS). Enlarged FAMSBASE provides a wide range of query keys, such as name of ORF (open reading frame), ORF keywords, Protein Data Bank (PDB) ID, PDB heterogen atoms and sequence similarity. Heterogen atoms in PDB include cofactors, ligands and other factors that interact with proteins, and are a good starting point for analyzing interactions between proteins and other molecules. The data may also work as a template for drug design. The present number of ORFs with protein 3D models in FAMSBASE is 183 805, and the database includes an average of three models for each ORF. FAMSBASE is available at http://famsbase.bio.nagoya-u.ac.jp/famsbase/.

Animals↗