Search PubMed⌕ Search

Biomedical subjects

Kei Yura

Publications and source records attributed to Kei Yura.

12 recordsLinked to original sources

Coverage of whole proteome by structural genomics observed through protein homology modeling database.

We have been developing FAMSBASE, a protein homology-modeling database of whole ORFs predicted from genome sequences. The latest update of FAMSBASE ( http://daisy.nagahama-i-bio.ac.jp/Famsbase/ ), which is based on the protein three-dimensional (3D) structures released by November 2003, contains modeled 3D structures for 368,724 open reading frames (ORFs) derived from genomes of 276 species, namely 17 archaebacterial, 130 eubacterial, 18 eukaryotic and 111 phage genomes. Those 276 genomes are predicted to have 734,193 ORFs in total and the current FAMSBASE contains protein 3D structure of approximately 50% of the ORF products. However, cases that a modeled 3D structure covers the whole part of an ORF product are rare. When portion of an ORF with 3D structure is compared in three kingdoms of life, in archaebacteria and eubacteria, approximately 60% of the ORFs have modeled 3D structures covering almost the entire amino acid sequences, however, the percentage falls to about 30% in eukaryotes. When annual differences in the number of ORFs with modeled 3D structure are calculated, the fraction of modeled 3D structures of soluble protein for archaebacteria is increased by 5%, and that for eubacteria by 7% in the last 3 years. Assuming that this rate would be maintained and that determination of 3D structures for predicted disordered regions is unattainable, whole soluble protein model structures of prokaryotes without the putative disordered regions will be in hand within 15 years. For eukaryotic proteins, they will be in hand within 25 years. The 3D structures we will have at those times are not the 3D structure of the entire proteins encoded in single ORFs, but the 3D structures of separate structural domains. Measuring or predicting spatial arrangements of structural domains in an ORF will then be a coming issue of structural genomics.

Amino Acid Sequence↗

Amino acid residue doublet propensity in the protein-RNA interface and its application to RNA interface prediction.

Protein-RNA interactions play essential roles in a number of regulatory mechanisms for gene expression such as RNA splicing, transport, translation and post-transcriptional control. As the number of available protein-RNA complex 3D structures has increased, it is now possible to statistically examine protein-RNA interactions based on 3D structures. We performed computational analyses of 86 representative protein-RNA complexes retrieved from the Protein Data Bank. Interface residue propensity, a measure of the relative importance of different amino acid residues in the RNA interface, was calculated for each amino acid residue type (residue singlet interface propensity). In addition to the residue singlet propensity, we introduce a new residue-based propensity, which gives a measure of residue pairing preferences in the RNA interface of a protein (residue doublet interface propensity). The residue doublet interface propensity contains much more information than the sum of two singlet propensities alone. The prediction of the RNA interface using the two types of propensities plus a position-specific multiple sequence profile can achieve a specificity of about 80%. The prediction method was then applied to the 3D structure of two mRNA export factors, TAP (Mex67) and UAP56 (Sub2). The prediction enables us to point out candidate RNA interfaces, part of which are consistent with previous experimental studies and may contribute to elucidation of atomic mechanisms of mRNA export.

Amino Acids↗

Large-scale identification and characterization of alternative splicing variants of human gene transcripts using 56,419 completely sequenced and manually annotated full-length cDNAs.

We report the first genome-wide identification and characterization of alternative splicing in human gene transcripts based on analysis of the full-length cDNAs. Applying both manual and computational analyses for 56,419 completely sequenced and precisely annotated full-length cDNAs selected for the H-Invitational human transcriptome annotation meetings, we identified 6877 alternative splicing genes with 18 297 different alternative splicing variants. A total of 37,670 exons were involved in these alternative splicing events. The encoded protein sequences were affected in 6005 of the 6877 genes. Notably, alternative splicing affected protein motifs in 3015 genes, subcellular localizations in 2982 genes and transmembrane domains in 1348 genes. We also identified interesting patterns of alternative splicing, in which two distinct genes seemed to be bridged, nested or having overlapping protein coding sequences (CDSs) of different reading frames (multiple CDS). In these cases, completely unrelated proteins are encoded by a single locus. Genome-wide annotations of alternative splicing, relying on full-length cDNAs, should lay firm groundwork for exploring in detail the diversification of protein function, which is mediated by the fast expanding universe of alternative splicing variants.

Alternative Splicing↗

Alternative splicing in human transcriptome: functional and structural influence on proteins.

Alternative splicing is a molecular mechanism that produces multiple proteins from a single gene, and is thought to produce variety in proteins translated from a limited number of genes. Here we analyzed how alternative splicing produced variety in protein structure and function, by using human full-length cDNAs on the assumption that all of the alternatively spliced mRNAs were translated to proteins. We found that the length of alternatively spliced amino acid sequences, in most cases, fell into a size shorter than that of average protein domain. We evaluated comprehensively the presumptive three-dimensional structures of the alternatively spliced products to assess the impact of alternative splicing on gene function. We found that more than half of the products encoded proteins which were involved in signal transduction, transcription and translation, and more than half of alternatively spliced regions comprised interaction sites between proteins and their binding partners, including substrates, DNA/RNA, and other proteins. Intriguingly, 67% of the alternatively spliced isoforms showed significant alterations to regions of the protein structural core, which likely resulted in large conformational change. Based on those findings, we speculate that there are a large number of cases that alternative splicing modulates protein networks through significant alteration in protein conformation.

Alternative Splicing↗

Newly sequenced eRF1s from ciliates: the diversity of stop codon usage and the molecular surfaces that are important for stop codon interactions.

The genetic code of nuclear genes in some ciliates was found to differ from that of other organisms in the assignment of UGA, UAG, and UAA codons, which are normally assigned as stop codons. In some ciliate species, the universal stop codons UAA and UAG instead encode glutamine. In some other ciliates, the universal stop codon UGA appears to be translated as cysteine or tryptophan. Eukaryotic release factor 1 (eRF1) is a key protein in stop codon recognition, thus, the protein is believed to play an important role in the stop codon reassignment in ciliates. We have cloned, sequenced, and analyzed the cDNA of eRF1 from four ciliate species of three different classes: Karyorelictea (Loxodes striatus), Heterotrichea (Blepharisma musculus), and Litostomatea (Didinium nasutum, Dileptus margaritifer). Phylogenetic analysis of these eRF1s supports the hypothesis that the genetic code in ciliates has deviated independently several times from the universal genetic code, and that different ciliate eRF1s may have undergone different processes to change the codon specificity. Using computational methods, we have also suggested areas on the surface of eRF1s that are important for stop codon recognition in ciliate eRF1s.

Amino Acid Sequence↗

Integrative annotation of 21,037 human genes validated by full-length cDNA clones.

The human genome sequence defines our inherent biological potential; the realization of the biology encoded therein requires knowledge of the function of each gene. Currently, our knowledge in this area is still limited. Several lines of investigation have been used to elucidate the structure and function of the genes in the human genome. Even so, gene prediction remains a difficult task, as the varieties of transcripts of a gene may vary to a great extent. We thus performed an exhaustive integrative characterization of 41,118 full-length cDNAs that capture the gene transcripts as complete functional cassettes, providing an unequivocal report of structural and functional diversity at the gene level. Our international collaboration has validated 21,037 human gene candidates by analysis of high-quality full-length cDNA clones through curation using unified criteria. This led to the identification of 5,155 new gene candidates. It also manifested the most reliable way to control the quality of the cDNA clones. We have developed a human gene database, called the H-Invitational Database (H-InvDB; http://www.h-invitational.jp/). It provides the following: integrative annotation of human genes, description of gene structures, details of novel alternative splicing isoforms, non-protein-coding RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein three-dimensional structure, mapping of known single nucleotide polymorphisms (SNPs), identification of polymorphic microsatellite repeats within human genes, and comparative results with mouse full-length cDNAs. The H-InvDB analysis has shown that up to 4% of the human genome sequence (National Center for Biotechnology Information build 34 assembly) may contain misassembled or missing regions. We found that 6.5% of the human gene candidates (1,377 loci) did not have a good protein-coding open reading frame, of which 296 loci are strong candidates for non-protein-coding RNA genes. In addition, among 72,027 uniquely mapped SNPs and insertions/deletions localized within human genes, 13,215 nonsynonymous SNPs, 315 nonsense SNPs, and 452 indels occurred in coding regions. Together with 25 polymorphic microsatellite repeats present in coding regions, they may alter protein structure, causing phenotypic effects or resulting in disease. The H-InvDB platform represents a substantial contribution to resources needed for the exploration of human biology and pathology.

Alternative Splicing↗

Structure and function of a family 10 beta-xylanase chimera of Streptomyces olivaceoviridis E-86 FXYN and Cellulomonas fimi Cex.

The catalytic domain of xylanases belonging to glycoside hydrolase family 10 (GH10) can be divided into 22 modules (M1 to M22; Sato, Y., Niimura, Y., Yura, K., and Go, M. (1999) Gene (Amst.) 238, 93-101). Inspection of the crystal structure of a GH10 xylanase from Streptomyces olivaceoviridis E-86 (SoXyn10A) revealed that the catalytic domain of GH10 xylanases can be dissected into two parts, an N-terminal larger region and C-terminal smaller region, by the substrate binding cleft, corresponding to the module border between M14 and M15. It has been suggested that the topology of the substrate binding clefts of GH10 xylanases are not conserved (Charnock, S. J., Spurway, T. D., Xie, H., Beylot, M. H., Virden, R., Warren, R. A. J., Hazlewood, G. P., and Gilbert, H. J. (1998) J. Biol. Chem. 273, 32187-32199). To facilitate a greater understanding of the structure-function relationship of the substrate binding cleft of GH10 xylanases, a chimeric xylanase between SoXyn10A and Xyn10A from Cellulomonas fimi (CfXyn10A) was constructed, and the topology of the hybrid substrate binding cleft established. At the three-dimensional level, SoXyn10A and CfXyn10A appear to possess 5 subsites, with the amino acid residues comprising subsites -3 to +1 being well conserved, although the +2 subsites are quite different. Biochemical analyses of the chimeric enzyme along with SoXyn10A and CfXyn10A indicated that differences in the structure of subsite +2 influence bond cleavage frequencies and the catalytic efficiency of xylooligosaccharide hydrolysis. The hybrid enzyme constructed in this study displays fascinating biochemistry, with an interesting combination of properties from the parent enzymes, resulting in a low production of xylose.

Catalytic Domain↗

Het-PDB Navi.: a database for protein-small molecule interactions.

The genomes of more than 100 species have been sequenced, and the biological functions of encoded proteins are now actively being researched. Protein function is based on interactions between proteins and other molecules. One approach to assuming protein function based on genomic sequence is to predict interactions between an encoded protein and other molecules. As a data source for such predictions, knowledge regarding known protein-small molecule interactions needs to be compiled. We have, therefore, surveyed interactions between proteins and other molecules in Protein Data Bank (PDB), the protein three-dimensional (3D) structure database. Among 20,685 entries in PDB (April, 2003), 4,189 types of small molecules were found to interact with proteins. Biologically relevant small molecules most often found in PDB were metal ions, such as calcium, zinc, and magnesium. Sugars and nucleotides were the next most common. These molecules are known to act as cofactors for enzymes and/or stabilizers of proteins. In each case of interactions between a protein and small molecule, we found preferred amino acid residues at the interaction sites. These preferences can be the basis for predicting protein function from genomic sequence and protein 3D structures. The data pertaining to these small molecules were collected in a database named Het-PDB Navi., which is freely available at http://daisy.nagahama-i-bio.ac.jp/golab/hetpdbnavi.html and linked to the official PDB home page.

Adenosine Triphosphate↗

Novel types of two-domain multi-copper oxidases: possible missing links in the evolution.

An analysis of the genome sequence database revealed novel types of two-domain multi-copper oxidases. The two-domain proteins have the conspicuous combination of blue-copper and inter-domain trinuclear copper binding residues, which is common in ceruloplasmin and ascorbate oxidase but not in nitrite reductase, and therefore are considered to retain the characteristics of the plausible ancestral form of ceruloplasmin and ascorbate oxidase. A possible evolutionary relationship of these proteins is proposed.

Amino Acid Sequence↗

Enlarged FAMSBASE: protein 3D structure models of genome sequences for 41 species.

Enlarged FAMSBASE is a relational database of comparative protein structure models for the whole genome of 41 species, presented in the GTOP database. The models are calculated by Full Automatic Modeling System (FAMS). Enlarged FAMSBASE provides a wide range of query keys, such as name of ORF (open reading frame), ORF keywords, Protein Data Bank (PDB) ID, PDB heterogen atoms and sequence similarity. Heterogen atoms in PDB include cofactors, ligands and other factors that interact with proteins, and are a good starting point for analyzing interactions between proteins and other molecules. The data may also work as a template for drug design. The present number of ORFs with protein 3D models in FAMSBASE is 183 805, and the database includes an average of three models for each ORF. FAMSBASE is available at http://famsbase.bio.nagoya-u.ac.jp/famsbase/.

Animals↗

Highly divergent actins from karyorelictean, heterotrich, and litostome ciliates.

We have cloned, sequenced, and characterized cDNA of actins from five ciliate species of three different classes of the phylum Ciliophora: Karyorelictea (Loxodes striatus), Heterotrichea (Blepharisma japonicum, Blepharisma musculus), and Litostomatea (Didinium nasutum, Dileptus margaritifer). Loxodes striatus uses UGA as the stop codon and has numerous in-frame UAA and UAG, which are translated into glutamine. The other four species use UAA as the stop codon and have no in-frame UAG nor UGA. The putative amino acid sequences of the newly determined actin genes were found to be highly divergent as expected from previous findings of other ciliate actins. These sequences were also highly divergent from other ciliate actins, indicating that actin genes are highly diverse even within the phylum Ciliophora. Phylogenetic analysis showed high evolutionary rate of ciliate actins. Our results suggest that the evolutionary rate was accelerated because of the differences in molecular interactions.

Actins↗