Search PubMed⌕ Search

Biomedical subjects

Tadashi Imanishi

Publications and source records attributed to Tadashi Imanishi.

At least 19 recordsLinked to original sources

Curated genome annotation of Oryza sativa ssp. japonica and comparative genome analysis with Arabidopsis thaliana.

We present here the annotation of the complete genome of rice Oryza sativa L. ssp. japonica cultivar Nipponbare. All functional annotations for proteins and non-protein-coding RNA (npRNA) candidates were manually curated. Functions were identified or inferred in 19,969 (70%) of the proteins, and 131 possible npRNAs (including 58 antisense transcripts) were found. Almost 5000 annotated protein-coding genes were found to be disrupted in insertional mutant lines, which will accelerate future experimental validation of the annotations. The rice loci were determined by using cDNA sequences obtained from rice and other representative cereals. Our conservative estimate based on these loci and an extrapolation suggested that the gene number of rice is approximately 32,000, which is smaller than previous estimates. We conducted comparative analyses between rice and Arabidopsis thaliana and found that both genomes possessed several lineage-specific genes, which might account for the observed differences between these species, while they had similar sets of predicted functional domains among the protein sequences. A system to control translational efficiency seems to be conserved across large evolutionary distances. Moreover, the evolutionary process of protein-coding genes was examined. Our results suggest that natural selection may have played a role for duplicated genes in both species, so that duplication was suppressed or favored in a manner that depended on the function of a gene.

Arabidopsis↗

Cancer-related mutations in BRCA1-BRCT cause long-range structural changes in protein-protein binding sites: a molecular dynamics study.

Cancer-associated mutations in the BRCT domain of BRCA1 (BRCA1-BRCT) abolish its tumor suppressor function by disrupting interactions with other proteins such as BACH1. Many cancer-related mutations do not cause sufficient destabilization to lead to global unfolding under physiological conditions, and thus abrogation of function probably is due to localized structural changes. To explore the reasons for mutation-induced loss of function, the authors performed molecular dynamics simulations on three cancer-associated mutants, A1708E, M1775R, and Y1853ter, and on the wild type and benign M1652I mutant, and compared the structures and fluctuations. Only the cancer-associated mutants exhibited significant backbone structure differences from the wild-type crystal structure in BACH1-binding regions, some of which are far from the mutation sites. Backbone differences of the A1708E mutant from the liganded wild type structure in these regions are much larger than those of the unliganded wild type X-ray or molecular dynamics structures. These BACH1-binding regions of the cancer-associated mutants also exhibited increases in their fluctuation magnitudes compared with the same regions in the wild type and M1562I mutant, as quantified by quasiharmonic analysis. Several of the regions of increased fluctuation magnitude correspond to correlated motions of residues in contact that provide a continuous path of fluctuating amino acids in contact from the A1708E and Y1853ter mutation sites to the BACH1-binding sites with altered structure and dynamics. The increased fluctuations in the disease-related mutants suggest an increase in vibrational entropy in the unliganded state that could result in a larger entropy loss in the disease-related mutants upon binding BACH1 than in the wild type. To investigate this possibility, vibrational entropies of the A1708E and wild type in the free state and bound to a BACH1-derived phosphopeptide were calculated using quasiharmonic analysis, to determine the binding entropy difference DeltaDeltaS between the A1708E mutant and the wild type. DeltaDeltaS was determined to be -4.0 cal mol(-1) K(-1), with an uncertainty of 2 cal mol(-1) K(-1); that is, the entropy loss upon binding the peptide is 4.0 cal mol(-1) K(-1) greater for the A1708E mutant, corresponding to an entropic contribution to the DeltaDeltaG of binding (-TDeltaDeltaS) 1.1 kcal mol(-1) more positive for the mutant. The observed differences in structure, flexibility, and entropy of binding likely are responsible for abolition of BACH1 binding, and illustrate that many disease- related mutations could have very long-range effects. The methods described here have potential for identifying correlated motions responsible for other long-range effects of deleterious mutations.

Algorithms↗

H-DBAS: alternative splicing database of completely sequenced and manually annotated full-length cDNAs based on H-Invitational.

The Human-transcriptome DataBase for Alternative Splicing (H-DBAS) is a specialized database of alternatively spliced human transcripts. In this database, each of the alternative splicing (AS) variants corresponds to a completely sequenced and carefully annotated human full-length cDNA, one of those collected for the H-Invitational human-transcriptome annotation meeting. H-DBAS contains 38,664 representative alternative splicing variants (RASVs) in 11,744 loci, in total. The data is retrievable by various features of AS, which were annotated according to manual annotations, such as by patterns of ASs, consequently invoked alternations in the encoded amino acids and affected protein motifs, GO terms, predicted subcellular localization signals and transmembrane domains. The database also records recently identified very complex patterns of AS, in which two distinct genes seemed to be bridged, nested or degenerated (multiple CDS): in all three cases, completely unrelated proteins are encoded by a single locus. By using AS Viewer, each AS event can be analyzed in the context of full-length cDNAs, enabling the user's empirical understanding of the relation between AS event and the consequent alternations in the encoded amino acid sequences together with various kinds of affected protein motifs. H-DBAS is accessible at http://jbirc.jbic.or.jp/h-dbas/.

Alternative Splicing↗

Frequent emergence and functional resurrection of processed pseudogenes in the human and mouse genomes.

Despite the wide distribution of processed pseudogenes in mammalian genomes, such as those of human and mouse, relatively little is known about their roles in genomic evolution. While gene duplications are recognized as one of the major driving forces in genome evolution, processed pseudogenes, which are retrotransposed copies of mRNAs, have been regarded as junk or selfish DNA for a long time. In order to elucidate the quantitative and qualitative contribution of processed pseudogenes to the mammalian genome evolution, we attempted to detect processed pseudogenes by extensively mapping the mRNAs to both the human and mouse genomes, and then we estimated the rate of their emergence. As a result, we revealed that the rate of pseudogene emergence was about 1-2% per gene per million years, which was as high as the rate (0.9%) of gene duplication in the human genome, although the rate of pseudogene emergence was found to drastically decrease in the hominid lineage. Furthermore, 1% of the processed pseudogenes seemed to be reinvigorated by post-retrotransposition transcription, many of them preserving the intact coding regions. Since the expression patterns of transcribed pseudogenes in various tissues were quite different between human and mouse, their emergence might have led to species-specific evolution. Our results indicate that the generation of processed pseudogenes was not wholly futile but instead has been an indispensable resource, driving dynamic evolution of the mammalian genomes.

Amino Acid Sequence↗

Large-scale identification and characterization of alternative splicing variants of human gene transcripts using 56,419 completely sequenced and manually annotated full-length cDNAs.

We report the first genome-wide identification and characterization of alternative splicing in human gene transcripts based on analysis of the full-length cDNAs. Applying both manual and computational analyses for 56,419 completely sequenced and precisely annotated full-length cDNAs selected for the H-Invitational human transcriptome annotation meetings, we identified 6877 alternative splicing genes with 18 297 different alternative splicing variants. A total of 37,670 exons were involved in these alternative splicing events. The encoded protein sequences were affected in 6005 of the 6877 genes. Notably, alternative splicing affected protein motifs in 3015 genes, subcellular localizations in 2982 genes and transmembrane domains in 1348 genes. We also identified interesting patterns of alternative splicing, in which two distinct genes seemed to be bridged, nested or having overlapping protein coding sequences (CDSs) of different reading frames (multiple CDS). In these cases, completely unrelated proteins are encoded by a single locus. Genome-wide annotations of alternative splicing, relying on full-length cDNAs, should lay firm groundwork for exploring in detail the diversification of protein function, which is mediated by the fast expanding universe of alternative splicing variants.

Alternative Splicing↗

TACT: Transcriptome Auto-annotation Conducting Tool of H-InvDB.

Transcriptome Auto-annotation Conducting Tool (TACT) is a newly developed web-based automated tool for conducting functional annotation of transcripts by the integration of sequence similarity searches and functional motif predictions. We developed the TACT system by integrating two kinds of similarity searches, FASTY and BLASTX, against protein sequence databases, UniProtKB (Swiss-Prot/TrEMBL) and RefSeq, and a unified motif prediction program, InterProScan, into the ORF-prediction pipeline originally designed for the 'H-Invitational' human transcriptome annotation project. This system successively applies these constituent programs to an mRNA sequence in order to predict the most plausible ORF and the function of the protein encoded. In this study, we applied the TACT system to 19 574 non-redundant human transcripts registered in H-InvDB and evaluated its predictive power by the degree of agreement with human-curated functional annotation in H-InvDB. As a result, the TACT system could assign functional description to 12 559 transcripts (64.2%), the remainder being hypothetical proteins. Furthermore, the overall agreement of functional annotation with H-InvDB, including those transcripts annotated as hypothetical proteins, was 83.9% (16 432/19 574). These results show that the TACT system is useful for functional annotation and that the prediction of ORFs and protein functions is highly accurate and close to the results of human curation. TACT is freely available at http://www.jbirc.aist.go.jp/tact/.

Amino Acid Motifs↗

Alternative splicing in human transcriptome: functional and structural influence on proteins.

Alternative splicing is a molecular mechanism that produces multiple proteins from a single gene, and is thought to produce variety in proteins translated from a limited number of genes. Here we analyzed how alternative splicing produced variety in protein structure and function, by using human full-length cDNAs on the assumption that all of the alternatively spliced mRNAs were translated to proteins. We found that the length of alternatively spliced amino acid sequences, in most cases, fell into a size shorter than that of average protein domain. We evaluated comprehensively the presumptive three-dimensional structures of the alternatively spliced products to assess the impact of alternative splicing on gene function. We found that more than half of the products encoded proteins which were involved in signal transduction, transcription and translation, and more than half of alternatively spliced regions comprised interaction sites between proteins and their binding partners, including substrates, DNA/RNA, and other proteins. Intriguingly, 67% of the alternatively spliced isoforms showed significant alterations to regions of the protein structural core, which likely resulted in large conformational change. Based on those findings, we speculate that there are a large number of cases that alternative splicing modulates protein networks through significant alteration in protein conformation.

Alternative Splicing↗

The Rice Annotation Project Database (RAP-DB): hub for Oryza sativa ssp. japonica genome information.

With the completion of the rice genome sequencing, a standardized annotation is necessary so that the information from the genome sequence can be fully utilized in understanding the biology of rice and other cereal crops. An annotation jamboree was held in Japan with the aim of annotating and manually curating all the genes in the rice genome. Here we present the Rice Annotation Project Database (RAP-DB), which has been developed to provide access to the annotation data. The RAP-DB has two different types of annotation viewers, BLAST and BLAT search, and other useful features. By connecting the annotations to other rice genomics data, such as full-length cDNAs and Tos17 mutant lines, the RAP-DB serves as a hub for rice genomics. All of the resources can be accessed through http://rapdb.lab.nig.ac.jp/.

Databases, Nucleic Acid↗

A web tool for comparative genomics: G-compass.

In order to assist the progression of comparative genomics, we have developed a new web-based tool, named G-compass, for browsing and analysis of genome alignments. G-compass utilizes 829,311 pieces of genome alignments between human and mouse that were originally produced for this tool. The quality of the genome alignment set was evaluated by using several statistics. As a result, the alignment set is found to cover approximately 17% of the human genome and 82% of the annotated exons. The averages of nucleotide sequence identity and sequence length are 71.2% and 673.6 bp, respectively. In comparison with public data, it appeared that our data is more expansive and possesses greater genome coverage. G-compass incorporates unique functions such as window analysis of individual alignments. Furthermore, with G-compass and the joint help of H-InvDB, we were able to find highly conserved genomic segments and a human specific antisense transcript candidate, demonstrating that G-compass is useful for facilitating biological discoveries. G-compass is publicly accessible on the WWW at http://www.jbirc.aist.go.jp/g-compass/.

Animals↗

Investigation of protein functions through data-mining on integrated human transcriptome database, H-Invitational database (H-InvDB).

H-Invitational Database (H-InvDB; ) is a human transcriptome database, containing integrative annotation of 41,118 full-length cDNA clones originated from 21,037 loci. H-InvDB is a product of the H-Invitational project, an international collaboration to systematically and functionally validate human genes by analysis of a unique set of high quality full-length cDNA clones using automatic annotation and human curation under unified criteria. Here, 19,574 proteins encoded by these cDNAs were classified into 11,709 function-known and 7865 function-unknown hypothetical proteins by similarity with protein databases and motif prediction (InterProScan). The proportion of "hypothetical proteins" in H-InvDB was as high as 40.4%. In this study, we thus conducted data-mining in H-InvDB with the aim of assigning advanced functional annotations to those hypothetical proteins. First, by data-mining in the H-InvDB version of GTOP, we identified 337 SCOP domains within 7865 H-Inv hypothetical proteins. Second, by data-mining of predicted subcellular localization by SOSUI and TMHMM in H-InvDB, we found 1032 transmembrane proteins within H-Inv hypothetical proteins. These results clearly demonstrate that structural prediction is effective for functional annotation of proteins with unknown functions. All the data in H-InvDB are shown in two main views, the cDNA view and the Locus view, and five auxiliary databases with web-based viewers; DiseaseInfo Viewer, H-ANGEL, Clustering Viewer, G-integra and TOPO Viewer; the data also are provided as flat files and XML files. The data consists of descriptions of their gene structures, novel alternative splicing isoforms, functional RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein 3D structure, mapping of SNPs and microsatellite repeat motifs in relation with orphan diseases, gene expression profiling, and comparisons with mouse full-length cDNAs in the context of molecular evolution. This unique integrative platform for conducting in silico data-mining represents a substantial contribution to resources required for the exploration of human biology and pathology.

Amino Acid Sequence↗

Whole genome association study of rheumatoid arthritis using 27 039 microsatellites.

A major goal of current human genome-wide studies is to identify the genetic basis of complex disorders. However, the availability of an unbiased, reliable, cost efficient and comprehensive methodology to analyze the entire genome for complex disease association is still largely lacking or problematic. Therefore, we have developed a practical and efficient strategy for whole genome association studies of complex diseases by charting the human genome at 100 kb intervals using a collection of 27,039 microsatellites and the DNA pooling method in three successive genomic screens of independent case-control populations. The final step in our methodology consists of fine mapping of the candidate susceptible DNA regions by single nucleotide polymorphisms (SNPs) analysis. This approach was validated upon application to rheumatoid arthritis, a destructive joint disease affecting up to 1% of the population. A total of 47 candidate regions were identified. The top seven loci, withstanding the most stringent statistical tests, were dissected down to individual genes and/or SNPs on four chromosomes, including the previously known 6p21.3-encoded Major Histocompatibility Complex gene, HLA-DRB1. Hence, microsatellite-based genome-wide association analysis complemented by end stage SNP typing provides a new tool for genetic dissection of multifactorial pathologies including common diseases.

Arthritis, Rheumatoid↗

Comparative genomics of bidirectional gene pairs and its implications for the evolution of a transcriptional regulation system.

Arrangement of genes in the human genome was not considered to be ordered like those of prokaryotes, as in many cases genes appeared to be randomly distributed across the genome. However, by focusing on the closely located adjacent gene pairs, it was recently suggested that the bidirectional pairs were enriched in the human genome and these pairs tended to be coexpressed by sharing promoter sequences. We compared this biased organization found in the human genome with those in the genomes of nine other eukaryotes to reveal when and how the biased organization had evolved using a total of 122,945 adjacent gene pairs. As a result, we found that the biased organization was found only in mammals, and not in other eukaryotes. Interestingly, we found that many of these genes in the bidirectional arrangement were not mammalian specific genes but conserved among various animals. Further analyses revealed that the bidirectional arrangement of these pairs had arisen by utilizing already-existing genes in the lineage leading to mammals recently, no earlier than the vertebrate-ascidian divergence. Since the novel bidirectional arrangement could result in novel co-regulated transcription, our results here provide evidence that shows how a transcriptional regulation system has evolved through changes in the genome organization, especially in the lineage leading to humans.

Animals↗

Length variation of CAG/CAA triplet repeats in 50 genes among 16 inbred mouse strains.

CAG repeats coding for poly-glutamines have been studied by many groups as repeat length variations contributes to differences in protein function and disease outcome. In this study, we systematically searched public databases for genes carrying CAG repeats. For the genes obtained, we experimentally analyzed variations of length and the purity of the repeats in 62 loci among 16 inbred mouse strains, including wild-derived and laboratory strains. We found that length was conserved in 50% of the loci, especially among wild-derived strains. Of 496 polymorphic repeat alleles, 78% were uninterrupted and 22% were interrupted with non-CAG codons. Interruptions tended to occur in longer repeats and all repeats of greater length than 23 were interrupted. Although interruptions can act as suppressors for the expansion of CAG repeats, we found that the occurrence of the interruptions depended on the length of the CAG repeats. Furthermore, most poly-glutamines examined in this study existed in human orthologous genes, reflecting the functional significance of poly-glutamines in proteins.

Alleles↗

The Human Anatomic Gene Expression Library (H-ANGEL), the H-Inv integrative display of human gene expression across disparate technologies and platforms.

The Human Anatomic Gene Expression Library (H-ANGEL) is a resource for information concerning the anatomical distribution and expression of human gene transcripts. The tool contains protein expression data from multiple platforms that has been associated with both manually annotated full-length cDNAs from H-InvDB and RefSeq sequences. Of the H-Inv predicted genes, 18 897 have associated expression data generated by at least one platform. H-ANGEL utilizes categorized mRNA expression data from both publicly available and proprietary sources. It incorporates data generated by three types of methods from seven different platforms. The data are provided to the user in the form of a web-based viewer with numerous query options. H-ANGEL is updated with each new release of cDNA and genome sequence build. In future editions, we will incorporate the capability for expression data updates from existing and new platforms. H-ANGEL is accessible at http://www.jbirc.aist.go.jp/hinv/h-angel/.

Database Management Systems↗

Novel algorithm for automated genotyping of microsatellites.

Microsatellites or short tandem repeats (STRs) are abundant in the human genome with easily assayed polymorphisms, providing powerful genetic tools for mapping both Mendelian and complex traits. Microsatellite genotyping requires detection of the products of polymerase chain reaction (PCR) amplification by electrophoresis, and analysis of the peak data for discrimination of the true allele. A high-throughput genotyping approach requires computer-based automation at both the detection and analysis phases. In order to achieve this, complicated peak patterns from individual alleles must be interpreted in order to assign alleles. Previous methods consider limited types of noise peaks and cannot provide enough accuracy. By pattern recognition of various types of noise peaks, such as stutter peaks and additional peaks, we have achieved an overall average accuracy of 94% for allele calling in our actual data. Our algorithm is crucial for a high-throughput genotyping system for microsatellite markers by reducing manual editing and human errors.

Algorithms↗

Integrative annotation of 21,037 human genes validated by full-length cDNA clones.

The human genome sequence defines our inherent biological potential; the realization of the biology encoded therein requires knowledge of the function of each gene. Currently, our knowledge in this area is still limited. Several lines of investigation have been used to elucidate the structure and function of the genes in the human genome. Even so, gene prediction remains a difficult task, as the varieties of transcripts of a gene may vary to a great extent. We thus performed an exhaustive integrative characterization of 41,118 full-length cDNAs that capture the gene transcripts as complete functional cassettes, providing an unequivocal report of structural and functional diversity at the gene level. Our international collaboration has validated 21,037 human gene candidates by analysis of high-quality full-length cDNA clones through curation using unified criteria. This led to the identification of 5,155 new gene candidates. It also manifested the most reliable way to control the quality of the cDNA clones. We have developed a human gene database, called the H-Invitational Database (H-InvDB; http://www.h-invitational.jp/). It provides the following: integrative annotation of human genes, description of gene structures, details of novel alternative splicing isoforms, non-protein-coding RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein three-dimensional structure, mapping of known single nucleotide polymorphisms (SNPs), identification of polymorphic microsatellite repeats within human genes, and comparative results with mouse full-length cDNAs. The H-InvDB analysis has shown that up to 4% of the human genome sequence (National Center for Biotechnology Information build 34 assembly) may contain misassembled or missing regions. We found that 6.5% of the human gene candidates (1,377 loci) did not have a good protein-coding open reading frame, of which 296 loci are strong candidates for non-protein-coding RNA genes. In addition, among 72,027 uniquely mapped SNPs and insertions/deletions localized within human genes, 13,215 nonsynonymous SNPs, 315 nonsense SNPs, and 452 indels occurred in coding regions. Together with 25 polymorphic microsatellite repeats present in coding regions, they may alter protein structure, causing phenotypic effects or resulting in disease. The H-InvDB platform represents a substantial contribution to resources needed for the exploration of human biology and pathology.

Alternative Splicing↗

Characteristic beta-globin gene cluster haplotypes of Evenkis and Oroqens in north China.

Haplotype frequencies of the beta-globin gene cluster were estimated for 114 Evenkis and 81 Oroqens from northeast China, and their characteristics were compared with those in Japanese, Koreans, and three Colombian Amerindian groups of South America (Wayuu, Kamsa, and Inga tribes). A major 5' subhaplotype (5' to the delta-globin gene) was + - - - - in Evenkis, whereas + - - - -, - + + - +, and - + - + + were the major subhaplotypes in Oroqens. One possible candidate for an ancestral 5' subhaplotype, - - - - -, was found in one Evenki (0.5%) and three Oroqen chromosomes (2.0%). They were observed as heterozygous forms for + ---- and -----. Major haplotypes were +-----+, + -----+-, and + - - - - + + in Evenkis, whereas they were +-----+,-++-+-+, +----+-, and -+-++-+ in Oroqens. The lowest Nei's genetic distance values of Evenkis or Oroqens based on the 5' subhaplotype frequency distributions were observed in relation to the Wayuu or Koreans, respectively, but those of Evenkis and Oroqens based on the haplotype frequency distributions were found in relation to Koreans.

Asian People↗