Search PubMed⌕ Search

Biomedical subjects

Takashi Gojobori

Publications and source records attributed to Takashi Gojobori.

At least 19 recordsLinked to original sources

Curated genome annotation of Oryza sativa ssp. japonica and comparative genome analysis with Arabidopsis thaliana.

We present here the annotation of the complete genome of rice Oryza sativa L. ssp. japonica cultivar Nipponbare. All functional annotations for proteins and non-protein-coding RNA (npRNA) candidates were manually curated. Functions were identified or inferred in 19,969 (70%) of the proteins, and 131 possible npRNAs (including 58 antisense transcripts) were found. Almost 5000 annotated protein-coding genes were found to be disrupted in insertional mutant lines, which will accelerate future experimental validation of the annotations. The rice loci were determined by using cDNA sequences obtained from rice and other representative cereals. Our conservative estimate based on these loci and an extrapolation suggested that the gene number of rice is approximately 32,000, which is smaller than previous estimates. We conducted comparative analyses between rice and Arabidopsis thaliana and found that both genomes possessed several lineage-specific genes, which might account for the observed differences between these species, while they had similar sets of predicted functional domains among the protein sequences. A system to control translational efficiency seems to be conserved across large evolutionary distances. Moreover, the evolutionary process of protein-coding genes was examined. Our results suggest that natural selection may have played a role for duplicated genes in both species, so that duplication was suppressed or favored in a manner that depended on the function of a gene.

Arabidopsis↗

Cancer-related mutations in BRCA1-BRCT cause long-range structural changes in protein-protein binding sites: a molecular dynamics study.

Cancer-associated mutations in the BRCT domain of BRCA1 (BRCA1-BRCT) abolish its tumor suppressor function by disrupting interactions with other proteins such as BACH1. Many cancer-related mutations do not cause sufficient destabilization to lead to global unfolding under physiological conditions, and thus abrogation of function probably is due to localized structural changes. To explore the reasons for mutation-induced loss of function, the authors performed molecular dynamics simulations on three cancer-associated mutants, A1708E, M1775R, and Y1853ter, and on the wild type and benign M1652I mutant, and compared the structures and fluctuations. Only the cancer-associated mutants exhibited significant backbone structure differences from the wild-type crystal structure in BACH1-binding regions, some of which are far from the mutation sites. Backbone differences of the A1708E mutant from the liganded wild type structure in these regions are much larger than those of the unliganded wild type X-ray or molecular dynamics structures. These BACH1-binding regions of the cancer-associated mutants also exhibited increases in their fluctuation magnitudes compared with the same regions in the wild type and M1562I mutant, as quantified by quasiharmonic analysis. Several of the regions of increased fluctuation magnitude correspond to correlated motions of residues in contact that provide a continuous path of fluctuating amino acids in contact from the A1708E and Y1853ter mutation sites to the BACH1-binding sites with altered structure and dynamics. The increased fluctuations in the disease-related mutants suggest an increase in vibrational entropy in the unliganded state that could result in a larger entropy loss in the disease-related mutants upon binding BACH1 than in the wild type. To investigate this possibility, vibrational entropies of the A1708E and wild type in the free state and bound to a BACH1-derived phosphopeptide were calculated using quasiharmonic analysis, to determine the binding entropy difference DeltaDeltaS between the A1708E mutant and the wild type. DeltaDeltaS was determined to be -4.0 cal mol(-1) K(-1), with an uncertainty of 2 cal mol(-1) K(-1); that is, the entropy loss upon binding the peptide is 4.0 cal mol(-1) K(-1) greater for the A1708E mutant, corresponding to an entropic contribution to the DeltaDeltaG of binding (-TDeltaDeltaS) 1.1 kcal mol(-1) more positive for the mutant. The observed differences in structure, flexibility, and entropy of binding likely are responsible for abolition of BACH1 binding, and illustrate that many disease- related mutations could have very long-range effects. The methods described here have potential for identifying correlated motions responsible for other long-range effects of deleterious mutations.

Algorithms↗

Rate of evolution in brain-expressed genes in humans and other primates.

Brain-expressed genes are known to evolve slowly in mammals. Nevertheless, since brains of higher primates have evolved rapidly, one might expect acceleration in DNA sequence evolution in their brain-expressed genes. In this study, we carried out full-length cDNA sequencing on the brain transcriptome of an Old World monkey (OWM) and then conducted three-way comparisons among (i) mouse, OWM, and human, and (ii) OWM, chimpanzee, and human. Although brain-expressed genes indeed appear to evolve more rapidly in species with more advanced brains (apes > OWM > mouse), a similar lineage effect is observable for most other genes. The broad inclusion of genes in the reference set to represent the genomic average is therefore critical to this type of analysis. Calibrated against the genomic average, the rate of evolution among brain-expressed genes is probably lower (or at most equal) in humans than in chimpanzee and OWM. Interestingly, the trend of slow evolution in coding sequence is no less pronounced among brain-specific genes, vis-à-vis brain-expressed genes in general. The human brain may thus differ from those of our close relatives in two opposite directions: (i) faster evolution in gene expression, and (ii) a likely slowdown in the evolution of protein sequences. Possible explanations and hypotheses are discussed.

Animals↗

Gene cluster analysis method identifies horizontally transferred genes with high reliability and indicates that they provide the main mechanism of operon gain in 8 species of gamma-Proteobacteria.

The formation mechanism of operons remains unresolved: operons may form by rearrangements within a genome or by acquisition of genes from other species, that is, horizontal gene transfer (HGT). One hindrance to its elucidation is the unavailability of a method to accurately identify HGT, although it is generally considered to occur. It is critically important first to select horizontally transferred (HT) genes reliably and then to determine the extent to which HGT is involved in operon formation. For this purpose, we considered indels in terms of gene clusters instead of individual genes and chose candidates of HT genes in 8 species of Escherichia, Shigella, and Salmonella based on the minimization of indels. To select a benchmark set of positively HT genes against which we can evaluate the candidate set, we devised another procedure using intergenetic alignments. Comparison with the benchmark set demonstrated the absence of a significant number of false positives in the candidate set, showing the high reliability of the method. Analyses of Escherichia coli K-12 operons revealed that although approximately 20 operons were probably gained from the last common ancestor of the 8 gamma-proteobacteria, deletion of intervening genes accounts for the formation of no operons, whereas horizontal transfer expanded 2 operons and introduced 4 entire operons. Based on these observations and reasoning, we suggest that the main mechanism of operon gain is HGT rather than intragenomic rearrangements. We propose that genes with related essential functions tend to reside in conserved operons, whereas genes in nonconserved operons mostly confer slight advantage to the organisms and frequently undergo horizontal transfer and decay. HT genes constitute at least 5.5% of the genes in the 8 species and approximately 45% of which originate from other gamma-proteobacteria. Genes involved in viral functions and mobile and extrachromosomal element functions are HT more often than expected. This finding indicates frequent mediation of HGT by bacteriophages. On the other hand, not only informational genes (those involved in transcription, translation, and related processes) but also operational genes (those involved in housekeeping) are HT less frequently than expected.

Cluster Analysis↗

A reduction in selective immune pressure during the course of chronic hepatitis C correlates with diminished biochemical evidence of hepatic inflammation.

It is considered that selection pressure exerted by the host immune response during early HCV infection might influence the outcome of that infection particularly as it relates to persistence or clearance of the agent. However, it is unclear whether positive selection pressure plays a role in determining the severity of hepatitis C during the course of persistent HCV infection. To address the evolutionary mechanism by which HCV escapes from the host immune response and to assess the relationship between viral evolution and hepatic inflammation, we determined 57 sequences (3-5 serial samples per patient) from 5 individuals with persistent HCV infection of genotype 1a who were under long-term follow-up ranging from 15.6 to 21.6 years. We applied a novel method to estimate serial alternations of selective pressure against the HCV enveloped region and compared this to fluctuation in transaminase level over time. Positive selection pressure was reduced over time postinfection, as evidenced by a reduction in nonsynonymous substitutions in the later phase of infection. Furthermore, serum transaminase, as a measure of inflammatory necrosis of hepatocytes, was reduced in parallel with decreased positive selection pressure. These results suggest that during persistent HCV infection, the virus faces diminished immune pressure over time, either from mutation to an immune resistant sequence or from immunologic exhaustion, and that this diminished immune attack is reflected in diminished inflammatory activity. This observation may be applicable to other viruses characterized by a slow rate of disease progression.

Aged↗

Exploration and grading of possible genes from 183 bacterial strains by a common protocol to identification of new genes: Gene Trek in Prokaryote Space (GTPS).

A large number of complete microorganism genomes has been sequenced and submitted to the public database and then incorporated into our complete genome database, Genome Information Broker (GIB, http://gib.genes.nig.ac.jp/). However, when comparative genomics is carried out, researchers must be aware that there are protein-coding genes not confirmed by homology or motif search and that reliable protein-coding genes are missing. Therefore, we developed a protocol (Gene Trek in Prokaryote Space, GTPS) for finding possible protein-coding genes in bacterial genomes. GTPS assigns a degree of reliability to predicted protein-coding genes. We first systematically applied the protocol to the complete genomes of all 123 bacterial species and strains that were publicly available as of July 2003, and then to those of 183 species and strains available as of September 2004. We found a number of incorrect genes and several new ones in the genome data in question. We also found a way to estimate the total number of orthologous genes in the bacterial world.

Bacteria↗

H-DBAS: alternative splicing database of completely sequenced and manually annotated full-length cDNAs based on H-Invitational.

The Human-transcriptome DataBase for Alternative Splicing (H-DBAS) is a specialized database of alternatively spliced human transcripts. In this database, each of the alternative splicing (AS) variants corresponds to a completely sequenced and carefully annotated human full-length cDNA, one of those collected for the H-Invitational human-transcriptome annotation meeting. H-DBAS contains 38,664 representative alternative splicing variants (RASVs) in 11,744 loci, in total. The data is retrievable by various features of AS, which were annotated according to manual annotations, such as by patterns of ASs, consequently invoked alternations in the encoded amino acids and affected protein motifs, GO terms, predicted subcellular localization signals and transmembrane domains. The database also records recently identified very complex patterns of AS, in which two distinct genes seemed to be bridged, nested or degenerated (multiple CDS): in all three cases, completely unrelated proteins are encoded by a single locus. By using AS Viewer, each AS event can be analyzed in the context of full-length cDNAs, enabling the user's empirical understanding of the relation between AS event and the consequent alternations in the encoded amino acid sequences together with various kinds of affected protein motifs. H-DBAS is accessible at http://jbirc.jbic.or.jp/h-dbas/.

Alternative Splicing↗

Frequent emergence and functional resurrection of processed pseudogenes in the human and mouse genomes.

Despite the wide distribution of processed pseudogenes in mammalian genomes, such as those of human and mouse, relatively little is known about their roles in genomic evolution. While gene duplications are recognized as one of the major driving forces in genome evolution, processed pseudogenes, which are retrotransposed copies of mRNAs, have been regarded as junk or selfish DNA for a long time. In order to elucidate the quantitative and qualitative contribution of processed pseudogenes to the mammalian genome evolution, we attempted to detect processed pseudogenes by extensively mapping the mRNAs to both the human and mouse genomes, and then we estimated the rate of their emergence. As a result, we revealed that the rate of pseudogene emergence was about 1-2% per gene per million years, which was as high as the rate (0.9%) of gene duplication in the human genome, although the rate of pseudogene emergence was found to drastically decrease in the hominid lineage. Furthermore, 1% of the processed pseudogenes seemed to be reinvigorated by post-retrotransposition transcription, many of them preserving the intact coding regions. Since the expression patterns of transcribed pseudogenes in various tissues were quite different between human and mouse, their emergence might have led to species-specific evolution. Our results indicate that the generation of processed pseudogenes was not wholly futile but instead has been an indispensable resource, driving dynamic evolution of the mammalian genomes.

Amino Acid Sequence↗

DDBJ working on evaluation and classification of bacterial genes in INSDC.

DNA Data Bank of Japan (DDBJ) (http://www.ddbj.nig.ac.jp) newly collected and released 12,927,184 entries or 13,787,688,598 bases in the period from July 2005 to June 2006. The released data contain honeybee expressed sequence tags (ESTs), re-examined and re-annotated complete genome data of Escherichia coli K-12 W3110, medaka WGS and human MGA. We also systematically evaluated and classified the genes in the complete bacterial genomes submitted to the International Nucleotide Sequence Database Collaboration (INSDC, http://insdc.org) that is composed of DDBJ, EMBL Bank and GenBank. The examination and classification selected 557,000 genes as reliable ones among all the bacterial genes predicted by us.

Animals↗

Degeneration after sexual differentiation in hydra and its relevance to the evolution of aging.

Aging occurs in most multicellular animals, yet some primitive animals do not show any sign of aging. This raises the following question: How have metazoans acquired the trait of aging in the course of evolution? Comparative studies of various species have provided a clue to this question by showing that sexually reproducing organisms predominantly undergo aging. The evolutionary theory "pleiotropy" also postulates aging as a price for facilitating the reproduction in the early life stage of an organism. For investigating the association between sexual reproduction and aging, a sexual phase-inducible organism in a laboratory would be suitable. One of such organisms is hydra, a genus of Cnidaria. Asexual hydra has been considered to be immortal, but there is the possibility that hydra undergoes aging after sexual reproduction. To search for signs of aging in hydra, we studied sexually differentiated Hydra oligactis at the individual and cellular levels. As a result, we found a significant decline in the capacities for food capture, contractile movements, and reproduction. More importantly, we discovered an exponential increase in the mortality rate of the population. These observations suggest that the degenerative process in H. oligactis represents the aging process. Furthermore, we found that the number of germ cells increased, whereas the number of somatic cells concomitantly decreased. The observed change of the cell composition is thus consistent with the "pleiotropy" theory of aging.

Animals↗

Radical amino acid change versus positive selection in the evolution of viral envelope proteins.

To detect positive selection in protein-coding sequence evolution, the ratio of the nonsynonymous to synonymous substitution rate (K(A)/K(S)) is commonly used. When this ratio is higher than 1, positive selection on nonsynonymous changes is considered to have occurred. However, the question of what kinds of amino acid change are likely to be involved in positive selection has not been well studied, though intuitively it seems that radical changes frequently occur in positively selected changes. To address this question, we examined chemically radical and conservative replacements in the evolution of hepatitis C virus (HCV) protein sequences. In the envelope region, 34 positively and 440 negatively selected sites were identified by the K(A)/K(S) ratio. Radical and conservative changes were compared between the two types of selected sites using two methods. First, the numbers of radical and conservative replacements were counted at the positively and negatively selected sites according to three kinds of chemical classifications. In all three classifications, the resulting ratios of the two numbers were not statistically different for the two types of selected sites (P>0.05). Second, the distribution of chemical changes was compared between the two types of selected sites using two kinds of chemical distances. The distributions of the two chemical distances were not statistically different for the two types of selected sites (P>0.05). These results indicate that the ratio of chemically radical and conservative changes is similar for positively and negatively selected sites in the envelope protein of HCV or, in other words, there is no correlation between radical change and positive selection in the evolution of this protein.

Amino Acid Substitution↗

Differential evolutionary rates of duplicated genes in protein interaction network.

In the network of protein-protein interactions (PPIs), a loss and gain of the partnering proteins can cause drastic changes of network formation during evolution. With the aim of examining the evolutionary effects of the loss and gain of the partnering proteins on PPIs, we examined a relationship between evolutionary rates and losses and/or gains of PPIs for duplicated gene pairs encoding proteins involved in the PPI network. For duplicated pairs, which provided us with a unique opportunity of making fair comparisons of the genes with the same initial condition, we found that the evolutionary rate of the protein with more PPI partners is much slower than that of the other with fewer PPI partners. Moreover, when the ratio of evolutionary rates (faster rate/slower rate) was computed for each of the duplicated pairs, the ratio for the duplicated pair sharing any PPI partners was significantly lower than that for the pair sharing no PPI partners. These results indicate that the duplicated gene pairs differentiate through the losses and/or gains of the PPI partners, resulting in a change in their evolutionary rates. In particular, we point out that the PPI losses for the duplicated gene products that are involved in the functional classes of 'transcription' and 'protein fate' have an impact on their evolutionary rates more than the PPI losses for others.

Biological Evolution↗

Large-scale identification and characterization of alternative splicing variants of human gene transcripts using 56,419 completely sequenced and manually annotated full-length cDNAs.

We report the first genome-wide identification and characterization of alternative splicing in human gene transcripts based on analysis of the full-length cDNAs. Applying both manual and computational analyses for 56,419 completely sequenced and precisely annotated full-length cDNAs selected for the H-Invitational human transcriptome annotation meetings, we identified 6877 alternative splicing genes with 18 297 different alternative splicing variants. A total of 37,670 exons were involved in these alternative splicing events. The encoded protein sequences were affected in 6005 of the 6877 genes. Notably, alternative splicing affected protein motifs in 3015 genes, subcellular localizations in 2982 genes and transmembrane domains in 1348 genes. We also identified interesting patterns of alternative splicing, in which two distinct genes seemed to be bridged, nested or having overlapping protein coding sequences (CDSs) of different reading frames (multiple CDS). In these cases, completely unrelated proteins are encoded by a single locus. Genome-wide annotations of alternative splicing, relying on full-length cDNAs, should lay firm groundwork for exploring in detail the diversification of protein function, which is mediated by the fast expanding universe of alternative splicing variants.

Alternative Splicing↗

TACT: Transcriptome Auto-annotation Conducting Tool of H-InvDB.

Transcriptome Auto-annotation Conducting Tool (TACT) is a newly developed web-based automated tool for conducting functional annotation of transcripts by the integration of sequence similarity searches and functional motif predictions. We developed the TACT system by integrating two kinds of similarity searches, FASTY and BLASTX, against protein sequence databases, UniProtKB (Swiss-Prot/TrEMBL) and RefSeq, and a unified motif prediction program, InterProScan, into the ORF-prediction pipeline originally designed for the 'H-Invitational' human transcriptome annotation project. This system successively applies these constituent programs to an mRNA sequence in order to predict the most plausible ORF and the function of the protein encoded. In this study, we applied the TACT system to 19 574 non-redundant human transcripts registered in H-InvDB and evaluated its predictive power by the degree of agreement with human-curated functional annotation in H-InvDB. As a result, the TACT system could assign functional description to 12 559 transcripts (64.2%), the remainder being hypothetical proteins. Furthermore, the overall agreement of functional annotation with H-InvDB, including those transcripts annotated as hypothetical proteins, was 83.9% (16 432/19 574). These results show that the TACT system is useful for functional annotation and that the prediction of ORFs and protein functions is highly accurate and close to the results of human curation. TACT is freely available at http://www.jbirc.aist.go.jp/tact/.

Amino Acid Motifs↗

Alternative splicing in human transcriptome: functional and structural influence on proteins.

Alternative splicing is a molecular mechanism that produces multiple proteins from a single gene, and is thought to produce variety in proteins translated from a limited number of genes. Here we analyzed how alternative splicing produced variety in protein structure and function, by using human full-length cDNAs on the assumption that all of the alternatively spliced mRNAs were translated to proteins. We found that the length of alternatively spliced amino acid sequences, in most cases, fell into a size shorter than that of average protein domain. We evaluated comprehensively the presumptive three-dimensional structures of the alternatively spliced products to assess the impact of alternative splicing on gene function. We found that more than half of the products encoded proteins which were involved in signal transduction, transcription and translation, and more than half of alternatively spliced regions comprised interaction sites between proteins and their binding partners, including substrates, DNA/RNA, and other proteins. Intriguingly, 67% of the alternatively spliced isoforms showed significant alterations to regions of the protein structural core, which likely resulted in large conformational change. Based on those findings, we speculate that there are a large number of cases that alternative splicing modulates protein networks through significant alteration in protein conformation.

Alternative Splicing↗

Rapid evolution of major histocompatibility complex class I genes in primates generates new disease alleles in humans via hitchhiking diversity.

A plausible explanation for many MHC-linked diseases is lacking. Sequencing of the MHC class I region (coding units or full contigs) in several human and nonhuman primate haplotypes allowed an analysis of single nucleotide variations (SNV) across this entire segment. This diversity was not evenly distributed. It was rather concentrated within two gene-rich clusters. These were each centered, but importantly not limited to, the antigen-presenting HLA-A and HLA-B/-C loci. Rapid evolution of MHC-I alleles, as evidenced by an unusually high number of haplotype-specific (hs) and hypervariable (hv) (which could not be traced to a single species or haplotype) SNVs within the classical MHC-I, seems to have not only hitchhiked alleles within nearby genes, but also hitchhiked deleterious mutations in these same unrelated loci. The overrepresentation of a fraction of these hvSNV (hv1SNV) along with hsSNV, as compared to those that appear to have been maintained throughout primate evolution (trans-species diversity; tsSNV; included within hv2SNV) tends to establish that the majority of the MHC polymorphism is de novo (species specific). This is most likely reminiscent of the fact that these hsSNV and hv1SNV have been selected in adaptation to the constantly evolving microbial antigenic repertoire.

Alleles↗

magp4 gene may contribute to the diversification of cichlid morphs and their speciation.

Lake Victoria harbors more than 300 species of cichlid fish, which are adapted to a variety of ecological niches with various morphological species-specific features. However, it is believed that these species arose explosively within the last 14,000 years and transcripts among Lake Victoria cichlid species are almost identical in sequence. These data prompted us to develop a DNA chip assay to compare patterns of gene expression among cichlid species. We prepared a DNA chip spotted with 6240 elements derived from cichlid expressed sequence tag (EST) clones and successfully characterized gene expression differences between the cichlid species Haplochromis chilotes and Haplochromis sp. "rockkribensis". We identified 14 transcripts that were differentially expressed between these species at an early developmental stage, 15 days post-fertilization (dpf), and several were further analyzed using quantitative real-time PCR (qPCR). One of these differentially expressed transcripts was a homolog of microfibril-associated glycoprotein 4 (magp4), a putative causative gene for the human inherited disease, Smith-Magenis syndrome (SMS), for which facial defects are among the phenotypic features. Further analysis of magp4 expression showed that magp4 was expressed in the jaw portion of cichlid fry and that expression profiles between Haplochromis chilotes and Haplochromis sp. "rockkribensis" differed during development. These data suggest that the differential expression of a gene associated with human cranial morphogenesis may be involved in the diversification of cichlid jaw morphs.

Animals↗

The evolutionary rate of a protein is influenced by features of the interacting partners.

Rates of protein evolution are thought to be influenced by features of protein-protein interaction (PPI). However, the most important features of interaction for determining the evolutionary rate are poorly understood. Here, we consider four categories for PPIs in Saccharomyces cerevisiae. Properties we consider are the extent to which proteins interact with proteins of the same function or different function (DF) and the extent to which these interactions involve connections in the dense part or sparse part (SP) of a PPI network. Our findings are that proteins with DF-SP interactions evolve at the slowest rate of all the proteins examined.

Cluster Analysis↗