Search PubMed⌕ Search

Biomedical subjects

Toshinori Endo

Publications and source records attributed to Toshinori Endo.

4 recordsLinked to original sources

Integrative annotation of 21,037 human genes validated by full-length cDNA clones.

The human genome sequence defines our inherent biological potential; the realization of the biology encoded therein requires knowledge of the function of each gene. Currently, our knowledge in this area is still limited. Several lines of investigation have been used to elucidate the structure and function of the genes in the human genome. Even so, gene prediction remains a difficult task, as the varieties of transcripts of a gene may vary to a great extent. We thus performed an exhaustive integrative characterization of 41,118 full-length cDNAs that capture the gene transcripts as complete functional cassettes, providing an unequivocal report of structural and functional diversity at the gene level. Our international collaboration has validated 21,037 human gene candidates by analysis of high-quality full-length cDNA clones through curation using unified criteria. This led to the identification of 5,155 new gene candidates. It also manifested the most reliable way to control the quality of the cDNA clones. We have developed a human gene database, called the H-Invitational Database (H-InvDB; http://www.h-invitational.jp/). It provides the following: integrative annotation of human genes, description of gene structures, details of novel alternative splicing isoforms, non-protein-coding RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein three-dimensional structure, mapping of known single nucleotide polymorphisms (SNPs), identification of polymorphic microsatellite repeats within human genes, and comparative results with mouse full-length cDNAs. The H-InvDB analysis has shown that up to 4% of the human genome sequence (National Center for Biotechnology Information build 34 assembly) may contain misassembled or missing regions. We found that 6.5% of the human gene candidates (1,377 loci) did not have a good protein-coding open reading frame, of which 296 loci are strong candidates for non-protein-coding RNA genes. In addition, among 72,027 uniquely mapped SNPs and insertions/deletions localized within human genes, 13,215 nonsynonymous SNPs, 315 nonsense SNPs, and 452 indels occurred in coding regions. Together with 25 polymorphic microsatellite repeats present in coding regions, they may alter protein structure, causing phenotypic effects or resulting in disease. The H-InvDB platform represents a substantial contribution to resources needed for the exploration of human biology and pathology.

Alternative Splicing↗

Standardized phylogenetic tree: a reference to discover functional evolution.

Functional evolution is often driven by positive natural selection. Although it is thought to be rare in evolution at the molecular level, its effects may be observed as the accelerated evolutionary rates. Therefore one of the effective ways to identify functional evolution is to identify accelerated evolution. Many methods have been developed to test the statistical significance of the accelerated evolutionary rate by comparison with the appropriate reference rate. The rates of synonymous substitution are one of the most useful and popular references, especially for large-scale analyses. On the other hand, these rates are applicable only to a limited evolutionary time period because they saturate quickly--i.e., multiple substitutions happen frequently because of the lower functional constraint. The relative rate test is an alternative method. This technique has an advantage in terms of the saturation effect but is not sufficiently powerful when the evolutionary rate differs considerably among phylogenetic lineages. For the aim to provide a universal reference tree, we propose a method to construct a standardized tree which serves as the reference for accelerated evolutionary rate. The method is based upon multiple molecular phylogenies of single genes with the aim of providing higher reliability. The tree has averaged and normalized branch lengths with standard deviations for statistical neutrality limits. The standard deviation also suggests the reliability level of the branch order. The resulting tree serves as a reference tree for the reliability level of the branch order and the test of evolutionary rate acceleration even when some of the species lineages show an accelerated evolutionary rate for most of their genes due to bottlenecking and other effects.

Animals↗

The genome sequence and structure of rice chromosome 1.

The rice species Oryza sativa is considered to be a model plant because of its small genome size, extensive genetic map, relative ease of transformation and synteny with other cereal crops. Here we report the essentially complete sequence of chromosome 1, the longest chromosome in the rice genome. We summarize characteristics of the chromosome structure and the biological insight gained from the sequence. The analysis of 43.3 megabases (Mb) of non-overlapping sequence reveals 6,756 protein coding genes, of which 3,161 show homology to proteins of Arabidopsis thaliana, another model plant. About 30% (2,073) of the genes have been functionally categorized. Rice chromosome 1 is (G + C)-rich, especially in its coding regions, and is characterized by several gene families that are dispersed or arranged in tandem repeats. Comparison with a draft sequence indicates the importance of a high-quality finished sequence.

Arabidopsis↗

Do introns favor or avoid regions of amino acid conservation?

Are intron positions correlated with regions of high amino acid conservation? For a set of ancient conserved proteins, with intronless prokaryotic but intron-containing eukaryotic homologs, multiple sequence alignments identified residues invariant throughout evolution. Intron positions between codons show no preferences. However, introns lying after the first base of a codon prefer conserved regions, markedly in glycines. Because glycines are in excess in conserved regions, this behavior could reflect phase-one introns entering glycine residues randomly in the ancestral sequences. Examination of intron positions within codons of evolutionarily invariable amino acids showed that roughly 50% of these introns are bordered by guanines at both 5'- and 3'-ends, 25% have a G only before the intron, and 5% have a G only after the intron, whereas about 20% are bordered by nonguanine bases.

Alternative Splicing↗