Search PubMed⌕ Search

Biomedical subjects

Rintaro Saito

Publications and source records attributed to Rintaro Saito.

At least 19 recordsLinked to original sources

Noise-reduction filtering for accurate detection of replication termini in bacterial genomes.

Bacterial chromosomes are highly polarized in their nucleotide composition through mutational selection related to replication. Using compositional skews such as the GC skew, replication origin and terminus can be predicted in silico by observing the shift points. However, the genome sequence is affected by myriad functional requirements and selection on numerous subgenomic features, and elimination of this "noise" should lead to better predictions. Here, we present a noise-reduction approach that uses low-pass filtering through Fast Fourier transform coupled with cumulative skew graphs. It increases the prediction accuracy of the replication termini compared with previously documented methods based on genomic base composition.

Bacteria↗

Inferring rules of Escherichia coli translational efficiency using an artificial neural network.

Although the machinery for translation initiation in Escherichia coli is very complicated, the translational efficiency has been reported to be predictable from upstream oligonucleotide sequences. Conventional models have difficulties in their generalization ability and prediction nonlinearity and in their ability to deal with a variety of input attributions. To address these issues, we employed structural learning by artificial neural networks to infer general rules for translational efficiency. The correlation between translational activities measured by biological experiments and those predicted by our method in the test data was significant (r=0.78), and our method uncovered underlying rules of translational activities and sequence patterns from the obtained skeleton structure. The significant rules for predicting translational efficiency were (1) G- and A-rich oligonucleotide sequences, resembling the Shine-Dalgarno sequence, at positions -10 to -7; (2) first base A in the initiation codon; (3) transport/binding or amino acid metabolism gene function; (4) high binding energy between mRNA and 16S rRNA at positions -15 to -5. An additional inferred novel rule was that C at position -1 increases translational efficiency. When our model was applied to the entire genomic sequence of E. coli, translational activities of genes for metabolism and translational were significantly high.

Base Sequence↗

Comparative analysis of base correlations in 5' untranslated regions of various species.

Translational initiation signals, such as Shine-Dalgarno (SD) sequences in bacteria and Kozak consensus sequences in vertebrates, direct ribosomes to initiate protein synthesis from mRNAs. Investigating sequence characteristics of these signals is important, particularly to infer translational initiation mechanisms. Although various statistical analyses of translational initiation signals have been done, few have focused on base correlations that assess base dependencies in the signal sequences. We used relative entropy and mutual information to analyze base conservation and correlation, respectively, in the 5' UTRs of various species. In eukaryotes, we found peaks of relative entropy at -3 from the translational start site but no peak of mutual information at that position, indicating that the base at that position (known as the core base of the Kozak sequence) is well conserved but not correlated with neighboring bases and thus functions as a single base. We observed unexpected peaks of mutual information between positions -2 and -1 in most eukaryotes. Surprisingly these base correlation also occurred in some bacteria and archaea, although there were no base preferences at neither position. Various dinucleotide patterns existed at these positions, and the correlation between bases at -2 and -1 may be relevant to the context of translational initiation. Because dinucleotide patterns of correlated pairs of nucleotides at -2 and -1 were not unique within respective organisms, the correlation could not be found when analyzing single-nucleotide conservation. Therefore, mutual information allowed us to discover signals that were not found by simply analyzing base conservation.

5' Untranslated Regions↗

Large-scale identification of protein-protein interaction of Escherichia coli K-12.

Protein-protein interactions play key roles in protein function and the structural organization of a cell. A thorough description of these interactions should facilitate elucidation of cellular activities, targeted-drug design, and whole cell engineering. A large-scale comprehensive pull-down assay was performed using a His-tagged Escherichia coli ORF clone library. Of 4339 bait proteins tested, partners were found for 2667, including 779 of unknown function. Proteins copurifying with hexahistidine-tagged baits on a Ni2+-NTA column were identified by MALDI-TOF MS (matrix-assisted laser desorption ionization time of flight mass spectrometry). An extended analysis of these interacting networks by bioinformatics and experimentation should provide new insights and novel strategies for E. coli systems biology.

Escherichia coli K12↗

Prediction of non-coding and antisense RNA genes in Escherichia coli with Gapped Markov Model.

A new mathematical index was developed to identify and characterize non-coding RNA (ncRNA) genes encoded within the Escherichia coli (E. coli) genome. It was designated the GMMI (Gapped Markov Model Index) and used to evaluate sequence patterns located at the separate positions of consensus sequences, codon biases and/or possible RNA structures on the basis of the Markov model. The GMMI was able to separate a set of known mRNA sequences from a mixture of ncRNAs including tRNAs and rRNAs. Consequently, the GMMI was employed to predict novel ncRNA candidates. At the beginning, possible transcription units were extracted from the E. coli genome using consensus sequences for the sigma70 promoter and the rho-independent terminator. Then, these units were evaluated by using the GMMI. This identified 133 candidate ncRNAs, which contain 29 previously annotated small RNA genes and 46 possible antisense ncRNAs. Furthermore 12 transcripts (including five antisense RNAs) were confirmed according to the expression analysis. These data suggests that the expression of small antisense RNAs might be more common than previously thought in the E. coli genome.

Computational Biology↗

Computational analysis of microRNA targets in Caenorhabditis elegans.

MicroRNAs (miRNAs) are endogenous approximately 22-nucleotide (nt) non-coding RNAs that post-transcriptionally regulate the expression of target genes via hybridization to target mRNA. Using known pairs of miRNA and target mRNA in Caenorhabditis elegans, we first performed computational analysis for specific hybridization patterns between these two RNAs. We counted the numbers of perfectly complementary dinucleotide sequences and calculated the free energy within complementary base pairs of each dinucleotide, observed by sliding a 2-nt window along all nucleotides of the miRNA-mRNA duplex. We confirmed not only strong base pairing within the 5' region of miRNAs (nts 1-8) in C. elegans, but also the required mismatch within the central region (nt 9 or nt 10), and we found weak binding within the 3' region (nts 13-14). We also predicted 687 possible miRNA target transcripts, many of which are thought to be involved in C. elegans development, by combining the above mentioned hybridization tendency with the following analyses: (1) prediction of the miRNA-mRNA duplex with free-energy minimization; (2) identification of the complementary pattern within the miRNA-mRNA duplex; (3) conservation of target sites between C. elegans and C. briggsae, a related soil nematode; and (4) extraction of mRNA candidates with multiple target sites. Rigorous tests using shuffled miRNA controls supported these predictions. Our results suggest that miRNAs recognize their target mRNAs by their hybridization pattern and that many target mRNAs may be regulated through a combination of several specific miRNA target sites in C. elegans.

3' Untranslated Regions↗

A problem in multivariate analysis of codon usage data and a possible solution.

Multivariate analyses are often used to identify major trends of variation in synonymous codon usage among genes. These analyses need to be performed on properly normalized codon usage data to avoid biases masking this synonymous variation, i.e., gene length, amino acid usage, and codon degeneracy; however, previous studies have failed to do so. In this paper, we demonstrate that the use of alternative normalized data (called 'relative adaptiveness' in the literature) can avoid all these biases and furthermore, can identify more trends of variation among genes, including GC-ending codon usage, GT-ending codon usage, and gene expression level.

Amino Acids↗

The 'weighted sum of relative entropy': a new index for synonymous codon usage bias.

Shannon entropy from information theory has been applied to estimate the degree of deviation from equal usage of synonymous codons; however, previous attempts have failed to take into account all three aspects of amino acid usage, i.e. (i) the number of distinct amino acids, (ii) their relative frequencies, and (iii) their degree of codon degeneracy. A new index taking into account all of these aspects is proposed. The index, designated as the 'weighted sum of relative entropy' (E(w)), is defined as the sum of the relative entropy of each amino acid weighted by its relative frequency in the sequence. In this paper, we demonstrate that E(w) allows us to avoid some amino acid usage biases and can yield results contradictory to those obtained by previous methods.

Algorithms↗

A new role for expressed pseudogenes as ncRNA: regulation of mRNA stability of its homologous coding gene.

We have earlier generated a mutant mouse in a course of making a transgenic line that exhibited interesting heterozygote phenotypes, which exhibited failure to thrive, severe bone deformities, and polycystic kidneys. This mutant mouse provided a clue to uncover a unique role of expressed pseudogenes. In this mutant the transgene was integrated into the vicinity of the expressing pseudogene of Makorin1 called Makorin1-p1. This insertion reduced transcription of the Makorin1-p1, resulting in destabilization of the Makorin1 mRNA in trans via a cis-acting RNA decay element within the 5' region of Makorin1 that is homologous between Makorin1 and Makorin1-p1. These findings demonstrate a novel and specific regulatory role of an expressed pseudogene as well as functional significance for noncoding RNAs. Next, we developed an original algorithm to determine how many pseudogenes are expressed. Based on our examination 2-3% of human processed pseudogenes are expressed using the most strict criteria. Interestingly, the mouse has a much smaller proportion of expressed pseudogenes (0.5-1%). Pseudogenes are functionally less constrained, and have accumulated more mutations than translated genes. If they have some functions in gene regulation, this property would allow more rapid functional diversification than protein-coding genes. In addition, some genetic phenomena that exhibit incomplete penetrance might be attributed to "mutation" or "variation" of pseudogenes.

Animals↗

Computational analysis of stop codon readthrough in D.melanogaster.

MOTIVATION: Readthrough is an unusual process in which a stop codon is misread or skipped. Recently it has been shown that some translation is regulated by the readthrough reactions although the complete mechanism is not clear. Therefore, the discovery of 'readthrough genes' is important for further investigation of their cellular roles, which may provide additional insights into the mechanism of translational regulation. RESULTS: We constructed a system that lists candidates of readthrough genes based on the existence of a 'protein motif' at the 3' untranslated region (UTR). Using this system, we extracted 85 candidates from 4082 nucleic acid sequences of Drosophila melanogaster in GenBank database. The sequences of these candidates had a slightly more stable secondary structure and different base preferences compared to the non-candidates. As these features are known to have an effect on readthrough events, we would like to suggest that these candidates contain actual readthrough genes. AVAILABILITY: Source code of the system is available upon request.

Algorithms↗

Collection, mapping, and annotation of over 28,000 cDNA clones from japonica rice.

We collected and completely sequenced 28,469 full-length complementary DNA clones from Oryza sativa L. ssp. japonica cv. Nipponbare. Through homology searches of publicly available sequence data, we assigned tentative protein functions to 21,596 clones (75.86%). Mapping of the cDNA clones to genomic DNA revealed that there are 19,000 to 20,500 transcription units in the rice genome. Protein informatics analysis against the InterPro database revealed the existence of proteins presented in rice but not in Arabidopsis. Sixty-four percent of our cDNAs are homologous to Arabidopsis proteins.

Alternative Splicing↗

Construction of reliable protein-protein interaction networks with a new interaction generality measure.

MOTIVATION: Recent screening techniques have made large amounts of protein-protein interaction data available, from which biologically important information such as the function of uncharacterized proteins, the existence of novel protein complexes, and novel signal-transduction pathways can be discovered. However, experimental data on protein interactions contain many false positives, making these discoveries difficult. Therefore computational methods of assessing the reliability of each candidate protein-protein interaction are urgently needed. RESULTS: We developed a new 'interaction generality' measure (IG2) to assess the reliability of protein-protein interactions using only the topological properties of their interaction-network structure. Using yeast protein-protein interaction data, we showed that reliable protein-protein interactions had significantly lower IG2 values than less-reliable interactions, suggesting that IG2 values can be used to evaluate and filter interaction data to enable the construction of reliable protein-protein interaction networks.

Binding Sites↗

Global insights into protein complexes through integrated analysis of the reliable interactome and knockout lethality.

We performed an integrated computational analysis of data derived from a comprehensive set of protein-protein interactions (interactome) and a phenotype dataset on lethality in Saccharomyces cerevisiae. For the analysis, we selected reliable interactome data using our previous 'interaction generality,' a computational approach to assess reliability of interactions. Those efforts gave clear evidence that proteins with lethal phenotypes in knockout studies (lethal proteins) may interact with each other to form functional protein complexes to perform their cellular roles. However, our analysis indicates that interactions between lethal proteins are rather restricted to the same cellular pathway or function, and it is quite unlikely that they interact with other lethal proteins functioning in different cellular roles. Furthermore, our results allowed us predictions on the functions of thus far uncharacterized lethal proteins with an estimated 93% accuracy. Thus, the analysis described in here can provide global insights into the biological features of the protein complexes.

Computational Biology↗

Identification of putative noncoding RNAs among the RIKEN mouse full-length cDNA collection.

With the sequencing and annotation of genomes and transcriptomes of several eukaryotes, the importance of noncoding RNA (ncRNA)-RNA molecules that are not translated to protein products-has become more evident. A subclass of ncRNA transcripts are encoded by highly regulated, multi-exon, transcriptional units, are processed like typical protein-coding mRNAs and are increasingly implicated in regulation of many cellular functions in eukaryotes. This study describes the identification of candidate functional ncRNAs from among the RIKEN mouse full-length cDNA collection, which contains 60,770 sequences, by using a systematic computational filtering approach. We initially searched for previously reported ncRNAs and found nine murine ncRNAs and homologs of several previously described nonmouse ncRNAs. Through our computational approach to filter artifact-free clones that lack protein coding potential, we extracted 4280 transcripts as the largest-candidate set. Many clones in the set had EST hits, potential CpG islands surrounding the transcription start sites, and homologies with the human genome. This implies that many candidates are indeed transcribed in a regulated manner. Our results demonstrate that ncRNAs are a major functional subclass of processed transcripts in mammals.

Animals↗

Inferring higher functional information for RIKEN mouse full-length cDNA clones with FACTS.

FACTS (Functional Association/Annotation of cDNA Clones from Text/Sequence Sources) is a semiautomated knowledge discovery and annotation system that integrates molecular function information derived from sequence analysis results (sequence inferred) with functional information extracted from text. Text-inferred information was extracted from keyword-based retrievals of MEDLINE abstracts and by matching of gene or protein names to OMIM, BIND, and DIP database entries. Using FACTS, we found that 47.5% of the 60,770 RIKEN mouse cDNA FANTOM2 clone annotations were informative for text searches. MEDLINE queries yielded molecular interaction-containing sentences for 23.1% of the clones. When disease MeSH and GO terms were matched with retrieved abstracts, 22.7% of clones were associated with potential diseases, and 32.5% with GO identifiers. A significant number (23.5%) of disease MeSH-associated clones were also found to have a hereditary disease association (OMIM Morbidmap). Inferred neoplastic and nervous system disease represented 49.6% and 36.0% of disease MeSH-associated clones, respectively. A comparison of sequence-based GO assignments with informative text-based GO assignments revealed that for 78.2% of clones, identical GO assignments were provided for that clone by either method, whereas for 21.8% of clones, the assignments differed. In contrast, for OMIM assignments, only 28.5% of clones had identical sequence-based and text-based OMIM assignments. Sequence, sentence, and term-based functional associations are included in the FACTS database (http://facts.gsc.riken.go.jp/), which permits results to be annotated and explored through web-accessible keyword and sequence search interfaces. The FACTS database will be a critical tool for investigating the functional complexity of the mouse transcriptome, cDNA-inferred interactome (molecular interactions), and pathome (pathologies).

Animals↗

CDS annotation in full-length cDNA sequence.

The identification of coding sequences (CDS) is an important step in the functional annotation of genes. CDS prediction for mammalian genes from genomic sequence is complicated by the vast abundance of intergenic sequence in the genome, and provides little information about how different parts of potential CDS regions are expressed. In contrast, mammalian gene CDS prediction from cDNA sequence offers obvious advantages, yet encounters a different set of complexities when performed on high-throughput cDNA (HTC) sequences, such as the set of 60,770 cDNAs isolated from full-length enriched libraries of the FANTOM2 project. We developed a CDS annotation strategy that uses a variety of different CDS prediction programs to annotate the CDS regions of FANTOM2 cDNAs. These include rsCDS, which uses sequence similarity to known proteins; ProCrest; Longest-ORF and Truncated-ORF, which are ab initio based predictors; and finally, DECODER and NCBI CDS predictor, which use a combination of both principles. Aided by graphical displays of these CDS prediction results in the context of other sequence similarity results for each cDNA, FANTOM2 CDS inspection by curators and follow-up quality control procedures resulted in high quality CDS predictions for a total of 14,345 FANTOM2 clones.

Animals↗

Targeting a complex transcriptome: the construction of the mouse full-length cDNA encyclopedia.

We report the construction of the mouse full-length cDNA encyclopedia,the most extensive view of a complex transcriptome,on the basis of preparing and sequencing 246 libraries. Before cloning,cDNAs were enriched in full-length by Cap-Trapper,and in most cases,aggressively subtracted/normalized. We have produced 1,442,236 successful 3'-end sequences clustered into 171,144 groups, from which 60,770 clones were fully sequenced cDNAs annotated in the FANTOM-2 annotation. We have also produced 547,149 5' end reads,which clustered into 124,258 groups. Altogether, these cDNAs were further grouped in 70,000 transcriptional units (TU),which represent the best coverage of a transcriptome so far. By monitoring the extent of normalization/subtraction, we define the tentative equivalent coverage (TEC),which was estimated to be equivalent to >12,000,000 ESTs derived from standard libraries. High coverage explains discrepancies between the very large numbers of clusters (and TUs) of this project,which also include non-protein-coding RNAs,and the lower gene number estimation of genome annotations. Altogether,5'-end clusters identify regions that are potential promoters for 8637 known genes and 5'-end clusters suggest the presence of almost 63,000 transcriptional starting points. An estimate of the frequency of polyadenylation signals suggests that at least half of the singletons in the EST set represent real mRNAs. Clones accounting for about half of the predicted TUs await further sequencing. The continued high-discovery rate suggests that the task of transcriptome discovery is not yet complete.

Animals↗