Search PubMed⌕ Search

Biomedical subjects

Shinichi Morishita

Publications and source records attributed to Shinichi Morishita.

At least 19 recordsLinked to original sources

Approximating edit distances between complex tandem repeats efficiently.

MOTIVATION: Extended tandem repeats (TRs) have been associated with 60 or more diseases over the past 30 years. Although most TRs have single repeat units (or motifs), complex TRs with different units have recently been correlated with some brain disorders. Of note, a population-scale analysis shows that complex TRs at one locus can be divergent, and different units are often expanded between individuals. To understand the evolution of high TR diversity, it is informative to visualize a phylogenetic tree. To do this, we need to measure the edit distance between pairs of complex TRs by considering duplication and contraction of units created by replication slippage. However, traditional rigorous algorithms for this purpose are computationally expensive. RESULTS: We here propose an efficient heuristic algorithm to estimate the edit distance with duplication and contraction of units (EDDC, for short). We select a set of frequent units that occur in given complex TRs, encode each unit as a single symbol, compress a TR into an optimal series of unit symbols that partially matches the original TR with the minimum Levenshtein distance, and estimate the EDDC between a pair of complex TRs from their compressed forms. Using substantial synthetic benchmark datasets, we demonstrate that the estimated EDDC is highly correlated with the accurate EDDC, with a Pearson correlation coefficient of >0.983, while the heuristic algorithm achieves orders of magnitude performance speedup. AVAILABILITY AND IMPLEMENTATION: The software program hEDDC that implements the proposed algorithm is available at https://github.com/Ricky-pon/hEDDC (DOI: 10.5281/zenodo.14732958).

Algorithms↗

A large-scale full-length cDNA analysis to explore the budding yeast transcriptome.

We performed a large-scale cDNA analysis to explore the transcriptome of the budding yeast Saccharomyces cerevisiae. We sequenced two cDNA libraries, one from the cells exponentially growing in a minimal medium and the other from meiotic cells. Both libraries were generated by using a vector-capping method that allows the accurate mapping of transcription start sites (TSSs). Consequently, we identified 11,575 TSSs associated with 3,638 annotated genomic features, including 3,599 ORFs, to suggest that most yeast genes have two or more TSSs. In addition, we identified 45 previously undescribed introns, including those affecting current ORF annotations and those spliced alternatively. Furthermore, the analysis revealed 667 transcription units in the intergenic regions and transcripts derived from antisense strands of 367 known features. We also found that 348 ORFs carry TSSs in their 3'-halves to generate sense transcripts starting from inside the ORFs. These results indicate that the budding yeast transcriptome is considerably more complex than previously thought, and it shares many recently revealed characteristics with the transcriptomes of mammals and other higher eukaryotes. Thus, the genome-wide active transcription that generates novel classes of transcripts appears to be an intrinsic feature of the eukaryotic cells. The budding yeast will serve as a versatile model for the studies on these aspects of transcriptome, and the full-length cDNA clones can function as an invaluable resource in such studies.

5' Untranslated Regions↗

PrimerStation: a highly specific multiplex genomic PCR primer design server for the human genome.

PrimerStation (http://ps.cb.k.u-tokyo.ac.jp) is a web service that calculates primer sets guaranteeing high specificity against the entire human genome. To achieve high accuracy, we used the hybridization ratio of primers in liquid solution. Calculating the status of sequence hybridization in terms of the stringent hybridization ratio is computationally costly, and no web service checks the entire human genome and returns a highly specific primer set calculated using a precise physicochemical model. To shorten the response time, we precomputed candidates for specific primers using a massively parallel computer with 100 CPUs (SunFire 15 K) about 3 months in advance. This enables PrimerStation to search and output qualified primers interactively. PrimerStation can select highly specific primers suitable for multiplex PCR by seeking a wider temperature range that minimizes the possibility of cross-reaction. It also allows users to add heuristic rules to the primer design, e.g. the exclusion of single nucleotide polymorphisms (SNPs) in primers, the avoidance of poly(A) and CA-repeats in the PCR products, and the elimination of defective primers using the secondary structure prediction. We performed several tests to verify the PCR amplification of randomly selected primers for ChrX, and we confirmed that the primers amplify specific PCR products perfectly.

DNA Primers↗

Expression profiling of muscles from Fukuyama-type congenital muscular dystrophy and laminin-alpha 2 deficient congenital muscular dystrophy; is congenital muscular dystrophy a primary fibrotic disease?

Fukuyama-type congenital muscular dystrophy (FCMD) and laminin-alpha2 deficient congenital muscular dystrophy (MDC1A) are congenital muscular dystrophies (CMDs) and they both are categorized into the same clinical entity of muscular dystrophy as Duchenne muscular dystrophy (DMD). All three disorders share a common etiologic defect in the dystrophin-glycoprotein complex, which connects muscle structural proteins with the extracellular basement membrane. To investigate the pathophysiology of these CMDs, we generated microarray gene expression profiles of skeletal muscle from patients in various clinical stages. Despite diverse pathological changes, the correlation coefficient of overall gene expression among these samples was considerably high. We performed a multi-dimensional statistical analysis, the Distillation, to extract determinant genes that distinguish CMD muscle from normal controls. Up-regulated genes were primarily extracellular matrix (ECM) components, whereas down-regulated genes included structural components of mature muscle. These observations reflect active interstitial fibrosis with less active regeneration of muscle cell components in the CMDs, characteristics that are clearly distinct from those of DMD. Although the severity of fibrosis varied among the specimens tested, ECM gene expression was consistently high without substantial changes through the clinical course. Further, in situ hybridization showed more prominent ECM gene expression on muscle cells than on interstitial tissue cells, suggesting that ECM components are induced by regeneration process rather than by 'dystrophy.' These data imply that the etiology of FCMD and MDC1A differs from that of the chronic phase of classical muscular dystrophy, and the major pathophysiologic change in CMDs might instead result from primary active fibrosis.

Child↗

Evaluation of image processing programs for accurate measurement of budding and fission yeast morphology.

To study the cellular functions of gene products, various yeast morphological mutants have been investigated. To describe yeast morphology objectively, we have developed image processing programs for budding and fission yeast. The programs, named CalMorph for budding yeast and F-CalMorph for fission yeast, directly process microscopic images and generate quantitative data about yeast cell shape, nuclear shape and location, and actin distribution. Using CalMorph, we can easily and quickly obtain various quantitative data reproducibly. To study the utility and reliability of CalMorph, we evaluated its data in three ways: (1) The programs extracted three-dimensional bud information from two-dimensional digital images with a low error rate (<1%). (2) The absolute values of the diameters of manufactured fluorescent beads calculated with CalMorph were very close to those given in the manufacturer's data sheet. (3) The programs generated reproducible data consistent with that obtained by hand. Based on these results, we determined that CalMorph could monitor yeast morphological changes accompanied by the progression of the cell cycle. We discuss the potential of the CalMorph series as a novel tool for the analysis of yeast cell morphology.

Cell Division↗

Comparative analysis of chimpanzee and human Y chromosomes unveils complex evolutionary pathway.

The mammalian Y chromosome has unique characteristics compared with the autosomes or X chromosomes. Here we report the finished sequence of the chimpanzee Y chromosome (PTRY), including 271 kb of the Y-specific pseudoautosomal region 1 and 12.7 Mb of the male-specific region of the Y chromosome. Greater sequence divergence between the human Y chromosome (HSAY) and PTRY (1.78%) than between their respective whole genomes (1.23%) confirmed the accelerated evolutionary rate of the Y chromosome. Each of the 19 PTRY protein-coding genes analyzed had at least one nonsynonymous substitution, and 11 genes had higher nonsynonymous substitution rates than synonymous ones, suggesting relaxation of selective constraint, positive selection or both. We also identified lineage-specific changes, including deletion of a 200-kb fragment from the pericentromeric region of HSAY, expansion of young Alu families in HSAY and accumulation of young L1 elements and long terminal repeat retrotransposons in PTRY. Reconstruction of the common ancestral Y chromosome reflects the dynamic changes in our genomes in the 5-6 million years since speciation.

Animals↗

High-dimensional and large-scale phenotyping of yeast mutants.

One of the most powerful techniques for attributing functions to genes in uni- and multicellular organisms is comprehensive analysis of mutant traits. In this study, systematic and quantitative analyses of mutant traits are achieved in the budding yeast Saccharomyces cerevisiae by investigating morphological phenotypes. Analysis of fluorescent microscopic images of triple-stained cells makes it possible to treat morphological variations as quantitative traits. Deletion of nearly half of the yeast genes not essential for growth affects these morphological traits. Similar morphological phenotypes are caused by deletions of functionally related genes, enabling a functional assignment of a locus to a specific cellular pathway. The high-dimensional phenotypic analysis of defined yeast mutant strains provides another step toward attributing gene function to all of the genes in the yeast genome.

Actins↗

dsCheck: highly sensitive off-target search software for double-stranded RNA-mediated RNA interference.

Off-target effects are one of the most serious problems in RNA interference (RNAi). Here, we present dsCheck (http://dsCheck.RNAi.jp/), web-based online software for estimating off-target effects caused by the long double-stranded RNA (dsRNA) used in RNAi studies. In the biochemical process of RNAi, the long dsRNA is cleaved by Dicer into short-interfering RNA (siRNA) cocktails. The software simulates this process and investigates individual 19 nt substrings of the long dsRNA. Subsequently, the software promptly enumerates a list of potential off-target gene candidates based on the order of off-target effects using its novel algorithm, which significantly improves both the efficiency and the sensitivity of the homology search. The website not only provides a rigorous off-target search to verify previously designed dsRNA sequences but also presents 'off-target minimized' dsRNA design, which is essential for reliable experiments in RNAi-based functional genomics.

Algorithms↗

Data mining tools for the Saccharomyces cerevisiae morphological database.

For comprehensive understanding of precise morphological changes resulting from loss-of-function mutagenesis, a large collection of 1,899,247 cell images was assembled from 91,71 micrographs of 4782 budding yeast disruptants of non-lethal genes. All the cell images were processed computationally to measure approximately 500 morphological parameters in individual mutants. We have recently made this morphological quantitative data available to the public through the Saccharomyces cerevisiae Morphological Database (SCMD). Inspecting the significance of morphological discrepancies between the wild type and the mutants is expected to provide clues to uncover genes that are relevant to the biological processes producing a particular morphology. To facilitate such intensive data mining, a suite of new software tools for visualizing parameter value distributions was developed to present mutants with significant changes in easily understandable forms. In addition, for a given group of mutants associated with a particular function, the system automatically identifies a combination of multiple morphological parameters that discriminates a mutant group from others significantly, thereby characterizing the function effectively. These data mining functions are available through the World Wide Web at http://scmd.gi.k.u-tokyo.ac.jp/.

Computer Graphics↗

5'SAGE: 5'-end Serial Analysis of Gene Expression database.

To comprehensively identify transcription start sites and the frequencies of individual mRNAs in human cell libraries, a method of 5' end Serial Analysis of Gene Expression (SAGE) was developed recently, which makes it possible to collect a large amount of start site information, and subsequently, we have established a related database server called 5'SAGE. This database displays the observed frequencies of individual 5' end SAGE tags and previously unknown transcription start sites in the promoter regions, introns and intergenic regions of known genes. 5'SAGE will be useful for analyzing promoter regions and start site variation in different tissues, and is freely available at http://5sage.gi.k.u-tokyo.ac.jp/.

5' Flanking Region↗

Cloning and expression of a brain-specific putative UDP-GalNAc: polypeptide N-acetylgalactosaminyltransferase gene.

We isolated a rat cDNA clone and its human orthologue, which are most homologous to UDP-GalNAc: polypeptide N-acetylgalactosaminyltransferase 9, by homology-based PCR from brain. Nucleotide sequence analysis of these putative GalNAc-transferases (designated pt-GalNAc-T) showed that they contained structural features characteristic of the GalNAc-transferase family. It was also found that human pt-GalNAc-T was identical to the gene WBSCR17, which is reported to be in the critical region of patients with Williams-Beuren Syndrome, a neurodevelopmental disorder, and to be predominantly expressed in brain and heart. In order to investigate the expression of pt-GalNAc-T in brain in more detail, we first examined that of human pt-GalNAc-T by Northern blot analysis and found the expression of the 5.0-kb mRNA to be most abundant in cerebral cortex with somewhat less abundant in cellebellum. The expression of rat pt-GalNAc-T was investigated more extensively. The brain-specific expression of 2.0-kb and 5.0-kb transcripts was demonstrated by Northern blot analysis. In situ hybridization in the adult brain revealed high levels of expression in cerebellum, hippocampus, thalamus, and cerebral cortex. Moreover, observation at high magnification revealed the expression to be associated with neurons, but not with glial cells. Analysis of the rat embryos also demonstrated that rat pt-GalNAc-T was expressed in the nervous system, including in the diencephalons, cerebellar primordium, and dorsal root ganglion. However, recombinant human pt-GalNAc-T, which was expressed in insect cells, did not glycosylate several peptides derived from mammalian mucins, suggesting that it may have a strict substrate specificity. The brain-specific expression of pt-GalNAc-T suggested its involvement in brain development, through O-glycosylation of proteins in the neurons.

Amino Acid Sequence↗

Accelerated off-target search algorithm for siRNA.

MOTIVATION: Designing highly effective short interfering RNA (siRNA) sequences with maximum target-specificity for mammalian RNA interference (RNAi) is one of the hottest topics in molecular biology. The relationship between siRNA sequences and RNAi activity has been studied extensively to establish rules for selecting highly effective sequences. However, there is a pressing need to compute siRNA sequences that minimize off-target silencing effects efficiently and to match any non-targeted sequences with mismatches. RESULTS: The enumeration of potential cross-hybridization candidates is non-trivial, because siRNA sequences are short, ca. 19 nt in length, and at least three mismatches with non-targets are required. With at least three mismatches, there are typically four or five contiguous matches, so that a BLAST search frequently overlooks off-target candidates. By contrast, existing accurate approaches are expensive to execute; thus we need to develop an accurate, efficient algorithm that uses seed hashing, the pigeonhole principle, and combinatorics to identify mismatch patterns. Tests show that our method can list potential cross-hybridization candidates for any siRNA sequence of selected human gene rapidly, outperforming traditional methods by orders of magnitude in terms of computational performance. AVAILABILITY: http://design.RNAi.jp CONTACT: yamada@cb.k.u-tokyo.ac.jp.

Algorithms↗

Dynactin is involved in a checkpoint to monitor cell wall synthesis in Saccharomyces cerevisiae.

Checkpoint controls ensure the completion of cell cycle events with high fidelity in the correct order. Here we show the existence of a novel checkpoint that ensures coupling of cell wall synthesis and mitosis. In response to a defect in cell wall synthesis, S. cerevisiae cells arrest the cell-cycle before spindle pole body separation. This arrest results from the regulation of the M-phase cyclin Clb2p at the transcriptional level through the transcription factor Fkh2p. Components of the dynactin complex are required to achieve the G2 arrest whilst keeping cells highly viable. Thus, the dynactin complex has a function in a checkpoint that monitors cell wall synthesis.

Cell Cycle↗

5'-end SAGE for the analysis of transcriptional start sites.

Identification of the mRNA start site is essential in establishing the full-length cDNA sequence of a gene and analyzing its promoter region, which regulates gene expression. Here we describe the development of a 5'-end serial analysis of gene expression (5' SAGE) that can be used to globally identify transcriptional start sites and the frequency of individual mRNAs. Of the 25,684 5' SAGE tags in the HEK293 human cell library, 19,893 matched to the human genome. Among 15,448 tags in one locus of the genome, 85.8%-96.1% of the 5' SAGE tags were assigned within -500 to +200 nt of mRNA start sites using the RefSeq, UniGene and DBTSS databases. This technique should facilitate 5'-end transcriptome analysis in a variety of cells and tissues.

5' Flanking Region↗

siDirect: highly effective, target-specific siRNA design software for mammalian RNA interference.

siDirect (http://design.RNAi.jp/) is a web-based online software system for computing highly effective small interfering RNA (siRNA) sequences with maximum target-specificity for mammalian RNA interference (RNAi). Highly effective siRNA sequences are selected using novel guidelines that were established through an extensive study of the relationship between siRNA sequences and RNAi activity. Our efficient software avoids off-target gene silencing to enumerate potential cross-hybridization candidates that the widely used BLAST search may overlook. The website accepts an arbitrary sequence as input and quickly returns siRNA candidates, providing a wide scope of applications in mammalian RNAi, including systematic functional genomics and therapeutic gene silencing.

Algorithms↗

Constrained clusters of gene expression profiles with pathological features.

MOTIVATION: Gene expression profiles should be useful in distinguishing variations in disease, since they reflect accurately the status of cells. The primary clustering of gene expression reveals the genotypes that are responsible for the proximity of members within each cluster, while further clustering elucidates the pathological features of the individual members of each cluster. However, since the first clustering process and the second classification step, in which the features are associated with clusters, are performed independently, the initial set of clusters may omit genes that are associated with pathologically meaningful features. Therefore, it is important to devise a way of identifying gene expression clusters that are associated with pathological features. RESULTS: We present the novel technique of 'itemset constrained clustering' (IC-Clustering), which computes the optimal cluster that maximizes the interclass variance of gene expression between groups, which are divided according to the restriction that only divisions that can be expressed using common features are allowed. This constraint automatically labels each cluster with a set of pathological features which characterize that cluster. When applied to liver cancer datasets, IC-Clustering revealed informative gene expression clusters, which could be annotated with various pathological features, such as 'tumor' and 'man', or 'except tumor' and 'normal liver function'. In contrast, the k-means method overlooked these clusters.

Algorithms↗