Search PubMed⌕ Search

Biomedical subjects

Jun Adachi

Publications and source records attributed to Jun Adachi.

11 recordsLinked to original sources

MAPU: Max-Planck Unified database of organellar, cellular, tissue and body fluid proteomes.

Mass spectrometry (MS)-based proteomics has become a powerful technology to map the protein composition of organelles, cell types and tissues. In our department, a large-scale effort to map these proteomes is complemented by the Max-Planck Unified (MAPU) proteome database. MAPU contains several body fluid proteomes; including plasma, urine, and cerebrospinal fluid. Cell lines have been mapped to a depth of several thousand proteins and the red blood cell proteome has also been analyzed in depth. The liver proteome is represented with 3200 proteins. By employing high resolution MS and stringent validation criteria, false positive identification rates in MAPU are lower than 1:1000. Thus MAPU datasets can serve as reference proteomes in biomarker discovery. MAPU contains the peptides identifying each protein, measured masses, scores and intensities and is freely available at http://www.mapuproteome.com using a clickable interface of cell or body parts. Proteome data can be queried across proteomes by protein name, accession number, sequence similarity, peptide sequence and annotation information. More than 4500 mouse and 2500 human proteins have already been identified in at least one proteome. Basic annotation information and links to other public databases are provided in MAPU and we plan to add further analysis tools.

Animals↗

Development and validation of the psychosomatic scale for atopic dermatitis in adults.

Psychosocial factors play an important role in the course of adult atopic dermatitis (AD). Nevertheless, AD patients are rarely treated for their psychosomatic concerns. The purpose of the present study was to develop and validate a brief self-rating scale for adult AD in order to aid dermatologists in evaluating psychosocial factors during the course of AD. A preliminary scale assessing stress-induced exacerbation, the secondary psychosocial burden, and attitude toward treatment was developed and administered to 187 AD patients (82 male, 105 female, aged 28.4 +/- 7.8, 13-61). Severity of skin lesions and improvement with standard dermatological treatment were assessed by both the dermatologist and the participant. Measures of anxiety and depression were also determined. In addition, psychosomatic evaluations were made according to the Psychosomatic Diagnostic Criteria for AD. Factor analysis resulted in the development of a 12-item scale (The Psychosomatic Scale for Atopic Dermatitis; PSS-AD) consisting of three factors: (i) exacerbation triggered by stress; (ii) disturbances due to AD; and (iii) ineffective control. Internal consistency indicated by Cronbach's alpha coefficient was 0.86 for the entire measure, 0.82 for (i), 0.81 for (ii), and 0.77 for (iii), verifying the acceptable reliability of PSS-AD. Patients with psychosomatic problems had higher PSS-AD scores than those without. PSS-AD scores were positively associated with the severity of the skin lesions, anxiety and depression. The scores were negatively associated with improvement during dermatological treatments. In conclusion, PSS-AD is a simple and reliable measure of the psychosomatic pathology of adult AD patients. It may be useful in dermatological practice for screening patients who would benefit from psychological or psychiatric interventions.

Adolescent↗

ANGLE: a sequencing errors resistant program for predicting protein coding regions in unfinished cDNA.

In the process of making full-length cDNA, predicting protein coding regions helps both in the preliminary analysis of genes and in any succeeding process. However, unfinished cDNA contains artifacts including many sequencing errors, which hinder the correct evaluation of coding sequences. Especially, predictions of short sequences are difficult because they provide little information for evaluating coding potential. In this paper, we describe ANGLE, a new program for predicting coding sequences in low quality cDNA. To achieve error-tolerant prediction, ANGLE uses a machine-learning approach, which makes better expression of coding sequence maximizing the use of limited information from input sequences. Our method utilizes not only codon usage, but also protein structure information which is difficult to be used for stochastic model-based algorithms, and optimizes limited information from a short segment when deciding coding potential, with the result that predictive accuracy does not depend on the length of an input sequence. The performance of ANGLE is compared with ESTSCAN on four dataset each of them having a different error rate (one frame-shift error or one substitution error per 200-500 nucleotides) and on one dataset which has no error. ANGLE outperforms ESTSCAN by 9.26% in average Matthews's correlation coefficient on short sequence dataset (< 1000 bases). On long sequence dataset, ANGLE achieves comparable performance.

Algorithms↗

The human urinary proteome contains more than 1500 proteins, including a large proportion of membrane proteins.

BACKGROUND: Urine is a desirable material for the diagnosis and classification of diseases because of the convenience of its collection in large amounts; however, all of the urinary proteome catalogs currently being generated have limitations in their depth and confidence of identification. Our laboratory has developed methods for the in-depth characterization of body fluids; these involve a linear ion trap-Fourier transform (LTQ-FT) and a linear ion trap-orbitrap (LTQ-Orbitrap) mass spectrometer. Here we applied these methods to the analysis of the human urinary proteome. RESULTS: We employed one-dimensional sodium dodecyl sulfate polyacrylamide gel electrophoresis and reverse phase high-performance liquid chromatography for protein separation and fractionation. Fractionated proteins were digested in-gel or in-solution, and digests were analyzed with the LTQ-FT and LTQ-Orbitrap at parts per million accuracy and with two consecutive stages of mass spectrometric fragmentation. We identified 1543 proteins in urine obtained from ten healthy donors, while essentially eliminating false-positive identifications. Surprisingly, nearly half of the annotated proteins were membrane proteins according to Gene Ontology (GO) analysis. Furthermore, extracellular, lysosomal, and plasma membrane proteins were enriched in the urine compared with all GO entries. Plasma membrane proteins are probably present in urine by secretion in exosomes. CONCLUSION: Our analysis provides a high-confidence set of proteins present in human urinary proteome and provides a useful reference for comparing datasets obtained using different methodologies. The urinary proteome is unexpectedly complex and may prove useful in biomarker discovery in the future.

Chromatography, High Pressure Liquid↗

Comparison of gene expression patterns between 2,3,7,8-tetrachlorodibenzo-p-dioxin and a natural arylhydrocarbon receptor ligand, indirubin.

Indirubin is a natural arylhydrocarbon receptor (AhR) ligand isolated from human urine. We previously reported that it was more potent than the prototypical ligand, 2,3,7,8-tetrachlorodibenzo-p-dioxin (TCDD) in a yeast assay system. Here we compared gene expression changes in HepG2 cells exposed to 10 nM of indirubin or TCDD using nylon-membrane-based cDNA arrays with 1176 genes to elucidate the toxic differences at the transcriptional level. The gene expression profiles for TCDD and indirubin were very similar. The number of up-regulated genes (fold change > or =2.0) was 11 and 4 and the number of down-regulated genes (fold change < or =0.5) was 17 and 21 in TCDD-treated and indirubin-treated cells, respectively. Cytochrome P450 (CYP) 1A1, 1A2, 19A1, insulin-like growth factor binding protein 1 (IGFBP1), and IGFBP10 were confirmed to be up-regulated using real-time reverse transcription polymerase chain reaction. CYP1A1 and CYP1A2 mRNAs were induced by as little as 1 pM of indirubin, whereas they were not induced by 10 pM of TCDD. In the time-course experiment, CYP1A1 mRNA was induced by indirubin transiently. Indirubin was also metabolized by CYP1A1 and lost its ligand activity. Indirubin would appear to be a good substrate of CYP1A1 given its low dissociation constant. Our results suggest that indirubin rapidly activates its own metabolism via AhR-mediated induction of CYP1A1 and this characteristic is consistent with the notion that indirubin is a physiological ligand of AhR.

Animals↗

Collection, mapping, and annotation of over 28,000 cDNA clones from japonica rice.

We collected and completely sequenced 28,469 full-length complementary DNA clones from Oryza sativa L. ssp. japonica cv. Nipponbare. Through homology searches of publicly available sequence data, we assigned tentative protein functions to 21,596 clones (75.86%). Mapping of the cDNA clones to genomic DNA revealed that there are 19,000 to 20,500 transcription units in the rice genome. Protein informatics analysis against the InterPro database revealed the existence of proteins presented in rice but not in Arabidopsis. Sixty-four percent of our cDNAs are homologous to Arabidopsis proteins.

Alternative Splicing↗

Identification of putative noncoding RNAs among the RIKEN mouse full-length cDNA collection.

With the sequencing and annotation of genomes and transcriptomes of several eukaryotes, the importance of noncoding RNA (ncRNA)-RNA molecules that are not translated to protein products-has become more evident. A subclass of ncRNA transcripts are encoded by highly regulated, multi-exon, transcriptional units, are processed like typical protein-coding mRNAs and are increasingly implicated in regulation of many cellular functions in eukaryotes. This study describes the identification of candidate functional ncRNAs from among the RIKEN mouse full-length cDNA collection, which contains 60,770 sequences, by using a systematic computational filtering approach. We initially searched for previously reported ncRNAs and found nine murine ncRNAs and homologs of several previously described nonmouse ncRNAs. Through our computational approach to filter artifact-free clones that lack protein coding potential, we extracted 4280 transcripts as the largest-candidate set. Many clones in the set had EST hits, potential CpG islands surrounding the transcription start sites, and homologies with the human genome. This implies that many candidates are indeed transcribed in a regulated manner. Our results demonstrate that ncRNAs are a major functional subclass of processed transcripts in mammals.

Animals↗

Impact of alternative initiation, splicing, and termination on the diversity of the mRNA transcripts encoded by the mouse transcriptome.

We analyzed the FANTOM2 clone set of 60,770 RIKEN full-length mouse cDNA sequences and 44,122 public mRNA sequences. We developed a new computational procedure to identify and classify the forms of splice variation evident in this data set and organized the results into a publicly accessible database that can be used for future expression array construction, structural genomics, and analyses of the mechanism and regulation of alternative splicing. Statistical analysis shows that at least 41% and possibly as much as 60% of multiexon genes in mouse have multiple splice forms. Of the transcription units with multiple splice forms, 49% contain transcripts in which the apparent use of an alternative transcription start (stop) is accompanied by alternative splicing of the initial (terminal) exon. This implies that alternative transcription may frequently induce alternative splicing. The fact that 73% of all exons with splice variation fall within the annotated coding region indicates that most splice variation is likely to affect the protein form. Finally, we compared the set of constitutive (present in all transcripts) exons with the set of cryptic (present only in some transcripts) exons and found statistically significant differences in their length distributions, the nucleotide distributions around their splice junctions, and the frequencies of occurrence of several short sequence motifs.

Alternative Splicing↗

CDS annotation in full-length cDNA sequence.

The identification of coding sequences (CDS) is an important step in the functional annotation of genes. CDS prediction for mammalian genes from genomic sequence is complicated by the vast abundance of intergenic sequence in the genome, and provides little information about how different parts of potential CDS regions are expressed. In contrast, mammalian gene CDS prediction from cDNA sequence offers obvious advantages, yet encounters a different set of complexities when performed on high-throughput cDNA (HTC) sequences, such as the set of 60,770 cDNAs isolated from full-length enriched libraries of the FANTOM2 project. We developed a CDS annotation strategy that uses a variety of different CDS prediction programs to annotate the CDS regions of FANTOM2 cDNAs. These include rsCDS, which uses sequence similarity to known proteins; ProCrest; Longest-ORF and Truncated-ORF, which are ab initio based predictors; and finally, DECODER and NCBI CDS predictor, which use a combination of both principles. Aided by graphical displays of these CDS prediction results in the context of other sequence similarity results for each cDNA, FANTOM2 CDS inspection by curators and follow-up quality control procedures resulted in high quality CDS predictions for a total of 14,345 FANTOM2 clones.

Animals↗

Connecting sequence and biology in the laboratory mouse.

The Mouse Genome Sequencing Consortium and the RIKEN Genome Exploration Research grouphave generated large sets of sequence data representing the mouse genome and transcriptome, respectively. These data provide a valuable foundation for genomic research. The challenges for the informatics community are how to integrate these data with the ever-expanding knowledge about the roles of genes and gene products in biological processes, and how to provide useful views to the scientific community. Public resources, such as the National Center for Biotechnology Information (NCBI; http://www.ncbi.nih.gov), and model organism databases, such as the Mouse Genome Informatics database (MGI; http://www.informatics.jax.org), maintain the primary data and provide connections between sequence and biology. In this paper, we describe how the partnership of MGI and NCBI LocusLink contributes to the integration of sequence and biology, especially in the context of the large-scale genome and transcriptome data now available for the laboratory mouse. In particular, we describe the methods and results of integration of 60,770 FANTOM2 mouse cDNAs with gene records in the databases of MGI and LocusLink.

Animals↗