Search PubMed⌕ Search

Biomedical subjects

Piero Carninci

Publications and source records attributed to Piero Carninci.

At least 55 records · Page 3Linked to original sources

Comparative analysis of plant and animal calcium signal transduction element using plant full-length cDNA data.

We obtained 32K full-length cDNA sequence data from the rice full-length cDNA project and performed a homology search against NCBI GenBank data. We have also searched homologs of Arabidopsis and other plants' genes with the databases. Comparative analysis of calcium ion transport proteins revealed that the genes specific for muscle and nerve calcium signal transduction systems (VDCC, IP3 receptor, ryanodine receptor) are very different in animals and plants. In contrast, Ca elements with basic functions in cell responses (CNGC, iGlu receptor, Ca(2+)ATPase, Ca2+/Na(+)-K+ ion exchanger) are basically conserved between plants and animals. We also performed comparative analyses of calcium ion binding and/or controlling signal transduction proteins. Many genes specific for muscle and nerve tissue do not exist in plants. However, calcium ion signal transduction genes of basic functions of cell homeostasis and responses were well conserved; plants have developed a calcium ion interacting system that is more direct than in animals. Many species of plants have specifically modified calcium ion binding proteins (CPK, CRK), Ca2+/phospholipid-binding domains, and calcium storage proteins.

Animals↗

Integrative annotation of 21,037 human genes validated by full-length cDNA clones.

The human genome sequence defines our inherent biological potential; the realization of the biology encoded therein requires knowledge of the function of each gene. Currently, our knowledge in this area is still limited. Several lines of investigation have been used to elucidate the structure and function of the genes in the human genome. Even so, gene prediction remains a difficult task, as the varieties of transcripts of a gene may vary to a great extent. We thus performed an exhaustive integrative characterization of 41,118 full-length cDNAs that capture the gene transcripts as complete functional cassettes, providing an unequivocal report of structural and functional diversity at the gene level. Our international collaboration has validated 21,037 human gene candidates by analysis of high-quality full-length cDNA clones through curation using unified criteria. This led to the identification of 5,155 new gene candidates. It also manifested the most reliable way to control the quality of the cDNA clones. We have developed a human gene database, called the H-Invitational Database (H-InvDB; http://www.h-invitational.jp/). It provides the following: integrative annotation of human genes, description of gene structures, details of novel alternative splicing isoforms, non-protein-coding RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein three-dimensional structure, mapping of known single nucleotide polymorphisms (SNPs), identification of polymorphic microsatellite repeats within human genes, and comparative results with mouse full-length cDNAs. The H-InvDB analysis has shown that up to 4% of the human genome sequence (National Center for Biotechnology Information build 34 assembly) may contain misassembled or missing regions. We found that 6.5% of the human gene candidates (1,377 loci) did not have a good protein-coding open reading frame, of which 296 loci are strong candidates for non-protein-coding RNA genes. In addition, among 72,027 uniquely mapped SNPs and insertions/deletions localized within human genes, 13,215 nonsynonymous SNPs, 315 nonsense SNPs, and 452 indels occurred in coding regions. Together with 25 polymorphic microsatellite repeats present in coding regions, they may alter protein structure, causing phenotypic effects or resulting in disease. The H-InvDB platform represents a substantial contribution to resources needed for the exploration of human biology and pathology.

Alternative Splicing↗

Transcriptional profiling of genes responsive to abscisic acid and gibberellin in rice: phenotyping and comparative analysis between rice and Arabidopsis.

We collected and completely sequenced 32,127 full-length complementary DNA clones from Oryza sativa L. ssp. japonica cv. "Nipponbare." Mapping of these clones to genomic DNA revealed approximately 20,500 transcriptional units (TUs) in the rice genome. For each TU, we selected 60-mers using an algorithm that took into account some DNA conditions such as base composition and sequence complexity. Using in situ synthesis technology, we constructed oligonucleotide arrays with these TUs on glass slides. We targeted RNAs prepared from normally grown rice callus and from callus treated with abscisic acid (ABA) or gibberellin (GA). We identified 200 ABA-responsive and 301 GA-responsive genes, many of which had never before been annotated as ABA or GA responsive in other expression analysis. Comparison of these genes revealed antagonistic regulation of almost all by both hormones; these had previously been annotated as being responsible for protein storage and defense against pathogens. Comparison of the cis-elements of genes responsive to one or antagonistic to both hormones revealed that the antagonistic genes had cis-elements related to ABA and GA responses. The genes responsive to only one hormone were rich in cis-elements that supported ABA and GA responses. In a search for the phenotypes of mutants in which a retrotransposon was inserted in these hormone-responsive genes, we identified phenotypes related to seed formation or plant height, including sterility, vivipary, and dwarfism. In comparison of cis-elements for hormone response genes between rice and Arabidopsis thaliana, we identified cis-elements for dehydration-stress response as Arabidopsis specific and for protein storage as rice specific.

Abscisic Acid↗

Gene discovery in genetically labeled single dopaminergic neurons of the retina.

In the retina, dopamine plays a central role in neural adaptation to light. Progress in the study of dopaminergic amacrine (DA) cells has been limited because they are very few (450 in each mouse retina, 0.005% of retinal neurons). Here, we applied transgenic technology, single-cell global mRNA amplification, and cDNA microarray screening to identify transcripts present in DA cells. To profile gene expression in single neurons, we developed a method (SMART7) that combines a PCR-based initial step (switching mechanism at the 5' end of the RNA transcript or SMART) with T7 RNA polymerase amplification. Single-cell targets were synthesized from genetically labeled DA cells to screen the RIKEN 19k mouse cDNA microarrays. Seven hundred ninety-five transcripts were identified in DA cells at a high level of confidence, and expression of the most interesting genes was confirmed by immunocytochemistry. Twenty-one previously undescribed proteins were found in DA cells, including a chloride channel, receptors and other membrane glycoproteins, kinases, transcription factors, and secreted neuroactive molecules. Thirty-eight percent of transcripts were ESTs or coding for hypothetical proteins, suggesting that a large portion of the DA cell proteome is still uncharacterized. Because cryptochrome-1 mRNA was found in DA cells, immunocytochemistry was extended to other components of the circadian clock machinery. This analysis showed that DA cells contain the most common clock-related proteins.

Animals↗

Absolute expression values for mouse transcripts: re-annotation of the READ expression database by the use of CAGE and EST sequence tags.

The RIKEN expression array database (READ) provides comprehensive gene expression data for the mouse, which were obtained as relative values from microarray double-staining experiments with E17.5 mRNA as common reference. To assign absolute expression values for mouse transcripts within READ, we applied the E17.5 reference sample to CAGE (cap analysis of gene expression) and expressed sequence tag (EST) high-throughput tag sequencing. Newly assigned values within the READ database were validated by comparison to expression data from serial analysis of gene expression, CAGE and EST experiments. These experiments confirmed the great significance of the absolute expression values within the improved READ database. The new Absolute READ database on absolute expression data is available under.

Animals↗

Solution structure of the SEA domain from the murine homologue of ovarian cancer antigen CA125 (MUC16).

Human CA125, encoded by the MUC16 gene, is an ovarian cancer antigen widely used for a serum assay. Its extracellular region consists of tandem repeats of SEA domains. In this study we determined the three-dimensional structure of the SEA domain from the murine MUC16 homologue using multidimensional NMR spectroscopy. The domain forms a unique alpha/beta sandwich fold composed of two alpha helices and four antiparallel beta strands and has a characteristic turn named the TY-turn between alpha1 and alpha2. The internal mobility of the main chain is low throughout the domain. The residues that form the hydrophobic core and the TY-turn are fully conserved in all SEA domain sequences, indicating that the fold is common in the family. Interestingly, no other residues are conserved throughout the family. Thus, the sequence alignment of the SEA domain family was refined on the basis of the three-dimensional structure, which allowed us to classify the SEA domains into several subfamilies. The residues on the surface differ between these subfamilies, suggesting that each subfamily has a different function. In the MUC16 SEA domains, the conserved surface residues, Asn-10, Thr-12, Arg-63, Asp-75, Asp-112, Ser-115, and Phe-117, are clustered on the beta sheet surface, which may be functionally important. The putative epitope (residues 58-77) for anti-MUC16 antibodies is located around the beta2 and beta3 strands. On the other hand the tissue tumor marker MUC1 has a SEA domain belonging to another subfamily, and its GSVVV motif for proteolytic cleavage is located in the short loop connecting beta2 and beta3.

Amino Acid Sequence↗

Solution structure of a BolA-like protein from Mus musculus.

The BolA-like proteins are widely conserved from prokaryotes to eukaryotes. The BolA-like proteins seem to be involved in cell proliferation or cell-cycle regulation, but the molecular function is still unknown. Here we determined the structure of a mouse BolA-like protein. The overall topology is alphabetabetaalphaalphabetaalpha, in which beta(1) and beta(2) are antiparallel, and beta(3) is parallel to beta(2). This fold is similar to the class II KH fold, except for the absence of the GXXG loop, which is well conserved in the KH fold. The conserved residues in the BolA-like proteins are assembled on the one side of the protein.

Amino Acid Sequence↗

FREP: a database of functional repeats in mouse cDNAs.

The FREP database (http://facts.gsc.riken.go.jp/FREP/) contains 31 396 RepeatMasker-identified non-redundant variant repeat sequences derived from 16,527 mouse cDNAs with protein-coding potential. The repeats were computationally associated with potential effects on transcriptional variation, translation, protein function or involvement in disease to identify Functional REPeats (FREPs). FREPs are defined by the (i) occurrence of exon-exon boundaries in repeats, (ii) presence of polyadenylation sites in 3'UTR-located repeats, (iii) effect on translation, (iv) position in the protein- coding region or protein domains or (v) conditional association with disease MeSH terms. Currently the database contains 9261 (29.5%) inferred FREPs derived from 6861 (41.5%) mouse cDNAs. Integrated evidence of the functional assignments and dynamically generated sequence similarity search results support the exploration and annotation of functional, ancestral or taxon-specific repeats. Keyword and pre-selected feature searches (e.g. coding sequence-repeat or splice site-repeat relations) support intuitive database querying as well as the retrieval of repeat sequences. Integrated sequence search and alignment tools allow the analysis of known or identification of new functional repeat candidates. FREP is a unique resource for illuminating the role of transposons and repetitive sequences in shaping the coding part of the mouse transcriptome and for selecting the appropriate experimental model to study diseases with suspected repeat etiology contributions.

Animals↗

Identification of unique transcripts from a mouse full-length, subtracted inner ear cDNA library.

A small-scale full-length library construction approach was developed to facilitate production of a mouse full-length cDNA encyclopedia representing approximately 250 enriched, normalized, and/or subtracted cDNA libraries. One library produced using this approach was a subtracted adult mouse inner ear cDNA library (sIEa). The average size of the inserts was approximately 2.5 kb, with the majority ranging from 0.5 to 7.0 kb. From this library 22,574 sequence reads were obtained from 15,958 independent clones. Sequencing and chromosomal localization established 5240 clusters, with 1302 clusters being unique and 359 representing new ESTs. Our sIEa library contributed 56.1% of the 7773 nonredundant Unigene clusters associated with the four mouse inner ear libraries in the NCBI dbEST. Based on homologous chromosomal regions between human and mouse, we identified 1018 UniGene clusters associated with the deafness locus critical regions. Of these, 59 clusters were found only in our sIEa library and represented approximately 50% of the identified critical regions.

Amino Acid Sequence↗

The status, quality, and expansion of the NIH full-length cDNA project: the Mammalian Gene Collection (MGC).

The National Institutes of Health's Mammalian Gene Collection (MGC) project was designed to generate and sequence a publicly accessible cDNA resource containing a complete open reading frame (ORF) for every human and mouse gene. The project initially used a random strategy to select clones from a large number of cDNA libraries from diverse tissues. Candidate clones were chosen based on 5'-EST sequences, and then fully sequenced to high accuracy and analyzed by algorithms developed for this project. Currently, more than 11,000 human and 10,000 mouse genes are represented in MGC by at least one clone with a full ORF. The random selection approach is now reaching a saturation point, and a transition to protocols targeted at the missing transcripts is now required to complete the mouse and human collections. Comparison of the sequence of the MGC clones to reference genome sequences reveals that most cDNA clones are of very high sequence quality, although it is likely that some cDNAs may carry missense variants as a consequence of experimental artifact, such as PCR, cloning, or reverse transcriptase errors. Recently, a rat cDNA component was added to the project, and ongoing frog (Xenopus) and zebrafish (Danio) cDNA projects were expanded to take advantage of the high-throughput MGC pipeline.

Animals↗

Solution structure of the RWD domain of the mouse GCN2 protein.

GCN2 is the alpha-subunit of the only translation initiation factor (eIF2alpha) kinase that appears in all eukaryotes. Its function requires an interaction with GCN1 via the domain at its N-terminus, which is termed the RWD domain after three major RWD-containing proteins: RING finger-containing proteins, WD-repeat-containing proteins, and yeast DEAD (DEXD)-like helicases. In this study, we determined the solution structure of the mouse GCN2 RWD domain using NMR spectroscopy. The structure forms an alpha + beta sandwich fold consisting of two layers: a four-stranded antiparallel beta-sheet, and three side-by-side alpha-helices, with an alphabetabetabetabetaalphaalpha topology. A characteristic YPXXXP motif, which always occurs in RWD domains, forms a stable loop including three consecutive beta-turns that overlap with each other by two residues (triple beta-turn). As putative binding sites with GCN1, a structure-based alignment allowed the identification of several surface residues in alpha-helix 3 that are characteristic of the GCN2 RWD domains. Despite the apparent absence of sequence similarity, the RWD structure significantly resembles that of ubiquitin-conjugating enzymes (E2s), with most of the structural differences in the region connecting beta-strand 4 and alpha-helix 3. The structural architecture, including the triple beta-turn, is fundamentally common among various RWD domains and E2s, but most of the surface residues on the structure vary. Thus, it appears that the RWD domain is a novel structural domain for protein-binding that plays specific roles in individual RWD-containing proteins.

Amino Acid Sequence↗

CTAB-urea method purifies RNA from melanin for cDNA microarray analysis.

Melanin represents a major problem for the study of melanoma by microarrays since it is retained during RNA extraction and inhibits the enzymatic reactions used for probe preparation. Here we report a new method for cleaning RNA from melanin, based on the use of the cationic detergent cetyl-trimethylammonium bromide (CTAB)-urea for RNA precipitation. This method is easy to perform and has a low cost. Purified RNA is recovered with high quality and good yield. CTAB-urea treated RNA from highly pigmented melanoma cells can be successfully reverse transcribed and labeled to obtain probes which can be subsequently used in cDNA microarray experiments, giving consistent and reproducible results.

Biotechnology↗

Construction of a full-length cDNA library from young spikelets of hexaploid wheat and its characterization by large-scale sequencing of expressed sequence tags.

The polyploid nature of wheat is a key characteristic of the plant. Full-length complementary DNAs (cDNAs) provide essential information that can be used to annotate the genes and provide a functional analysis of these genes and their products. We constructed a full-length cDNA library derived from young spikelets of common wheat, and obtained 24056 expressed sequence tags (ESTs) from both ends of the cDNA clones. These ESTs were grouped into 3605 contigs using the phrap method, representing expressed loci from each of the three genomes. Using BLAST, 3605 contigs were grouped into 1902 gene clusters, showing that loci of the three genomes are not always expressed. A homology search of these gene clusters against a wheat EST database (15964 gene clusters) and a rice full-length cDNA database (21447 gene clusters) revealed that a quarter of the wheat full-length cDNAs were novel. A protein database of Arabidopsis was used to examine the functional classification of these gene clusters. The GC-content in the 5 -UTR region of wheat cDNAs was compared to that of rice. Forty-three genes (3.5% of wheat cDNAs homologous to those of rice) possessed distinct GC-content in the 5 -UTR region, suggesting different breeding behaviors of wheat and rice.

5' Untranslated Regions↗

Large-scale collection and characterization of promoters of human and mouse genes.

We report the generation and initial characterization of a large-scale collection of sequences of putative promoter regions (PPRs) of human and mouse genes. Based on our unique collection of 400,225 and 580,209 human and mouse full-length cDNAs, we determined exact transcriptional start sites (TSSs). Using positional information of the TSSs, we could retrieve adjacent sequences as PPRs for 8,793 and 6,875 human and mouse genes, respectively. The positions of the PPRs were 4 kb upstream to previously reported 5'-ends of cDNAs on average, demonstrating that full-length cDNA information is indispensable for this purpose. Among those PPRs supported by experimentally validated TSSs, 3,324 could be paired as mutually homologous genes between human and mouse and were used for the comprehensive comparative studies. The sequence identities in the proximal regions of the TSSs were 45% on average, and 22,794 putative transcription factor binding sites that are conserved between human and mouse were identified. The data resource created in the present work and the results of the sequences' initial characterization should lay the firm foundation for deciphering the transcriptional modulations of human genes. All the data were deposited and made available through a database for comparative studies, DBTSS.

Animals↗

Comprehensive analysis of NAC family genes in Oryza sativa and Arabidopsis thaliana.

The NAC domain was originally characterized from consensus sequences from petunia NAM and from Arabidopsis ATAF1, ATAF2, and CUC2. Genes containing the NAC domain (NAC family genes) are plant-specific transcriptional regulators and are expressed in various developmental stages and tissues. We performed a comprehensive analysis of NAC family genes in Oryza sativa (a monocot) and Arabidopsis thaliana (a dicot). We found 75 predicted NAC proteins in full-length cDNA data sets of O. sativa (28,469 clones) and 105 in putative genes (28,581 sequences) from the A. thaliana genome. NAC domains from both predicted and known NAC family proteins were classified into two groups and 18 subgroups by sequence similarity. There were a few differences in amino acid sequences in the NAC domains between O. sativa and A. thaliana. In addition, we found 13 common sequence motifs from transcriptional activation regions in the C-terminal regions of predicted NAC proteins. These motifs probably diverged having correlations with NAC domain structures. We discuss the relationship between the structure and function of the NAC family proteins in light of our results and the published data. Our results will aid further functional analysis of NAC family genes.

Arabidopsis↗

Genomics approach to abscisic acid- and gibberellin-responsive genes in rice.

We used an 8987-EST collection to construct a cDNA microarray system with various genomics information (full-length cDNA, expression profile, high accuracy genome sequence, phenotype, genetic map, and physical map) in rice. This array was used as a probe to hybridize target RNAs prepared from normally grown callus of rice and from callus treated for 6 hr or 3 days with the hormones abscisic acid (ABA) or gibberellin (GA). We identified 509 clones, including many clones that had never been annotated as ABA-or GA-responsive. These genes included not only ABA- or GA-responsive genes but also genes responsive to other physiological conditions such as pathogen infection, heat shock, and metal ion stress. Comparison of ABA- and GA-responsive genes revealed antagonistic regulation for these genes by both hormones except for one defense-related gene, thionin. The gene for thionin was up-regulated by both hormone treatments for 3 days. The upstream regions of all the genes that were regulated by both hormones had cis-elements for ABA and GA response. We performed a clustering analysis of genes regulated by both hormones and various expression profiles that showed three notable clusters (seed tissues, low temperature and sugar starvation, and thionin-gene related). A comparison of the cis-elements for hormone response genes between rice and Arabidopsis thaliana, we identified cis-elements for dehydration-stress response or for expression of amylase gene as Arabidopsis gene-specific or rice gene-specific, respectively.

Abscisic Acid↗

Antisense transcripts with rice full-length cDNAs.

BACKGROUND: Natural antisense transcripts control gene expression through post-transcriptional gene silencing by annealing to the complementary sequence of the sense transcript. Because many genome and mRNA sequences have become available recently, genome-wide searches for sense-antisense transcripts have been reported, but few plant sense-antisense transcript pairs have been studied. The Rice Full-Length cDNA Sequencing Project has enabled computational searching of a large number of plant sense-antisense transcript pairs. RESULTS: We identified sense-antisense transcript pairs from 32,127 full-length rice cDNA sequences produced by this project and public rice mRNA sequences by aligning the cDNA sequences with rice genome sequences. We discovered 687 bidirectional transcript pairs in rice, including sense-antisense transcript pairs. Both sense and antisense strands of 342 pairs (50%) showed homology to at least one expressed sequence tag other than that of the pair. Microarray analysis showed 82 pairs (32%) out of 258 pairs on the microarray were more highly expressed than the median expression intensity of 21,938 rice transcriptional units. Both sense and antisense strands of 594 pairs (86%) had coding potential. CONCLUSIONS: The large number of plant sense-antisense transcript pairs suggests that gene regulation by antisense transcripts occurs in plants and not only in animals. On the basis of our results, experiments should be carried out to analyze the function of plant antisense transcripts.

DNA, Antisense↗