Search PubMed⌕ Search

Biomedical subjects

Piero Carninci

Publications and source records attributed to Piero Carninci.

At least 73 records · Page 4Linked to original sources

Cap analysis gene expression for high-throughput analysis of transcriptional starting point and identification of promoter usage.

We introduce cap analysis gene expression (CAGE), which is based on preparation and sequencing of concatamers of DNA tags deriving from the initial 20 nucleotides from 5' end mRNAs. CAGE allows high-throughout gene expression analysis and the profiling of transcriptional start points (TSP), including promoter usage analysis. By analyzing four libraries (brain, cortex, hippocampus, and cerebellum), we redefined more accurately the TSPs of 11-27% of the analyzed transcriptional units that were hit. The frequency of CAGE tags correlates well with results from other analyses, such as serial analysis of gene expression, and furthermore maps the TSPs more accurately, including in tissue-specific cases. The high-throughput nature of this technology paves the way for understanding gene networks via correlation of promoter usage and gene transcriptional factor expression.

Animals↗

A novel feature of microsatellites in plants: a distribution gradient along the direction of transcription.

A computer-based analysis was conducted to assess the characteristics of microsatellites in transcribed regions of rice and Arabidopsis. In addition, two mammals were simultaneously analyzed for a comparative analysis. Our analyses confirmed a novel plant-specific feature in which there is a gradient in microsatellite density along the direction of transcription. It was also confirmed that pyrimidine-rich microsatellites are found intensively near the transcription start sites, specifically in the two plants, but not in the mammals. Our results suggest that microsatellites located at high frequency in the 5'-flanking regions of plant genes can potentially act as factors in regulating gene expression.

5' Flanking Region↗

Empirical analysis of transcriptional activity in the Arabidopsis genome.

Functional analysis of a genome requires accurate gene structure information and a complete gene inventory. A dual experimental strategy was used to verify and correct the initial genome sequence annotation of the reference plant Arabidopsis. Sequencing full-length cDNAs and hybridizations using RNA populations from various tissues to a set of high-density oligonucleotide arrays spanning the entire genome allowed the accurate annotation of thousands of gene structures. We identified 5817 novel transcription units, including a substantial amount of antisense gene transcription, and 40 genes within the genetically defined centromeres. This approach resulted in completion of approximately 30% of the Arabidopsis ORFeome as a resource for global functional experimentation of the plant proteome.

Arabidopsis↗

Collection, mapping, and annotation of over 28,000 cDNA clones from japonica rice.

We collected and completely sequenced 28,469 full-length complementary DNA clones from Oryza sativa L. ssp. japonica cv. Nipponbare. Through homology searches of publicly available sequence data, we assigned tentative protein functions to 21,596 clones (75.86%). Mapping of the cDNA clones to genomic DNA revealed that there are 19,000 to 20,500 transcription units in the rice genome. Protein informatics analysis against the InterPro database revealed the existence of proteins presented in rice but not in Arabidopsis. Sixty-four percent of our cDNAs are homologous to Arabidopsis proteins.

Alternative Splicing↗

Genetic control of the innate immune response.

BACKGROUND: Susceptibility to infectious diseases is directed, in part, by the interaction between the invading pathogen and host macrophages. This study examines the influence of genetic background on host-pathogen interactions, by assessing the transcriptional responses of macrophages from five inbred mouse strains to lipopolysaccharide (LPS), a major determinant of responses to gram-negative microorganisms. RESULTS: The mouse strains examined varied greatly in the number, amplitude and rate of induction of genes expressed in response to LPS. The response was attenuated in the C3H/HeJlpsd strain, which has a mutation in the LPS receptor Toll-like receptor 4 (TLR4). Variation between mouse strains allowed clustering into early (C57Bl/6J and DBA/2J) and delayed (BALB/c and C3H/ARC) transcriptional phenotypes. There was no clear correlation between gene induction patterns and variation at the Bcg locus (Slc11A1) or propensity to bias Th1 versus Th2 T cell activation responses. CONCLUSION: Macrophages from each strain responded to LPS with unique gene expression profiles. The variation apparent between genetic backgrounds provides insights into the breadth of possible inflammatory responses, and paradoxically, this divergence was used to identify a common transcriptional program that responds to TLR4 signalling, irrespective of genetic background. Our data indicates that many additional genetic loci control the nature and the extent of transcriptional responses promoted by a single pathogen-associated molecular pattern (PAMP), such as LPS.

Animals↗

Comparative genomics of Physcomitrella patens gametophytic transcriptome and Arabidopsis thaliana: implication for land plant evolution.

The mosses and flowering plants diverged >400 million years ago. The mosses have haploid-dominant life cycles, whereas the flowering plants are diploid-dominant. The common ancestors of land plants have been inferred to be haploid-dominant, suggesting that genes used in the diploid body of flowering plants were recruited from the genes used in the haploid body of the ancestors during the evolution of land plants. To assess this evolutionary hypothesis, we constructed an EST library of the moss Physcomitrella patens, and compared the moss transcriptome to the genome of Arabidopsis thaliana. We constructed full-length enriched cDNA libraries from auxin-treated, cytokinin-treated, and untreated gametophytes of P. patens, and sequenced both ends of >40,000 clones. These data, together with the mRNA sequences in the public databases, were assembled into 15,883 putative transcripts. Sequence comparisons of A. thaliana and P. patens showed that at least 66% of the A. thaliana genes had homologues in P. patens. Comparison of the P. patens putative transcripts with all known proteins, revealed 9,907 putative transcripts with high levels of similarity to vascular plant genes, and 850 putative transcripts with high levels of similarity to other organisms. The haploid transcriptome of P. patens appears to be quite similar to the A. thaliana genome, supporting the evolutionary hypothesis. Our study also revealed that a number of genes are moss specific and were lost in the flowering plant lineage.

Arabidopsis↗

Multiple tissue-specific promoters control expression of the murine tartrate-resistant acid phosphatase gene.

Tartrate-resistant acid phosphatase (TRAP) is highly expressed in osteoclasts and in a subset of tissue macrophages and dendritic cells. It is expressed at lower levels in the parenchymal cells of the liver, glomerular mesangial cells of the kidney and pancreatic acinar cells. We have identified novel TRAP mRNAs that differ in their 5'-untranslated region (5'-UTR) sequence, but align with the known murine TRAP mRNA from the first base of Exon 2. The novel 5'-UTRs represent alternative first exons located upstream of the known 5'-UTR. A similar genomic structure exists for the human TRAP gene with partial conservation of the exon and promoter sequences. Expression of the most distal 5'-UTR (Exon 1A) is restricted to adult bone and spleen tissue. Exon 1B is expressed primarily in tissues containing TRAP-positive non-haematopoietic cells. The known TRAP 5'-UTR (Exon 1C) is expressed in tissues characteristic of myeloid cell expression. In addition the Exon 1C promoter sequence is shown to comprise distinct transcription start regions, with an osteoclast-specific transcription initiation site identified downstream of a TATA-like element. Macrophages are shown to initiate transcription of the Exon 1C transcript from a purine-rich region located upstream of the osteoclast-specific transcription start point. The distinct expression patterns for each of the TRAP 5'-UTRs suggest that TRAP mRNA expression is regulated by the use of four alternative tissue- and cell-restricted promoters.

5' Untranslated Regions↗

Continued discovery of transcriptional units expressed in cells of the mouse mononuclear phagocyte lineage.

The current RIKEN transcript set represents a significant proportion of the mouse transcriptome but transcripts expressed in the innate and acquired immune systems are poorly represented. In the present study we have assessed the complexity of the transcriptome expressed in mouse macrophages before and after treatment with lipopolysaccharide, a global regulator of macrophage gene expression, using existing RIKEN 19K arrays. By comparison to array profiles of other cells and tissues, we identify a large set of macrophage-enriched genes, many of which have obvious functions in endocytosis and phagocytosis. In addition, a significant number of LPS-inducible genes were identified. The data suggest that macrophages are a complex source of mRNA for transcriptome studies. To assess complexity and identify additional macrophage expressed genes, cDNA libraries were created from purified populations of macrophage and dendritic cells, a functionally related cell type. Sequence analysis revealed a high incidence of novel mRNAs within these cDNA libraries. These studies provide insights into the depths of transcriptional complexity still untapped amongst products of inducible genes, and identify macrophage and dendritic cell populations as a starting point for sampling the inducible mammalian transcriptome.

Animals↗

Targeting a complex transcriptome: the construction of the mouse full-length cDNA encyclopedia.

We report the construction of the mouse full-length cDNA encyclopedia,the most extensive view of a complex transcriptome,on the basis of preparing and sequencing 246 libraries. Before cloning,cDNAs were enriched in full-length by Cap-Trapper,and in most cases,aggressively subtracted/normalized. We have produced 1,442,236 successful 3'-end sequences clustered into 171,144 groups, from which 60,770 clones were fully sequenced cDNAs annotated in the FANTOM-2 annotation. We have also produced 547,149 5' end reads,which clustered into 124,258 groups. Altogether, these cDNAs were further grouped in 70,000 transcriptional units (TU),which represent the best coverage of a transcriptome so far. By monitoring the extent of normalization/subtraction, we define the tentative equivalent coverage (TEC),which was estimated to be equivalent to >12,000,000 ESTs derived from standard libraries. High coverage explains discrepancies between the very large numbers of clusters (and TUs) of this project,which also include non-protein-coding RNAs,and the lower gene number estimation of genome annotations. Altogether,5'-end clusters identify regions that are potential promoters for 8637 known genes and 5'-end clusters suggest the presence of almost 63,000 transcriptional starting points. An estimate of the frequency of polyadenylation signals suggests that at least half of the singletons in the EST set represent real mRNAs. Clones accounting for about half of the predicted TUs await further sequencing. The continued high-discovery rate suggests that the task of transcriptome discovery is not yet complete.

Animals↗

Analysis of the mouse transcriptome for genes involved in the function of the nervous system.

We analyzed the mouse Representative Transcript and Protein Set for molecules involved in brain function. We found full-length cDNAs of many known brain genes and discovered new members of known brain gene families, including Family 3 G-protein coupled receptors, voltage-gated channels, and connexins. We also identified previously unknown candidates for secreted neuroactive molecules. The existence of a large number of unique brain ESTs suggests an additional molecular complexity that remains to be explored.A list of genes containing CAG stretches in the coding region represents a first step in the potential identification of candidates for hereditary neurological disorders.

Adenine↗

Subtraction of cap-trapped full-length cDNA libraries to select rare transcripts.

The normalization and subtraction of highly expressed cDNAs from relatively large tissues before cloning dramatically enhanced the gene discovery by sequencing for the mouse full-length cDNA encyclopedia, but these methods have not been suitable for limited RNA materials. To normalize and subtract full-length cDNA libraries derived from limited quantities of total RNA, here we report a method to subtract plasmid libraries excised from size-unbiased amplified lambda phage cDNA libraries that avoids heavily biasing steps such as PCR and plasmid library amplification. The proportion of full-length cDNAs and the gene discovery rate are high, and library diversity can be validated by in silico randomization.

Gene Expression Profiling↗

Generation and initial analysis of more than 15,000 full-length human and mouse cDNA sequences.

The National Institutes of Health Mammalian Gene Collection (MGC) Program is a multiinstitutional effort to identify and sequence a cDNA clone containing a complete ORF for each human and mouse gene. ESTs were generated from libraries enriched for full-length cDNAs and analyzed to identify candidate full-ORF clones, which then were sequenced to high accuracy. The MGC has currently sequenced and verified the full ORF for a nonredundant set of >9,000 human and >6,000 mouse genes. Candidate full-ORF clones for an additional 7,800 human and 3,500 mouse genes also have been identified. All MGC sequences and clones are available without restriction through public databases and clone distribution networks (see http:mgc.nci.nih.gov).

Algorithms↗

Initial sequencing and comparative analysis of the mouse genome.

The sequence of the mouse genome is a key informational tool for understanding the contents of the human genome and a key experimental tool for biomedical research. Here, we report the results of an international collaboration to produce a high-quality draft sequence of the mouse genome. We also present an initial comparative analysis of the mouse and human genomes, describing some of the insights that can be gleaned from the two sequences. We discuss topics including the analysis of the evolutionary forces shaping the size, structure and sequence of the genomes; the conservation of large-scale synteny across most of the genomes; the much lower extent of sequence orthology covering less than half of the genomes; the proportions of the genomes under selection; the number of protein-coding genes; the expansion of gene families related to reproduction and immunity; the evolution of proteins; and the identification of intraspecies polymorphism.

Animals↗

Monitoring the expression pattern of around 7,000 Arabidopsis genes under ABA treatments using a full-length cDNA microarray.

Full-length cDNAs are essential for functional analysis of plant genes. Recently, cDNA microarray analysis has been developed for quantitative analysis of global and simultaneous analysis of expression profiles. Microarray technology is a powerful tool for identifying genes induced by environmental stimuli or stress and for analyzing their expression profiles in response to environmental signals. We prepared an Arabidopsis full-length cDNA microarray containing around 7,000 independent full-length cDNA groups and analyzed the expression profiles of genes. The transcripts of 245, 299, 54 and 213 genes increased after abscisic acid (ABA), drought-, cold-, and salt-stress treatments, respectively, with inducibilities more than fivefold compared with those of control genes [corrected]. The cDNA microarray analysis showed that many ABA-inducible genes were induced after drought- and high-salinity-stress treatments, and that there is more crosstalk between drought and ABA responses than between ABA and cold responses. Among the ABA-inducible genes identified, we identified 22 transcription factor genes, suggesting that many transcriptional regulatory mechanisms exist in the ABA signal transduction pathways.

Abscisic Acid↗

Functional annotation of a full-length Arabidopsis cDNA collection.

Full-length complementary DNAs (cDNAs) are essential for the correct annotation of genomic sequences and for the functional analysis of genes and their products. We isolated 155,144 RIKEN Arabidopsis full-length (RAFL) cDNA clones. The 3'-end expressed sequence tags (ESTs) of 155,144 RAFL cDNAs were clustered into 14,668 nonredundant cDNA groups, about 60% of predicted genes. We also obtained 5' ESTs from 14,034 nonredundant cDNA groups and constructed a promoter database. The sequence database of the RAFL cDNAs is useful for promoter analysis and correct annotation of predicted transcription units and gene products. Furthermore, the full-length cDNAs are useful resources for analyses of the expression profiles, functions, and structures of plant proteins.

Arabidopsis↗

Mapping of 19032 mouse cDNAs on mouse chromosomes.

Finding genes by the positional candidate approach requires abundant cDNAs mapped to chromosomes. To provide such important information, we computationally mapped 19032 of our mouse cDNAs to mouse chromosomes by using data from public databases. We used 2 approaches. In the first, we integrated the mapping data of cDNAs on the human genome, known gene-related data, and comparative mapping data. From this, we calculated map positions on the mouse chromosomes. For this first approach, we developed a simple and powerful criterion to choose the correct map position from candidate positions in sequence homology searches. In the second approach, we related cDNAs to expressed sequence tags (EST) previously mapped in radiation hybrid experiments. We discuss improving the mapping by combining the 2 methods.

Animals↗

Monitoring the expression profiles of 7000 Arabidopsis genes under drought, cold and high-salinity stresses using a full-length cDNA microarray.

Full-length cDNAs are essential for functional analysis of plant genes in the post-sequencing era of the Arabidopsis genome. Recently, cDNA microarray analysis has been developed for quantitative analysis of global and simultaneous analysis of expression profiles. We have prepared a full-length cDNA microarray containing approximately 7000 independent, full-length cDNA groups to analyse the expression profiles of genes under drought, cold (low temperature) and high-salinity stress conditions over time. The transcripts of 53, 277 and 194 genes increased after cold, drought and high-salinity treatments, respectively, more than fivefold compared with the control genes. We also identified many highly drought-, cold- or high-salinity- stress-inducible genes. However, we observed strong relationships in the expression of these stress-responsive genes based on Venn diagram analysis, and found 22 stress-inducible genes that responded to all three stresses. Several gene groups showing different expression profiles were identified by analysis of their expression patterns during stress-responsive gene induction. The cold-inducible genes were classified into at least two gene groups from their expression profiles. DREB1A was included in a group whose expression peaked at 2 h after cold treatment. Among the drought, cold or high-salinity stress-inducible genes identified, we found 40 transcription factor genes (corresponding to approximately 11% of all stress-inducible genes identified), suggesting that various transcriptional regulatory mechanisms function in the drought, cold or high-salinity stress signal transduction pathways.

Arabidopsis↗

Inferring alternative splicing patterns in mouse from a full-length cDNA library and microarray data.

Although many studies on alternative splicing of specific genes have been reported in the literature, the general mechanism that regulates alternative splicing has not been clearly understood. In this study, we systematically aligned each pair of the 21,076 cDNA sequences of Mus musculus, searched for putative alternative splicing patterns, and constructed a list of potential alternative splicing sites. Two cDNAs are suspected to be alternatively spliced and originating from a common gene if they share most of their region with a high degree of sequence homology, but parts of the sequences are very distinctive or deleted in either cDNA. The list contains the following information: (1) tissue, (2) developmental stage, (3) sequences around splice sites, (4) the length of each gapped region, and (5) other comments. The list is available at http://www.bioinfo.sfc.keio.ac.jp/intron. Our results have predicted a number of unreported alternatively spliced genes, some of which are expressed only in a specific tissue or at a specific developmental stage.

Alternative Splicing↗