Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “RNA sequencing analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Exploiting conserved structure for faster annotation of non-coding RNAs without loss of accuracy.

MOTIVATION: Non-coding RNAs (ncRNAs)-functional RNA molecules not coding for proteins-are grouped into hundreds of families of homologs. To find new members of an ncRNA gene family in a large genome database, covariance models (CMs) are a useful statistical tool, as they use both sequence and RNA secondary structure information. Unfortunately, CM searches are slow. Previously, we introduced 'rigorous filters', which provably sacrifice none of CMs' accuracy, although often scanning much faster. A rigorous filter, using a profile hidden Markov model (HMM), is built based on the CM, and filters the genome database, eliminating sequences that provably could not be annotated as homologs. The CM is run only on the remainder. Some biologically important ncRNA families could not be scanned efficiently with this technique, largely due to the significance of conserved secondary structure relative to primary sequence in identifying these families. Current heuristic filters are also expected to perform poorly on such families. RESULTS: By augmenting profile HMMs with limited secondary structure information, we obtain rigorous filters that accelerate CM searches for virtually all known ncRNA families from the Rfam Database and tRNA models in tRNAscan-SE. These filters scan an 8 gigabase database in weeks instead of years, and uncover homologs missed by heuristic techniques to speed CM searches. AVAILABILITY: Software in development; contact the authors.

Algorithms↗

MoD Tools: regulatory motif discovery in nucleotide sequences from co-regulated or homologous genes.

Understanding the complex mechanisms regulating gene expression at the transcriptional and post-transcriptional levels is one of the greatest challenges of the post-genomic era. The MoD (MOtif Discovery) Tools web server comprises a set of tools for the discovery of novel conserved sequence and structure motifs in nucleotide sequences, motifs that in turn are good candidates for regulatory activity. The server includes the following programs: Weeder, for the discovery of conserved transcription factor binding sites (TFBSs) in nucleotide sequences from co-regulated genes; WeederH, for the discovery of conserved TFBSs and distal regulatory modules in sequences from homologous genes; RNAProfile, for the discovery of conserved secondary structure motifs in unaligned RNA sequences whose secondary structure is not known. In this way, a given gene can be compared with other co-regulated genes or with its homologs, or its mRNA can be analyzed for conserved motifs regulating its post-transcriptional fate. The web server thus provides researchers with different strategies and methods to investigate the regulation of gene expression, at both the transcriptional and post-transcriptional levels. Available at http://www.pesolelab.it/modtools/ and http://www.beacon.unimi.it/modtools/.

Binding Sites↗

Structure of a precursor to the yeast mitochondrial tRNAMetf. Implications for the function of the tRNA synthesis locus.

A transcript from the yeast mitochondrial tRNAMetf gene has been isolated from a petite deletion mutant, ND40. RNA sequence analysis demonstrates that it has a 5' unprocessed extension of 28 nucleotides, and capping experiments with guanylyltransferase reveal that the first nucleotide has a 5' di- or triphosphate. A comparison of the RNA sequence with the tRNAMetf gene sequence reported here shows that transcription initiation occurs at a sequence homologous to a nonanucleotide segment implicated as a promoter element of yeast mitochondrial DNA. The 3' end of this transcript is identical to the mature tRNAMetf and carries a CCA sequence. The transcript can be processed in vitro to yield a mature tRNAMetf and thus appears to be a bona fide tRNA precursor. This tRNA precursor accumulates in the absence of the mitochondrial tRNA synthesis locus whereas mature tRNAMetf can be made from the same gene in the presence of the locus. This data provides clear and convincing evidence that the synthesis locus codes for a 5' tRNA processing function.

Base Sequence↗

RNA editing in transcripts of the mitochondrial genes of the insect trypanosome Crithidia fasciculata.

With the aid of cDNA and RNA sequence analysis, we have determined to what extent transcripts of mitochondrial maxicircle genes of the insect trypanosome Crithidia fasciculata are altered by RNA editing, a novel mechanism of gene expression which operates via the insertion and deletion of uridine residues. Editing of cytochrome c oxidase (cox) subunit II and III transcripts and of maxicircle unidentified reading frame (MURF) 2 RNA is limited to a small section and results in the creation of a potential AUG translational initiation codon (coxIII, MURF2) or the removal of a frameshift (coxII). No differences with the genomic sequences were observed in the remainder of these RNAs. Surprisingly, NADH dehydrogenase subunit I transcripts were completely unedited in the coding region, implying that an AUG translational initiation codon is absent. The partial ribosomal RNA sequences determined also conform to the gene sequences. Together these results lead to the conclusion that the unusual sequences predicted by the protein and rRNA genes must indeed be present in the gene products. Editing also occurred in the poly(A) tail of RNAs from all protein genes, including those that are unedited in the coding region. The tails display a large variation in AU sequence motifs. Finally, some cDNAs contained sequences absent from both the DNA and the edited RNA. Some of these may represent intermediates in the RNA editing process. We argue, however, that long runs of T may be artefacts of cDNA synthesis.

Animals↗

Isolation and genomic analysis of the rat polymeric immunoglobulin receptor gene terminal domain and transcriptional control region.

The polymeric immunoglobulin receptor (pIgR) transports IgA and IgM across secretory epithelial cells and is essential in external immunity maintenance. We report here the structural characterization of the single-copy rat gene distributed over 30 kb of chromosomal DNA and analysis of its transcriptional control region. RNA sequencing and genomic analysis show a 5' terminal region originates at a major (+1) and a minor site producing an unusual 124-bp nontranslated exon I separated from a small 96-bp initiator ATG coding exon II by a 7.5-kb intron. The pIgR 5' region comprises a structured promoter with abundant helix-loop-helix (bHLH) cis elements positioned within an equivalent internal -70, -290, -528, and three centered at -745. The three latter bHLH elements each occur within 30-bp repeats at -690 to -780. Transient expression assays show a 1.3-kb 5' region is sufficient to drive expression in rat primary hepatocyte monolayer cultures, transformed human hepatic (HepG2) cells, and a mammary epithelial tumor cell line MCF-7, but is inactive in the rodent fibroblast 3T3 cell line. A minimal transcriptional promoter domain was deduced from sequentially deleted vectors revealing a +40 to -922 sequence to be sufficient for full activity. Further deletions within this region yield incremental losses in cis activity, indicating that multiple subregions comprise an extended transcriptional control region.

Animals↗

Novel aromatase transcripts from bovine placenta contain repeated sequence motifs.

Aromatase cytochrome P-450 (Aro) is the major enzyme of estrogen biosynthesis. The aim of the present investigation was the isolation and comparative sequence analysis of the bovine aromatase cytochrome P-450 transcript (bCyp19) from a placental lambda gt10 cDNA library. From three overlapping clones, a total sequence of 5180 bp could be derived, including two polyadenylation sites and signals located next to each other. As found in other species, the open reading frame (ORF) comprises 1509 bp and shows 87, 78 and 78% sequence homology to the coding areas of the human, rat and mouse genes, respectively. The 3'-untranslated region (UTR) of the bovine transcript is about 2-kb longer than that of the human gene (hCYP19). It contains homologous retroposon elements of the bovidae dimer family (BDF) at two different positions, and ends with a sequence motif which also occurs repeatedly within the bovine genome. The 5'-UTR isolated from placenta includes a new sequence upstream from exon II that was not found in cattle or other species so far. We conclude from our data that (i) as found in other species, bCyp19 is likely to be transcribed into different mRNA species, (ii) the bovine 3'-UTR was the target for multiple insertions of repeated sequence motifs, (iii) the unusual length of the bCyp19 transcript is mainly due to the long 3'-UTR, (iv) it includes sequences which are found in humans only on the genomic level, conceivably due to mutational inactivation of a primordial polyadenylation signal (PAS) and (v) the recently used, functional PAS is contributed by a downstream bovine repeat element.

Amino Acid Sequence↗

Application of tetranucleotide frequencies for the assignment of genomic fragments.

A basic problem of the metagenomic approach in microbial ecology is the assignment of genomic fragments to a certain species or taxonomic group, when suitable marker genes are absent. Currently, the (G + C)-content together with phylogenetic information and codon adaptation for functional genes is mostly used to assess the relationship of different fragments. These methods, however, can produce ambiguous results. In order to evaluate sequence-based methods for fragment identification, we extensively compared (G + C)-contents and tetranucleotide usage patterns of 9054 fosmid-sized genomic fragments generated in silico from 118 completely sequenced bacterial genomes (40 982 931 fragment pairs were compared in total). The results of this systematic study show that the discriminatory power of correlations of tetranucleotide-derived z-scores is by far superior to that of differences in (G + C)-content and provides reasonable assignment probabilities when applied to metagenome libraries of small diversity. Using six fully sequenced fosmid inserts from a metagenomic analysis of microbial consortia mediating the anaerobic oxidation of methane (AOM), we demonstrate that discrimination based on tetranucleotide-derived z-score correlations was consistent with corresponding data from 16S ribosomal RNA sequence analysis and allowed us to discriminate between fosmid inserts that were indistinguishable with respect to their (G + C)-contents.

Bacteria↗

Bacillus sp. WW3-SN6, a novel facultatively alkaliphilic bacterium isolated from the washwaters of edible olives.

A novel Gram-positive facultatively alkaliphilic, sporulating, rod-shaped bacterium, designated as WW3-SN6, has been isolated from the alkaline washwaters derived from the preparation of edible olives. The bacterium is nonmotile, and flagella are not observed. It is oxidase positive and catalase negative. The facultative alkaliphile grows from pH 7.0 to 10.5, with a broad optimum from pH 8.0 to 9.0. It could grow in up to 15% (w/v) NaCl, and over the temperature range from 4 degrees to 37 degrees C, with an optimum between 27 degrees and 32 degrees C: therefore, it is both halotolerant and psychrotolerant. The bacterium is sensitive to a range of beta-lactam, sulfonamide, and aminoglycoside antibiotics, but resistant to trimethoprim. The range of amino acids, sugars, and polyols utilized as growth substrates indicates that this alkaliphile is a heterotrophic bacterium. D(+)-glucose, D(+)-glucose-6-phosphate, D(+)-cellobiose, starch, or sucrose are the substrates best utilized. The major membrane lipids are phosphatidylglycerol and diphosphatidylglycerol, with smaller amounts of phosphatidylethanolamine and an unknown phospholipid. During growth at high pH, the proportion of phosphatidylglycerol is increased relative to phosphatidylethanolamine. The fatty acyl components in the membrane phospholipids are mainly branched chain, with 13-methyl tetradecanoic and 12-methyl tetradecanoic acids as the predominant components. The G + C content of the genomic DNA is 41.1 +/- 1.0 mol%. The results of 16S ribosomal RNA sequence analysis place this alkaliphilic bacterium in a cluster, together with an unnamed alkaliphilic Bacillus species (98.2% similarity).

Alkalies↗

Ribosome-binding sites on chloroplast rbcL and psbA mRNAs and light-induced initiation of D1 translation.

Chloroplast ribosome-binding sites were identified on the plastid rbcL and psbA mRNAs using toeprint analysis. The rbcL translation initiation domain is highly conserved and contains a prokaryotic Shine-Dalgarno (SD) sequence (GGAGG) located 4 to 12 nucleotides upstream of the initiator AUG. Toeprint analysis of rbcL mRNA associated with plastid polysomes revealed strong toeprint signals 15 nucleotides downstream from the AUG indicating ribosome binding at the translation initiation site. Escherichia coli 30S ribosomes generated similar toeprint signals when mixed with rbcL mRNA in the presence of initiator tRNA. These results indicate that plastid SD sequences are functional in chloroplast translation initiation. The psbA initiator region lacks a SD sequence within 12 nucleotides of the initiator AUG. However, toeprint analysis of soluble and membrane polysome-associated psbA mRNA revealed ribosomes bound to the initiator region. E. coli 30S ribosomes did not associate with the psbA translation initiation region. E. coli and chloroplast ribosomes bind to an upstream region which contains a conserved SD-like sequence. Therefore, translation initiation on psbA mRNA may involve the transient binding of chloroplast ribosomes to this upstream SD-like sequence followed by scanning to localize the initiator AUG. Illumination 8-day-old dark-grown barley seedlings caused an increase in polysome-associated psbA mRNA and the abundance of initiation complexes bound to psbA mRNA. These results demonstrate that light modulates D1 translation initiation in plastids of older dark-grown barley seedlings.

Bacterial Proteins↗

An algorithm for selection of functional siRNA sequences.

Randomly designed siRNA targeting different positions within the same mRNA display widely differing activities. We have performed a statistical analysis of 46 siRNA, identifying various features of the 19bp duplex that correlate significantly with functionality at the 70% knockdown level and verified these results against an independent data set of 34 siRNA recently reported by others. Features that consistently correlated positively with functionality across the two data sets included an asymmetry in the stability of the duplex ends (measured as the A/U differential of the three terminal basepairs at either end of the duplex) and the motifs S1, A6, and W19. The presence of the motifs U1 or G19 was associated with lack of functionality. A selection algorithm based on these findings strongly differentiated between the two functional groups of siRNA in both data sets and proved highly effective when used to design siRNA targeting new endogenous human genes.

Algorithms↗

Phylogenetic analysis of the spirochetes.

The 16S rRNA sequences were determined for species of Spirochaeta, Treponema, Borrelia, Leptospira, Leptonema, and Serpula, using a modified Sanger method of direct RNA sequencing. Analysis of aligned 16S rRNA sequences indicated that the spirochetes form a coherent taxon composed of six major clusters or groups. The first group, termed the treponemes, was divided into two subgroups. The first treponeme subgroup consisted of Treponema pallidum, Treponema phagedenis, Treponema denticola, a thermophilic spirochete strain, and two species of Spirochaeta, Spirochaeta zuelzerae and Spirochaeta stenostrepta, with an average interspecies similarity of 89.9%. The second treponeme subgroup contained Treponema bryantii, Treponema pectinovorum, Treponema saccharophilum, Treponema succinifaciens, and rumen strain CA, with an average interspecies similarity of 86.2%. The average interspecies similarity between the two treponeme subgroups was 84.2%. The division of the treponemes into two subgroups was verified by single-base signature analysis. The second spirochete group contained Spirochaeta aurantia, Spirochaeta halophila, Spirochaeta bajacaliforniensis, Spirochaeta litoralis, and Spirochaeta isovalerica, with an average similarity of 87.4%. The Spirochaeta group was related to the treponeme group, with an average similarity of 81.9%. The third spirochete group contained borrelias, including Borrelia burgdorferi, Borrelia anserina, Borrelia hermsii, and a rabbit tick strain. The borrelias formed a tight phylogenetic cluster, with average similarity of 97%. THe borrelia group shared a common branch with the Spirochaeta group and was closer to this group than to the treponemes. A single spirochete strain isolated fromt the shew constituted the fourth group. The fifth group was composed of strains of Serpula (Treponema) hyodysenteriae and Serpula (Treponema) innocens. The two species of this group were closely related, with a similarity of greater than 99%. Leptonema illini, Leptospira biflexa, and Leptospira interrogans formed the sixth and most deeply branching group. The average similarity within this group was 83.2%. This study represents the first demonstration that pathogenic and saprophytic Leptospira species are phylogenetically related. The division of the spirochetes into six major phylogenetic clusters was defined also by sequence signature elements. These signature analyses supported the conclusion that the spirochetes represent a monophylectic bacterial phylum.

Animals↗

Signal recognition particle receptor (SRPR) is downregulated in a rat model of cyclosporin A-induced gingival overgrowth.

Differential display is a powerful technique which can be used to identify those genes whose expression is altered between two or more tissues under investigation. We have applied differential display to a rat model of cyclosporin A-induced gingival overgrowth (CIGO) to identify genes which are differentially expressed as a result of drug treatment. Ten weanling Wistar rats were fed with a pelleted diet containing cyclosporin A (CsA) at 120 mg/kg for 10 d and then 200 mg/kg for a further 30 d prior to culling. Experimental rats were compared with 10 age/sex-matched rats on a control diet. Significant evidence of overgrowth was observed in the interdental papilla between the mandibular first and second molar teeth in the CsA group. Differential display was performed on total cellular RNA extracted from the mandibular buccal gingiva. A cDNA product was isolated which was underexpressed in the overgrowth tissue and demonstrated a 95% sequence homology to the human signal recognition particle receptor (Human Docking Protein). Preliminary studies indicate that this gene is also underexpressed in human CIGO tissue. The method of approach and the potential implications of our findings are discussed.

Animals↗

Inventory and analysis of the protein subunits of the ribonucleases P and MRP provides further evidence of homology between the yeast and human enzymes.

The RNases P and MRP are involved in tRNA and rRNA processing, respectively. Both enzymes in eukaryotes are composed of an RNA molecule and 9-12 protein subunits. Most of the protein subunits are shared between RNases P and MRP. We have here performed a computational analysis of the protein subunits in a broad range of eukaryotic organisms using profile-based searches and phylogenetic methods. A number of novel homologues were identified, giving rise to a more complete inventory of RNase P/MRP proteins. We present evidence of a relationship between fungal Pop8 and the protein subunit families Rpp14/Pop5 as well as between fungal Pop6 and metazoan Rpp25. These relationships further emphasize a structural and functional similarity between the yeast and human P/MRP complexes. We have also identified novel P and MRP RNAs and analysis of all available sequences revealed a K-turn motif in a large number of these RNAs. We suggest that this motif is a binding site for the Pop3/Rpp38 proteins and we discuss other structural features of the RNA subunit and possible relationships to the protein subunit repertoire.

Amino Acid Sequence↗

Expression of class I O-methyltransferase in healthy and TMV-infected tobacco.

Tobacco possesses two distinct classes of O-methyltransferases (OMTs; S-adenosyl-L-methionine:o-diphenol O-methyltransferases; EC 2.1.1.6). Here we report on the cloning and the expression pattern of the class I OMT that is specifically involved in lignin biosynthesis. Near-full-length cDNAs have been isolated from tobacco libraries constructed from leaf and stem poly(A)+ RNA. Sequence analysis demonstrated that the OMT I clones derived from two different mRNA species. The two types of OMT I mRNA were reverse transcribed from total RNA and the cDNAs were amplified by polymerase chain reaction and characterized by restriction analysis. The same proportion of the two transcripts was measured in stem tissue of healthy plants and in leaves reacting hypersensitively to tobacco mosaic virus, indicating a coordinate expression of the two OMT I genes. Consistently, genomic hybridization indicated the presence of two OMT I genes in the amphidiploid genome of tobacco. The pattern of expression of OMT I genes was studied by in situ mRNA hybridization. In stem, petiole, and root tissues, OMT I genes were found to be specifically expressed in vascular cells and epidermis. In healthy leaves OMT I mRNA was only detected in vascular strands, whereas, in leaves bearing tobacco mosaic virus-induced necrotic lesions, a particularly strong accumulation of the labeling was also localized in the upper and lower epidermis.

Amino Acid Sequence↗

The molecular genetic basis of Glanzmann's thrombasthenia in a gypsy population in France: identification of a new mutation on the alpha IIb gene.

Glanzmann's thrombasthenia is a rare inherited bleeding disorder caused by a qualitative or quantitative defect of platelet alpha IIb beta 3. We describe here a new mutation that is the molecular genetic basis of Glanzmann's thrombasthenia in two gypsy families. Our investigation was focused on the alpha IIb gene as a result of biochemical and immunologic analysis of patients' platelets showing undetectable alpha IIb but residual beta 3 levels. The entire alpha IIb cDNA was polymerase chain reaction (PCR) amplified using patients platelet RNA. Sequence analysis showed an 8-bp deletion located at the 3' end of exon 15. This deletion causes a reading-frame shift leading to a premature stop codon and the synthesis of a severely truncated form of alpha IIb. Genomic DNA study showed a G-->A substitution, the Gypsy mutation, at the splice donor site of intron 15. This mutation results in an abnormal splicing occurring at an alternative donor site located 8 bp upstream from the mutation. Based on those results, an allele-specific PCR analysis was developed to allow a rapid identification of the mutation in patients and potential carriers of the gypsy community. This PCR analysis can also be used for genetic counseling and antenatal diagnosis.

Alleles↗

An in vitro model of hepatitis C virion production.

The hepatitis C virus (HCV) is a major cause of liver disease worldwide. The understanding of the viral life cycle has been hampered by the lack of a satisfactory cell culture system. The development of the HCV replicon system has been a major advance, but the system does not produce virions. In this study, we constructed an infectious HCV genotype 1b cDNA between two ribozymes that are designed to generate the exact 5' and 3' ends of HCV. A second construct with a mutation in the active site of the viral RNA-dependent RNA polymerase (RdRp) was generated as a control. The HCV-ribozyme expression construct was transfected into Huh7 cells. Both HCV structural and nonstructural proteins were detected by immunofluorescence and Western blot. RNase protection assays showed positive- and negative-strand HCV RNA. Sequence analysis of the 5' and 3' ends provided further evidence of viral replication. Sucrose density gradient centrifugation of the culture medium revealed colocalization of HCV RNA and structural proteins in a fraction with the density of 1.16 g/ml, the putative density of HCV virions. Electron microscopy showed viral particles of approximately 50 nm in diameter. The level of HCV RNA in the culture medium was as high as 10 million copies per milliliter. The HCV-ribozyme construct with the inactivating mutation in the RdRp did not show evidence of viral replication, assembly, and release. This system supports the production and secretion of high-level HCV virions and extends the repertoire of tools available for the study of HCV biology.

Base Sequence↗

RNA structures and folding: from conventional to new issues in structure predictions.

Prediction and modeling of RNA structures has become an indispensable tool of biological research disciplines. Currently, reliable predictions require massive input of experimental data. Structure-forming elements are conventional base pairs, as well as a rapidly increasing repertoire of novel structural motifs. New developments extend structural analysis beyond the one-sequence/one-structure paradigm and allow questions that are relevant to molecular evolution to be answered.

Algorithms↗

Mapping of determinants required for the function of the HIV-1 env nuclear retention sequence.

Control of HIV-1 RNA processing and transport are critical to the successful replication of the virus. In previous work, we identified a region within the HIV-1 env that is involved in mediating nuclear retention of unspliced viral RNA. To define this sequence further and identify elements required for function, deletion mutagenesis was carried out. Progressive 5' and 3' deletions map the nuclear retention sequence (NRS) within the intron between nts 8281 and 8381. While deletion of sequences comprising the 3'ss had no effect, removal of the 5'ss resulted in cytoplasmic accumulation of unspliced RNA. Sequence analysis determined that the region corresponding to the NRS is highly conserved among HIV-1 strains. To evaluate whether this NRS interacts with cellular factors, RNA electrophoretic mobility shift assays (REMSA) were performed. We show that the NRS specifically interacts with cellular factors present in HeLa nuclear extracts, and, by UV crosslinking, correlates with the binding of a 49-kDa protein. Immunoprecipitation of the UV crosslinked products determined that this 49-kDa protein corresponds to hnRNP C.

Cell Nucleus↗