Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “RNA sequencing analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

RibAlign: a software tool and database for eubacterial phylogeny based on concatenated ribosomal protein subunits.

BACKGROUND: Until today, analysis of 16S ribosomal RNA (rRNA) sequences has been the de-facto gold standard for the assessment of phylogenetic relationships among prokaryotes. However, the branching order of the individual phlya is not well-resolved in 16S rRNA-based trees. In search of an improvement, new phylogenetic methods have been developed alongside with the growing availability of complete genome sequences. Unfortunately, only a few genes in prokaryotic genomes qualify as universal phylogenetic markers and almost all of them have a lower information content than the 16S rRNA gene. Therefore, emphasis has been placed on methods that are based on multiple genes or even entire genomes. The concatenation of ribosomal protein sequences is one method which has been ascribed an improved resolution. Since there is neither a comprehensive database for ribosomal protein sequences nor a tool that assists in sequence retrieval and generation of respective input files for phylogenetic reconstruction programs, RibAlign has been developed to fill this gap. RESULTS: RibAlign serves two purposes: First, it provides a fast and scalable database that has been specifically adapted to eubacterial ribosomal protein sequences and second, it provides sophisticated import and export capabilities. This includes semi-automatic extraction of ribosomal protein sequences from whole-genome GenBank and FASTA files as well as exporting aligned, concatenated and filtered sequence files that can directly be used in conjunction with the PHYLIP and MrBayes phylogenetic reconstruction programs. CONCLUSION: Up to now, phylogeny based on concatenated ribosomal protein sequences is hampered by the limited set of sequenced genomes and high computational requirements. However, hundreds of full and draft genome sequencing projects are on the way, and advances in cluster-computing and algorithms make phylogenetic reconstructions feasible even with large alignments of concatenated marker genes. RibAlign is a first step in this direction and may be particularly interesting to scientists involved in whole genome sequencing of representatives of new or sparsely studied eubacterial phyla. RibAlign is available at http://www.megx.net/ribalign.

Algorithms↗

Genome-wide analysis of polymerase III-transcribed Alu elements suggests cell-type-specific enhancer function.

Alu elements are one of the most successful families of transposons in the human genome. A portion of Alu elements is transcribed by RNA Pol III, whereas the remaining ones are part of Pol II transcripts. Because Alu elements are highly repetitive, it has been difficult to identify the Pol III-transcribed elements and quantify their expression levels. In this study, we generated high-resolution, long-genomic-span RAMPAGE data in 155 biosamples all with matching RNA-seq data and built an atlas of 17,249 Pol III-transcribed Alu elements. We further performed an integrative analysis on the ChIP-seq data of 10 histone marks and hundreds of transcription factors, whole-genome bisulfite sequencing data, ChIA-PET data, and functional data in several biosamples, and our results revealed that although the human-specific Alu elements are transcriptionally repressed, the older, expressed Alu elements may be exapted by the human host to function as cell-type-specific enhancers for their nearby protein-coding genes.

Alu Elements↗

Small RNA genes expressed from Staphylococcus aureus genomic and pathogenicity islands with specific expression among pathogenic strains.

Small RNA (sRNA) genes are expressed in all organisms, primarily as regulators of translation and message stability. We have developed comparative genomic approaches to identify sRNAs that are expressed by Staphylococcus aureus, the most common cause of hospital-acquired infections. This study represents an in-depth analysis of the RNome of a Gram-positive bacterium. A set of sRNAs candidates were identified in silico within intergenic regions, and their expression levels were monitored by using microarrays and confirmed by Northern blot hybridizations. Two sRNAs were also detected directly from purification and RNA sequence determination. In total, at least 12 sRNAs are expressed from the S. aureus genome, five from the core genome and seven from pathogenicity islands that confer virulence and antibiotic resistance. Three sRNAs are present in multiple (two to five) copies. For the sRNAs that are conserved throughout the bacterial phylogeny, their secondary structures were inferred by phylogenetic comparative methods. In vitro binding assays indicate that one sRNA encoded within a pathogenicity island is a trans-encoded antisense RNA regulating the expression of target genes at the posttranscriptional level. Some of these RNAs show large variations of expression among pathogenic strains, suggesting that they are involved in the regulation of staphylococcal virulence.

Base Pairing↗

The vasa locus in zebrafish: multiple RGG boxes from duplications.

Vasa is the most extensively characterized germ cell-specific marker in metazoans. We determined the sequence of the zebrafish vasa locus, which - together with the flanking regions - is 25 kb long. It contains 26 + 1 exons, with an average length of 83 bp. The 5' end of zebrafish vasa is rich in repeats; it includes a peculiar tandem repeat with three units spanning a total of eight exons, each unit containing several tandemly arranged RNA-binding motifs (RGG boxes). The presence of various, nonmicrosatellite type repeats in vasa seems to be universal among the species studied, and could have functional importance during the evolution of the gene, due to the increase in the number of RGG boxes located in these repeats. Analysis of vasa transcripts in zebrafish identified numerous isoforms resulting from alternative splicing (seven) and polyadenylation (four). Mapping to two radiation hybrid panels strengthened the position of the gene on LG10 of zebrafish.

Animals↗

Characterization of a Brome mosaic virus strain and its use as a vector for gene silencing in monocotyledonous hosts.

Virus-induced gene silencing (VIGS) is used to analyze gene function in dicotyledonous plants but less so in monocotyledonous plants (particularly rice and corn), partially due to the limited number of virus expression vectors available. Here, we report the cloning and modification for VIGS of a virus from Festuca arundinacea Schreb. (tall fescue) that caused systemic mosaic symptoms on barley, rice, and a specific cultivar of maize (Va35) under greenhouse conditions. Through sequencing, the virus was determined to be a strain of Brome mosaic virus (BMV). The virus was named F-BMV (F for Festuca), and genetic determinants that controlled the systemic infection of rice were mapped to RNAs 1 and 2 of the tripartite genome. cDNA from RNA 3 of the Russian strain of BMV (R-BMV) was modified to accept inserts from foreign genes. Coinoculation of RNAs 1 and 2 from F-BMV and RNA 3 from R-BMV expressing a portion of a plant gene to leaves of barley, rice, and maize plants resulted in visual silencing-like phenotypes. The visual phenotypes were correlated with decreased target host transcript levels in the corresponding leaves. The VIGS visual phenotype varied from maintained during silencing of actin 1 transcript expression to transient with incomplete penetration through affected tissue during silencing of phytoene desaturase expression. F-BMV RNA 3 was modified to allow greater accumulation of virus while minimizing virus pathogenicity. The modified vector C-BMV(A/G) (C for chimeric) was shown to be useful for VIGS. These BMV vectors will be useful for analysis of gene function in rice and maize for which no VIGS system is reported.

Bromovirus↗

Predicting the secondary structures and tertiary interactions of 211 group I introns in IE subgroup.

The large number of currently available group I intron sequences in the public databases provides opportunity for studying this large family of structurally complex catalytic RNA by large-scale comparative sequence analysis. In this study, the detailed secondary structures of 211 group I introns in the IE subgroup were manually predicted. The secondary structure-favored alignments showed that IE introns contain 14 conserved stems. The P13 stem formed by long-range base-pairing between P2.1 and P9.1 is conserved among IE introns. Sequence variations in the conserved core divide IE introns into three distinct minor subgroups, namely IE1, IE2 and IE3. Co-variation of the peripheral structural motifs with core sequences supports that the peripheral elements function in assisting the core structure folding. Interestingly, host-specific structural motifs were found in IE2 introns inserted at S516 position. Competitive base-pairing is found to be conserved at the junctions of all long-range paired regions, suggesting a possible mechanism of establishing long-range base-pairing during large RNA folding. These findings extend our knowledge of IE introns, indicating that comparative analysis can be a very good complement for deepening our understanding of RNA structure and function in the genomic era.

Base Pairing↗

Small subunit ribosomal RNA of Blastomyces dermatitidis: sequence and phylogenetic analysis.

We determined the small subunit (18S) ribosomal RNA sequence of the dimorphic fungus Blastomyces dermatitidis. The sequence was compared to that of fourteen other eukaryotic organisms, ten of which were higher fungi, and an evolutionary tree was constructed based on these sequences. B. dermatitidis aligned most closely with the Ascomycetes Neurospora crassa and Podospora anserina, in agreement with previous phylogenetic analysis based on morphological criteria. Phase-specific cDNA clones derived by reverse transcription of RNA isolated from the yeast and mycelial phases of B. dermatitidis were also sequenced. The 18S ribosome sequence was found to be the same in both phases. Heterogeneity was found at both the genomic and RNA level at position 1352.

Base Sequence↗

Analysis of Leishbuviridae from Trypanosomatids.

Over the last decade, considerable progress has been made in unraveling RNA virus diversity. This has contributed to our understanding of the evolution of these viruses, which include emerging zoonotic human pathogens. Current success has been greatly facilitated by the development of next-generation sequencing platforms instrumental for meta-transcriptomic studies. However, due to the rapid evolution of RNA viruses, there are numerous "blind spots" waiting to be explored; one of those is the RNA virome of unicellular eukaryotes. Here, we present the pipeline, which has been successfully used to characterize various types of RNA viruses, including Leishbuviridae (Bunyaviricetes, Hareavirales) in the parasitic flagellates of the family Trypanosomatidae. The pipeline relies on axenic in vitro cell culture and double-stranded RNA enrichment, followed by direct RNA-sequencing. A detailed procedure description starting from the initial total RNA preparation to the final assembly of the viral segments is provided.

High-Throughput Nucleotide Sequencing↗

P-RnaPredict--a parallel evolutionary algorithm for RNA folding: effects of pseudorandom number quality.

This paper presents a fully parallel version of RnaPredict, a genetic algorithm (GA) for RNA secondary structure prediction. The research presented here builds on previous work and examines the impact of three different pseudorandom number generators (PRNGs) on the GA's performance. The three generators tested are the C standard library PRNG RAND, a parallelized multiplicative congruential generator (MCG), and a parallelized Mersenne Twister (MT). A fully parallel version of RnaPredict using the Message Passing Interface (MPI) was implemented on a 128-node Beowulf cluster. The PRNG comparison tests were performed with known structures whose sequences are 118, 122, 468, 543, and 556 nucleotides in length. The effects of the PRNGs are investigated and the predicted structures are compared to known structures. Results indicate that P-RnaPredict demonstrated good prediction accuracy, particularly so for shorter sequences.

Algorithms↗

Detection and genotyping of SHV beta-lactamase variants by mass spectrometry after base-specific cleavage of in vitro-generated RNA transcripts.

Matrix-assisted laser desorption ionization-time of flight mass spectrometry (MALDI-TOF MS) after base-specific cleavage of PCR-amplified and in vitro-transcribed bla(SHV) genes was used for the identification and genotyping of SHV beta-lactamases. For evaluation, bla(SHV) stretches of 21 clinical Enterobacteriaceae isolates were PCR amplified using T7 promoter-tagged forward and reverse primers, respectively. In vitro transcripts were generated with T7 RNA and DNA polymerase in the presence of modified analogues replacing either CTP or UTP. Using RNase A, the in vitro transcripts were base-specifically cleaved at every "T" or "C" position. Resulting cleavage products were analyzed by MALDI-TOF MS, generating a characteristic signal pattern based on the fragment masses. All 21 individual SHV genes were identified unambiguously using reference sequences, and the results were in perfect concordance with those obtained by fluorescent dideoxy sequencing, which represents the current standard method. As multiple point mutations can be detected in a single assay and newly emerged mutations which are not yet described in public databases can be identified too, MALDI-TOF MS appears to be an ideal tool for analysis of sequence polymorphisms in resistance-associated gene loci.

Bacteriological Techniques↗

Sequence analysis and genome organisation of poinsettia mosaic virus (PnMV) reveal closer relationship to marafiviruses than to tymoviruses.

Sequence comparison and genome organisation of poinsettia mosaic virus (PnMV), a putative member of the tymoviruses, revealed a closer relationship to marafiviruses. The complete nucleotide sequence of PnMV was determined. The 6099-nt RNA genome encodes a putative 221-kDa polyprotein that lacks a stop codon between the replicase and the coat protein genes, as in most tymovirus RNAs. The genomic RNA has a poly(A) tail at its 3'-terminus in contrast to the tRNA-like structure found in the RNA of most tymoviruses, and no homology was observed to the conserved noncoding region of the tymoviral 3'-termini. The tymobox of PnMV, a 16-nt region of the subgenomic RNA (sgRNA) promoter shared by most tymoviruses, differs in 3 nt from the RNA sequence of tymoviruses but is identical to the sequence of marafiviruses. At least three sgRNAs were found in PnMV-infected Euphorbia pulcherrima and in isolated PnMV particles; one that is 650 nt long encodes the 21.4-kDa coat protein, and the others are about 3.5 and 1.7 kb and contain the 5'- and the 3'-terminal parts of genomic RNA, respectively. Like tymoviruses, PnMV particles sediment as top and bottom components. The particles of the top component contain the sgRNA (650 nt) encoding the coat protein, and those of bottom component contain both genomic and sgRNAs.

Amino Acid Sequence↗

Swine hepatitis E virus strains in Japan form four phylogenetic clusters comparable with those of Japanese isolates of human hepatitis E virus.

Japanese patients with sporadic acute hepatitis E are infected with polyphyletic strains of hepatitis E virus (HEV). Hepatitis E is considered a zoonotic disease. Thus far in Japan, only three strains of swine HEV have been identified and an antibody study for HEV antibodies has not been done on Japanese pigs. To determine the prevalence of swine HEV infection in Japan and the extent of genetic variation among Japanese swine HEV strains, we tested serum samples obtained from 2500 pigs from 2 to 6 months of age at 25 commercial swine farms in Japan for the presence of IgG antibodies to HEV and swine HEV RNA. Anti-HEV antibodies were detected in 1448 pigs (58 %). One-hundred-and-thirteen (15 %) of the 750 3-month-old pigs and 24 (13 %) of the 180 4-month-old pigs were positive for swine HEV RNA. The nucleotide sequence of a 412 bp region within open reading frame 2 of the 137 swine HEV isolates was determined. Sequence analyses revealed that the 137 isolates shared 76.6-100 % nucleotide sequence identities and were classifiable into genotype III (93 %) or IV (7 %) and that the isolates from the same farm were > or = 97.1 % similar to each other. Phylogenetic analysis showed that the Japanese swine and human HEV isolates segregated into four clusters, with the highest nucleotide identity being 94.4-100 % between swine and human isolates in each cluster. These results indicate that swine HEV is widespread in the Japanese swine population and further support the hypothesis that swine serve as reservoirs for HEV infection.

Animals↗

Effects of abasic sites and DNA single-strand breaks on prokaryotic RNA polymerases.

Abasic sites are thought to be the most frequently occurring cellular DNA damage and are generated spontaneously or as the result of chemical or radiation damage to DNA. In contrast to the wealth of information that exists on the effects of abasic sites on DNA polymerases, very little is known about how these lesions interact with RNA polymerases. An in vitro transcription system was used to determine the effects of abasic sites and single-strand breaks on transcriptional elongation. DNA templates were constructed containing single abasic sites or nicks placed at unique locations downstream from two different promoters and were transcribed by SP6 and Escherichia coli RNA polymerases. SP6 RNA polymerase is initially stalled at abasic sites with subsequent, efficient bypass of these lesions. E. coli RNA polymerase also bypassed abasic sites. In contrast, single-strand breaks introduced at abasic sites completely blocked the progression of both RNA polymerases. Sequence analysis of full-length transcripts revealed that SP6 and E. coli RNA polymerases insert primarily, if not exclusively, adenine residues opposite to abasic sites. This finding suggests that abasic sites may be highly mutagenic in vivo at the level of transcription.

Base Composition↗

Faster cyclic loess: normalizing RNA arrays via linear models.

MOTIVATION: Our goal was to develop a normalization technique that yields results similar to cyclic loess normalization and with speed comparable to quantile normalization. RESULTS: Fastlo yields normalized values similar to cyclic loess and quantile normalization and is fast; it is at least an order of magnitude faster than cyclic loess and approaches the speed of quantile normalization. Furthermore, fastlo is more versatile than both cyclic loess and quantile normalization because it is model-based. AVAILABILITY: The Splus/R function for fastlo normalization is available from the authors.

Algorithms↗

Molecular evolution of the major capsid protein VP1 of enterovirus 70.

Nucleotide sequences of the genome RNA encoding capsid protein VP1 (918 nucleotides) of 18 enterovirus 70 (EV70) isolates collected from various parts of the world in 1971 to 1981 were determined, and nucleotide substitutions among them were studied. The genetic distances between isolates were calculated by the pairwise comparison of nucleotide difference. Regression analysis of the genetic distances against time of isolation of the strains showed that the synonymous substitution rate was very high at 21.53 x 10(-3) substitution per nucleotide per year, while the nonsynonymous rate was extremely low at 0.32 x 10(-3) substitution per nucleotide per year. The rate estimated by the average value of synonymous and nonsynonymous substitutions (W.-H. Li, C.-C. Wu, and C.-C. Luo, Mol. Biol. Evol. 2:150-174, 1985) was 5.00 x 10(-3) substitution per nucleotide per year. Taking the average value of synonymous and nonsynonymous substitutions as genetic distances between isolates, the phylogenetic tree was inferred by the unweighted pairwise grouping method of arithmetic average and by the neighbor-joining method. The tree indicated that the virus had evolved from one focal place, and the time of emergence was estimated to be August 1967 +/- 15 months, 2 years before first recognition of the pandemic of acute hemorrhagic conjunctivitis. By superimposing every nucleotide substitution on the branches of the phylogenetic tree, we analyzed nucleotide substitution patterns of EV70 genome RNA. In synonymous substitutions, the proportion of transitions, i.e., C<==>U and G<==>A, was found to be extremely frequent in comparison with that reported on other viruses or pseudogenes. In addition, parallel substitutions (independent substitutions at the same nucleotide position on different branches, i.e., different isolates, of the tree) were frequently found in both synonymous and nonsynonymous substitutions. These frequent parallel substitutions and the low nonsynonymous substitution rate despite the very high synonymous substitution rate described above imply a strong restriction on nonsynonymous substitution sites of VP1, probably due to the requirement for maintaining the rigid icosahedral conformation of the virus.

Base Sequence↗