Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “RNA sequencing analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Identification of base-triples in RNA using comparative sequence analysis.

Comparative sequence analysis has proven to be a very efficient tool for the determination of RNA secondary structure and certain tertiary interactions. However, base-triples, an important RNA structural element, cannot be predicted accurately from sequence data. We show here that the poor base correlations observed at base-triple positions are the result of two factors. (1) Base covariation is not as strictly required in triples as it is in Watson-Crick pairs. (2) Base-triple structures are less conserved among homologous molecules. A particularity of known triple-helical regions is the presence of multiple base correlations that do not reflect direct pairing. We suggest that natural mutations in base-triples create structural changes that require compensatory mutations in adjacent base-pairs and triples to maintain the triple-helix conformation. On the basis of these observations, we devised two new measures of association that significantly enhance the base-triple signal in correlation studies. We evaluated correlations between base-pairs and single stranded bases, and correlations between adjacent base-pairs. Positions that score well in both analyses are the best triple candidates. This procedure correctly identifies triples, or interactions very close to the proposed triples, in type I and type II tRNAs and in the group I intron.

Base Composition↗

Microarray analysis of RNA processing and modification.

Most RNAs are processed from precursors by mechanisms that include covalent modifications, as well as the removal of flanking and intervening sequences. Traditional methods to detect RNA processing, such as Northern blotting, reverse-transcribed polymerase chain reaction and primer extension assays, are difficult to apply on a large scale. This chapter outlines several methods for analysis of the processing and modification of RNA using microarrays. These encompass protocols for the application of homemade microarrays and custom-designed commercial inkjet microarrays and are tailored for the large-scale analysis of processing of mRNA, including alternative splicing, as well as for the analysis of processing and modification of noncoding RNA. This chapter also describes practical aspects of microarray design, sample preparation, hybridization, and data analysis.

Animals↗

Interrelatedness of 5S RNA sequences investigated by correspondence analysis.

Correspondence analysis (a form of multivariate statistics) applied to 74 5S ribosomal RNA sequences indicates that the sequences are interrelated in a systematic, nonrandom fashion. Aligned sequences are represented as vectors in a 5N-dimensional space, where N is the number of base positions in the 5S RNA molecule. Mutually orthogonal directions (called factor axes) along which intersequence variance is greatest are defined in this hyperspace. Projection of the sequences onto planes defined by these factorial directions reveals clustering of species that is suggestive of phylogenetic relationships. For each factorial direction, correspondence analysis points to regions of "importance," i.e., those base positions at which the systematic changes occur that define that particular direction. In effect, the technique provides a rapid determination of group-specific signatures. In several instances, similarities between sequences are indicated that have only recently been inferred from visual base-to-base comparisons. These results suggest that correspondence analysis may provide a valuable starting point from which to uncover the patterns of change underlying the evolution of a macromolecule, such as 5S RNA.

Analysis of Variance↗

Microbial community dynamics in a humic lake: differential persistence of common freshwater phylotypes.

In an effort to better understand the factors contributing to patterns in freshwater bacterioplankton community composition and diversity, we coupled automated ribosomal intergenic spacer analysis (ARISA) to analysis of 16S ribosomal RNA (rRNA) gene sequences to follow the persistence patterns of 46 individual phylotypes over 3 years in Crystal Bog Lake. Additionally, we sought to identify linkages between the observed phylotype variations and known chemical and biological drivers. Sequencing of 16S rRNA genes obtained from the water column indicated the presence of phylotypes associated with the Actinobacteria, Bacteroidetes, Firmicutes, Proteobacteria, TM7 and Verrucomicrobia phyla, as well as phylotypes with unknown affiliation. Employment of the 16S rRNA gene/ARISA method revealed that specific phylotypes varied independently of the entire bacterial community dynamics. Actinobacteria, which were present on greater than 95% of sampling dates, did not share the large temporal variability of the other identified phyla. Examination of phylotype relative abundance patterns (inferred using ARISA fragment relative fluorescence) revealed a strong correlation between the dominant phytoplankton succession and the relative abundance patterns of the majority of individual phylotypes. Further analysis revealed covariation among unique phylotypes, which formed several distinct bacterial assemblages correlated with particular phytoplankton communities. These data indicate the existence of unique persistence patterns for different common freshwater phylotypes, which may be linked to the presence of dominant phytoplankton species.

Bacteria↗

Three-dimensional comparative modeling of RNA.

Comparative sequence analysis and ERNA-3D software were used to model the three-dimensional structure of the small domain of signal recognition particle RNA. RNA secondary structures were established by allowing only phylogenetically-supported base pairs. The folding of the RNA molecules was constrained further to include a well-supported pseudoknot. Helical sections were oriented coaxially where a continuous helical stack was formed in the RNA of another species. Finally, RNA helices were placed at distances that preserved the connectivity of the molecule with the smallest number of single-stranded nucleotide residues as identified from the aligned sequences. We show that the comparative three-dimensional structure modeling approach is an extremely powerful tool as it requires only a critical number of carefully aligned sequences.

Bacillus subtilis↗

Widespread occurrence of antisense transcription in the human genome.

An increasing number of eukaryotic genes are being found to have naturally occurring antisense transcripts. Here we study the extent of antisense transcription in the human genome by analyzing the public databases of expressed sequences using a set of computational tools designed to identify sense-antisense transcriptional units on opposite DNA strands of the same genomic locus. The resulting data set of 2,667 sense-antisense pairs was evaluated by microarrays containing strand-specific oligonucleotide probes derived from the region of overlap. Verification of specific cases by northern blot analysis with strand-specific riboprobes proved transcription from both DNA strands. We conclude that > or =60% of this data set, or approximately 1,600 predicted sense-antisense transcriptional units, are transcribed from both DNA strands. This indicates that the occurrence of antisense transcription, usually regarded as infrequent, is a very common phenomenon in the human genome. Therefore, antisense modulation of gene expression in human cells may be a common regulatory mechanism.

Algorithms↗

Sequence analysis of the entire RNA genome of a sweet potato chlorotic fleck virus isolate reveals that it belongs to a distinct carlavirus species.

Since the paucity of information on sweet potato chlorotic fleck virus (SPCFV) had precluded its classification, we have determined the complete nucleotide sequence of the single-stranded RNA genome of a Ugandan isolate of SPCFV. The genome is 9104 nucleotides long (excluding the poly(A) tail) and potentially includes six open reading frames (ORFs). Based on genomic organisation and sequence similarity, SPCFV appears to be a member of the genus Carlavirus (family Flexiviridae). However, SPCFV is distantly related to typical carlaviruses, as most of its putative gene products share amino acid sequence identities of <40% with those of typical carlaviruses. Its closest relative is melon yellowing-associated virus, a proposed carlavirus from Brazil, with which it shares ORF5 and ORF6 amino acid sequence identities of 61 and 46%, respectively.

Base Sequence↗

Nucleotide sequencing of S-RNA segment and sequence analysis of the nucleocapsid protein gene of the newly isolated Akabane virus PT-17 strain.

The nucleotide sequences of the S-RNA of Akabane viruses JaGAr-39, OBE-1, Iriki and the newly isolated PT-17 strains and the Aino virus were determined and compared. The results reveal that the S-RNAs of the four Akabane strains share 96.9% homology in nucleotide sequences. Only one amino acid difference out of the 233 amino acids of the nucleocapsid protein (N) and three amino acid differences in the 91 amino acids of the nonstructural protein (NSs) were found among the Akabane viruses. Amino acid sequences of N and NSs proteins of the Aino virus have approximately 80% identity as compared with the Akabane viruses. The results also demonstrate that the four Akabane viruses and the Aino virus can be clearly differentiated by RFLP (restriction fragments length polymorphism) analysis using RT-PCR generated nucleocapsid protein genes and digested with HaeIII and HindIII. The phylogenetic tree based on the UPGMA (Unweighted Pair Group Method with Arithmetic Mean) analysis of the sequences of nucleocapsid protein genes and the S-DNAs revealed that the newly isolated PT-17 strain is most closely related to Iriki strain, than the JaGAr-39 or OBE-1 strains.

Amino Acid Sequence↗

Sequence analysis of RNase MRP RNA reveals its origination from eukaryotic RNase P RNA.

RNase MRP is a eukaryote-specific endoribonuclease that generates RNA primers for mitochondrial DNA replication and processes precursor rRNA. RNase P is a ubiquitous endoribonuclease that cleaves precursor tRNA transcripts to produce their mature 5' termini. We found extensive sequence homology of catalytic domains and specificity domains between their RNA subunits in many organisms. In Candida glabrata, the internal loop of helix P3 is 100% conserved between MRP and P RNAs. The helix P8 of MRP RNA from microsporidia Encephalitozoon cuniculi is identical to that of P RNA. Sequence homology can be widely spread over the whole molecule of MRP RNA and P RNA, such as those from Dictyostelium discoideum. These conserved nucleotides between the MRP and P RNAs strongly support the hypothesis that the MRP RNA is derived from the P RNA molecule in early eukaryote evolution.

Base Sequence↗

Complete nucleotide sequence of the RNA-2 of grapevine deformation and Grapevine Anatolian ringspot viruses.

The nucleotide sequence of RNA-2 of Grapevine Anatolian ringspot virus (GARSV) and Grapevine deformation virus (GDefV), two recently described nepoviruses, has been determined. These RNAs are 3753 nt (GDefV) and 4607 nt (GARSV) in size and contain a single open reading frame encoding a polyprotein of 122 kDa (GDefV) and 150 kDa (GARSV). Full-length nucleotide sequence comparison disclosed 71-73% homology between GDefV RNA-2 and that of Grapevine fanleaf virus (GFLV) and Arabis mosaic virus (ArMV), and 62-64% homology between GARSV RNA-2 and that of Grapevine chrome mosaic virus (GCMV) and Tomato black ring virus (TBRV). As previously observed in other nepoviruses, the 5' non-coding regions of both RNAs are capable of forming stem-loop structures. Phylogenetic analysis of the three proteins encoded by RNA-2 (i.e. protein 2A, movement protein and coat protein) confirmed that GDefV and GARSV are distinct viruses which can be assigned as definitive species in subgroup A and subgroup B of the genus Nepovirus, respectively.

5' Untranslated Regions↗

PETScan: score-based genome-wide association analysis of RNA-Seq and ATAC-Seq data.

MOTIVATION: High-dimensional sequencing data, such as RNA-Seq for gene expression and ATAC-Seq for chromatin accessibility, are widely used in studying systems biology. Accessible chromatin allows transcription factors and regulatory elements to bind to DNA, thereby regulating transcription through the activation or repression of target genes. The association analysis of RNA-Seq and ATAC-Seq data provides insights into gene regulatory mechanisms. Most existing analytic tools exclusively focus on cis-associations, despite regulatory elements being able to physically interact with distant target genes. Furthermore, conventional approaches often utilize Pearson or Spearman correlations, which ignore the count-based nature of RNA-Seq data. RESULTS: To address these limitations, we introduce PETScan, a computationally efficient genome-wide PEak-Transcript Score-based association analysis, utilizing negative binomial models to better accommodate RNA-Seq data. We leverage score tests and matrix calculations for improved computational efficiency, and combine an empirical permutation method with genomic control to ensure valid p-value calculations in studies with limited sample sizes. In real-world datasets, PETScan achieved three orders of magnitude faster than Wald tests, while identifying similar significant gene-peak pairs. AVAILABILITY: The PETScan R package is available on GitHub at https://github.com/yajing-hao/PETScan.

Chromatin Immunoprecipitation Sequencing↗

Salmon pancreas disease virus, an alphavirus infecting farmed Atlantic salmon, Salmo salar L.

A 5.2-kb region at the 3' terminus of the salmon pancreas disease virus (SPDV) RNA genome has been cloned and sequenced. The nucleotide and predicted amino acid sequences show that SPDV shares considerable organizational and sequence identity to members of the genus alphavirus within the family Togaviridae. The SPDV structural proteins encoded by the 5.2-kb region contain a number of unique features when compared to other sequenced alphaviruses. Based on cleavage site homologies, the predicted sizes of the SPDV envelope glycoproteins E2 (438 aa) and E1 (461 aa) are larger than those of other alphaviruses, while the predicted size of the alphavirus 6K protein is 3.2 K (32 aa) in SPDV. The E2 and E1 proteins each carry one putative N-linked glycosylation site, with the site in E1 being found at a unique position. From amino acid sequence comparisons of the SPDV structural region with sequenced alphaviruses overall homology is uniform, ranging from 32 to 33%. While nucleotide sequence analysis of the 26S RNA junction region shows that SPDV is similar to other alphaviruses, analysis of the 3'-nontranslated region reveals that SPDV shows divergence in this region.

3' Untranslated Regions↗

Revised dinoflagellate phylogeny inferred from molecular analysis of large-subunit ribosomal RNA gene sequences.

The nucleotide sequence analysis of the PCR products corresponding to the variable large-subunit rRNA domains D1, D2, D9, and D10 from ten representative dinoflagellate species is reported. Species were selected among the main laboratory-grown dinoflagellate groups: Prorocentrales, Gymnodiniales, and Peridiniales which comprise a variety of morphological and ecological characteristics. The sequence alignments comprising up to 1,000 nucleotides from all ten species were employed to analyze the phylogenetic relationships among these dinoflagellates. Maximum parsimony and neighbor-joining trees were inferred from the data generated and subsequently tested by bootstrapping. Both the D1/D2 and the D9/D10 regions led to coherent trees in which the main class of dinoflagellates. Dinophyceae, is divided in three groups: prorocentroid, gymnodinioid, and peridinioid. An interesting outcome from the molecular phylogeny obtained was the uncertain emergence of Prorocentrum lima. The molecular results reported agreed with morphological classifications within Peridiniales but not with those of Prorocentrales and Gymnodiniales. Additionally, the sequence comparison analysis provided strong evidence to suggest that Alexandrium minutum and Alexandrium lusitanicum were synonymous species given the identical sequence they shared. Moreover, clone Gg1V, which was determined Gymnodinium catenatum based on morphological criteria, would correspond to a new species of the genus Gymnodinium as its sequence clearly differed from that obtained in G. catenatum. The sequence of the amplified fragments was demonstrated to be a valuable tool for phylogenetic and taxonomical analysis among these highly diversified species.

Animals↗

Perspectives on archaeal diversity, thermophily and monophyly from environmental rRNA sequences.

Phylogenetic analysis of ribosomal RNA sequences obtained from uncultivated organisms of a hot spring in Yellowstone National Park reveals several novel groups of Archaea, many of which diverged from the crenarchaeal line of descent prior to previously characterized members of that kingdom. Universal phylogenetic trees constructed with the addition of these sequences indicate monophyly of Archaea, with modest bootstrap support. The data also show a specific relationship between low-temperature marine Archaea and some hot spring Archaea. Two of the environmental sequences are enigmatic: depending upon the data set and analytical method used, these sequences branch deeply within the Crenarchaeota, below the bifurcation between Crenarchaeota and Euryarchaeota, or even as the sister group to Eukaryotes. If additional data confirm either of the latter two placements, then the organisms represented by these ribosomal RNA sequences would merit recognition as a new kingdom, provisionally named "Korarchaeota."

Archaea↗

ModiCal: A Targeted Calibration Workflow for Site-Specific m5C Validation by Nanopore Direct RNA Sequencing.

Accurate identification of RNA 5-methylcytidine (m5C) at the single-nucleotide resolution remains a central challenge in nanopore direct RNA sequencing (DRS). Current global scanning and modification-aware basecalling methods enable transcriptome-wide profiling but often yield high false-positive rates and lack site-specific accuracy. To address this, we repurposed ModiDeC, originally a de novo multimodification classifier, into a targeted, high-precision validation tool for RNA modification sites with prior biochemical knowledge. This was implemented through a three-step calibration workflow that alternates between biochemical and computational modules using the well-characterized m5C2278 site in 25S rRNA as a starting point. Baseline training uses short synthetic RNAs carrying either a methylated or unmodified C2278 as ground truth, followed by IVT-derived calibration and validation in methyltransferase knockout yeast. The baseline model accurately detected the bona fide m5C2278 site but initially produced off-target predictions. Iterative retraining with unmodified IVT signals progressively reduced and ultimately eliminated false positives while maintaining a strong signal at the bona fide site. The final model retained enzyme-dependent detection in wild-type versus knockout yeast and, when explicitly targeted, was also able to detect the second rRNA site, C2870, which remained invisible in the initial analysis. Application to native human prerRNA processing intermediates further resolved two distinct m5C deposition regimes on 28S rRNA, while generalization to dengue virus genomic RNA confirmed that the same calibration logic transfers across diverse RNA contexts. Together, this study establishes a reproducible and transferable framework that integrates biochemical validation with iterative neural network refinement, providing a route toward reliable site-specific m5C confirmation by nanopore direct RNA sequencing.

RNA Methylation↗

Identification of mycobacterial species by comparative sequence analysis of the RNA polymerase gene (rpoB).

For the differentiation and identification of mycobacterial species, the rpoB gene, encoding the beta subunit of RNA polymerase, was investigated. rpoB DNAs (342 bp) were amplified from 44 reference strains of mycobacteria and clinical isolates (107 strains) by PCR. The nucleotide sequences were directly determined (306 bp) and aligned by using the multiple alignment algorithm in the MegAlign package (DNASTAR) and the MEGA program. A phylogenetic tree was constructed by the neighbor-joining method. Comparative sequence analysis of rpoB DNAs provided the basis for species differentiation within the genus Mycobacterium. Slowly and rapidly growing groups of mycobacteria were clearly separated, and each mycobacterial species was differentiated as a distinct entity in the phylogenetic tree. Pathogenic Mycobacterium kansasii was easily differentiated from nonpathogenic M. gastri; this differentiation cannot be achieved by using 16S rRNA gene (rDNA) sequences. By being grouped into species-specific clusters with low-level sequence divergence among strains of the same species, all of the clinical isolates could be easily identified. These results suggest that comparative sequence analysis of amplified rpoB DNAs can be used efficiently to identify clinical isolates of mycobacteria in parallel with traditional culture methods and as a supplement to 16S rDNA gene analysis. Furthermore, in the case of M. tuberculosis, rifampin resistance can be simultaneously determined.

Amino Acid Sequence↗

Differences in Ibaraki virus RNA segment 3 sequences from three epidemics.

Phylogenetic tree and partial nucleotide sequence analysis of RNA segment 3 were conducted to compare the Ibaraki virus (IBAV) strains from three epidemics in Japan, and serotype 2 epizootic hemorrhagic disease virus strains isolated in Australia, Taiwan, and Canada. Each strain was classified relative to the Ibaraki disease (IBAD) epidemics, which occurred in 1959-1960, 1987, or 1997-1998. In particular, major variation of the gene was identified in the strains isolated after 1997 when a new type of IBAD with the abnormal birth was confirmed. Ibaraki viruses isolated in Japan were more closely related to Taiwanese and Australian strains based on genetics, while the Canadian strain was more distantly related.

Animals↗

Self-cleaving circular RNA associated with rice yellow mottle virus is the smallest viroid-like RNA.

We report the sequence, structural features, and self-cleaving activity of the small circular RNA (sc-RNA) associated with rice yellow mottle sobemovirus (RYMV). At 220 nucleotides, the RYMV sc-RNA represents the smallest naturally occurring viroid-like RNA currently documented in the literature. It is similar to other circular satellite RNAs (sat-RNAs) and viroids in being G-C-rich with a high level of self-complementarity. The predicted native structure is essentially a rod with one branched terminus. A region of the RYMV sc-RNA, constituting 24% of the sequence, exhibits 89% identity to the sat-RNA associated with the Australasian isolates of lucerne transient streak sobemovirus. This region is also structurally similar in all three RNAs in that it forms the left terminus of each rod. Dimeric runoff transcripts of cloned RYMV sc-RNA undergo efficient autocatalytic in vitro cleavage in the (+) but not the (-) polarity. Analysis of the (+) sequence indicates the presence of a hammerhead ribozyme resembling that of carnation small retroviroid-like RNA and the genomic satellite transcript of newt. Inefficient cleavage of (+) monomeric transcripts, and a short stem III in the hammerhead, are features consistent with a double-hammerhead mode of self-cleavage. The presence of sat-RNA and retroviroid-like structures within a single RNA suggests a possible role for the RYMV sc-RNA as an evolutionary intermediate between these subviral RNAs.

Base Sequence↗