Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “RNA sequencing analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

An efficient algorithm to compute the landscape of locally optimal RNA secondary structures with respect to the Nussinov-Jacobson energy model.

We make a novel contribution to the theory of biopolymer folding, by developing an efficient algorithm to compute the number of locally optimal secondary structures of an RNA molecule, with respect to the Nussinov-Jacobson energy model. Additionally, we apply our algorithm to analyze the folding landscape of selenocysteine insertion sequence (SECIS) elements from A. Bock (personal communication), hammerhead ribozymes from Rfam (Griffiths-Jones et al., 2003), and tRNAs from Sprinzl's database (Sprinzl et al., 1998). It had previously been reported that tRNA has lower minimum free energy than random RNA of the same compositional frequency (Clote et al., 2003; Rivas and Eddy, 2000), although the situation is less clear for mRNA (Seffens and Digby, 1999; Workman and Krogh, 1999; Cohen and Skienna, 2002),(1) which plays no structural role. Applications of our algorithm extend knowledge of the energy landscape differences between naturally occurring and random RNA. Given an RNA molecule a(1), ... , a(n) and an integer k > or = 0, a k-locally optimal secondary structure S is a secondary structure on a(1), ... , a(n) which has k fewer base pairs than the maximum possible number, yet for which no basepairs can be added without violation of the definition of secondary structure (e.g., introducing a pseudoknot). Despite the fact that the number numStr(k) of k-locally optimal structures for a given RNA molecule in general is exponential in n, we present an algorithm running in time O(n (4)) and space O(n (3)), which computes numStr(k) for each k. Structurally important RNA, such as SECIS elements, hammerhead ribozymes, and tRNA, all have a markedly smaller number of k-locally optimal structures than that of random RNA of the same dinucleotide frequency, for small and moderate values of k. This suggests a potential future role of our algorithm as a tool to detect noncoding RNA genes.

Algorithms↗

Design of hammerhead ribozymes to distinguish single base changes in substrate RNA.

Hammerhead ribozymes are attractive tools in antisense gene inactivation because of their catalytic cleavage of target molecules. High sequence discrimination should be possible, since the cleavage efficiencies were already significantly reduced, if single base changes in substrate RNA introduce mismatches next to the cleavage site. This was observed at the first innermost base pair in helix I and the two innermost base pairs in helix III. In addition to its position, the nature of the mismatch pair was important.

Autoradiography↗

Sequence-specific cleavage of hepatitis X RNA in cis and trans by novel monotarget and multitarget hammerhead motif-containing ribozymes.

We constructed two monoribozymes and a diribozyme against the conserved region of the X RNA of hepatitis B virus (HBV). All the ribozymes (Rzs) possessed sequence-specific cleavage activities under standard and simulated physiologic conditions. Specific cleavage was also obtained when the same Rzs were placed in cis configuration with respect to X gene in multiple combinations. Rz-expressing cells were able to specifically interfere with the functional expression of X RNA and protein production in a liver-specific cell line HepG2. Potential applications of these novel Rzs are discussed.

Base Sequence↗

Algorithmic approaches for identification of RNA editing sites.

Recently a number of groups have introduced computational methods for the detection of A-to-I RNA editing sites. These approaches have resulted in finding thousands of editing sites within the genomic repeats, as well as a few novel genetic recoding sites. We review these recent advancements, emphasizing the principles underlying the various methods used. Possible directions for extending these methods are discussed.

Algorithms↗

Decipher RNA isoform combinations from minigene splicing assays and massive parallel sequencing with MAGIC.

SUMMARY: Functional testing of RNA using minigene splicing assays is increasingly being realized to demonstrate the effects of variants on splicing. In complex cases, variant pathogenicity is assessed by Sanger sequencing, which can be time consuming and may be replaced by short read sequencing. Moreover, strategies based on long read sequencing of the amplified minigene construct are promising and allow the isoforms to be fully characterized. We introduce MAGIC, a user-friendly tool that first generates the artificial construction genome files required to then perform alignment, assembly and annotation of the isoforms obtained by either short or long read minigene splicing assay sequencing. AVAILABILITY AND IMPLEMENTATION: MAGIC is available at https://github.com/LBGC-CFB/MAGIC. Zenodo DOI: 10.5281/zenodo.17052752.

High-Throughput Nucleotide Sequencing↗

Markovian negentropies in bioinformatics. 1. A picture of footprints after the interaction of the HIV-1 Psi-RNA packaging region with drugs.

MOTIVATION: Many experts worldwide have highlighted the potential of RNA molecules as drug targets for the chemotherapeutic treatment of a range of diseases. In particular, the molecular pockets of RNA in the HIV-1 packaging region have been postulated as promising sites for antiviral action. The discovery of simpler methods to accurately represent drug-RNA interactions could therefore become an interesting and rapid way to generate models that are complementary to docking-based systems. RESULTS: The entropies of a vibrational Markov chain have been introduced here as physically meaningful descriptors for the local drug-nucleic acid complexes. A study of the interaction of the antibiotic Paromomycin with the packaging region of the RNA present in type-1 HIV has been carried out as an illustrative example of this approach. A linear discriminant function gave rise to excellent discrimination among 80.13% of interacting/non-interacting sites. More specifically, the model classified 36/45 nucleotides (80.0%) that interacted with paromomycin and, in addition, 85/106 (80.2%) footprinted (non-interacting) sites from the RNA viral sequence were recognized. The model showed a high Matthews' regression coefficient (C = 0.64). The Jackknife method was also used to assess the stability and predictability of the model by leaving out adenines, C, G, or U. Matthews' coefficients and overall accuracies for these approaches were between 0.55 and 0.68 and 75.8 and 82.7, respectively. On the other hand, a linear regression model predicted the local binding affinity constants between a specific nucleotide and the aforementioned antibiotic (R2 = 0.83,Q2 = 0.825). These kinds of models may play an important role either in the discovery of new anti-HIV compounds or in the elucidation of their mode of action. AVAILABILITY: On request from the corresponding author (humbertogd@cbq.uclv.edu.cu or humbertogd@navegalia.com).

Binding Sites↗

Evaluating the predictability of conformational switching in RNA.

MOTIVATION: There are various cases where the biological function of an RNA molecule involves a reversible change of conformation. paRNAss is a software approach to the prediction of such structural switching in RNA. It is based on three hypotheses about the secondary structure space of a switching RNA molecule that can be evaluated by RNA folding and structure comparison. In the positive case, the predicted structural switching must be verified experimentally. RESULTS: After reviewing the strategy used in paRNAss, we present recent improvements on the algorithmic level of the approach, and the results of an evaluation procedure, comprising 1500 RNA sequences. It could be shown that the paRNAss approach performs well on known examples for conformational switching in RNA. The overall number of positive predictions was small, whereas for human 3' UTRs, representing regulatory important regions, it was substantially higher than for arbitrary natural and random sequences. AVAILABILITY: paRNAss is available as a Web service at http://bibiserv.techfak.uni-bielefeld.de/parnass SUPPLEMENTARY INFORMATION: Detailed information on the analyses summarized in Table 1 can be found at http://bibiserv.techfak.uni-bielefeld.de/parnass/examples.html

Algorithms↗

Second eigenvalue of the Laplacian matrix for predicting RNA conformational switch by mutation.

MOTIVATION: Conformational switching in RNAs is thought to be of fundamental importance in several biological processes, including translational regulation, regulation of self-cleavage in viruses, protein biosynthesis and mRNA splicing. Current methods for detecting bi-stable RNAs that can lead to structural switching when triggered by an outside event rely on kinetics, energetics and properties of the combinatorial structure space of RNAs. Based on these properties, tools have been developed to predict whether a given sequence folds to a structure characterized by a bi-stable conformation, or to design multi-stable RNAs by an iterative algorithm. A useful addition is in developing a local procedure to prescribe, given an initial sequence, the least amount of mutations needed to drive the system into an optimal bi-stable conformation. RESULTS: We introduce a local procedure for predicting mutations, by generating and analyzing eigenvalue tables, that are capable of transforming the wild-type sequence into a bi-stable conformation. The method is independent of the folding algorithms but relies on their success. It can be used in conjunction with existing tools, as well as being incorporated into more general RNA prediction packages. We apply this procedure on three well-studied structures. First, the method is validated on the mutation leading to a conformational switch in the spliced leader RNA from Leptomonas collosoma, a mutation that has already been confirmed by an experiment. Second, the method is used to predict a mutation that can lead to a novel conformational switch in the P5abc subdomain of the group I intron ribozyme in Tetrahymena thermophila. Third, the method is applied on Hepatitis delta virus to predict mutations that transform the wild-type into a bi-stable conformation, a configuration assessed by calculating the free energies using folding prediction algorithms. The predictions in the final examples need to be verified experimentally, whereas the mutation predicted in the first example complies with the experiment. This supports the use of our proposed method on other known structures, as well as genetically engineered ones. AVAILABILITY: An eigenvalue application will be available in the near future attached to one of the existing tools.

Algorithms↗

AUG codons at the beginning of protein coding sequences are frequent in eukaryotic mRNAs with a suboptimal start codon context.

MOTIVATION: The translation start site plays an important role in the control of translation efficiency of eukaryotic mRNAs. However, mRNAs with a suboptimal context of start AUG codon are relatively abundant. It is likely that at least some mRNAs with suboptimal start codon context contain the other signals providing additional information for efficient AUG recognition. RESULTS: Frequency of AUG codons at the beginning of the coding part of eukaryotic mRNAs was analyzed in relation to the context of translation start codon. It was found that the observed downstream AUG content in the mRNAs with optimal start codon context was close to the expected value, whereas it was significantly higher in the mRNAs with a suboptimal context. It is likely that downstream AUG codons can often be utilized as additional start sites to increase translation rate of mRNAs with a suboptimal context of the annotated start codon and many eukaryotic proteins can be characterized by some N-end heterogeneity.

Adenine↗

Sequence to Structure (S2S): display, manipulate and interconnect RNA data from sequence to structure.

SUMMARY: Efficient RNA sequence manipulations (such as multiple alignments) need to be constrained by rules of RNA structure folding. The structural knowledge has increased dramatically in the last years with the accumulation of several large RNA structures similar to those of the bacterial ribosome subunits. However, no tool in the RNA community provides an easy way to link and integrate progress made at the sequence level using the available three-dimensional information. Sequence to Structure (S2S) proposes a framework in which an user can easily display, manipulate and interconnect heterogeneous RNA data, such as multiple sequence alignments, secondary and tertiary structures. S2S has been implemented using the Java language and has been developed and tested under UNIX systems, such as Linux and MacOSX. AVAILABILITY: S2S is available at http://bioinformatics.org/S2S/.

Algorithms↗

Local RNA base pairing probabilities in large sequences.

SUMMARY: The genome-wide search for non-coding RNAs requires efficient methods to compute and compare local secondary structures. Since the exact boundaries of such putative transcripts are typically unknown, arbitrary sequence windows have to be used in practice. Here we present a method for robustly computing the probabilities of local base pairs from long RNA sequences independent of the exact positions of the sequence window. AVAILABILITY: The program RNAplfold is part of the Vienna RNA Package and can be downloaded from http://www.tbi.univie.ac.at/RNA/.

Algorithms↗

Using mRNAs lengths to accurately predict the alternatively spliced gene products in Caenorhabditis elegans.

MOTIVATION: Computational gene prediction methods are an important component of whole genome analyses. While ab initio gene finders have demonstrated major improvements in accuracy, the most reliable methods are evidence-based gene predictors. These algorithms can rely on several different sources of evidence including predictions from multiple ab initio gene finders, matches to known proteins, sequence conservation and partial cDNAs to predict the final product. Despite the success of these algorithms, prediction of complete gene structures, especially for alternatively spliced products, remains a difficult task. RESULTS: LOCUS (Length Optimized Characterization of Unknown Spliceforms) is a new evidence-based gene finding algorithm which integrates a length-constraint into a dynamic programming-based framework for prediction of gene products. On a Caenorhabditis elegans test set of alternatively spliced internal exons, its performance exceeds that of current ab initio gene finders and in most cases can accurately predict the correct form of all the alternative products. As the length information used by the algorithm can be obtained in a high-throughput fashion, we propose that integration of such information into a gene-prediction pipeline is feasible and doing so may improve our ability to fully characterize the complete set of mRNAs for a genome. AVAILABILITY: LOCUS is available from http://ural.wustl.edu/software.html

Algorithms↗

INFO-RNA--a fast approach to inverse RNA folding.

MOTIVATION: The structure of RNA molecules is often crucial for their function. Therefore, secondary structure prediction has gained much interest. Here, we consider the inverse RNA folding problem, which means designing RNA sequences that fold into a given structure. RESULTS: We introduce a new algorithm for the inverse folding problem (INFO-RNA) that consists of two parts; a dynamic programming method for good initial sequences and a following improved stochastic local search that uses an effective neighbor selection method. During the initialization, we design a sequence that among all sequences adopts the given structure with the lowest possible energy. For the selection of neighbors during the search, we use a kind of look-ahead of one selection step applying an additional energy-based criterion. Afterwards, the pre-ordered neighbors are tested using the actual optimization criterion of minimizing the structure distance between the target structure and the mfe structure of the considered neighbor. We compared our algorithm to RNAinverse and RNA-SSD for artificial and biological test sets. Using INFO-RNA, we performed better than RNAinverse and in most cases, we gained better results than RNA-SSD, the probably best inverse RNA folding tool on the market. AVAILABILITY: www.bioinf.uni-freiburg.de?Subpages/software.html.

Algorithms↗

CONTRAfold: RNA secondary structure prediction without physics-based models.

MOTIVATION: For several decades, free energy minimization methods have been the dominant strategy for single sequence RNA secondary structure prediction. More recently, stochastic context-free grammars (SCFGs) have emerged as an alternative probabilistic methodology for modeling RNA structure. Unlike physics-based methods, which rely on thousands of experimentally-measured thermodynamic parameters, SCFGs use fully-automated statistical learning algorithms to derive model parameters. Despite this advantage, however, probabilistic methods have not replaced free energy minimization methods as the tool of choice for secondary structure prediction, as the accuracies of the best current SCFGs have yet to match those of the best physics-based models. RESULTS: In this paper, we present CONTRAfold, a novel secondary structure prediction method based on conditional log-linear models (CLLMs), a flexible class of probabilistic models which generalize upon SCFGs by using discriminative training and feature-rich scoring. In a series of cross-validation experiments, we show that grammar-based secondary structure prediction methods formulated as CLLMs consistently outperform their SCFG analogs. Furthermore, CONTRAfold, a CLLM incorporating most of the features found in typical thermodynamic models, achieves the highest single sequence prediction accuracies to date, outperforming currently available probabilistic and physics-based techniques. Our result thus closes the gap between probabilistic and thermodynamic models, demonstrating that statistical learning procedures provide an effective alternative to empirical measurement of thermodynamic parameters for RNA secondary structure prediction. AVAILABILITY: Source code for CONTRAfold is available at http://contra.stanford.edu/contrafold/.

Algorithms↗

Ribostral: an RNA 3D alignment analyzer and viewer based on basepair isostericities.

UNLABELLED: RNA atomic resolution structures have revealed the existance of different families of basepair interactions, each of which with its own isosteric sub-families. Ribostral (Ribonucleic Structural Aligner) is a user-friendly framework for analyzing, evaluating and viewing RNA sequence alignments with at least one available atomic resolution structure. It is the first of its kind that makes direct and easy- to-understand superposition of the isostericity matrices of basepairs observed in the structure onto sequence alignments, easily indicating allowed and unallowed substitutions at each BP position. Potential mistakes in the alignments can then be corrected using other sequence editing software. Ribostral has been developed and tested under Windows XP, and is capable of running on any PC or MAC platform with MATLAB 7.1 (SP3) or higher installed version. A stand-alone version is also available for the PC platform. AVAILABILITY: http://rna.bgsu.edu/ribostral.

Algorithms↗

Mining frequent stem patterns from unaligned RNA sequences.

MOTIVATION: In detection of non-coding RNAs, it is often necessary to identify the secondary structure motifs from a set of putative RNA sequences. Most of the existing algorithms aim to provide the best motif or few good motifs, but biologists often need to inspect all the possible motifs thoroughly. RESULTS: Our method RNAmine employs a graph theoretic representation of RNA sequences and detects all the possible motifs exhaustively using a graph mining algorithm. The motif detection problem boils down to finding frequently appearing patterns in a set of directed and labeled graphs. In the tasks of common secondary structure prediction and local motif detection from long sequences, our method performed favorably both in accuracy and in efficiency with the state-of-the-art methods such as CMFinder. AVAILABILITY: The software is available upon request.

Algorithms↗

Robust prediction of consensus secondary structures using averaged base pairing probability matrices.

MOTIVATION: Recent transcriptomic studies have revealed the existence of a considerable number of non-protein-coding RNA transcripts in higher eukaryotic cells. To investigate the functional roles of these transcripts, it is of great interest to find conserved secondary structures from multiple alignments on a genomic scale. Since multiple alignments are often created using alignment programs that neglect the special conservation patterns of RNA secondary structures for computational efficiency, alignment failures can cause potential risks of overlooking conserved stem structures. RESULTS: We investigated the dependence of the accuracy of secondary structure prediction on the quality of alignments. We compared three algorithms that maximize the expected accuracy of secondary structures as well as other frequently used algorithms. We found that one of our algorithms, called McCaskill-MEA, was more robust against alignment failures than others. The McCaskill-MEA method first computes the base pairing probability matrices for all the sequences in the alignment and then obtains the base pairing probability matrix of the alignment by averaging over these matrices. The consensus secondary structure is predicted from this matrix such that the expected accuracy of the prediction is maximized. We show that the McCaskill-MEA method performs better than other methods, particularly when the alignment quality is low and when the alignment consists of many sequences. Our model has a parameter that controls the sensitivity and specificity of predictions. We discussed the uses of that parameter for multi-step screening procedures to search for conserved secondary structures and for assigning confidence values to the predicted base pairs. AVAILABILITY: The C++ source code that implements the McCaskill-MEA algorithm and the test dataset used in this paper are available at http://www.ncrna.org/papers/McCaskillMEA/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

In vivo expression of the nucleolar group I intron-encoded I-dirI homing endonuclease involves the removal of a spliceosomal intron.

The Didymium iridis DiSSU1 intron is located in the nuclear SSU rDNA and has an unusual twin-ribozyme organization. One of the ribozymes (DiGIR2) catalyses intron excision and exon ligation. The other ribozyme (DiGIR1), which along with the endonuclease-encoding I-DirI open reading frame (ORF) is inserted in DiGIR2, carries out hydrolysis at internal processing sites (IPS1 and IPS2) located at its 3' end. Examination of the in vivo expression of DiSSU1 shows that after excision, DiSSU1 is matured further into the I-DirI mRNA by internal DiGIR1-catalysed cleavage upstream of the ORF 5' end, as well as truncation and polyadenylation downstream of the ORF 3' end. A spliceosomal intron, the first to be reported within a group I intron and the rDNA, is removed before the I-DirI mRNA associates with the polysomes. Taken together, our results imply that DiSSU1 uses a unique combination of intron-supplied ribozyme activity and adaptation to the general RNA polymerase II pathway of mRNA expression to allow a protein to be produced from the RNA polymerase I-transcribed rDNA.

Amoeba↗