Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “RNA structure”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

RNA secondary structure prediction using stochastic context-free grammars and evolutionary history.

MOTIVATION: Many computerized methods for RNA secondary structure prediction have been developed. Few of these methods, however, employ an evolutionary model, thus relevant information is often left out from the structure determination. This paper introduces a method which incorporates evolutionary history into RNA secondary structure prediction. The method reported here is based on stochastic context-free grammars (SCFGs) to give a prior probability distribution of structures. RESULTS: The phylogenetic tree relating the sequences can be found by maximum likelihood (ML) estimation from the model introduced here. The tree is shown to reveal information about the structure, due to mutation patterns. The inclusion of a prior distribution of RNA structures ensures good structure predictions even for a small number of related sequences. Prediction is carried out using maximum a posteriori estimation (MAP) estimation in a Bayesian approach. For small sequence sets, the method performs very well compared to current automated methods.

Algorithms↗

Statistical prediction of single-stranded regions in RNA secondary structure and application to predicting effective antisense target sites and beyond.

Single-stranded regions in RNA secondary structure are important for RNA-RNA and RNA-protein interactions. We present a probability profile approach for the prediction of these regions based on a statistical algorithm for sampling RNA secondary structures. For the prediction of phylogenetically-determined single-stranded regions in secondary structures of representative RNA sequences, the probability profile offers substantial improvement over the minimum free energy structure. In designing antisense oligonucleotides, a practical problem is how to select a secondary structure for the target mRNA from the optimal structure(s) and many suboptimal structures with similar free energies. By summarizing the information from a statistical sample of probable secondary structures in a single plot, the probability profile not only presents a solution to this dilemma, but also reveals 'well-determined' single-stranded regions through the assignment of probabilities as measures of confidence in predictions. In antisense application to the rabbit beta-globin mRNA, a significant correlation between hybridization potential predicted by the probability profile and the degree of inhibition of in vitro translation suggests that the probability profile approach is valuable for the identification of effective antisense target sites. Coupling computational design with DNA-RNA array technique provides a rational, efficient framework for antisense oligonucleotide screening. This framework has the potential for high-throughput applications to functional genomics and drug target validation.

Algorithms↗

An interactive framework for RNA secondary structure prediction with a dynamical treatment of constraints.

A novel approach aiding in the prediction of RNA secondary structures is presented. Although phylogenetic methods are the most successful at deriving RNA secondary structures, the are not applicable when the number of sequences or the sequence variability is too low. Methods based on energy minimization are therefore of great interest. However, some of the suboptimal RNA secondary structures computed with classic methods are unsaturated structures, i.e. some structures are included into others. Thus, the incorporation of constraints during the process of folding is not possible, while the incorporation of constraints before the process of folding often introduces a bias into the energy function. This paper describes a new procedure which allows for the incorporation of constraints before and during the process of RNA folding. SAPSSARN is an interactive program which offers a framework, both to specify a secondary structure through a set of folding constraints and to compute all the supoptimal saturated RNA secondary structures which satisfy all the folding constraints. At the start, it relies on the computation of the probabilities of pairing of each base with all others according to McCaskill's algorithm. The constraint satisfaction formulation of the problem deals dynamically with a chosen set of folding constraints and, finally, a search algorithm computes all the suboptimal saturated secondary structures which satisfy those folding constraints. Within such a framework, it is possible to test new ideas about RNA folding and secondary structures, including pseudoknots, can be computed. The program is illustrated with RNA sequences on which we obtained results in agreement with known structures by using a protocol which mimics the hierarchical folding of RNA molecules.

Algorithms↗

Model for folding and aggregation in RNA secondary structures.

We study the statistical mechanics of RNA secondary structures designed to have an attraction between two different types of structures as a model system for heteropolymer aggregation. The competition between the branching entropy of the secondary structure and the energy gained by pairing drives the RNA to undergo a "temperature independent" second order phase transition from a molten to an aggregated phase. The aggregated phase thus obtained has a macroscopically large number of contacts between different RNAs. The partition function scaling exponent for this phase is theta approximately 1/2 and the crossover exponent of the phase transition is nu approximately 5/3. The relevance of these calculations to the aggregation of biological molecules is discussed.

Models, Chemical↗

Hidden messages in the nef gene of human immunodeficiency virus type 1 suggest a novel RNA secondary structure.

The coexistence of multiple codes in the genome of human immunodeficiency virus type 1 (HIV-1) was analyzed. We explored factors constraining the variability of the virus genome primarily in relation to conserved RNA secondary structures overlapping coding sequences, and used a simple combination of algorithms for RNA secondary structure prediction based on the nearest-neighbor thermodynamic rules and a statistical approach. In our previous study, we applied this combination to a non- redundant data set of env nucleotide sequences, confirmed the conservative secondary structure of the rev-responsive element (RRE) and found a new RNA structure in the first conserved (C1) region of the env gene. In this study, we analyzed the variability of putative RNA secondary structures inside the nef gene of HIV-1 by applying these algorithms to a non-redundant data set of 104 nef sequences retrieved from the Los Alamos HIV database, and predicted the existence of a novel functional RNA secondary structure in the beta3/beta4 regions of nef. The predicted RNA fold in the beta3/beta4 region of nef appears in two forms with different loop sizes. The loop of the first fold consists of seven nucleotides (positions 494-500), with consensus UCAAGCU appearing in 79% of sequences. The other has a five-base loop (positions 495-499) with consensus CAAGC. The difference in size between these two loops may reflect the difference between respective counterparts in the hairpin recognition. This may also have an adaptive biological significance.

Algorithms↗

Prediction of consensus RNA secondary structures including pseudoknots.

Most functional RNA molecules have characteristic structures that are highly conserved in evolution. Many of them contain pseudoknots. Here, we present a method for computing the consensus structures including pseudoknots based on alignments of a few sequences. The algorithm combines thermodynamic and covariation information to assign scores to all possible base pairs, the base pairs are chosen with the help of the maximum weighted matching algorithm. We applied our algorithm to a number of different types of RNA known to contain pseudoknots. All pseudoknots were predicted correctly and more than 85 percent of the base pairs were identified.

Algorithms↗

Automatic RNA secondary structure determination with stochastic context-free grammars.

We have developed a method for predicting the common secondary structure of large RNA multiple alignments using only the information in the alignment. It uses a series of progressively more sensitive searches of the data in an iterative manner to discover regions of base pairing; the first pass examines the entire multiple alignment. The searching uses two methods to find base pairings. Mutual information is used to measure covariation between pairs of columns in the multiple alignment and a minimum length encoding method is used to detect column pairs with high potential to base pair. Dynamic programming is used to recover the optimal tree made up of the best potential base pairs and to create a stochastic context-free grammar. The information in the tree guides the next iteration of searching. The method is similar to the traditional comparative sequence analysis technique. The method correctly identifies most of the common secondary structure in 16S and 23S rRNA.

Algorithms↗

RNA secondary structure switching during DNA synthesis catalyzed by HIV-1 reverse transcriptase.

Changes in RNA secondary structure have been found to play important roles in translational regulation, protein synthesis, and mRNA splicing. In studies utilizing a 66 nucleotide RNA template with a stable hairpin structure, we have examined the effects of RNA secondary structure on HIV-1 reverse transcriptase activity. We identify several pause sites in the stem of the hairpin and show that these pause sites are correlated with the free energy of melting the next base pair in the stem. We also identify a pause site appearing in the loop of the hairpin and show that this is due to the rapid formation of a new hairpin structure occurring during the progress of DNA polymerization through the hairpin. The rapid change in RNA secondary structure to form the new hairpin selectively destabilizes the major hairpin and thereby accelerates the rate at which reverse transcriptase reads through RNA secondary structure.

Base Sequence↗

Structure, recognition and adaptive binding in RNA aptamer complexes.

Novel features of RNA structure, recognition and discrimination have been recently elucidated through the solution structural characterization of RNA aptamers that bind cofactors, aminoglycoside antibiotics, amino acids and peptides with high affinity and specificity. This review presents the solution structures of RNA aptamer complexes with adenosine monophosphate, flavin mononucleotide, arginine/citrulline and tobramycin together with an example of hydrogen exchange measurements of the base-pair kinetics for the AMP-RNA aptamer complex. A comparative analysis of the structures of these RNA aptamer complexes yields the principles, patterns and diversity associated with RNA architecture, molecular recognition and adaptive binding associated with complex formation.

Adenosine Monophosphate↗

Combinatorial properties of RNA secondary structures.

The secondary structure of an RNA molecule is of great importance and possesses influence, e.g., on the interaction of tRNA molecules with proteins or on the stabilization of mRNA molecules. The classification of secondary structures by means of their order proved useful with respect to numerous applications. In 1978, Waterman, who gave the first precise formal framework for the topic, suggested to determine the number a(n,p) of secondary structures of size n and given order p. Since then, no satisfactory result has been found. Based on an observation due to Viennot et al., we will derive generating functions for the secondary structures of order p from generating functions for binary tree structures with Horton-Strahler number p. These generating functions enable us to compute a precise asymptotic equivalent for a(n,p). Furthermore, we will determine the related number of structures when the number of unpaired bases shows up as an additional parameter. Our approach proves to be general enough to compute the average order of a secondary structure together with all the r-th moments and to enumerate substructures such as hairpins or bulges in dependence on the order of the secondary structures considered.

Algorithms↗

Expanded sequence dependence of thermodynamic parameters improves prediction of RNA secondary structure.

An improved dynamic programming algorithm is reported for RNA secondary structure prediction by free energy minimization. Thermodynamic parameters for the stabilities of secondary structure motifs are revised to include expanded sequence dependence as revealed by recent experiments. Additional algorithmic improvements include reduced search time and storage for multibranch loop free energies and improved imposition of folding constraints. An extended database of 151,503 nt in 955 structures? determined by comparative sequence analysis was assembled to allow optimization of parameters not based on experiments and to test the accuracy of the algorithm. On average, the predicted lowest free energy structure contains 73 % of known base-pairs when domains of fewer than 700 nt are folded; this compares with 64 % accuracy for previous versions of the algorithm and parameters. For a given sequence, a set of 750 generated structures contains one structure that, on average, has 86 % of known base-pairs. Experimental constraints, derived from enzymatic and flavin mononucleotide cleavage, improve the accuracy of structure predictions.

Algorithms↗

Structure and assembly of turnip crinkle virus. VI. Identification of coat protein binding sites on the RNA.

Structural studies of turnip crinkle virus have been extended to include the identification of high-affinity coat protein binding sites on the RNA genome. Virus was dissociated at elevated pH and ionic strength, and a ribonucleoprotein complex (rp-complex) was isolated by chromatography on Sephacryl S-200. Genomic RNA fragments in the rp-complex, resistant to RNase A and RNase T1 digestion and associated with tightly bound coat protein subunits, were isolated using coat-protein-specific antibodies. The identity of the protected fragments was determined by direct RNA sequencing. These approaches allowed us to study the specific RNA-protein interactions in the rp-complex obtained from dissociated virus particles. The location of one protected fragment downstream from the amber terminator codon in the first and largest of the three viral open reading frames suggests that the coat protein may play a role in the regulation of the expression of the polymerase gene. We have also identified an additional cluster of T1-protected fragments in the region of the coat protein gene that may represent further high-affinity sites involved in assembly recognition.

Base Sequence↗

CONTRAfold: RNA secondary structure prediction without physics-based models.

MOTIVATION: For several decades, free energy minimization methods have been the dominant strategy for single sequence RNA secondary structure prediction. More recently, stochastic context-free grammars (SCFGs) have emerged as an alternative probabilistic methodology for modeling RNA structure. Unlike physics-based methods, which rely on thousands of experimentally-measured thermodynamic parameters, SCFGs use fully-automated statistical learning algorithms to derive model parameters. Despite this advantage, however, probabilistic methods have not replaced free energy minimization methods as the tool of choice for secondary structure prediction, as the accuracies of the best current SCFGs have yet to match those of the best physics-based models. RESULTS: In this paper, we present CONTRAfold, a novel secondary structure prediction method based on conditional log-linear models (CLLMs), a flexible class of probabilistic models which generalize upon SCFGs by using discriminative training and feature-rich scoring. In a series of cross-validation experiments, we show that grammar-based secondary structure prediction methods formulated as CLLMs consistently outperform their SCFG analogs. Furthermore, CONTRAfold, a CLLM incorporating most of the features found in typical thermodynamic models, achieves the highest single sequence prediction accuracies to date, outperforming currently available probabilistic and physics-based techniques. Our result thus closes the gap between probabilistic and thermodynamic models, demonstrating that statistical learning procedures provide an effective alternative to empirical measurement of thermodynamic parameters for RNA secondary structure prediction. AVAILABILITY: Source code for CONTRAfold is available at http://contra.stanford.edu/contrafold/.

Algorithms↗

Effect of quinacrine on nuclear structure and RNA synthesis in cultured rat hepatocytes.

The effects of quinacrine, an antimetabolite which intercalates into DNA, on the ultrastructure of interphase nuclei and on RNA turnover were studied in primary cultures of rat hepatocytes. Procedures included ultrastructural cytochemical staining for ribonucleoprotein and DNA, autoradiography, and measurement of labeled uridine uptake and incorporation. Addition to the culture medium of a nontoxic dose (10 microM for 30 min) reduces the net accumulation of labeled uridine in RNA. This involves first heterogeneous RNA and then ribosomal RNA since their structural precursors, interchromatin fibrils and nucleolar fibrils, respectively, diminish in that order. Intranucleolar chromatin retracts, and perinucleolar chromatin becomes unusually condensed. A toxic dose (50 microM for 30 min) produces greater inhibition of tritiated uridine incorporation in RNA. This precedes and is not due to a drop in uridine uptake into the cells. Toxic doses produce unusually large clusters of interchromatin granules which are embedded in an unusual dense material which stains positively for ribonucleoprotein. Three regions of the chromatin are altered. (a) Perinuclear condensed chromatin retracts from the nuclear envelope, remaining attached by short DNA-containing bridges. (b) The normally dispersed nucleoplasmic chromatin condenses into a stainable network which retracts centrifugally. (c) Perinucleolar chromatin becomes a network of small highly condensed masses or bands interconnected by fibrils which are either decondensed or stretched. These alterations in chromatin structure probably form the basis of quinacrine-impaired nuclear metabolism.

Animals↗

Analysis of an RNA pseudoknot structure by CD spectroscopy.

The RNA PK5 (GCGAUUUCUGACCGCUUUUUUGUCAG) forms a pseudoknotted structure at low temperatures and a hairpin containing an A.C opposition at higher temperatures (J. Mol. Biol. 214, 455-470 (1990)). CD and absorption spectra of PK5 were measured at several temperatures. A basis set of spectra were fit to the spectra of PK5 using a method that can provide estimates of the numbers of A.U, G.C, and G.U base pairs as well as the number of each of 11 nearest-neighbor base pairs in an RNA (Biopolymers 31, 373-384 (1991)). The fits were close, indicating that PK5 retained the A conformation in the pseudoknot structure and that the fitting technique is not hindered by pseudoknots or A.C oppositions. The results from the analysis were consistent with the pseudoknotted structure at low temperatures and with the hairpin structure at higher temperatures. We concluded that the method of spectral analysis should be useful for determining the secondary structures of other RNAs containing pseudoknots and A.C oppositions.

Base Composition↗

Evolutionarily conserved structural elements are critical for processing of Internal Transcribed Spacer 2 from Saccharomyces cerevisiae precursor ribosomal RNA.

Structural features of Internal Transcribed Spacer 2 (ITS2) important for the correct and efficient removal of this spacer from Saccharomyces cerevisiae pre-rRNA were identified by in vivo mutational analysis based upon phylogenetic comparison with its counterparts from four different yeast species. Compatibility between ITS2 structure and the S. cerevisiae processing machinery was found to have been maintained over only a short evolutionary distance, in contrast to the situation for ITS1. Nevertheless, cis-acting elements required for correct and efficient processing are confined predominantly to those regions of the spacer that show the highest degree of evolutionary conservation. Mutation or deletion of each of these regions severely reduced production of mature 26 S, but not 17 S rRNA, mainly by impeding processing of the 29 SB precursor. In some cases, however, conversion of 29SA into 29 SB pre-rRNA also appeared to be affected. Deletion of non-conserved segments, on the other hand, caused little or no disturbance in processing. Surprisingly, some combinations of such individually neutral deletions had a severe negative effect on the removal of ITS2, suggesting a requirement for a higher-order structure of ITS2. Finally, even structural alterations of ITS2 that did not noticeably affect processing, significantly reduced the growth rate of cells that exclusively express the mutant rDNA units. We take this as further evidence for a direct role of ITS2 in the formation of fully functional 60 S ribosomal subunits.

Base Sequence↗

An algorithm for comparing RNA secondary structures and searching for similar substructures.

To access the functional informations carried by RNA molecules at the level of their secondary structure interactions, we propose a comparison method based on a tree edit algorithm which takes into account the tree structure of RNA foldings. Any secondary structure is translated into a tree involving all its elementary substructures; then a shorter condensed tree is built in which any unbranched helix interspersed with bulges and interior loops is taken as a single node. This method includes several parameters: a comparison matrix between structural units, gap penalties, and the scoring between nodes of the condensed trees. Their effects have been analysed using as a model a rapidly divergent domain of the large ribosomal RNA, for which structural variation during evolution is well known. This method allows one to recognize precisely, in large target molecules, definite substructures that present with the query molecules only a limited set of closely related secondary structure features; it is still efficient if intervening features, which can correspond to insertion/deletion of entire stem regions, separate such structural elements. When coupled with a hierarchical clustering algorithm, this method is suitable for classifying RNA molecules according to their secondary structure homologies.

Algorithms↗

Detailed mapping of RNA secondary structures in core and NS5B-encoding region sequences of hepatitis C virus by RNase cleavage and novel bioinformatic prediction methods.

There is accumulating evidence from bioinformatic studies that hepatitis C virus (HCV) possesses extensive RNA secondary structure in the core and NS5B-encoding regions of the genome. Recent functional studies have defined one such stem-loop structure in the NS5B region as an essential cis-acting replication element (CRE). A program was developed (STRUCTUR_DIST) that analyses multiple rna-folding patterns predicted by mfold to determine the evolutionary conservation of predicted stem-loop structures and, by a new method, to analyse frequencies of covariant sites in predicted RNA folding between HCV genotypes. These novel bioinformatic methods have been combined with enzymic mapping of RNA transcripts from the core and NS5B regions to precisely delineate the RNA structures that are present in these genomic regions. Together, these methods predict the existence of multiple, often juxtaposed stem-loops that are found in all HCV genotypes throughout both regions, as well as several strikingly conserved single-stranded regions, one of which coincides with a region of the genome to which ribosomal access is required for translation initiation. Despite the existence of marked sequence conservation between genotypes in the HCV CRE and single-stranded regions, there was no evidence for comparable suppression of variability at either synonymous or non-synonymous sites in the other predicted stem-loop structures. The configuration and genetic variability of many of these other NS5B and core structures is perhaps more consistent with their involvement in genome-scale ordered RNA structure, a structural configuration of the genomes of many positive-stranded RNA viruses that is associated with host persistence.

Computational Biology↗