Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “RNA structure”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

A graph theoretical approach for predicting common RNA secondary structure motifs including pseudoknots in unaligned sequences.

MOTIVATION: RNA structure motifs contained in mRNAs have been found to play important roles in regulating gene expression. However, identification of novel RNA regulatory motifs using computational methods has not been widely explored. Effective tools for predicting novel RNA regulatory motifs based on genomic sequences are needed. RESULTS: We present a new method for predicting common RNA secondary structure motifs in a set of functionally or evolutionarily related RNA sequences. This method is based on comparison of stems (palindromic helices) between sequences and is implemented by applying graph-theoretical approaches. It first finds all possible stable stems in each sequence and compares stems pairwise between sequences by some defined features to find stems conserved across any two sequences. Then by applying a maximum clique finding algorithm, it finds all significant stems conserved across at least k sequences. Finally, it assembles in topological order all possible compatible conserved stems shared by at least k sequences and reports a number of the best assembled stem sets as the best candidate common structure motifs. This method does not require prior structural alignment of the sequences and is able to detect pseudoknot structures. We have tested this approach on some RNA sequences with known secondary structures, in which it is capable of detecting the real structures completely or partially correctly and outperforms other existing programs for similar purposes. AVAILABILITY: The algorithm has been implemented in C++ in a program called comRNA, which is available at http://ural.wustl.edu/softwares.html

Algorithms↗

Computer-aided prediction of RNA secondary structures.

A brief survey of computer algorithms that have been developed to generate predictions of the secondary structures of RNA molecules is presented. Two particular methods are described in some detail. The first utilizes a thermodynamic energy minimization algorithm that takes into account the likelihood that short-range folding tends to be favored over long-range interactions. The second utilizes an interactive computer graphic modelling algorithm that enables the user to consider thermodynamic criteria as well as structural data obtained by nuclease susceptibility, chemical reactivity and phylogenetic studies. Examples of structures for prokaryotic 16S and 23S ribosomal RNAs, several eukaryotic 5S ribosomal RNAs and rabbit beta-globin messenger RNA are presented as case studies in order to describe the two techniques. Anm argument is made for integrating the two approaches presented in this paper, enabling the user to generate proposed structures using thermodynamic criteria, allowing interactive refinement of these structures through the application of experimentally derived data.

Animals↗

Method for predicting RNA secondary structure.

We report a method for predicting the most stable secondary structure of RNA from its primary sequence of nucleotides. The technique consists of a series of three computer programs interfaced to take the nucleotide sequence of any RNA and (a) list all possible helical regions, using modified Watson-Crick base-pairing rules; (b) create all possible secondary structures by forming permutations of compatible helical regions; and (c)evaluate each structure for total free energy of formation from a completely extended chain. A free energy distribution and the base-by-base bonding interactions of each possible structure are catalogued by the system and are readily available for examination. The method has been applied to 62 tRNA sequences. The total free-energy of the predicted most stable structures ranged from -19 to -41 kcal/mole (-22 to -49 kJ/mole). The number of structures created was also highly sequence-dependent and ranged from 200 to 13,000. In nearly all cases the cloverleaf is predicted to be the structure with the lowest free energy of formation.

Base Sequence↗

Analytical description of finite size effects for RNA secondary structures.

The ensemble of RNA secondary structures of uniform sequences is studied analytically. We calculate the partition function for very long sequences and discuss how the crossover length, beyond which asymptotic scaling laws apply, depends on thermodynamic parameters. For realistic choices of parameters this length can be much longer than natural RNA molecules. This has to be taken into account when applying asymptotic theory to interpret experiments or numerical results.

Base Pairing↗

Prediction of structured non-coding RNAs in the genomes of the nematodes Caenorhabditis elegans and Caenorhabditis briggsae.

We present a survey for non-coding RNAs and other structured RNA motifs in the genomes of Caenorhabditis elegans and Caenorhabditis briggsae using the RNAz program. This approach explicitly evaluates comparative sequence information to detect stabilizing selection acting on RNA secondary structure. We detect 3,672 structured RNA motifs, of which only 678 are known non-translated RNAs (ncRNAs) or clear homologs of known C. elegans ncRNAs. Most of these signals are located in introns or at a distance from known protein-coding genes. With an estimated false positive rate of about 50% and a sensitivity on the order of 50%, we estimate that the nematode genomes contain between 3,000 and 4,000 RNAs with evolutionary conserved secondary structures. Only a small fraction of these belongs to the known RNA classes, including tRNAs, snoRNAs, snRNAs, or microRNAs. A relatively small class of ncRNA candidates is associated with previously observed RNA-specific upstream elements.

Animals↗

An algorithm for detecting homologues of known structured RNAs in genomes.

Distinct RNA structures are frequently involved in a wide-range of functions in various biological mechanisms. The three dimensional RNA structures solved by X-ray crystallography and various well-established RNA phylogenetic structures indicate that functional RNAs have characteristic RNA structural motifs represented by specific combinations of base pairings and conserved nucleotides in the loop region. Discovery of well-ordered RNA structures and their homologues in genome-wide searches will enhance our ability to detect the RNA structural motifs and help us to highlight their association with functional and regulatory RNA elements. We present here a novel computer algorithm, HomoStRscan, that takes a single RNA sequence with its secondary structure to search for homologous-RNAs in complete genomes. This novel algorithm completely differs from other currently used search algorithms of homologous structures or structural motifs. For an arbitrary segment (or window) given in the target sequence, that has similar size to the query sequence, HomoStRscan finds the most similar structure to the input query structure and computes the maximal similarity score (MSS) between the two structures. The homologousRNA structures are then statistically inferred from the MSS distribution computed in the target genome. The method provides a flexible, robust and fine search tool for any homologous structural RNAs.

Algorithms↗

Graphical exploratory data analysis of RNA secondary structure dynamics predicted by the massively parallel genetic algorithm.

Studies indicate that RNA may enter intermediate and multiple conformational states, which may impact gene expression and molecular function. It is known that the biologically functional states of RNA molecules may not correspond to their minimum energy conformations, that kinetic barriers may trap the molecule in a local minimum, that folding often occurs during transcription, and that cases exist in which a molecule will transition between one or more functional conformations. Thus, methods for simulating the folding pathway and dynamic behavior of an RNA molecule are important for the prediction of RNA structure and its associated functions. We have developed several data mining techniques guided by interactive visualization tools associated with our massively parallel genetic algorithm for RNA/DNA secondary structure prediction, MPGAfold, and StructureLab analysis workbench. Most of the methods and tools are also applicable to dynamic programming algorithm (DPA) folding data analysis. When applied to MPGAfold results these methodologies are used to determine the significant intermediate and final structures associated with co-transcriptional and full length RNA folding. Since the genetic algorithm is essentially stochastic, multiple runs are required to develop a consensus understanding of an RNA structure. The interactive visualizations facilitate interpretation of results from sequential or full length individual MPGAfold runs, final results of multiple folding runs, including multiple population sizes, and the results from multiple RNA sequences of one family. This paper describes several of these techniques and shows how they are used to help solve this highly combinatoric problem.

Algorithms↗

Stochastic modeling of RNA pseudoknotted structures: a grammatical approach.

MOTIVATION: Modeling RNA pseudoknotted structures remains challenging. Methods have previously been developed to model RNA stem-loops successfully using stochastic context-free grammars (SCFG) adapted from computational linguistics; however, the additional complexity of pseudoknots has made modeling them more difficult. Formally a context-sensitive grammar is required, which would impose a large increase in complexity. RESULTS: We introduce a new grammar modeling approach for RNA pseudoknotted structures based on parallel communicating grammar systems (PCGS). Our new approach can specify pseudoknotted structures, while avoiding context-sensitive rules, using a single CFG synchronized with a number of regular grammars. Technically, the stochastic version of the grammar model can be as simple as an SCFG. As with SCFG, the new approach permits automatic generation of a single-RNA structure prediction algorithm for each specified pseudoknotted structure model. This approach also makes it possible to develop full probabilistic models of pseudoknotted structures to allow the prediction of consensus structures by comparative analysis and structural homology recognition in database searches.

Algorithms↗

A comparative method for finding and folding RNA secondary structures within protein-coding regions.

Existing computational methods for RNA secondary-structure prediction tacitly assume RNA to only encode functional RNA structures. However, experimental studies have revealed that some RNA sequences, e.g. compact viral genomes, can simultaneously encode functional RNA structures as well as proteins, and evidence is accumulating that this phenomenon may also be found in Eukaryotes. We here present the first comparative method, called RNA-DECODER, which explicitly takes the known protein-coding context of an RNA-sequence alignment into account in order to predict evolutionarily conserved secondary-structure elements, which may span both coding and non-coding regions. RNA-DECODER employs a stochastic context-free grammar together with a set of carefully devised phylogenetic substitution-models, which can disentangle and evaluate the different kinds of overlapping evolutionary constraints which arise. We show that RNA-DECODER's parameters can be automatically trained to successfully fold known secondary structures within the HCV genome. We scan the genomes of HCV and polio virus for conserved secondary-structure elements, and analyze performance as a function of available evolutionary information. On known secondary structures, RNA-DECODER shows a sensitivity similar to the programs MFOLD, PFOLD and RNAALIFOLD. When scanning the entire genomes of HCV and polio virus for structure elements, RNA-DECODER's results indicate a markedly higher specificity than MFOLD, PFOLD and RNAALIFOLD.

Codon↗

Relation between genomic and capsid structures in RNA viruses.

We described a new computer program for calculation of RNA secondary structure. Calculation of 20 viral RNAs with this program showed that genomes of the icosahedral capsid viruses had higher folding probabilities than those of the helical capsid viruses. As this explains virus assembly quite well, the information of capsid structure must be imprinted not only in the capsid protein structures but also in the base sequence of the whole genome. We compared folding probability of the original sequence with that of the random sequence in which base composition was the same as the original. All the actual genomes of RNA viruses were more folded than the corresponding random sequences, even though most transcripts of chromosomal genes tended to be less folded. The data can be related to encapsidation of viral genomes. It was thus suggested that there exists a relation between actual sequences and random sequences with the same base ratios, and that the base ratio itself has some evolutional meaning.

Animals↗

Evolutionary change in 5S RNA secondary structure and a phylogenic tree of 54 5S RNA species.

Secondary structure models of 54 5S RNA species are constructed based on the comparative analyses of their primary structure. All 5S RNAs examined have essentially the same secondary structure. However, there are revealing characteristic differences between eukaryotic and prokaryotic types. The prokaryotic 5S RNAs may be further classified into two types, one having 120 nucleotides (120-N type) and another having 116 (116-N type). A possible mechanism for the conversion of the prokaryotic 116-N type to the 120-N type 5S RNAs (or vice versa) is discussed on the basis of their nucleotide alignments. Finally, by comparing the nucleotide alignments, we propose a phylogenic tree of the 54 5S RNA species.

Animals↗

Statistical and Bayesian approaches to RNA secondary structure prediction.

Prediction of RNA secondary structure is a fundamental problem in computational structural biology. For several decades, free energy minimization has been the most popular method for prediction from a single sequence. In recent years, the McCaskill algorithm for computation of partition function and base-pair probabilities has become increasingly appreciated. This paradigm-shifting work has inspired the developments of extended partition function algorithms, statistical sampling and clustering, and application of Bayesian statistical inference. The performance of thermodynamics-based methods is limited by thermodynamic rules and parameters. However, further improvements may come from statistical estimates derived from structural databases for thermodynamics parameters with weak or little experimental data. The Bayesian inference approach appears to be promising in this context.

Algorithms↗

Widespread selection for local RNA secondary structure in coding regions of bacterial genes.

Redundancy of the genetic code dictates that a given protein can be encoded by a large collection of distinct mRNA species, potentially allowing mRNAs to simultaneously optimize desirable RNA structural features in addition to their protein-coding function. To determine whether natural mRNAs exhibit biases related to local RNA secondary structure, a new randomization procedure was developed, DicodonShuffle, which randomizes mRNA sequences while preserving the same encoded protein sequence, the same codon usage, and the same dinucleotide composition as the native message. Genes from 10 of 14 eubacterial species studied and one eukaryote, the yeast Saccharomyces cerevisiae, exhibited statistically significant biases in favor of local RNA structure as measured by folding free energy. Several significant associations suggest functional roles for mRNA structure, including stronger secondary structure bias in the coding regions of intron-containing yeast genes than in intronless genes, and significantly higher folding potential in polycistronic messages than in monocistronic messages in Escherichia coli. Potential secondary structure generally increased in genes from the 5' to the 3' end of E. coli operons, and secondary structure potential was conserved in homologous Salmonella typhi operons. These results are interpreted in terms of possible roles of RNA structures in RNA processing, regulation of mRNA stability, and translational control.

Computational Biology↗

[Influence of ionic strength on RNA-polymerase structure].

Chromatography of RNA polymerase holoenzyme preincubated under different ionic strength conditions on the DNA agarose column was studied. Ratio of two peaks identified to be core and holoenzyme was analysed. In the range of 0.15 to 0.05 M KCl the relative content of the holoenzyme peak gradually decreased from 100 to 50%. At the same time a peak of free sigma-subunit appeared as detected by the chromatography on DNA agarose gel A-1.5 m. The dissociation of half of the sigma-subunit amount occured within the enzyme dimer-monomer transition range. The results suggest that the dimerization follows the equation: E sigma + E sigma in equilibrium with E2 sigma. Reconstitution of the RNA polymerase holoenzyme from purified core enzyme and sigma-subunit was also studied by the same method. Reconstitution did not occur at a low ionic strength (0--0.1 M KCl), but takes place at ionic strength of 0.2 M or higher. Possible function of the dimerisation of the enzyme in search of promoter site and regulation of RNA synthesis is discussed.

DNA↗

Prediction of RNA secondary structure, including pseudoknotting, by computer simulation.

A computer program is presented which determines the secondary structure of linear RNA molecules by simulating a hypothetical process of folding. This process implies the concept of 'nucleation centres', regions in RNA which locally trigger the folding. During the simulation, the RNA is allowed to fold into pseudoknotted structures, unlike all other programs predicting RNA secondary structure. The simulation uses published, experimentally determined free energy values for nearest neighbour base pair stackings and loop regions, except for new extrapolated values for loops larger than seven nucleotides. The free energy value for a loop arising from pseudoknot formation is set to a single, estimated value of 4.2 kcal/mole. Especially in the case of long RNA sequences, our program appears superior to other secondary structure predicting programs described so far, as tests on tRNAs, the LSU intron of Tetrahymena thermophila and a number of plant viral RNAs show. In addition, pseudoknotted structures are often predicted successfully. The program is written in mainframe APL and is adapted to run on IBM compatible PCs, Atari ST and Macintosh personal computers. On an 8 MHz 8088 standard PC without coprocessor, using STSC APL, it folds a sequence of 700 nucleotides in one and a half hour.

Algorithms↗

Analysis of internal loops within the RNA secondary structure in almost quadratic time.

MOTIVATION: Evaluating all possible internal loops is one of the key steps in predicting the optimal secondary structure of an RNA molecule. The best algorithm available runs in time O(L(3)), L is the length of the RNA. RESULTS: We propose a new algorithm for evaluating internal loops, its run-time is O(M(*)log(2)L), M < L(2) is a number of possible nucleotide pairings. We created a software tool Afold which predicts the optimal secondary structure of RNA molecules of lengths up to 28 000 nt, using a computer with 2 Gb RAM. We also propose algorithms constructing sets of conditionally optimal multi-branch loop free (MLF) structures, e.g. the set that for every possible pairing (x, y) contains an optimal MLF structure in which nucleotides x and y form a pair. All the algorithms have run-time O(M(*)log(2)L).

Algorithms↗