Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “RNA sequencing analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

The massively parallel genetic algorithm for RNA folding: MIMD implementation and population variation.

A massively parallel Genetic Algorithm (GA) has been applied to RNA sequence folding on three different computer architectures. The GA, an evolution-like algorithm that is applied to a large population of RNA structures based on a pool of helical stems derived from an RNA sequence, evolves this population in parallel. The algorithm was originally designed and developed for a 16384 processor SIMD (Single Instruction Multiple Data) MasPar MP-2. More recently it has been adapted to a 64 processor MIMD (Multiple Instruction Multiple Data) SGI ORIGIN 2000, and a 512 processor MIMD CRAY T3E. The MIMD version of the algorithm raises issues concerning RNA structure data-layout and processor communication. In addition, the effects of population variation on the predicted results are discussed. Also presented are the scaling properties of the algorithm from the perspective of the number of physical processors utilized and the number of virtual processors (RNA structures) operated upon.

Algorithms↗

PseudoViewer: automatic visualization of RNA pseudoknots.

MOTIVATION: Several algorithms have been developed for drawing RNA secondary structures, however none of these can be used to draw RNA pseudoknot structures. In the sense of graph theory, a drawing of RNA secondary structures is a tree, whereas a drawing of RNA pseudoknots is a graph with inner cycles within a pseudoknot as well as possible outer cycles formed between a pseudoknot and other structural elements. Thus, RNA pseudoknots are more difficult to visualize than RNA secondary structures. Since no automatic method for drawing RNA pseudoknots exists, visualizing RNA pseudoknots relies on significant amount of manual work and does not yield satisfactory results. The task of visualizing RNA pseudoknots by hand becomes more challenging as the size and complexity of the RNA pseudoknots increase. RESULTS: We have developed a new representation and an algorithm for drawing H-type pseudoknots with RNA secondary structures. Compared to existing representations of H-type pseudoknots, the new representation ensures uniform and clear drawings with no edge crossing for any H-type pseudoknots. To the best of our knowledge, this is the first algorithm for automatically drawing RNA pseudoknots with RNA secondary structures. The algorithm has been implemented in a Java program, which can be executed on any computing system. Experimental results demonstrate that the algorithm generates an aesthetically pleasing drawing of all H-type pseudoknots. The results have also shown that the drawing has high readability, enabling the user to quickly and easily recognize the whole RNA structure as well as the pseudoknots themselves.

Algorithms↗

Silent DNA: speaking RNA language?

The sequence of silent DNA in the human genome (intergenic spacers, introns and synonymous codon positions of protein-coding genes) was found here to have the higher thermostability of corresponding RNA/RNA and RNA/DNA duplexes as compared with randomized sequence. This difference increased with elevation of GC content. The revealed effect was not due to correlation of RNA/RNA and RNA/DNA thermostabilities with thermostability of the DNA/DNA duplex, which, on the contrary, was lower than in the randomized sequence and lagged behind the elevation of GC content. The same picture was observed in the genomes of other warm-blooded vertebrates but not in the lower organisms. This finding suggests that RNA-RNA and RNA-DNA interactions could be involved in the putative function of silent DNA.

Animals↗

Prediction of locally stable RNA secondary structures for genome-wide surveys.

MOTIVATION: Recently novel classes of functional RNAs, most prominently the miRNAs have been discovered, strongly suggesting that further types of functional RNAs are still hidden in the recently completed genomic DNA sequences. Only few techniques are known, however, to survey genomes for such RNA genes. When sufficiently similar sequences are not available for comparative approaches the only known remedy is to search directly for structural features. RESULTS: We present here efficient algorithms for computing locally stable RNA structures at genome-wide scales. Both the minimum energy structure and the complete matrix of base pairing probabilities can be computed in theta(N x L2) time and theta(N + L2) memory in terms of the length N of the genome and the size L of the largest secondary structure motifs of interest. In practice, the 100 Mb of the complete genome of Caenorhabditis elegans can be folded within about half a day on a modern PC with a search depth of L = 100. This is sufficient example for a survey for miRNAs. AVAILABILITY: The software described in this contribution will be available for download at http://www.tbi.univie.ac.at/~ivo/RNA/ as part of the Vienna RNA Package.

Algorithms↗

A memory-efficient algorithm for multiple sequence alignment with constraints.

MOTIVATION: Recently, the concept of the constrained sequence alignment was proposed to incorporate the knowledge of biologists about structures/functionalities/consensuses of their datasets into sequence alignment such that the user-specified residues/nucleotides are aligned together in the computed alignment. The currently developed programs use the so-called progressive approach to efficiently obtain a constrained alignment of several sequences. However, the kernels of these programs, the dynamic programming algorithms for computing an optimal constrained alignment between two sequences, run in (gamman2) memory, where gamma is the number of the constraints and n is the maximum of the lengths of sequences. As a result, such a high memory requirement limits the overall programs to align short sequences only. RESULTS: We adopt the divide-and-conquer approach to design a memory-efficient algorithm for computing an optimal constrained alignment between two sequences, which greatly reduces the memory requirement of the dynamic programming approaches at the expense of a small constant factor in CPU time. This new algorithm consumes only O(alphan) space, where alpha is the sum of the lengths of constraints and usually alpha << n in practical applications. Based on this algorithm, we have developed a memory-efficient tool for multiple sequence alignment with constraints. AVAILABILITY: http://genome.life.nctu.edu.tw/MUSICME.

Algorithms↗

Memory efficient folding algorithms for circular RNA secondary structures.

BACKGROUND: A small class of RNA molecules, in particular the tiny genomes of viroids, are circular. Yet most structure prediction algorithms handle only linear RNAs. The most straightforward approach is to compute circular structures from 'internal' and 'external' substructures separated by a base pair. This is incompatible, however, with the memory-saving approach of the Vienna RNA Package which builds a linear RNA structure from shorter (internal) structures only. RESULT: Here we describe how circular secondary structures can be obtained without additional memory requirements as a kind of 'post-processing' of the linear structures. AVAILABILITY: The circular folding algorithm is implemented in the current version of the of RNAfold program of the Vienna RNA Package, which can be downloaded from http://www.tbi.univie.ac.at/RNA/

Algorithms↗

Isolation and sequence determination of the 3'-terminal regions of isotopically labelled RNA molecules.

The method which was developed for the selective isolation of 3'-terminal polynucleotides from large RNA molecules on columns of cellulose derivatives containing covalently bound dihydroxyboryl groups has been modified and adapted for use on radioactively labelled RNAs. The 3'-terminal polynucleotide fragments which result from specific ribonuclease digestion of isotopically detectable quantities of RNA can be selectively obtained in both high yield and purity by the modified procedure and can be subsequently analyzed by standard electrophoretic and chromatographic techniques. In addition, when the extent of enzymatic fragmentation of the RNA is controlled, the procedure permits the selective isolation of discrete "sets" of fragments of variable chain length, all of which derive from the 3'-terminus of the RNA molecule. These overlapping polynucleotides can be used directly to obtain extensive sequence information regarding the primary structure in the 3'-region of the RNA.

Autoradiography↗

Sequencing RNA by a combination of exonuclease digestion and uridine specific chemical cleavage using MALDI-TOF.

The determination of DNA sequences by partial exonuclease digestion followed by Matrix-Assisted Laser Desorption Time of Flight Mass Spectrometry (MALDI-TOF) is a well established method. When the same procedure is applied to RNA, difficulties arise due to the small (1 Da) mass difference between the nucleotides U and C, which makes unambiguous assignment difficult using a MALDI-TOF instrument. Here we report our experiences with sequence specific endonucleases and chemical methods followed by MALDI-TOF to resolve these sequence ambiguities. We have found chemical methods superior to endonucleases both in terms of correct specificity and extent of sequence coverage. This methodology can be used in combination with exonuclease digestion to rapidly assign RNA sequences.

Aniline Compounds↗

GPRM: A genetic programming approach to finding common RNA secondary structure elements.

RNA molecules play an important role in many biological activities. Knowing its secondary structure can help us better understand the molecule's ability to function. The methods for RNA structure determination have traditionally been implemented through biochemical, biophysical and phylogenetic analyses. As the advance of computer technology, an increasing number of computational approaches have recently been developed. They have different goals and apply various algorithms. For example, some focus on secondary structure prediction for a single sequence; some aim at finding a global alignment of multiple sequences. Some predict the structure based on free energy minimization; some make comparative sequence analyses to determine the structure. In this paper, we describe how to correctly use GPRM, a genetic programming approach to finding common secondary structure elements in a set of unaligned coregulated or homologous RNA sequences. GPRM can be accessed at http://bioinfo.cis.nctu.edu.tw/service/gprm/.

Internet↗

Integrated databanks access and sequence/structure analysis services at the PBIL.

The World Wide Web server of the PBIL (Pôle Bioinformatique Lyonnais) provides on-line access to sequence databanks and to many tools of nucleic acid and protein sequence analyses. This server allows to query nucleotide sequence banks in the EMBL and GenBank formats and protein sequence banks in the SWISS-PROT and PIR formats. The query engine on which our data bank access is based is the ACNUC system. It allows the possibility to build complex queries to access functional zones of biological interest and to retrieve large sequence sets. Of special interest are the unique features provided by this system to query the data banks of gene families developed at the PBIL. The server also provides access to a wide range of sequence analysis methods: similarity search programs, multiple alignments, protein structure prediction and multivariate statistics. An originality of this server is the integration of these two aspects: sequence retrieval and sequence analysis. Indeed, thanks to the introduction of re-usable lists, it is possible to perform treatments on large sets of data. The PBIL server can be reached at: http://pbil.univ-lyon1.fr.

Databases, Genetic↗

Paradigms for computational nucleic acid design.

The design of DNA and RNA sequences is critical for many endeavors, from DNA nanotechnology, to PCR-based applications, to DNA hybridization arrays. Results in the literature rely on a wide variety of design criteria adapted to the particular requirements of each application. Using an extensively studied thermodynamic model, we perform a detailed study of several criteria for designing sequences intended to adopt a target secondary structure. We conclude that superior design methods should explicitly implement both a positive design paradigm (optimize affinity for the target structure) and a negative design paradigm (optimize specificity for the target structure). The commonly used approaches of sequence symmetry minimization and minimum free-energy satisfaction primarily implement negative design and can be strengthened by introducing a positive design component. Surprisingly, our findings hold for a wide range of secondary structures and are robust to modest perturbation of the thermodynamic parameters used for evaluating sequence quality, suggesting the feasibility and ongoing utility of a unified approach to nucleic acid design as parameter sets are refined further. Finally, we observe that designing for thermodynamic stability does not determine folding kinetics, emphasizing the opportunity for extending design criteria to target kinetic features of the energy landscape.

Algorithms↗

ILM: a web server for predicting RNA secondary structures with pseudoknots.

The ILM web server provides a web interface to two algorithms, iterated loop matching and maximum weighted matching, for efficiently predicting RNA secondary structures with pseudoknots. The algorithms can utilize either thermodynamic or comparative information or both, and thus can work on both aligned and individual sequences. Predicted secondary structures are presented in several formats compatible with a variety of existing visualization tools. The service can be accessed at http://cic.cs.wustl.edu/RNA/.

Algorithms↗

RNALOSS: a web server for RNA locally optimal secondary structures.

RNAomics, analogous to proteomics, concerns aspects of the secondary and tertiary structure, folding pathway, kinetics, comparison, function and regulation of all RNA in a living organism. Given recently discovered roles played by micro RNA, small interfering RNA, riboswitches, ribozymes, etc., it is important to gain insight into the folding process of RNA sequences. We describe the web server RNALOSS, which provides information about the distribution of locally optimal secondary structures, that possibly form kinetic traps in the folding process. The tool RNALOSS may be useful in designing RNA sequences which not only have low folding energy, but whose distribution of locally optimal secondary structures would suggest rapid and robust folding. Website: http://clavius.bc.edu/~clotelab/RNALOSS/.

Algorithms↗

Kinefold web server for RNA/DNA folding path and structure prediction including pseudoknots and knots.

The Kinefold web server provides a web interface for stochastic folding simulations of nucleic acids on second to minute molecular time scales. Renaturation or co-transcriptional folding paths are simulated at the level of helix formation and dissociation in agreement with the seminal experimental results. Pseudoknots and topologically 'entangled' helices (i.e. knots) are efficiently predicted taking into account simple geometrical and topological constraints. To encourage interactivity, simulations launched as immediate jobs are automatically stopped after a few seconds and return adapted recommendations. Users can then choose to continue incomplete simulations using the batch queuing system or go back and modify suggested options in their initial query. Detailed output provide (i) a series of low free energy structures, (ii) an online animated folding path and (iii) a programmable trajectory plot focusing on a few helices of interest to each user. The service can be accessed at http://kinefold.curie.fr/.

Algorithms↗

DINAMelt web server for nucleic acid melting prediction.

The DINAMelt web server simulates the melting of one or two single-stranded nucleic acids in solution. The goal is to predict not just a melting temperature for a hybridized pair of nucleic acids, but entire equilibrium melting profiles as a function of temperature. The two molecules are not required to be complementary, nor must the two strand concentrations be equal. Competition among different molecular species is automatically taken into account. Calculations consider not only the heterodimer, but also the two possible homodimers, as well as the folding of each single-stranded molecule. For each of these five molecular species, free energies are computed by summing Boltzmann factors over every possible hybridized or folded state. For temperatures within a user-specified range, calculations predict species mole fractions together with the free energy, enthalpy, entropy and heat capacity of the ensemble. Ultraviolet (UV) absorbance at 260 nm is simulated using published extinction coefficients and computed base pair probabilities. All results are available as text files and plots are provided for species concentrations, heat capacity and UV absorbance versus temperature. This server is connected to an active research program and should evolve as new theory and software are developed. The server URL is http://www.bioinfo.rpi.edu/applications/hybrid/.

Algorithms↗

Predicting candidate genomic sequences that correspond to synthetic functional RNA motifs.

Riboswitches and RNA interference are important emerging mechanisms found in many organisms to control gene expression. To enhance our understanding of such RNA roles, finding small regulatory motifs in genomes presents a challenge on a wide scale. Many simple functional RNA motifs have been found by in vitro selection experiments, which produce synthetic target-binding aptamers as well as catalytic RNAs, including the hammerhead ribozyme. Motivated by the prediction of Piganeau and Schroeder [(2003) Chem. Biol., 10, 103-104] that synthetic RNAs may have natural counterparts, we develop and apply an efficient computational protocol for identifying aptamer-like motifs in genomes. We define motifs from the sequence and structural information of synthetic aptamers, search for sequences in genomes that will produce motif matches, and then evaluate the structural stability and statistical significance of the potential hits. Our application to aptamers for streptomycin, chloramphenicol, neomycin B and ATP identifies 37 candidate sequences (in coding and non-coding regions) that fold to the target aptamer structures in bacterial and archaeal genomes. Further energetic screening reveals that several candidates exhibit energetic properties and sequence conservation patterns that are characteristic of functional motifs. Besides providing candidates for experimental testing, our computational protocol offers an avenue for expanding natural RNA's functional repertoire.

Algorithms↗

Reverse Sanger sequencing of RNA by MALDI-TOF mass spectrometry after solid phase purification.

Several DNA/RNA sequencing strategies have been developed using matrix-assisted laser desorption ionization mass spectrometry (MALDI-MS). In the reverse Sanger sequencing approach alpha-thiophosphate-containing NTPs are employed. Sequencing ladders are produced by the subsequent exonuclease cleavage, which is inhibited by the alpha-S-NTP at the 3' terminus. Here the reverse Sanger sequencing of RNA is described. The stability of RNA during the UV-MALDI process is higher relative to DNA, and RNA can be easily synthesized by transcription using bacteriophage RNA polymerase. alpha-S-rNTP was added to the reaction in a ratio of 1:3 to the native rNTPs and was incorporated statistically by the RNA polymerase. Four separate sequence ladders were produced, to avoid the problem of the only 1u mass difference between uridine and cytidine. However, it was shown that RNA transcription does not produce homogeneous transcripts. Therefore isolation of the full-length transcript is required to attain a non-ambiguous interpretation of cleavage spectra. This is achieved by the exclusive immobilization of the full-length transcript on a solid phase. The full-length transcripts were hybridized to magnetic beads, coated with short universal sequences, complementary to the in vitro RNA. After purification and isolation the RNA full-length transcript is cleaved by snake venom phosphodiesterase (SVP) and the obtained sequence ladder is analyzed by MALDI-MS.

RNA↗

AtMRD1 and AtMRU1, two novel genes with altered mRNA levels in the methionine over-accumulating mto1-1 mutant of Arabidopsis thaliana.

The mto1-1 mutant of Arabidopsis thaliana over-accumulates soluble methionine (Met) up to 40-fold higher than that in its Col-0 wild type. In order to identify genes regulated by altered Met concentrations, microarray analysis of gene expression in young rosettes and developing siliques of the mto1-1 mutant were performed. Expression of selected genes was then examined in detail in three developmental stages of the mto1-1 mutant using a combination of Northern hybridisation analysis and real-time PCR. Eight genes were identified that had altered mRNA accumulation levels in the mto1-1 mutant compared to that in wild-type plants. Three of the genes have known roles in plant development unrelated to amino acid biosynthesis. One other gene up-regulated specifically in mto1-1 rosettes shared similarity with the embryo-specific protein 3 (ATS3). Two novel genes, referred to as AtMRD1 and AtMRU1, were also identified that were expressed in a developmental manner in wild-type Col-0 and do not share sequence similarity with genes of known function. AtMRD1 was strongly down-regulated in both rosette and young silique tissues of the mto1-1 mutant. AtMRU1 was up-regulated approximately 3-fold in young mto1-1 rosettes and exhibited a developmental response to the mto1-1 mutation.

Amino Acid Sequence↗