Search PubMed⌕ Search

Biomedical subjects

David H Mathews

Publications and source records attributed to David H Mathews.

15 recordsLinked to original sources

A set of nearest neighbor parameters for predicting the enthalpy change of RNA secondary structure formation.

A complete set of nearest neighbor parameters to predict the enthalpy change of RNA secondary structure formation was derived. These parameters can be used with available free energy nearest neighbor parameters to extend the secondary structure prediction of RNA sequences to temperatures other than 37 degrees C. The parameters were tested by predicting the secondary structures of sequences with known secondary structure that are from organisms with known optimal growth temperatures. Compared with the previous set of enthalpy nearest neighbor parameters, the sensitivity of base pair prediction improved from 65.2 to 68.9% at optimal growth temperatures ranging from 10 to 60 degrees C. Base pair probabilities were predicted with a partition function and the positive predictive value of structure prediction is 90.4% when considering the base pairs in the lowest free energy structure with pairing probability of 0.99 or above. Moreover, a strong correlation is found between the predicted melting temperatures of RNA sequences and the optimal growth temperatures of the host organism. This indicates that organisms that live at higher temperatures have evolved RNA sequences with higher melting temperatures.

Algorithms↗

Interpreting oligonucleotide microarray data to determine RNA secondary structure: application to the 3' end of Bombyx mori R2 RNA.

A method to deduce RNA secondary structure on the basis of data from microarrays of 2'-O-methyl RNA 9-mers immobilized in agarose film on glass slides is tested with a 249 nucleotide RNA from the 3' end of the R2 retrotransposon from Bombyx mori. Various algorithms incorporating binding data and free-energy minimization calculations were compared for interpreting the data to provide possible secondary structures. Two different methods give structures with 100 and 87% of the base pairs determined by sequence comparison. In contrast, structures predicted by free-energy minimization alone by Mfold and RNAstructure contain 52 and 72% of the known base pairs, respectively. This combination of high throughput microarray techniques with algorithms using free-energy calculations has potential to allow for fast determination of RNA secondary structure. It should also facilitate the design of antisense and siRNA oligonucleotides.

3' Untranslated Regions↗

Nearest neighbor parameters for Watson-Crick complementary heteroduplexes formed between 2'-O-methyl RNA and RNA oligonucleotides.

Results from optical melting studies of Watson-Crick complementary heteroduplexes formed between 2'-O-methyl RNA and RNA oligonucleotides are used to determine nearest neighbor thermodynamic parameters for predicting the stabilities of such duplexes. The results are consistent with the physical model assumed by the individual nearest neighbor-hydrogen bonding model, which contains terms for helix initiation, base pair stacking and base pair composition. The sequence dependence is similar to that for Watson-Crick complementary RNA/RNA duplexes, which suggests that the sequence dependence may also be similar to that for other backbones that favor A-form RNA conformations.

Base Pairing↗

Prediction of RNA secondary structure by free energy minimization.

RNA secondary structure is often predicted from sequence by free energy minimization. Over the past two years, advances have been made in the estimation of folding free energy change, the mapping of secondary structure and the implementation of computer programs for structure prediction. The trends in computer program development are: efficient use of experimental mapping of structures to constrain structure prediction; use of statistical mechanics to improve the fidelity of structure prediction; inclusion of pseudoknots in secondary structure prediction; and use of two or more homologous sequences to find a common structure.

Base Sequence↗

Detection of non-coding RNAs on the basis of predicted secondary structure formation free energy change.

BACKGROUND: Non-coding RNAs (ncRNAs) have a multitude of roles in the cell, many of which remain to be discovered. However, it is difficult to detect novel ncRNAs in biochemical screens. To advance biological knowledge, computational methods that can accurately detect ncRNAs in sequenced genomes are therefore desirable. The increasing number of genomic sequences provides a rich dataset for computational comparative sequence analysis and detection of novel ncRNAs. RESULTS: Here, Dynalign, a program for predicting secondary structures common to two RNA sequences on the basis of minimizing folding free energy change, is utilized as a computational ncRNA detection tool. The Dynalign-computed optimal total free energy change, which scores the structural alignment and the free energy change of folding into a common structure for two RNA sequences, is shown to be an effective measure for distinguishing ncRNA from randomized sequences. To make the classification as a ncRNA, the total free energy change of an input sequence pair can either be compared with the total free energy changes of a set of control sequence pairs, or be used in combination with sequence length and nucleotide frequencies as input to a classification support vector machine. The latter method is much faster, but slightly less sensitive at a given specificity. Additionally, the classification support vector machine method is shown to be sensitive and specific on genomic ncRNA screens of two different Escherichia coli and Salmonella typhi genome alignments, in which many ncRNAs are known. The Dynalign computational experiments are also compared with two other ncRNA detection programs, RNAz and QRNA. CONCLUSION: The Dynalign-based support vector machine method is more sensitive for known ncRNAs in the test genomic screens than RNAz and QRNA. Additionally, both Dynalign-based methods are more sensitive than RNAz and QRNA at low sequence pair identities. Dynalign can be used as a comparable or more accurate tool than RNAz or QRNA in genomic screens, especially for low-identity regions. Dynalign provides a method for discovering ncRNAs in sequenced genomes that other methods may not identify. Significant improvements in Dynalign runtime have also been achieved.

Algorithms↗

The RNA Ontology Consortium: an open invitation to the RNA community.

The aim of the RNA Ontology Consortium (ROC) is to create an integrated conceptual framework-an RNA Ontology (RO)-with a common, dynamic, controlled, and structured vocabulary to describe and characterize RNA sequences, secondary structures, three-dimensional structures, and dynamics pertaining to RNA function. The RO should produce tools for clear communication about RNA structure and function for multiple uses, including the integration of RNA electronic resources into the Semantic Web. These tools should allow the accurate description in computer-interpretable form of the coupling between RNA architecture, function, and evolution. The purposes for creating the RO are, therefore, (1) to integrate sequence and structural databases; (2) to allow different computational tools to interoperate; (3) to create powerful software tools that bring advanced computational methods to the bench scientist; and (4) to facilitate precise searches for all relevant information pertaining to RNA. For example, one initial objective of the ROC is to define, identify, and classify RNA structural motifs described in the literature or appearing in databases and to agree on a computer-interpretable definition for each of these motifs. To achieve these aims, the ROC will foster communication and promote collaboration among RNA scientists by coordinating frequent face-to-face workshops to discuss, debate, and resolve difficult conceptual issues. These meeting opportunities will create new directions at various levels of RNA research. The ROC will work closely with the PDB/NDB structural databases and the Gene, Sequence, and Open Biomedical Ontology Consortia to integrate the RO with existing biological ontologies to extend existing content while maintaining interoperability.

Databases, Genetic↗

Revolutions in RNA secondary structure prediction.

RNA structure formation is hierarchical and, therefore, secondary structure, the sum of canonical base-pairs, can generally be predicted without knowledge of the three-dimensional structure. Secondary structure prediction algorithms evolved from predicting a single, lowest free energy structure to their current state where statistics can be determined from the thermodynamic ensemble. This article reviews the free energy minimization technique and the salient revolutions in the dynamic programming algorithm methods for secondary structure prediction. Emphasis is placed on highlighting the recently developed method, which statistically samples structures from the complete Boltzmann ensemble.

Algorithms↗

Nudged elastic band calculation of minimal energy paths for the conformational change of a GG non-canonical pair.

The nudged elastic band (NEB) technique has been implemented in AMBER to calculate low-energy paths for conformational changes. A novel simulated annealing protocol that does not require an initial hypothesis for the path is used to sample low-energy paths. This was used to study the conformational change of an RNA cis Watson-Crick/Hoogsteen GG non-canonical pair, with one G syn around the glycosidic bond and the other anti. A previous solution structure, determined by NMR-constrained modeling, demonstrated that the GG pairs change from (syn)G-(anti)G to (anti)G-(syn)G in the context of duplex r(GCAGGCGUGC) on the millisecond timescale. The set of low-energy paths found by NEB show that each G flips independently around the glycosidic bond, with the anti G flipping to syn first. Guanine bases flip without opening adjacent base-pairs by protruding into the major groove, accommodated by a transient change by the ribose to C2'-exo sugar pucker. Hydrogen bonds between bases and the backbone, which lower the energetic barrier to flipping, are observed along the path. The results show the plasticity of RNA base-pairs in helices, which is important for biological processes, including mismatch repair, protein recognition, and translation. The modeling of the GG conformational change also demonstrates that NEB can be used to discover non-trivial paths for macromolecules and therefore NEB can be used as an exploratory method for predicting putative conformational change paths.

Base Pairing↗

The influence of locked nucleic acid residues on the thermodynamic properties of 2'-O-methyl RNA/RNA heteroduplexes.

The influence of locked nucleic acid (LNA) residues on the thermodynamic properties of 2'-O-methyl RNA/RNA heteroduplexes is reported. Optical melting studies indicate that LNA incorporated into an otherwise 2'-O-methyl RNA oligonucleotide usually, but not always, enhances the stabilities of complementary duplexes formed with RNA. Several trends are apparent, including: (i) a 3' terminal U LNA and 5' terminal LNAs are less stabilizing than interior and other 3' terminal LNAs; (ii) most of the stability enhancement is achieved when LNA nucleotides are separated by at least one 2'-O-methyl nucleotide; and (iii) the effects of LNA substitutions are approximately additive when the LNA nucleotides are separated by at least one 2'-O-methyl nucleotide. An equation is proposed to approximate the stabilities of complementary duplexes formed with RNA when at least one 2'-O-methyl nucleotide separates LNA nucleotides. The sequence dependence of 2'-O-methyl RNA/RNA duplexes appears to be similar to that of RNA/RNA duplexes, and preliminary nearest-neighbor free energy increments at 37 degrees C are presented for 2'-O-methyl RNA/RNA duplexes. Internal mismatches with LNA nucleotides significantly destabilize duplexes with RNA.

Base Pair Mismatch↗

Predicting a set of minimal free energy RNA secondary structures common to two sequences.

MOTIVATION: Function derives from structure, therefore, there is need for methods to predict functional RNA structures. RESULTS: The Dynalign algorithm, which predicts the lowest free energy secondary structure common to two unaligned RNA sequences, is extended to the prediction of a set of low-energy structures. Dot plots can be drawn to show all base pairs in structures within an energy increment. Dynalign predicts more well-defined structures than structure prediction using a single sequence; in 5S rRNA sequences, the average number of base pairs in structures with energy within 20% of the lowest energy structure is 317 using Dynalign, but 569 using a single sequence. Structure prediction with Dynalign can also be constrained according to experiment or comparative analysis. The accuracy, measured as sensitivity and positive predictive value, of Dynalign is greater than predictions with a single sequence. AVAILABILITY: Dynalign can be downloaded at http://rna.urmc.rochester.edu

Algorithms↗

Incorporating chemical modification constraints into a dynamic programming algorithm for prediction of RNA secondary structure.

A dynamic programming algorithm for prediction of RNA secondary structure has been revised to accommodate folding constraints determined by chemical modification and to include free energy increments for coaxial stacking of helices when they are either adjacent or separated by a single mismatch. Furthermore, free energy parameters are revised to account for recent experimental results for terminal mismatches and hairpin, bulge, internal, and multibranch loops. To demonstrate the applicability of this method, in vivo modification was performed on 5S rRNA in both Escherichia coli and Candida albicans with 1-cyclohexyl-3-(2-morpholinoethyl) carbodiimide metho-p-toluene sulfonate, dimethyl sulfate, and kethoxal. The percentage of known base pairs in the predicted structure increased from 26.3% to 86.8% for the E. coli sequence by using modification constraints. For C. albicans, the accuracy remained 87.5% both with and without modification data. On average, for these sequences and a set of 14 sequences with known secondary structure and chemical modification data taken from the literature, accuracy improves from 67% to 76%. This enhancement primarily reflects improvement for three sequences that are predicted with <40% accuracy on the basis of energetics alone. For these sequences, inclusion of chemical modification constraints improves the average accuracy from 28% to 78%. For the 11 sequences with <6% pseudoknotted base pairs, structures predicted with constraints from chemical modification contain on average 84% of known canonical base pairs.

Algorithms↗

Secondary structure models of the 3' untranslated regions of diverse R2 RNAs.

The RNA structure of the 3' untranslated region (UTR) of the R2 retrotransposable element is recognized by the R2-encoded reverse transcriptase in a reaction called target primed reverse transcription (TPRT). To provide insight into structure-function relationships important for TPRT, we have created alignments that reveal the secondary structure for 22 Drosophila and five silkmoth 3' UTR R2 sequences. In addition, free energy minimization has been used to predict the secondary structure for the 3' UTR R2 RNA of Forficula auricularia. The predicted structures for Bombyx mori and F. auricularia are consistent with chemical modification data obtained with beta-ethoxy-alpha-ketobutyraldehyde (kethoxal), dimethyl sulfate, and 1-cyclohexyl-3-(2-morpholinoethyl)carbodiimide metho-p-toluene sulfonate. The structures appear to have common helices that are likely important for function.

3' Untranslated Regions↗

Using an RNA secondary structure partition function to determine confidence in base pairs predicted by free energy minimization.

A partition function calculation for RNA secondary structure is presented that uses a current set of nearest neighbor parameters for conformational free energy at 37 degrees C, including coaxial stacking. For a diverse database of RNA sequences, base pairs in the predicted minimum free energy structure that are predicted by the partition function to have high base pairing probability have a significantly higher positive predictive value for known base pairs. For example, the average positive predictive value, 65.8%, is increased to 91.0% when only base pairs with probability of 0.99 or above are considered. The quality of base pair predictions can also be increased by the addition of experimentally determined constraints, including enzymatic cleavage, flavin mono-nucleotide cleavage, and chemical modification. Predicted secondary structures can be color annotated to demonstrate pairs with high probability that are therefore well determined as compared to base pairs with lower probability of pairing.

Algorithms↗

Dynalign: an algorithm for finding the secondary structure common to two RNA sequences.

With the rapid increase in the size of the genome sequence database, computational analysis of RNA will become increasingly important in revealing structure-function relationships and potential drug targets. RNA secondary structure prediction for a single sequence is 73 % accurate on average for a large database of known secondary structures. This level of accuracy provides a good starting point for determining a secondary structure either by comparative sequence analysis or by the interpretation of experimental studies. Dynalign is a new computer algorithm that improves the accuracy of structure prediction by combining free energy minimization and comparative sequence analysis to find a low free energy structure common to two sequences without requiring any sequence identity. It uses a dynamic programming construct suggested by Sankoff. Dynalign, however, restricts the maximum distance, M, allowed between aligned nucleotides in the two sequences. This makes the calculation tractable because the complexity is simplified to O(M(3)N(3)), where N is the length of the shorter sequence. The accuracy of Dynalign was tested with sets of 13 tRNAs, seven 5 S rRNAs, and two R2 3' UTR sequences. On average, Dynalign predicted 86.1 % of known base-pairs in the tRNAs, as compared to 59.7 % for free energy minimization alone. For the 5 S rRNAs, the average accuracy improves from 47.8 % to 86.4 %. The secondary structure of the R2 3' UTR from Drosophila takahashii is poorly predicted by standard free energy minimization. With Dynalign, however, the structure predicted in tandem with the sequence from Drosophila melanogaster nearly matches the structure determined by comparative sequence analysis.

3' Untranslated Regions↗

Experimentally derived nearest-neighbor parameters for the stability of RNA three- and four-way multibranch loops.

Algorithms for predicting RNA secondary structure require approximations for the free energies of multibranch loops, also called junctions. The stabilities of 62 RNA duplexes with three- and four-way multibranch loops were determined by optical melting. To account for the observed sequence dependence, a revised loop free-energy approximation is proposed that accounts for the strain in three-way junctions with fewer than two unpaired nucleotides, penalizes asymmetry in the distribution of unpaired nucleotides, and gives a bonus for four-way loops relative to three-way loops. Parameters for this equation were determined by linear regression.

Algorithms↗