Search PubMed⌕ Search

Biomedical subjects

Paul P Gardner

Publications and source records attributed to Paul P Gardner.

9 recordsLinked to original sources

Exploring genomic dark matter: a critical assessment of the performance of homology search methods on noncoding RNA.

Homology search is one of the most ubiquitous bioinformatic tasks, yet it is unknown how effective the currently available tools are for identifying noncoding RNAs (ncRNAs). In this work, we use reliable ncRNA data sets to assess the effectiveness of methods such as BLAST, FASTA, HMMer, and Infernal. Surprisingly, the most popular homology search methods are often the least accurate. As a result, many studies have used inappropriate tools for their analyses. On the basis of our results, we suggest homology search strategies using the currently available tools and some directions for future development.

Algorithms↗

Identification of miRNA targets with stable isotope labeling by amino acids in cell culture.

miRNAs are small noncoding RNAs that regulate gene expression. We have used stable isotope labeling by amino acids in cell culture (SILAC) to investigate the effect of miRNA-1 on the HeLa cell proteome. Expression of 12 out of 504 investigated proteins was repressed by miRNA-1 transfection. This repressed set of genes significantly overlaps with miRNA-1 regulated genes that have been identified with DNA array technology and are predicted by computational methods. Moreover, we find that the 3'-untranslated region for the repressed set are enriched in miRNA-1 complementary sites. Our findings demonstrate that SILAC can be used for miRNA target identification and that one highly expressed miRNA can regulate the levels of many different proteins.

3' Untranslated Regions↗

A hidden Markov model approach for determining expression from genomic tiling micro arrays.

BACKGROUND: Genomic tiling micro arrays have great potential for identifying previously undiscovered coding as well as non-coding transcription. To-date, however, analyses of these data have been performed in an ad hoc fashion. RESULTS: We present a probabilistic procedure, ExpressHMM, that adaptively models tiling data prior to predicting expression on genomic sequence. A hidden Markov model (HMM) is used to model the distributions of tiling array probe scores in expressed and non-expressed regions. The HMM is trained on sets of probes mapped to regions of annotated expression and non-expression. Subsequently, prediction of transcribed fragments is made on tiled genomic sequence. The prediction is accompanied by an expression probability curve for visual inspection of the supporting evidence. We test ExpressHMM on data from the Cheng et al. (2005) tiling array experiments on ten Human chromosomes. Results can be downloaded and viewed from our web site. CONCLUSION: The value of adaptive modelling of fluorescence scores prior to categorisation into expressed and non-expressed probes is demonstrated. Our results indicate that our adaptive approach is superior to the previous analysis in terms of nucleotide sensitivity and transfrag specificity.

Algorithms↗

A comparison of RNA folding measures.

BACKGROUND: In the last few decades there has been a great deal of discussion concerning whether or not noncoding RNA sequences (ncRNAs) fold in a more well-defined manner than random sequences. In this paper, we investigate several existing measures for how well an RNA sequence folds, and compare the behaviour of these measures over a large range of Rfam ncRNA families. Such measures can be useful in, for example, identifying novel ncRNAs, and indicating the presence of alternate RNA foldings. RESULTS: Our analysis shows that ncRNAs, but not mRNAs, in general have lower minimal free energy (MFE) than random sequences with the same dinucleotide frequency. Moreover, even when the MFE is significant, many ncRNAs appear to not have a unique fold, but rather several alternative folds, at least when folded in silico. Furthermore, we find that the six investigated measures are correlated to varying degrees. CONCLUSION: Due to the correlations between the different measures we find that it is sufficient to use only two of them in RNA folding studies, one to test if the sequence in question has lower energy than a random sequence with the same dinucleotide frequency (the Z-score) and the other to see if the sequence has a unique fold (the average base-pair distance, D).

Models, Molecular↗

A benchmark of multiple sequence alignment programs upon structural RNAs.

To date, few attempts have been made to benchmark the alignment algorithms upon nucleic acid sequences. Frequently, sophisticated PAM or BLOSUM like models are used to align proteins, yet equivalents are not considered for nucleic acids; instead, rather ad hoc models are generally favoured. Here, we systematically test the performance of existing alignment algorithms on structural RNAs. This work was aimed at achieving the following goals: (i) to determine conditions where it is appropriate to apply common sequence alignment methods to the structural RNA alignment problem. This indicates where and when researchers should consider augmenting the alignment process with auxiliary information, such as secondary structure and (ii) to determine which sequence alignment algorithms perform well under the broadest range of conditions. We find that sequence alignment alone, using the current algorithms, is generally inappropriate <50-60% sequence identity. Second, we note that the probabilistic method ProAlign and the aging Clustal algorithms generally outperform other sequence-based algorithms, under the broadest range of applications.

Algorithms↗

A comprehensive comparison of comparative RNA structure prediction approaches.

BACKGROUND: An increasing number of researchers have released novel RNA structure analysis and prediction algorithms for comparative approaches to structure prediction. Yet, independent benchmarking of these algorithms is rarely performed as is now common practice for protein-folding, gene-finding and multiple-sequence-alignment algorithms. RESULTS: Here we evaluate a number of RNA folding algorithms using reliable RNA data-sets and compare their relative performance. CONCLUSIONS: We conclude that comparative data can enhance structure prediction but structure-prediction-algorithms vary widely in terms of both sensitivity and selectivity across different lengths and homologies. Furthermore, we outline some directions for future research.

Algorithms↗

Optimal alphabets for an RNA world.

Experiments have shown that the canonical AUCG genetic alphabet is not the only possible nucleotide alphabet. In this work we address the question 'is the canonical alphabet optimal?' We make the assumption that the genetic alphabet was determined in the RNA world. Computational tools are used to infer the RNA secondary structure (shape) from a given RNA sequence, and statistics from RNA shapes are gathered with respect to alphabet size. Then, simulations based upon the replication and selection of fixed-sized RNA populations are used to investigate the effect of alternative alphabets upon RNA's ability to step through a fitness landscape. These results show that for a low copy fidelity the canonical alphabet is fitter than two-, six- and eight-letter alphabets. In higher copy-fidelity experiments, six-letter alphabets outperform the four-letter alphabets, suggesting that the canonical alphabet is indeed a relic of the RNA world.

Animals↗

A search for H/ACA snoRNAs in yeast using MFE secondary structure prediction.

MOTIVATION: Noncoding RNA genes produce functional RNA molecules rather than coding for proteins. One such family is the H/ACA snoRNAs. Unlike the related C/D snoRNAs these have resisted automated detection to date. RESULTS: We develop an algorithm to screen the yeast genome for novel H/ACA snoRNAs. To achieve this, we introduce some new methods for facilitating the search for noncoding RNAs in genomic sequences which are based on properties of predicted minimum free-energy (MFE) secondary structures. The algorithm has been implemented and can be generalized to enable screening of other eukaryote genomes. We find that use of primary sequence alone is insufficient for identifying novel H/ACA snoRNAs. Only the use of secondary structure filters reduces the number of candidates to a manageable size. From genomic context, we identify three strong H/ACA snoRNA candidates. These together with a further 47 candidates obtained by our analysis are being experimentally screened.

Algorithms↗

Sequence diversity and functional conservation of the origin of replication in lactococcal prolate phages.

Prolate or c2-like phages are a large homologous group of viruses that infect the bacterium Lactococcus lactis. In a collection of 122 prolate phages, three distinct, non-cross-hybridizing groups of origins of DNA replication were found. The nonconserved sequence was confined to the template for an untranslated transcript, P(E)1-T, 300 to 400 nucleotides in length, while the flanking sequences were conserved. All three origin types, despite the low sequence homology, have the same functional characteristics: they express abundant P(E)1-T transcripts and can function as origins of plasmid replication in the absence of phage proteins. Using chimeric constructs, we showed that hybrids of two nonhomologous origin sequences failed to function as replication origins, suggesting that preservation of a particular secondary structure of the P(E)1-T transcript is required for replication. This is the first systematic survey of the sequence and function of origins of replication in a group of lactococcal phages.

Bacteriophages↗