Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “RNA sequencing analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

EGASP: the human ENCODE Genome Annotation Assessment Project.

BACKGROUND: We present the results of EGASP, a community experiment to assess the state-of-the-art in genome annotation within the ENCODE regions, which span 1% of the human genome sequence. The experiment had two major goals: the assessment of the accuracy of computational methods to predict protein coding genes; and the overall assessment of the completeness of the current human genome annotations as represented in the ENCODE regions. For the computational prediction assessment, eighteen groups contributed gene predictions. We evaluated these submissions against each other based on a 'reference set' of annotations generated as part of the GENCODE project. These annotations were not available to the prediction groups prior to the submission deadline, so that their predictions were blind and an external advisory committee could perform a fair assessment. RESULTS: The best methods had at least one gene transcript correctly predicted for close to 70% of the annotated genes. Nevertheless, the multiple transcript accuracy, taking into account alternative splicing, reached only approximately 40% to 50% accuracy. At the coding nucleotide level, the best programs reached an accuracy of 90% in both sensitivity and specificity. Programs relying on mRNA and protein sequences were the most accurate in reproducing the manually curated annotations. Experimental validation shows that only a very small percentage (3.2%) of the selected 221 computationally predicted exons outside of the existing annotation could be verified. CONCLUSION: This is the first such experiment in human DNA, and we have followed the standards established in a similar experiment, GASP1, in Drosophila melanogaster. We believe the results presented here contribute to the value of ongoing large-scale annotation projects and should guide further experimental methods when being scaled up to the entire human genome sequence.

Alternative Splicing↗

Performance assessment of promoter predictions on ENCODE regions in the EGASP experiment.

BACKGROUND: This study analyzes the predictions of a number of promoter predictors on the ENCODE regions of the human genome as part of the ENCODE Genome Annotation Assessment Project (EGASP). The systems analyzed operate on various principles and we assessed the effectiveness of different conceptual strategies used to correlate produced promoter predictions with the manually annotated 5' gene ends. RESULTS: The predictions were assessed relative to the manual HAVANA annotation of the 5' gene ends. These 5' gene ends were used as the estimated reference transcription start sites. With the maximum allowed distance for predictions of 1,000 nucleotides from the reference transcription start sites, the sensitivity of predictors was in the range 32% to 56%, while the positive predictive value was in the range 79% to 93%. The average distance mismatch of predictions from the reference transcription start sites was in the range 259 to 305 nucleotides. At the same time, using transcription start site estimates from DBTSS and H-Invitational databases as promoter predictions, we obtained a sensitivity of 58%, a positive predictive value of 92%, and an average distance from the annotated transcription start sites of 117 nucleotides. In this experiment, the best performing promoter predictors were those that combined promoter prediction with gene prediction. The main reason for this is the reduced promoter search space that resulted in smaller numbers of false positive predictions. CONCLUSION: The main finding, now supported by comprehensive data, is that the accuracy of human promoter predictors for high-throughput annotation purposes can be significantly improved if promoter prediction is combined with gene prediction. Based on the lessons learned in this experiment, we propose a framework for the preparation of the next similar promoter prediction assessment.

Computational Biology↗

Direct selection of trans-acting ligase ribozymes by in vitro compartmentalization.

We have used a compartmentalized in vitro selection method to directly select for ligase ribozymes that are capable of acting on and turning over separable oligonucleotide substrates. Starting from a degenerate pool, we selected a trans-acting variant of the Bartel class I ligase which statistically may have been the only active variant in the starting pool. The isolation of this sequence from the population suggests that this selection method is extremely robust at selecting optimal ribozymes and should, therefore, prove useful for the selection and optimization of other trans-acting nucleic acid catalysts capable of multiple turnover catalysis.

Base Pairing↗

Finding specific RNA motifs: function in a zeptomole world?

We have developed a new method for estimating the abundance of any modular (piecewise) RNA motif within a longer random region. We have used this method to estimate the size of the active motifs available to modern SELEX experiments (picomoles of unique sequences) and to a plausible RNA World (zeptomoles of unique sequences: 1 zmole = 602 sequences). Unexpectedly, activities such as specific isoleucine binding are almost certainly present in zeptomoles of molecules, and even ribozymes such as self-cleavage motifs may appear (depending on assumptions about the minimal structures). The number of specified nucleotides is not the only important determinant of a motif's rarity: The number of modules into which it is divided, and the details of this division, are also crucial. We propose three maxims for easily isolated motifs: the Maxim of Minimization, the Maxim of Multiplicity, and the Maxim of the Median. These maxims together state that selected motifs should be small and composed of as many separate, equally sized modules as possible. For evenly divided motifs with four modules, the largest accessible activity in picomole scale (1-1000 pmole) pools of length 100 is about 34 nucleotides; while for zeptomole scale (1-1000 zmole) pools it is about 20 specific nucleotides (50% probability of occurrence). This latter figure includes some ribozymes and aptamers. Consequently, an RNA metabolism apparently could have begun with only zeptomoles of RNA molecules.

Animals↗

BayesFold: rational 2 degrees folds that combine thermodynamic, covariation, and chemical data for aligned RNA sequences.

BayesFold is a Web application that folds an alignment of closely related sequences and evaluates hypotheses about their shared structure. It uses Bayes's Theorem to combine information from several sources, including chemical mapping (if available), thermodynamic folding, and observed sequence variations. Its method provides a rational basis for integrating results, even when these methods conflict. On a gapped alignment of 86 tRNAPhe sequences each 77 bases long, BayesFold takes 31 sec to perform the calculations; the best structure contained 95% of the base pairs in the true structure, and the true structure was ranked second. Notably, similar results come from random samples of only 10 sequences from the alignment (running time 3 sec), suggesting that remarkably few sequences are required for good results. In contrast, folding single sequences with BayesFold produced structures 9.6 bp different, or with the Vienna package, 13.4 bp different, from the true structure. Similar results were obtained for other families of tRNAs. We especially recommend BayesFold for alignments of 3-50 closely related sequences, such as the sequence families frequently found in SELEX. In addition to providing a convenient way to explore the effects of each of the criteria on the plausibility of different structures, BayesFold also makes it easy to produce publication-quality secondary-structure graphics. The Web interface, available at http://bayes.colorado.edu/fold/, includes the flexibility to thread any of the sequences (or the consensus sequence) through any of the structures, including the one judged most probable.

Algorithms↗

A mutational analysis of U12-dependent splice site dinucleotides.

Introns spliced by the U12-dependent minor spliceosome are divided into two classes based on their splice site dinucleotides. The /AU-AC/ class accounts for about one-third of U12-dependent introns in humans, while the /GU-AG/ class accounts for the other two-thirds. We have investigated the in vivo and in vitro splicing phenotypes of mutations in these dinucleotide sequences. A 5' A residue can splice to any 3' residue, although C is preferred. A 5' G residue can splice to 3' G or U residues with a preference for G. Little or no splicing was observed to 3' A or C residues. A 5' U or C residue is highly deleterious for U12-dependent splicing, although some combinations, notably 5' U to 3' U produced detectable spliced products. The dependence of 3' splice site activity on the identity of the 5' residue provides evidence for communication between the first and last nucleotides of the intron. Most mutants in the second position of the 5' splice site and the next to last position of the 3' splice site were defective for splicing. Double mutants of these residues showed no evidence of communication between these nucleotides. Varying the distance between the branch site and the 3' splice site dinucleotide in the /GU-AG/ class showed that a somewhat larger range of distances was functional than for the /AU-AC/ class. The optimum branch site to 3' splice site distance of 11-12 nucleotides appears to be the same for both classes.

Animals↗

Theoretical and experimental selection parameters for HBV-directed antisense RNA are related to increased RNA-RNA annealing.

Annealing kinetics of antisense species against two different target regions of the hepatitis B virus (HBV) were measured by kinetic in vitro selection. Individual association rates were related to energies calculated for local sequence segments and predicted structures of the complete pregenomic target RNA. A relationship between the presence of external loops and joint sequences with fast pairing was observed whereas internal loops did not favor fast RNA-RNA annealing. The findings were used to predict a fast-annealing HBV-directed antisense oligodeoxyribonucleotide that turned out to pair with its target RNA at an association rate constant of k=9.2 x 10(4) M(-1) s(-1), which is substantially faster than the annealing rates of artificial antisense RNA so far included in in vitro selection assays.

Computer Simulation↗

Predicting RNA structure using mutual information.

BACKGROUND: With the ever-increasing number of sequenced RNAs and the establishment of new RNA databases, such as the Comparative RNA Web Site and Rfam, there is a growing need for accurately and automatically predicting RNA structures from multiple alignments. Since RNA secondary structure is often conserved in evolution, the well known, but underused, mutual information measure for identifying covarying sites in an alignment can be useful for identifying structural elements. This article presents MIfold, a MATLAB toolbox that employs mutual information, or a related covariation measure, to display and predict conserved RNA secondary structure (including pseudoknots) from an alignment. RESULTS: We show that MIfold can be used to predict simple pseudoknots, and that the performance can be adjusted to make it either more sensitive or more selective. We also demonstrate that the overall performance of MIfold improves with the number of aligned sequences for certain types of RNA sequences. In addition, we show that, for these sequences, MIfold is more sensitive but less selective than the related RNAalifold structure prediction program and is comparable with the COVE structure prediction package. CONCLUSION: MIfold provides a useful supplementary tool to programs such as RNA Structure Logo, RNAalifold and COVE, and should be useful for automatically generating structural predictions for databases such as Rfam.

Algorithms↗

Identification and sequence determination of mRNAs detected in dormant (diapausing) Aedes triseriatus mosquito embryos.

Many insects survive adverse climatic conditions in a dormant state known as diapause. In this study, we identified and sequenced several mRNAs in diapausing Aedes triseriatus mosquito embryos. Using reverse-transcription PCR and 5' RACE, we identified a 995-nucleotide cDNA that encodes a 259-amino acid protein of unknown function. This putative protein displays strong sequence similarity to Drosophila melanogaster (95%), human (87%), Caenorhabditis elegans (86%) and yeast (81%) counterparts. The second identified full-length cDNA consists of 624 nucleotides and encodes a 174-amino acid protein of unknown function. This putative protein displays significant sequence similarity to D. melanogaster (68%), human (59%), plant (57%) and yeast (49%) counterparts. We also detected a number of cDNA fragments that exhibited significant sequence similarity to a mitochondrial cytochrome C oxidase subunit, human N33 protein (a potential human prostate tumor suppressor), 18S and 28S ribosomal RNAs, protein disulfide-isomerase, and guanine nucleotide-binding protein.

Aedes↗

Management of a measles outbreak among Old World nonhuman primates.

BACKGROUND AND PURPOSE: A measles outbreak in a facility housing Old World nonhuman primates developed over a 2-month period in 1996, providing an opportunity to study the epidemiology of this highly infectious disease in an animal-handling setting. METHODS: Serum and urine specimens were collected from monkeys housed in the room where the initial measles cases were identified, other monkeys with suspicious measles-like signs, and employees working in the affected areas. Serum specimens were tested for measles virus-specific IgG and IgM antibodies, and urine specimens were tested for measles virus by virus isolation or reverse transcriptase-polymerase chain reaction (RT-PCR). RESULTS: A total of 94 monkeys in two separate facilities had evidence of an acute measles infection. The outbreak was caused by a wild-type virus that had been associated with recent human cases of acute measles in the United States; however, an investigation was unable to identify the original source of the outbreak. Quarantine and massive vaccination helped to control further spread of infection. CONCLUSIONS: Results emphasize the value of having a measles control plan in place that includes a preventive measles vaccination program involving human and nonhuman primates to decrease the likelihood of a facility outbreak.

Animals↗

[Acute transverse myelitis caused by enterovirus].

Enteroviruses comprise a group of commonly encountered small RNA viruses with genetic similarities. We report a case of transverse myelitis associated with an enterovirus infection. By use of PCR, enterovirus-specific RNA sequences were detected in the cerebrospinal fluid and in swabs from the throat and the rectum at the time of admission. Routine cultures and serology gave no other explanation for the clinical condition. The patient was treated with intravenous immunoglobulin and steroids and improved gradually. She was fully recovered 18 months after onset of disease.

Child, Preschool↗

[Self-aggregation of the structural protein encoded by rice ragged stunt Oryzavirus genome segment 8].

Rice ragged stunt oryzavirus (RRSV) is a member of the genus oryzavirus within the family Reoviridae. Its genome consists of ten segments of dsRNA. The functions of all products encoded by these viral genome segments, except one encoded by S9, have not yet been elucidated. In the present study, the ORF of S 8 of RRSV-Philippines isolate was sequenced and expressed in E. coli. The 67 kD product of S8 could be self-cleaved into two fragments with molecular weights of 43 kD and 26 kD. Western blotting indicated that both 67 kD and 43 kD products were major structural proteins of the virus. It was also found that the 67 kD protein could self aggregate into aggregates with higher sedimentation rate in sucrose gradients during centrifugation. Moreover, the self-aggregation process could be accelerated by the complex of S6 product and genome dsRNAs of RRSV. These results suggest that the S8 products, 67 kD or 43 kD, may be the structural components of the viral inner-capsid.

Amino Acid Sequence↗