Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “RNA sequencing analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Serotype-specific identification of polioviruses by PCR using primers containing mixed-base or deoxyinosine residues at positions of codon degeneracy.

We have developed a method for determining the serotypes of poliovirus isolates by PCR. Three sets of serotype-specific antisense PCR-initiating primers (primers seroPV1A, seroPV2A, and seroPV3A) were designed to pair with codons of VP1 amino acid sequences that are conserved within but that differ across serotypes. The sense polarity primers (primers seroPV1S, seroPV2S, and seroPV3S) matched codons of more conserved capsid sequences. The primers contain mixed-base and deoxyinosine residues to compensate for the high rate of degeneracy of the targeted codons. The serotypes of all polioviruses tested (48 vaccine-related isolates and 110 diverse wild isolates) were correctly identified by PCR with the serotype-specific primers. None of the genomic sequences of 49 nonpolio enterovirus reference strains were amplified under equivalent reaction conditions with any of the three primer sets. These primers are useful for the rapid screening of poliovirus isolates and for determining the compositions of cultures containing mixtures of poliovirus serotypes.

Amino Acid Sequence↗

Genetic changes in hepatitis delta virus from acutely and chronically infected woodchucks.

A woodchuck-derived hepatitis delta virus (HDV) inoculum was created by transfection of a genotype I HDV cDNA clone directly into the liver of a woodchuck that was chronically infected with woodchuck hepatitis virus. All woodchucks receiving this inoculum became positive for HDV RNA in serum, and 67% became chronically infected, similar to the rate of chronic HDV infection in humans. Analysis of HDV sequences obtained at 73 weeks postinfection indicated that changes had occurred at a rate of 0.5% per year; many of these modifications were consistent with editing by host RNA adenosine deaminase. The appearance of sequence changes, which were not evenly distributed on the genome, was correlated with the course of HDV infection. A limited number of modifications occurred in the consensus sequence of the viral genome that altered the sequence of the hepatitis delta antigen (HDAg). All chronically infected animals examined exhibited these changes 73 weeks following infection, but at earlier times, only one of the HDV carriers exhibited consensus sequence substitutions. On the other hand, sequence modifications in animals that eventually recovered from HDV infection were apparent after 27 weeks. The data are consistent with a model in which HDV sequence changes are selected by host immune responses. Chronic HDV infection in woodchucks may result from a delayed and weak immune response that is limited to a small number of epitopes on HDAg.

Acute Disease↗

An eight-nucleotide sequence in the potato virus X 3' untranslated region is required for both host protein binding and viral multiplication.

Gel retardation and UV-cross-linking techniques were used to demonstrate that two tobacco proteins, with approximate molecular masses of 28 and 32 kDa, bind to a site within the 3' region of potato virus X (PVX) genomic RNA. The protein binding is specific, in that a 50-fold excess of unlabeled probe prevents formation of the complexes but no reduction is observed with a 2,000-fold molar excess of yeast tRNA. Complex formation is inhibited by poly(U) but is relatively unaffected by poly(A), poly(G), or poly(C-I). PVX RNA-host protein complex formation occurs in vitro at salt concentrations up to 400 mM. Deletion mapping indicates that the proteins bind within the 3' untranslated region (UTR) of PVX genomic RNA and that an 8-nucleotide U-rich sequence (5'-UAUUUUCU) is required for the binding. Deletion of the 8-nucleotide U-rich region from the 3' UTR of a sensitive PVX reporter virus that carries the luciferase gene in place of the PVX coat protein gene results in a more than 70,000-fold reduction in luciferase expression in tobacco protoplasts. RNA probes carrying the sequence GCGC in place of the central four contiguous uridines of the 8-nucleotide U-rich motif fail to bind host protein at detectable levels, and the same mutation, when introduced into the PVX reporter virus, eliminates viral multiplication. Mutations of 1 or 2 nucleotides within the same four uridines reduced both binding of host proteins and replication of reporter virus. These results indicate that the 8-nucleotide U-rich motif within the PVX 3' UTR is important for some aspect of viral multiplication and suggest that host protein binding plays a role in the process.

Gene Deletion↗

Molecular mechanisms of coxsackievirus persistence in chronic inflammatory myopathy: viral RNA persists through formation of a double-stranded complex without associated genomic mutations or evolution.

Enterovirus infection and persistence have been implicated in the pathogenesis of certain chronic muscle diseases. In vitro studies suggest that persistent enteroviruses mutate, evolving into forms that are less lytic and display altered tropism, but it is less clear whether these mechanisms operate in vivo. In this study, persistent coxsackievirus RNA from the muscle of mice afflicted with chronic inflammatory myopathy (CIM) was characterized and compared with RNA from a virus that had established a persistent infection of G8 mouse myoblasts for 30 passages in vitro. Competitive strand-specific reverse transcription-PCR and susceptibility to RNase I treatment revealed that plus- and minus-strand viral RNAs were present at nearly equivalent levels in muscle and that they persisted in a double-stranded conformation. All regions of the viral genome persisted and were amplified as a series of seven overlapping fragments. Restriction endonuclease fingerprinting coupled with sequencing indicated that there was no evolution of the viral genome associated with its persistence in muscle. This contrasted with the productive persistent infection that was established in myoblast cultures, where plus-strand RNA predominated and persistent virus developed distinct mutations. In vitro persistence proceeded by a carrier culture mechanism and was completely dependent on production of infectious virus, since persistent viral RNA was not detected in cultures subjected to antibody-mediated curing. These experiments demonstrate that persistence of coxsackievirus RNA in muscle is not facilitated by distinct genetic changes in the virus that give rise to replication-defective forms but occurs primarily through production of stable double-stranded RNA that is produced as the acute viral infection resolves. The data suggest a mechanism for coxsackievirus persistence in myofibers and perhaps other nondividing cells whereby cells that survive infection could harbor persistent viral RNA for extended times without producing detectable levels of infectious virus.

Animals↗

Pathogenesis of borna disease virus: granulocyte fractions of psychiatric patients harbor infectious virus in the absence of antiviral antibodies.

Borna disease virus (BDV) causes acute and persistent infections in various vertebrates. During recent years, BDV-specific serum antibodies, BDV antigen, and BDV-specific nucleic acid were found in humans suffering from psychiatric disorders. Furthermore, viral antigen was detected in human autopsy brain tissue by immunohistochemical staining. Whether BDV infection can be associated with psychiatric disorders is still a matter of debate; no direct evidence has ever been presented. In the present study we report on (i) the detection of BDV-specific nucleic acid in human granulocyte cell fraction from three different psychiatric patients and (ii) the isolation of infectious BDV from these cells obtained from a patient with multiple psychiatric disorders. In leukocyte preparations other than granulocytes, either no BDV RNA was detected or positive PCR results were obtained only if there was at least 20% contamination with granulocytes. Parts of the antigenome of the isolated virus were sequenced, demonstrating the close relationship to the prototype BDV strains (He/80 and strain V) as well as to other human virus sequences. Our data provide strong evidence that cells in the granulocyte fraction represent the major if not the sole cell type harboring BDV-specific nucleic acid in human blood and contain infectious virus. In contrast to most other reports of putative human isolates, where sequences are virtually identical to those of the established laboratory strains, this isolate shows divergence in the region previously defined as variable in BDV from naturally infected animals.

Antibodies, Viral↗

Molecular and biological characterization of deformed wing virus of honeybees (Apis mellifera L.).

Deformed wing virus (DWV) of honeybees (Apis mellifera) is closely associated with characteristic wing deformities, abdominal bloating, paralysis, and rapid mortality of emerging adult bees. The virus was purified from diseased insects, and its genome was cloned and sequenced. The genomic RNA of DWV is 10,140 nucleotides in length and contains a single large open reading frame encoding a 328-kDa polyprotein. The coding sequence is flanked by a 1,144-nucleotide 5' nontranslated leader sequence and a 317-nucleotide 3' nontranslated region, followed by a poly(A) tail. The three major structural proteins, VP1 (44 kDa), VP2 (32 kDa), and VP3 (28 kDa), were identified, and their genes were mapped to the N-terminal section of the polyprotein. The C-terminal part of the polyprotein contains sequence motifs typical of well-characterized picornavirus nonstructural proteins: an RNA helicase, a chymotrypsin-like 3C protease, and an RNA-dependent RNA polymerase. The genome organization, capsid morphology, and sequence comparison data indicate that DWV is a member of the recently established genus Iflavirus.

Amino Acid Sequence↗

CFTR transcripts are undetectable in lymphocytes and respiratory epithelial cells of a CF patient homozygous for the nonsense mutation R553X.

In order to analyse the influence of the nonsense mutation R553X on CFTR gene expression, transcripts from epithelial cells and lymphocytes were examined from nine subjects (one CF patient homozygous for R553X, one CF patient compound heterozygous for R553X/delta F508, four CF carriers heterozygous for R553X, one CF carrier with the genotype delta F508/N, and two uncharacterized normal adults). After reverse transcription of the region from exons 10 to 13 to cDNA, fragments of the expected size were amplified from all heterozygous and normal subjects. In three subjects an additional alternatively spliced product was observed, which was found to contain a termination codon. In repeated experiments it was not possible to detect any CFTR mRNA in cells derived from the R553X homozygous patient. Furthermore, in subjects heterozygous for R553X we could not detect by hybridisation with a specific oligonucleotide probe and direct sequencing any CFTR mRNA derived from the R553X allele. However, the wild type product was present in all of these subjects. Our results support the view that nonsense mutations in the CFTR gene can lead to a reduction or absence of cytoplasmic CFTR mRNA.

Adolescent↗

Data mining tools for biological sequences.

We describe a methodology, as well as some related data mining tools, for analyzing sequence data. The methodology comprises three steps: (a) generating candidate features from the sequences, (b) selecting relevant features from the candidates, and (c) integrating the selected features to build a system to recognize specific properties in sequence data. We also give relevant techniques for each of these three steps. For generating candidate features, we present various types of features based on the idea of k-grams. For selecting relevant features, we discuss signal-to-noise, t-statistics, and entropy measures, as well as a correlation-based feature selection method. For integrating selected features, we use machine learning methods, including C4.5, SVM, and Naive Bayes. We illustrate this methodology on the problem of recognizing translation initiation sites. We discuss how to generate and select features that are useful for understanding the distinction between ATG sites that are translation initiation sites and those that are not. We also discuss how to use such features to build reliable systems for recognizing translation initiation sites in DNA sequences.

Artificial Intelligence↗

Local sequence-structure motifs in RNA.

Ribonuclic acid (RNA) enjoys increasing interest in molecular biology; despite this interest fundamental algorithms are lacking, e.g. for identifying local motifs. As proteins, RNA molecules have a distinctive structure. Therefore, in addition to sequence information, structure plays an important part in assessing the similarity of RNAs. Furthermore, common sequence-structure features in two or several RNA molecules are often only spatially local, where possibly large parts of the molecules are dissimilar. Consequently, we address the problem of comparing RNA molecules by computing an optimal local alignment with respect to sequence and structure information. While local alignment is superior to global alignment for identifying local similarities, no general local sequence-structure alignment algorithms are currently known. We suggest a new general definition of locality for sequence-structure alignments that is biologically motivated and efficiently tractable. To show the former, we discuss locality of RNA and prove that the defined locality means connectivity by atomic and non-atomic bonds. To show the latter, we present an efficient algorithm for the newly defined pairwise local sequence-structure alignment (lssa) problem for RNA. For molecules of lengthes n and m, the algorithm has worst-case time complexity of O(n2 x m2 x max(n,m)) and a space complexity of only O(n x m). An implementation of our algorithm is available at http://www.bio.inf.uni-jena.de. Its runtime is competitive with global sequence-structure alignment.

Algorithms↗

Complete nucleotide sequence and the genome organization of Patchouli mild mosaic virus RNA1.

The nucleotide sequence of the RNA1 of Patchouli mild mosaic virus (PatMMV) has been determined. It contains 5,957 nucleotides excluding the 3'-terminal poly(A) tail and contains a single long open reading frame (ORF) of 5,613 nucleotides extending from nucleotide 235 to 5847. The predicted polyprotein encoded by the long ORF is 1,870 amino acids in length with a molecular weight of 210 kD. The conserved residues of RNA-dependent RNA polymerase, cysteine protease, purine NTP-binding domain and a cofactor for viral protease are present in a 210-kD polyprotein. As PatMMV RNA showed high sequence identity (81-97%) with BBWV-2 RNA, PaMMV may be one strain of BBWV-2.

5' Untranslated Regions↗

Evolutionary models for insertions and deletions in a probabilistic modeling framework.

BACKGROUND: Probabilistic models for sequence comparison (such as hidden Markov models and pair hidden Markov models for proteins and mRNAs, or their context-free grammar counterparts for structural RNAs) often assume a fixed degree of divergence. Ideally we would like these models to be conditional on evolutionary divergence time. Probabilistic models of substitution events are well established, but there has not been a completely satisfactory theoretical framework for modeling insertion and deletion events. RESULTS: I have developed a method for extending standard Markov substitution models to include gap characters, and another method for the evolution of state transition probabilities in a probabilistic model. These methods use instantaneous rate matrices in a way that is more general than those used for substitution processes, and are sufficient to provide time-dependent models for standard linear and affine gap penalties, respectively. Given a probabilistic model, we can make all of its emission probabilities (including gap characters) and all its transition probabilities conditional on a chosen divergence time. To do this, we only need to know the parameters of the model at one particular divergence time instance, as well as the parameters of the model at the two extremes of zero and infinite divergence. I have implemented these methods in a new generation of the RNA genefinder QRNA (eQRNA). CONCLUSION: These methods can be applied to incorporate evolutionary models of insertions and deletions into any hidden Markov model or stochastic context-free grammar, in a pair or profile form, for sequence modeling.

Algorithms↗

Identification of consensus RNA secondary structures using suffix arrays.

BACKGROUND: The identification of a consensus RNA motif often consists in finding a conserved secondary structure with minimum free energy in an ensemble of aligned sequences. However, an alignment is often difficult to obtain without prior structural information. Thus the need for tools to automate this process. RESULTS: We present an algorithm called Seed to identify all the conserved RNA secondary structure motifs in a set of unaligned sequences. The search space is defined as the set of all the secondary structure motifs inducible from a seed sequence. A general-to-specific search allows finding all the motifs that are conserved. Suffix arrays are used to enumerate efficiently all the biological palindromes as well as for the matching of RNA secondary structure expressions. We assessed the ability of this approach to uncover known structures using four datasets. The enumeration of the motifs relies only on the secondary structure definition and conservation only, therefore allowing for the independent evaluation of scoring schemes. Twelve simple objective functions based on free energy were evaluated for their potential to discriminate native folds from the rest. CONCLUSION: Our evaluation shows that 1) support and exclusion constraints are sufficient to make an exhaustive search of the secondary structure space feasible. 2) The search space induced from a seed sequence contains known motifs. 3) Simple objective functions, consisting of a combination of the free energy of matching sequences, can generally identify motifs with high positive predictive value and sensitivity to known motifs.

Algorithms↗

Impact of RNA structure on the prediction of donor and acceptor splice sites.

BACKGROUND: gene identification in genomic DNA sequences by computational methods has become an important task in bioinformatics and computational gene prediction tools are now essential components of every genome sequencing project. Prediction of splice sites is a key step of all gene structural prediction algorithms. RESULTS: we sought the role of mRNA secondary structures and their information contents for five vertebrate and plant splice site datasets. We selected 900-nucleotide sequences centered at each (real or decoy) donor and acceptor sites, and predicted their corresponding RNA structures by Vienna software. Then, based on whether the nucleotide is in a stem or not, the conventional four-letter nucleotide alphabet was translated into an eight-letter alphabet. Zero-, first- and second-order Markov models were selected as the signal detection methods. It is shown that applying the eight-letter alphabet compared to the four-letter alphabet considerably increases the accuracy of both donor and acceptor site predictions in case of higher order Markov models. CONCLUSION: Our results imply that RNA structure contains important data and future gene prediction programs can take advantage of such information.

Algorithms↗

Efficient pairwise RNA structure prediction and alignment using sequence alignment constraints.

BACKGROUND: We are interested in the problem of predicting secondary structure for small sets of homologous RNAs, by incorporating limited comparative sequence information into an RNA folding model. The Sankoff algorithm for simultaneous RNA folding and alignment is a basis for approaches to this problem. There are two open problems in applying a Sankoff algorithm: development of a good unified scoring system for alignment and folding and development of practical heuristics for dealing with the computational complexity of the algorithm. RESULTS: We use probabilistic models (pair stochastic context-free grammars, pairSCFGs) as a unifying framework for scoring pairwise alignment and folding. A constrained version of the pairSCFG structural alignment algorithm was developed which assumes knowledge of a few confidently aligned positions (pins). These pins are selected based on the posterior probabilities of a probabilistic pairwise sequence alignment. CONCLUSION: Pairwise RNA structural alignment improves on structure prediction accuracy relative to single sequence folding. Constraining on alignment is a straightforward method of reducing the runtime and memory requirements of the algorithm. Five practical implementations of the pairwise Sankoff algorithm - this work (Consan), David Mathews' Dynalign, Ian Holmes' Stemloc, Ivo Hofacker's PMcomp, and Jan Gorodkin's FOLDALIGN - have comparable overall performance with different strengths and weaknesses.

Algorithms↗

Prediction and verification of microRNA targets by MovingTargets, a highly adaptable prediction method.

BACKGROUND: MicroRNAs (miRNAs) mediate a form of translational regulation in animals. Hundreds of animal miRNAs have been identified, but only a few of their targets are known. Prediction of miRNA targets for translational regulation is challenging, since the interaction with the target mRNA usually occurs via incomplete and interrupted base pairing. Moreover, the rules that govern such interactions are incompletely defined. RESULTS: MovingTargets is a software program that allows a researcher to predict a set of miRNA targets that satisfy an adjustable set of biological constraints. We used MovingTargets to identify a high-likelihood set of 83 miRNA targets in Drosophila, all of which adhere to strict biological constraints. We tested and verified 3 of these predictions in cultured cells, including a target for the Drosophila let-7 homolog. In addition, we utilized the flexibility of MovingTargets by relaxing the biological constraints to identify and validate miRNAs targeting tramtrack, a gene also known to be subject to translational control dependent on the RNA binding protein Musashi. CONCLUSION: MovingTargets is a flexible tool for the accurate prediction of miRNA targets in Drosophila. MovingTargets can be used to conduct a genome-wide search of miRNA targets using all Drosophila miRNAs and potential targets, or it can be used to conduct a focused search for miRNAs targeting a specific gene. In addition, the values for a set of biological constraints used to define a miRNA target are adjustable, allowing the software to incorporate the rules used to characterize a miRNA target as these rules are experimentally determined and interpreted.

3' Untranslated Regions↗

Prediction and identification of Arabidopsis thaliana microRNAs and their mRNA targets.

BACKGROUND: A class of eukaryotic non-coding RNAs termed microRNAs (miRNAs) interact with target mRNAs by sequence complementarity to regulate their expression. The low abundance of some miRNAs and their time- and tissue-specific expression patterns make experimental miRNA identification difficult. We present here a computational method for genome-wide prediction of Arabidopsis thaliana microRNAs and their target mRNAs. This method uses characteristic features of known plant miRNAs as criteria to search for miRNAs conserved between Arabidopsis and Oryza sativa. Extensive sequence complementarity between miRNAs and their target mRNAs is used to predict miRNA-regulated Arabidopsis transcripts. RESULTS: Our prediction covered 63% of known Arabidopsis miRNAs and identified 83 new miRNAs. Evidence for the expression of 25 predicted miRNAs came from northern blots, their presence in the Arabidopsis Small RNA Project database, and massively parallel signature sequencing (MPSS) data. Putative targets functionally conserved between Arabidopsis and O. sativa were identified for most newly identified miRNAs. Independent microarray data showed that the expression levels of some mRNA targets anti-correlated with the accumulation pattern of their corresponding regulatory miRNAs. The cleavage of three target mRNAs by miRNA binding was validated in 5' RACE experiments. CONCLUSIONS: We identified new plant miRNAs conserved between Arabidopsis and O. sativa and report a wide range of transcripts as potential miRNA targets. Because MPSS data are generated from polyadenylated RNA molecules, our results suggest that at least some miRNA precursors are polyadenylated at certain stages. The broad range of putative miRNA targets indicates that miRNAs participate in the regulation of a variety of biological processes.

Arabidopsis↗

A computational investigation of kinetoplastid trans-splicing.

Trans-splicing is an unusual process in which two separate RNA strands are spliced together to yield a mature mRNA. We present a novel computational approach which has an overall accuracy of 82% and can predict 92% of known trans-splicing sites. We have applied our method to chromosomes 1 and 3 of Leishmania major, with high-confidence predictions for 85% and 88% of annotated genes respectively. We suggest some extensions of our method to other systems.

Animals↗