Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “RNA sequencing analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

seq2ribo: structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context. AVAILABILITY: seq2ribo is available at https://github.com/Kingsford-Group/seq2ribo.

Machine Learning↗

Discovering well-ordered folding patterns in nucleotide sequences.

MOTIVATION: Growing evidence demonstrates that local well-ordered structures are closely correlated with cis-acting elements in the post-transcriptional regulation of gene expression. The prediction of a well-ordered folding sequence (WFS) in genomic sequences is very helpful in the determination of local RNA elements with structure-dependent functions in mRNAs. RESULTS: In this study, the quality of local WFS is assessed by the energy difference (E(diff)) between the free energies of the global minimal structure folded in the segment and its corresponding optimal restrained structure (ORS). The ORS is an optimal structure under the condition in which none of the base-pairs in the global minimal structure is allowed to form. Those WFSs in HIV-1 mRNA, various ferritin mRNAs and genomic sequences containing let-7 RNA gene were searched by a novel method, ed_scan. Our results indicate that the detected WFSs are coincident with known Rev response element in HIV-1 mRNA, iron-responsive elements in ferritin mRNAs and small let-7 RNAs in Caenorhabditis elegans, Caenorhabditis briggsae and Drosophila melanogaster genomic sequences. Statistical significance of the WFS is addressed by a quantitative measure Zscr(e) that is a z-score of E(diff) and extensive random simulations. We suggest that WFSs with high statistical significance have structural roles involving their sequence information. AVAILABILITY: The source code of ed_scan is available via anonymous ftp as ftp://ftp.ncifcrf.gov/pub/users/shuyun/scan/ed_scan.tar.

Algorithms↗

MicroRNA identification based on sequence and structure alignment.

MOTIVATION: MicroRNAs (miRNA) are approximately 22 nt long non-coding RNAs that are derived from larger hairpin RNA precursors and play important regulatory roles in both animals and plants. The short length of the miRNA sequences and relatively low conservation of pre-miRNA sequences restrict the conventional sequence-alignment-based methods to finding only relatively close homologs. On the other hand, it has been reported that miRNA genes are more conserved in the secondary structure rather than in primary sequences. Therefore, secondary structural features should be more fully exploited in the homologue search for new miRNA genes. RESULTS: In this paper, we present a novel genome-wide computational approach to detect miRNAs in animals based on both sequence and structure alignment. Experiments show this approach has higher sensitivity and comparable specificity than other reported homologue searching methods. We applied this method on Anopheles gambiae and detected 59 new miRNA genes. AVAILABILITY: This program is available at http://bioinfo.au.tsinghua.edu.cn/miralign. SUPPLEMENTARY INFORMATION: Supplementary information is available at http://bioinfo.au.tsinghua.edu.cn/miralign/supplementary.htm.

Algorithms↗

Inferring global levels of alternative splicing isoforms using a generative model of microarray data.

MOTIVATION: Alternative splicing (AS) is a frequent step in metozoan gene expression whereby the exons of genes are spliced in different combinations to generate multiple isoforms of mature mRNA. AS functions to enrich an organism's proteomic complexity and regulates gene expression. Despite its importance, the mechanisms underlying AS and its regulation are not well understood, especially in the context of global gene expression patterns. We present here an algorithm referred to as the Generative model for the Alternative Splicing Array Platform (GenASAP) that can predict the levels of AS for thousands of exon skipping events using data generated from custom microarrays. GenASAP uses Bayesian learning in an unsupervised probability model to accurately predict AS levels from the microarray data. GenASAP is capable of learning the hybridization profiles of microarray data, while modeling noise processes and missing or aberrant data. GenASAP has been successfully applied to the global discovery and analysis of AS in mammalian cells and tissues. RESULTS: GenASAP was applied to data obtained from a custom microarray designed for the monitoring of 3126 AS events in mouse cells and tissues. The microarray design included probes specific for exon body and junction sequences formed by the splicing of exons. Our results show that GenASAP provides accurate predictions for over one-third of the total events, as verified by independent RT-PCR assays. SUPPLEMENTARY INFORMATION: http://www.psi.toronto.edu/GenASAP.

Algorithms↗

A stimulatory RNA associated with RecBCD enzyme.

RecBCD enzyme acts in the major pathway of homologous recombination of linear DNA in Escherichia coli. The enzyme unwinds DNA and is an ATP-dependent double-strand and single-strand exonuclease and a single-strand endonuclease; it acts at Chi recombination hotspots (5'-GCTGGTGG-3') to produce a recombinogenic single-stranded DNA 3'-end. We found that a small RNA with a unique sequence of approximately 24 nt was tightly bound to RecBCD enzyme and co-purified with it. When added to native enzyme this RNA, but not four others, increased DNA unwinding and Chi nicking activities of the enzyme. In seven similarly active enzyme preparations the molar ratio of RNA molecules to RecBCD enzyme molecules ranged from 0.2 to <0.008. These results suggest that, although this unique RNA is not an essential enzyme subunit, it has a biological role in stimulating RecBCD enzyme activity.

DNA Helicases↗

Pfold: RNA secondary structure prediction using stochastic context-free grammars.

RNA secondary structures are important in many biological processes and efficient structure prediction can give vital directions for experimental investigations. Many available programs for RNA secondary structure prediction only use a single sequence at a time. This may be sufficient in some applications, but often it is possible to obtain related RNA sequences with conserved secondary structure. These should be included in structural analyses to give improved results. This work presents a practical way of predicting RNA secondary structure that is especially useful when related sequences can be obtained. The method improves a previous algorithm based on an explicit evolutionary model and a probabilistic model of structures. Predictions can be done on a web server at http://www.daimi.au.dk/~compbio/pfold.

Algorithms↗

Riboswitch finder--a tool for identification of riboswitch RNAs.

We describe a dedicated RNA motif search program and web server to identify RNA riboswitches. The Riboswitch finder analyses a given sequence using the web interface, checks specific sequence elements and secondary structure, calculates and displays the energy folding of the RNA structure and runs a number of tests including this information to determine whether high-sensitivity riboswitch motifs (or variants) according to the Bacillus subtilis type are present in the given RNA sequence. Batch-mode determination (all sequences input at once and separated by FASTA format) is also possible. The program has been implemented and is available both as local software for in-house installation and as a web server at http://www.biozentrum.uni-wuerzburg.de/bioinformatik/Riboswitch/.

Allosteric Regulation↗

Prediction of CsrA-regulating small RNAs in bacteria and their experimental verification in Vibrio fischeri.

The role of small RNAs as critical components of global regulatory networks has been highlighted by several recent studies. An important class of such small RNAs is represented by CsrB and CsrC of Escherichia coli, which control the activity of the global regulator CsrA. Given the critical role played by CsrA in several bacterial species, an important problem is the identification of CsrA-regulating small RNAs. In this paper, we develop a computer program (CSRNA_FIND) designed to locate potential CsrA-regulating small RNAs in bacteria. Using CSRNA_FIND to search the genomes of bacteria having homologs of CsrA, we identify all the experimentally known CsrA-regulating small RNAs and also make predictions for several novel small RNAs. We have verified experimentally our predictions for two CsrA-regulating small RNAs in Vibrio fischeri. As more genomes are sequenced, CSRNA_FIND can be used to locate the corresponding small RNAs that regulate CsrA homologs. This work thus opens up several avenues of research in understanding the mode of CsrA regulation through small RNAs in bacteria.

Aliivibrio fischeri↗

The terminal balls characteristic of eukaryotic rRNA transcription units in chromatin spreads are rRNA processing complexes.

When spread chromatin is visualized by electron microscopy, active rRNA genes have a characteristic Christmas tree appearance: From a DNA "trunk" extend closely packed "branches" of nascent transcripts whose ends are decorated with terminal "balls." These terminal balls have been known for more than two decades, are shown in most biology textbooks, and are reported in hundreds of papers, yet their nature has remained elusive. Here, we show that a rRNA-processing signal in the 5'-external transcribed spacer (ETS) of the Xenopus laevis ribosomal primary transcript forms a large, processing-related complex with factors of the Xenopus oocyte, analogous to 5' ETS processing complexes found in other vertebrate cell types. Using mutant rRNA genes, we find that the same rRNA residues are required for this biochemically defined complex formation and for terminal ball formation, analyzed electron microscopically after injection of these cloned genes into Xenopus oocytes. This, plus other presented evidence, implies that rRNA terminal balls in Xenopus, and by inference, also in the multitude of other species where they have been observed, are the ultrastructural visualization of an evolutionarily conserved 5' ETS processing complex that forms on the nascent rRNA.

Animals↗

23S rRNA domain V, a fragment that can be specifically methylated in vitro by the ErmSF (TlrA) methyltransferase.

The DNA sequence that encodes 23S rRNA domain V of Bacillus subtilis, nucleotides 2036 to 2672 (C. J. Green, G. C. Stewart, M. A. Hollis, B. S. Vold, and K. F. Bott, Gene 37:261-266, 1985), was cloned and used as a template from which to transcribe defined domain V RNA in vitro. The RNA transcripts served as a substrate in vitro for specific methylation of B. subtilis adenine 2085 (adenine 2058 in Escherichia coli 23S rRNA) by the ErmSF methyltransferase, an enzyme that confers resistance to the macrolide-lincosamide-streptogramin B group of antibiotics on Streptomyces fradiae NRRL 2702, the host from which it was cloned. Thus, neither RNA sequences belonging to domains other than V nor the association of 23S rRNA with ribosomal proteins is needed for the specific methylation of adenine that confers resistance to the macrolide-lincosamide-streptogramin B group of antibiotics.

Adenine↗

Development of a cDNA array for chicken gene expression analysis.

BACKGROUND: The application of microarray technology to functional genomic analysis in the chicken has been limited by the lack of arrays containing large numbers of genes. RESULTS: We have produced cDNA arrays using chicken EST collections generated by BBSRC, University of Delaware and the Fred Hutchinson Cancer Research Center. From a total of 363,838 chicken ESTs representing 24 different adult or embryonic tissues, a set of 11,447 non-redundant ESTs were selected and added to an existing collection of clones (4,162) from immune tissues and a chicken bursal cell line (DT40). Quality control analysis indicates there are 13,007 useable features on the array, including 160 control spots. The array provides broad coverage of mRNAs expressed in many tissues; in addition, clones with expression unique to various tissues can be detected. CONCLUSIONS: A chicken multi-tissue cDNA microarray with 13,007 features is now available to academic researchers from http://genomics@fhcrc.org. Sequence information for all features on the array is in GenBank, and clones can be readily obtained. Targeted users include researchers in comparative and developmental biology, immunology, vaccine and agricultural technology. These arrays will be an important resource for the entire research community using the chicken as a model.

Animals↗

Beyond the proteome: non-coding regulatory RNAs.

A variety of RNA molecules have been found over the last 20 years to have a remarkable range of functions beyond the well-known roles of messenger, ribosomal and transfer RNAs. Here, we present a general categorization of all non-coding RNAs and briefly discuss the ones that affect transcription, translation and protein function.

Animals↗

Long-distance RNA-RNA interactions between terminal elements and the same subset of internal elements on the potato virus X genome mediate minus- and plus-strand RNA synthesis.

Potexvirus genomes contain conserved terminal elements that are complementary to multiple internal octanucleotide elements. Both local sequences and structures at the 5' terminus and long-distance interactions between this region and internal elements are important for accumulation of potato virus X (PVX) plus-strand RNA in vivo. In this study, the role of the conserved hexanucleotide motif within SL3 of the 3' NTR and internal conserved octanucleotide elements in minus-strand RNA synthesis was analyzed using both a template-dependent, PVX RNA-dependent RNA polymerase (RdRp) extract and a protoplast replication system. Template analyses in vitro indicated that 3' terminal templates of 850 nucleotides (nt), but not 200 nt, supported efficient, minus-strand RNA synthesis. Mutational analyses of the longer templates indicated that optimal transcription requires the hexanucleotide motif in SL3 within the 3' NTR and the complementary CP octanucleotide element 747 nt upstream. Additional experiments to disrupt interactions between one or more internal conserved elements and the 3' hexanucleotide element showed that long-distance interactions were necessary for minus-strand RNA synthesis both in vitro and in vivo. Additionally, multiple internal octanucleotide elements could serve as pairing partners with the hexanucleotide element in vivo. These cis-acting elements and interactions correlate in several ways to those previously observed for plus-strand RNA accumulation in vivo, suggesting that dynamic interactions between elements at both termini and the same subset of internal octanucleotide elements are required for both minus- and plus-strand RNA synthesis and potentially other aspects of PVX replication.

Base Sequence↗

Characterization of two distinct RNA domains that regulate translation of the Drosophila gypsy retroelement.

The genomic RNA of the gypsy retroelement from Drosophila melanogaster exhibits features similar to other retroviral RNAs because its 5' untranslated (5' UTR) region is unusually long (846 nucleotides) and potentially highly structured. Our initial aim was to search for an internal ribosome entry site (IRES) element in the 5' UTR of the gypsy genomic RNA by using various monocistronic and bicistronic RNAs in the rabbit reticulocyte lysate (RRL) system and in cultured cells. Results reported here show that two functionally distinct and independent RNA domains control the production of gypsy encoded proteins. The first domain corresponds to the 5' UTR of the env subgenomic RNA and exhibits features of an efficient IRES (IRES(E)) both in the reticulocyte lysate and in cells. The second RNA domain that encompasses the gypsy insulator can function as an IRES in the rabbit reticulocyte lysate but strongly represses translation in cultured cells. Taken together, these results suggest that expression of the gypsy encoded proteins from the genomic and subgenomic RNAs can be regulated at the level of translation.

5' Untranslated Regions↗

Transcriptome analysis of mouse stem cells and early embryos.

Understanding and harnessing cellular potency are fundamental in biology and are also critical to the future therapeutic use of stem cells. Transcriptome analysis of these pluripotent cells is a first step towards such goals. Starting with sources that include oocytes, blastocysts, and embryonic and adult stem cells, we obtained 249,200 high-quality EST sequences and clustered them with public sequences to produce an index of approximately 30,000 total mouse genes that includes 977 previously unidentified genes. Analysis of gene expression levels by EST frequency identifies genes that characterize preimplantation embryos, embryonic stem cells, and adult stem cells, thus providing potential markers as well as clues to the functional features of these cells. Principal component analysis identified a set of 88 genes whose average expression levels decrease from oocytes to blastocysts, stem cells, postimplantation embryos, and finally to newborn tissues. This can be a first step towards a possible definition of a molecular scale of cellular potency. The sequences and cDNA clones recovered in this work provide a comprehensive resource for genes functioning in early mouse embryos and stem cells. The nonrestricted community access to the resource can accelerate a wide range of research, particularly in reproductive and regenerative medicine.

Animals↗

Pre-mRNA secondary structure prediction aids splice site prediction.

Accurate splice site prediction is a critical component of any computational approach to gene prediction in higher organisms. Existing approaches generally use sequence-based models that capture local dependencies among nucleotides in a small window around the splice site. We present evidence that computationally predicted secondary structure of moderate length pre-mRNA subsequencies contains information that can be exploited to improve acceptor splice site prediction beyond that possible with conventional sequence-based approaches. Both decision tree and support vector machine classifiers, using folding energy and structure metrics characterizing helix formation near the splice site, achieve a 5-10% reduction in error rate with a human data set. Based on our data, we hypothesize that acceptors preferentially exhibit short helices at the splice site.

Computer Simulation↗

STR2: a structure to string approach for locating G-box riboswitch shapes in pre-selected genes.

Traditional sequence-based search methods such as BLAST and FASTA can be used to identify sequence similarities. Recently, there is a growing interest in performing RNA shape similarity searches inside selected genes to locate RNA structure motifs that are known to possess functionally important roles. For example, in the newly discovered RNA genetic control elements called "riboswitches", the box domain is known to be highly conserved among various bacterial species in both its nucleotide composition and shape. However, in non-bacterial species, shape conservation is likely to become more important than sequence conservation when searching for riboswitch patterns. For this purpose, we present an approach tailored for detecting RNA shape similarities. We extend the Structure to String (ST R2) method that was initially proposed to locate shape similarities in proteins to identify predicted secondary structures of RNAs. The ST R2 for RNAs is a translation of a secondary structure to a string of characters, after which known sequence-based search algorithms with an efficient implementation are being used. We validate that the ST R2 succeeds to locate G-box riboswitches in prokaryotes, as expected. Subsequently we show running examples when attempting to detect G-box riboswitch candidates in eukaryotes.

Algorithms↗