Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “RNA sequencing analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

In vivo dissection of cis-acting determinants for plastid RNA editing.

Substitutional RNA editing changes single C nucleotides in higher plant chloroplast transcripts into U residues. To determine the cis-acting sequence elements involved in plastid RNA editing, we constructed a series of chloroplast transformation vectors harboring selected editing sites of the tobacco ndhB transcript in a chimeric context. The constructs were inserted into the tobacco plastid genome by biolistic transformation leading to the production of stable chimeric RNAs. Analysis of RNA editing revealed unexpected differences in the size of the essential cis elements or in their distance from the editing site. Flanking sequences of identical size direct virtually complete editing for one pair of editing sites, partial editing for a second and no editing at all for a third pair of sites. Serial 5' and 3' deletions allowed us to define the cis-acting elements more precisely and to identify a sequence element essential for editing site recognition. In addition, a single nucleotide substitution immediately upstream of an editing position was introduced. This mutation was found drastically and selectively to reduce the editing efficiency of the downstream editing site, demonstrating that position -1 is important for either site recognition or catalysis. Our results indicate that the editing of adjacent sites is likely to be mechanistically coupled. In no case did the presence in the plastome of the additional editing sites have any effect on the editing efficiency of the endogenous ndhB sites, indicating that the availability of site-specific trans-acting factors is not rate limiting.

Base Sequence↗

RNAmute: RNA secondary structure mutation analysis tool.

BACKGROUND: RNAMute is an interactive Java application that calculates the secondary structure of all single point mutations, given an RNA sequence, and organizes them into categories according to their similarity with respect to the wild type predicted structure. The secondary structure predictions are performed using the Vienna RNA package. Several alternatives are used for the categorization of single point mutations: Vienna's RNAdistance based on dot-bracket representation, as well as tree edit distance and second eigenvalue of the Laplacian matrix based on Shapiro's coarse grain tree graph representation. RESULTS: Selecting a category in each one of the processed tables lists all single point mutations belonging to that category. Selecting a mutation displays a graphical drawing of the single point mutation and the wild type, and includes basic information such as associated energies, representations and distances. RNAMute can be used successfully with very little previous experience and without choosing any parameter value alongside the initial RNA sequence. The package runs under LINUX operating system. CONCLUSION: RNAMute is a user friendly tool that can be used to predict single point mutations leading to conformational rearrangements in the secondary structure of RNAs. In several cases of substantial interest, notably in virology, a point mutation may lead to a loss of important functionality such as the RNA virus replication and translation initiation because of a conformational rearrangement in the secondary structure.

Animals↗

Structure, chromosomal localization, and expression of the Drosophila melanogaster gene encoding sepiapterin reductase.

We have isolated and characterized a Drosophila melanogaster gene encoding the sepiapterin reductase (SR). The gene does not have introns. The 5'- and 3'-RACE analysis, which determined the transcription start point (tsp) and polyadenylation site, respectively, showed that the gene produces single mRNA species. The potential promoter region lacks distinct TATAAA or CCAAT box consensus sequences. RNA blot analysis revealed that the gene encodes a 1.4kb transcript that could be detected throughout development and in both heads and bodies of adults. The Drosophila SR gene maps to 15A on the X chromosome.

Alcohol Oxidoreductases↗

Clustering of RNA secondary structures with application to messenger RNAs.

There is growing evidence of translational gene regulation at the mRNA level, and of the important roles of RNA secondary structure in these regulatory processes. Because mRNAs likely exist in a population of structures, the popular free energy minimization approach may not be well suited to prediction of mRNA structures in studies of post-transcriptional regulation. Here, we describe an alternative procedure for the characterization of mRNA structures, in which structures sampled from the Boltzmann-weighted ensemble of RNA secondary structures are clustered. Based on a random sample of full-length human mRNAs, we find that the minimum free energy (MFE) structure often poorly represents the Boltzmann ensemble, that the ensemble often contains multiple structural clusters, and that the centroids of a small number of structural clusters more effectively characterize the ensemble. We show that cluster-level characteristics and statistics are statistically reproducible. In a comparison between mRNAs and structural RNAs, similarity is observed for the number of clusters and the energy gap between the MFE structure and the sampled ensemble. However, for structural RNAs, there are more high-frequency base-pairs in both the Boltzmann ensemble and the clusters, and the clusters are more compact. The clustering features have been incorporated into the Sfold software package for nucleic acid folding and design.

Base Pairing↗

Transfer of Methanolobus siciliae to the genus Methanosarcina, naming it Methanosarcina siciliae, and emendation of the genus Methanosarcina.

A sequence analysis of the 16S rRNA of Methanolobus siciliae T4/M(T) (T = type strain) showed that this strain is closely related to members of the genus Methanosarcina, especially Methanosarcina acetivorans C2A(T). Methanolobus siciliae T4/M(T) and HI350 were morphologically more similar to members of the genus Methanosarcina than to members of the genus Methanolobus in that they both formed massive cell aggregates with pseudosarcinae. Thus, we propose that Methanolobus siciliae should be transferred to the genus Methanosarcina as Methanosarcina siciliae.

Methanosarcina↗

SPA: a probabilistic algorithm for spliced alignment.

Recent large-scale cDNA sequencing efforts show that elaborate patterns of splice variation are responsible for much of the proteome diversity in higher eukaryotes. To obtain an accurate account of the repertoire of splice variants, and to gain insight into the mechanisms of alternative splicing, it is essential that cDNAs are very accurately mapped to their respective genomes. Currently available algorithms for cDNA-to-genome alignment do not reach the necessary level of accuracy because they use ad hoc scoring models that cannot correctly trade off the likelihoods of various sequencing errors against the probabilities of different gene structures. Here we develop a Bayesian probabilistic approach to cDNA-to-genome alignment. Gene structures are assigned prior probabilities based on the lengths of their introns and exons, and based on the sequences at their splice boundaries. A likelihood model for sequencing errors takes into account the rates at which misincorporation, as well as insertions and deletions of different lengths, occurs during sequencing. The parameters of both the prior and likelihood model can be automatically estimated from a set of cDNAs, thus enabling our method to adapt itself to different organisms and experimental procedures. We implemented our method in a fast cDNA-to-genome alignment program, SPA, and applied it to the FANTOM3 dataset of over 100,000 full-length mouse cDNAs and a dataset of over 20,000 full-length human cDNAs. Comparison with the results of four other mapping programs shows that SPA produces alignments of significantly higher quality. In particular, the quality of the SPA alignments near splice boundaries and SPA's mapping of the 5' and 3' ends of the cDNAs are highly improved, allowing for more accurate identification of transcript starts and ends, and accurate identification of subtle splice variations. Finally, our splice boundary analysis on the human dataset suggests the existence of a novel non-canonical splice site that we also find in the mouse dataset. The SPA software package is available at http://www.biozentrum.unibas.ch/personal/nimwegen/cgi-bin/spa.cgi.

Algorithms↗

Approaches to sequence analysis of 125I-labeled RNA.

A method is described for the initial steps of sequence analysis of RNase T1-and pancreatic RN-ase-resistant oligonucleotides of RNA containing cytidylate residues labeled in vitro with 125I. In many cases an oligonucleotide sequence can be deduced from a consideration of (i) its relative position in the two-dimensional fingerprint (with DEAE thin layer homochromatographic second dimension), (ii) its electrophoretic mobility on DEAE paper at pH 1.9, and (iii) identification of its products of further enzymatic digestion by comparison with a set of marker oligonucleotides. Additional methods including analysis of oligonucleotides following chemical blocking of uridylate residues with CMCT and analysis of products of incomplete enzymatic digestion are also discussed.

Base Sequence↗

Further evidence that the failure to cleave the aminopropeptide of type I procollagen is the cause of Ehlers-Danlos syndrome type VII.

Dermal fibroblasts from a Chinese Ehlers-Danlos syndrome type VII patient synthesized approximately equal amounts of normal pro-alpha 2(I) chains of type I procollagen and abnormal ones with electrophoretic mobility of pN alpha 2(I) chains, in which the amino-propeptide (N-propeptide) was retained. Reverse-transcriptase PCR analysis of the proband's RNA showed outsplicing of the 54 base exon 6 in half of the pro-alpha 2(I) mRNAs. Exon 6 encodes 18 amino acids of the N-telopeptide which contains the procollagen N-proteinase cleavage site and a cross-link precursor lysine. Loss of these sequences would result in failure to cleave the amino-propeptide of pro-alpha 2(I) and the accumulation of pN-alpha 2(I) chains. Nucleotide sequencing analyses of the proband's COL1A2 gene showed the presence of a T to C transition at position +2 of intron 6 in one allele and the proband is heterozygous for the defect. This mutation which destroyed the consensus GT dinucleotide at the 5' splice donor site of the intron is responsible for the loss of exon 6 by exon skipping. Electron microscopic analysis of the patient's dermis showed the presence of abnormal collagen I fibrils of irregular diameter and circularity. This mutation in COL1A2 in an EDS VII patient is the first reported case in the Chinese population and is identical to one reported for another EDS-VII (Libyan) patient. The occurrence of an identical mutation in two probands of different ethnic origin is direct evidence that the mutant genotype is the cause of the EDS VII phenotype.(ABSTRACT TRUNCATED AT 250 WORDS)

Base Sequence↗

Partial characterization of hepatitis A viruses from three intermediate passage levels of a series resulting in adaptation to growth in cell culture and attenuation of virulence.

Viruses from passages 9/10, 21, and 32 of a serially passaged human isolate of hepatitis A virus, strain HM-175, were partially sequenced and compared for their abilities to grow in cell cultures and to cause disease in primates. Viruses from all passages grew more efficiently in cell culture than did the wild-type virulent virus from which they were derived, and all displayed some degree of attenuation of virulence for primates. Within the 5' noncoding region and the 2B2C region of the HAV genome, passage 9/10 virus differed in sequence from wild-type at a single and novel position in the 2C gene, while the sequence of the passage 32 virus was almost identical to that of a fully attenuated passage 35 virus. Passage 21 viruses were found to consist of a mixture of viruses which included all but two of the 13 mutations present in the sequenced regions of the virus from passage 32.

Animals↗

Comprehensive aligned sequence construction for automated design of effective probes (CASCADE-P) using 16S rDNA.

MOTIVATION: Prokaryotic organisms have been identified utilizing the sequence variation of the 16S rRNA gene. Variations steer the design of DNA probes for the detection of taxonomic groups or specific organisms. The long-term goal of our project is to create probe arrays capable of identifying 16S rDNA sequences in unknown samples. This necessitated the authentication, categorization and alignment of the >75 000 publicly available '16S' sequences. Preferably, the entire process should be computationally administrated so the aligned collection could periodically absorb 16S rDNA sequences from the public records. A complete multiple sequence alignment would provide a foundation for computational probe selection and facilitates microbial taxonomy and phylogeny. RESULTS: Here we report the alignment and similarity clustering of 62 662 16S rDNA sequences and an approach for designing effective probes for each cluster. A novel alignment compression algorithm, NAST (Nearest Alignment Space Termination), was designed to produce the uniform multiple sequence alignment referred to as the prokMSA. From the prokMSA, 9020 Operational Taxonomic Units (OTUs) were found based on transitive sequence similarities. An automated approach to probe design was straightforward using the prokMSA clustered into OTUs. As a test case, multiple probes were computationally picked for each of the 27 OTUs that were identified within the Staphylococcus Group. The probes were incorporated into a customized microarray and were able to correctly categorize Staphylococcus aureus and Bacillus anthracis into their correct OTUs. Although a successful probe picking strategy is outlined, the main focus of creating the prokMSA was to provide a comprehensive, categorized, updateable 16S rDNA collection useful as a foundation for any probe selection algorithm.

Algorithms↗

CART variance stabilization and regularization for high-throughput genomic data.

MOTIVATION: mRNA expression data obtained from high-throughput DNA microarrays exhibit strong departures from homogeneity of variances. Often a complex relationship between mean expression value and variance is seen. Variance stabilization of such data is crucial for many types of statistical analyses, while regularization of variances (pooling of information) can greatly improve overall accuracy of test statistics. RESULTS: A Classification and Regression Tree (CART) procedure is introduced for variance stabilization as well as regularization. The CART procedure adaptively clusters genes by variances. Using both local and cluster wide information leads to improved estimation of population variances which improves test statistics. Whereas making use of cluster wide information allows for variance stabilization of data. AVAILABILITY: Sufficient details for our CART procedure are given so that the interested reader can program the method for themselves. The algorithm is also accessible within the Java software package BAMarray(TM), which is freely available to non-commercial users at www.bamarray.com. CONTACT: hemant.ishwaran@gmail.com.

Artificial Intelligence↗

Nucleotide sequence analysis of precursor 5S RNA from Bacillus licheniformis.

The complete nucleotide sequences of the various precursor 5S RNA species occurring in Bacillus licheniformis have been elucidated. The B. licheniformis precursors contain a 5'-precursor-specific segment of 95 nucleotides which is four times as long as the corresponding segment of the p5S RNAs from the closely related strains B. subtilis (Sogin, M.L., Pace, N.R., Rosenberg, M., Weissman, S.M. (1976) J. Biol. Chem. 251, 3480-3488) and Bacillus Q (Stiekema, W.J., Raué, H.A., Planta, R.J. (1980) Nucl. Ac. Res. 8, 2193-2211). However, fourteen of the sixteen nucleotides at the 5'-end are identical in the precursors from all three strains. These conserved nucleotides can form a stem and loop structure which is likely to play an important role in the biosynthesis of 5S RNA. Extension secondary and tertiary structure is present in the 5'-precursor-specific segment as concluded from the results of digestion with RNAase T1 both of the isolated segment and the intact precursors. No sequence homology exists between the 3'-precursor-specific segments of the B. licheniformis precursors and those of the other two strains except for a stretch of U residues at the 3-terminus. This stretch of U residues is not immediately preceded by a hairpin loop, however, as expected for a transcription termination signal (20). The question whether the precursors have already undergone processing at the 3'-end, therefore, remains open. The total number of genetically distinct precursor species in B. licheniformis is at least five and at most ten. Most likely each ribosomal RNA cistron produces a separate p5S RNA as is also the case in Bacillus Q.

Bacillus↗

Characterization of METTL3/14-mediated m6A modification in human transcriptome using Nanopore direct RNA sequencing.

Post-transcriptional RNA modifications modulate diverse aspects of RNA metabolism. N6-methyladenosine (m6A), one of the most abundant internal RNA modifications, is deposited by the core methyltransferase complex, METTL3 and METTL14. Oxford Nanopore Technologies (ONT) platform permits direct, single RNA molecule sequencing while preserving native modifications. However, without rigorous benchmarking, the accuracy and reproducibility of modification detection remain uncertain. Here, we leveraged ONT to comprehensively profile bona fide m6A modifications in cellular RNAs at single-nucleotide resolution by integrating two direct RNA sequencing chemistries (RNA002 and RNA004) with the m6Anet and Dorado modification-detection models. We independently depleted METTL3 and METTL14 in human cells and rigorously validated modification calls through several assays and independent orthogonal methods (GLORI and miCLIP). We find that Dorado detected a higher number of m6A events and enabled simultaneous detection of other RNA modifications (5-methylcytosine, pseudouridine, and inosine). Pairing Dorado with an in vitro transcribed, unmodified control under stringent filtering, we provide compelling evidence supporting a global reduction in m6A sites and stoichiometry within coding sequences and across genes, particularly in highly modified genes and sites, and at consensus DRACH motifs. We report a differential and complex regulation of modified transcripts, accompanied by a global reduction in poly(A) tail length. Notably, METTL3 and METTL14 depletion produced distinct transcript-specific effects, supporting non-redundant roles within the m6A writer complex. Together, our study illustrates a notable advancement of ONT capabilities and establishes a robust transcriptome-wide framework for RNA modification detection, thereby laying the groundwork for exploring the contribution of METTL3/METTL14 to cellular functions and disease.

Humans↗

Comparison of RNA expression profiles based on maize expressed sequence tag frequency analysis and micro-array hybridization.

Assembly of 73,000 expressed sequence tags (ESTs) representing multiple organs and developmental stages of maize (Zea mays) identified approximately 22,000 tentative unique genes (TUGs) at the criterion of 95% identity. Based on sequence similarity, overlap between any two of nine libraries with more than 3,000 ESTs ranged from 4% to 20% of the constituent TUGs. The most abundant ESTs were recovered from only one or a minority of the libraries, and only 26 EST contigs had members from all nine EST sets (presumably representing ubiquitously expressed genes). For several examples, ESTs for different members of gene families were detected in distinct organs. To study this further, two types of micro-array slides were fabricated, one containing 5,534 ESTs from 10- to 14-d-old endosperm, and the other 4,844 ESTs from immature ear, estimated to represent about 2,800 and 2,500 unique genes, respectively. Each array type was hybridized with fluorescent cDNA targets prepared from endosperm and immature ear poly(A(+)) RNA. Although the 10- to 14-d-old postpollination endosperm TUGs showed only 12% overlap with immature ear TUGs, endosperm target hybridized with 94% of the ear TUGs, and ear target hybridized with 57% of the endosperm TUGs. Incomplete EST sampling of low-abundance transcripts contributes to an underestimate of shared gene expression profiles. Reassembly of ESTs at the criterion of 90% identity suggests how cross hybridization among gene family members can overestimate the overlap in genes expressed in micro-array hybridization experiments.

Contig Mapping↗

Role of the DIS hairpin in replication of human immunodeficiency virus type 1.

The virion-associated genome of human immunodeficiency virus type 1 consists of a noncovalently linked dimer of two identical, unspliced RNA molecules. A hairpin structure within the untranslated leader transcript is postulated to play a role in RNA dimerization through base pairing of the autocomplementary loop sequences. This hairpin motif with the palindromic loop sequence is referred to as the dimer initiation site (DIS), and the type of interaction is termed loop-loop kissing. Detailed phylogenetic analysis of the DIS motifs in different human and simian immunodeficiency viruses revealed conservation of the hairpin structure with a 6-mer palindrome in the loop, despite considerable sequence divergence. This finding supports the loop-loop kissing mechanism. To test this possibility, proviral genomes with mutations in the DIS palindrome were constructed. The appearance of infectious virus upon transfection into SupT1 T cells was delayed for the DIS mutants compared with that obtained by transfection of the wild-type provirus (pLAI), confirming that this RNA motif plays an important role in virus replication. Surprisingly, the RNA genome extracted from mutant virions was found to be fully dimeric and to have a normal thermal stability. These results indicate that the DIS motif is not essential for human immunodeficiency virus type 1 RNA dimerization and suggest that DIS base pairing does not contribute to the stability of the mature RNA dimer. Instead, we measured a reduction in the amount of viral RNA encapsidated in the mutant virions, suggesting a role of the DIS motif in RNA packaging. This result correlates with the idea that the processes of RNA dimerization and packaging are intrinsically linked, and we propose that DIS pairing is a prerequisite for RNA packaging.

Genome, Viral↗

Multiple transcription initiation sites, alternative splicing, and differential polyadenylation contribute to the complexity of human neurofibromatosis 2 transcripts.

Northern blot analysis has shown that the human neurofibromatosis type 2 (NF2) cDNA hybridizes to multiple RNA species. To examine whether these hybridizing RNA species represent NF2 transcripts, we cloned the complete NF2 cDNA by a combination of techniques: 5' and 3' rapid amplification of cDNA ends, RT-PCR, and searching and sequencing the NF2-related cDNA clones from the IMAGE consortium. We showed that human NF2 transcripts initiate at multiple positions. Analogous to those reported previously, NF2 transcripts undergo alternative splicing in the coding exons. We isolated eight alternatively spliced NF2 cDNA isoforms, including one that contains a new exon termed exon 2', which potentially could encode proteins of different sizes. We assembled the overlapping cDNA fragments, and the longest NF2 cDNA, containing all 17 exons, consists of 6067 nucleotides, which is consistent with the size of the major RNA species hybridized to the NF2 probe. The cDNA has a 425-nucleotide 5' untranslated region upstream from the ATG start codon, and a long 3' untranslated region of 3869 nucleotides. We also isolated two shorter NF2 cDNAs that were terminated by different polyadenylation signal sequences, which indicates that differential usage of multiple polyadenylation sites also contributes to the complexity of human NF2 transcripts. By reference to the transcription initiation site mapped, we analyzed the 5' flanking sequence of the human NF2 gene. Transient transfection analysis in human 293 kidney, SK-N-AS neuroblastoma, and NT2/D1 teratocarcinoma cells with NF2 promoter-luciferase chimeric constructs revealed a core promoter region extending 400 base pairs from the major transcription initiation site. Although multiple regions are required for full promoter activity, a site-directed mutagenesis experiment identified a GC-rich sequence (position -58 to -46), which could be bound by transcription factor Sp1, as a positive cis-acting regulatory element. Cotransfection studies in Drosophila melanogaster SL2 cells showed that Sp1 could activate the NF2 promoter through the GC-rich sequence.

Alternative Splicing↗

Sequence analysis of the leader RNA of two porcine coronaviruses: transmissible gastroenteritis virus and porcine respiratory coronavirus.

The leader RNA sequence was determined for two pig coronaviruses, transmissible gastroenteritis virus (TGEV), and porcine respiratory coronavirus (PRCV). Primer extension, of a synthetic oligonucleotide complementary to the 5' end of the nucleoprotein gene of TGEV was used to produce a single-stranded DNA copy of the leader RNA from the nucleoprotein mRNA species from TGEV and PRCV, the sequences of which were determined by Maxam and Gilbert cleavage. Northern blot analysis, using a synthetic oligonucleotide complementary to the leader RNA, showed that the leader RNA sequence was present on all of the subgenomic mRNA species. The porcine coronavirus leader RNA sequences were compared to each other and to published coronavirus leader RNA sequences. Sequence homologies and secondary structure similarities were identified that may play a role in the biological function of these RNA sequences.

Amino Acid Sequence↗

Identification of RBP binding sites using RNA deaminases.

RNA-binding proteins (RBPs) are critical regulators of gene expression and RNA processing. Identification of their binding sites has important implications for their physiological and disease-related functions. Crosslinking and immunoprecipitation, followed by sequencing (CLIP-seq) and its derivatives, are the most commonly used methods to identify RBP binding sites, but are laborious and require a large amount of starting material. Recent advancements harnessing RNA deaminases in fusion to any RBP of interest, allow for the profiling of RBP binding sites from low-input samples in simpler procedures. Among these efforts, we developed STAMP (Surveying Targets by APOBEC-Mediated Profiling), which efficiently detects RBP-RNA interactions. This chapter describes the detailed protocol for the STAMP method, including plasmid construction, delivery and sorting, library preparation and bioinformatic data analysis.

RNA-Binding Proteins↗