Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “RNA sequencing analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Taurine as a constituent of mitochondrial tRNAs: new insights into the functions of taurine and human mitochondrial diseases.

Taurine (2-aminoethanesulphonic acid), a naturally occurring, sulfur-containing amino acid, is found at high concentrations in mammalian plasma and tissues. Although taurine is involved in a variety of processes in humans, it has never been found as a component of a protein or a nucleic acid, and its precise biochemical functions are not fully understood. Here, we report the identification of two novel taurine-containing modified uridines (5-taurinomethyluridine and 5-taurinomethyl-2-thiouridine) in human and bovine mitochondrial tRNAs. Our work further revealed that these nucleosides are synthesized by the direct incorporation of taurine supplied to the medium. This is the first reported evidence that taurine is a constituent of biological macromolecules, unveiling the prospect of obtaining new insights into the functions and subcellular localization of this abundant amino acid. Since modification of these taurine-containing uridines has been found to be lacking in mutant mitochondrial tRNAs for Leu(UUR) and Lys from pathogenic cells of the mitochondrial encephalomyopathies MELAS and MERRF, respectively, our findings will considerably deepen our understanding of the molecular pathogenesis of mitochondrial encephalomyopathic diseases.

Anticodon↗

Categorization and characterization of transcript-confirmed constitutively and alternatively spliced introns and exons from human.

By spliced alignment of human DNA and transcript sequence data we constructed a data set of transcript-confirmed exons and introns from 2793 genes, 796 of which (28%) were seen to have multiple isoforms. We find that over one-third of human exons can translate in more than one frame, and that this is highly correlated with G+C content. Introns containing adenosine at donor site position +3 (A3), rather than guanosine (G3), are more common in low G+C regions, while the converse is true in high G+C regions. These two classes of introns are shown to have distinct lengths, consensus sequences and correlations among splice signals, leading to the hypothesis that A3 donor sites are associated with exon definition, and G3 donor sites with intron definition. Minor classes of introns, including GC-AG, U12-type GT-AG, weak, and putative AG-dependant introns are identified and characterized. Cassette exons are more prevalent in low G+C regions, while exon isoforms are more prevalent in high G+C regions. Cassette exon events outnumber other alternative events, while exon isoform events involve truncation twice as often as extension, and occur at acceptor sites twice as often as at donor sites. Alternative splicing is usually associated with weak splice signals, and in a majority of cases, preserves the coding frame. The reported characteristics of constitutive and alternative splice signals, and the hypotheses offered regarding alternative splicing and genome organization, have important implications for experimental research into RNA processing. The 'AltExtron' data sets are available at http://www.bit.uq.edu.au/altExtron/ and http://www.ebi.ac.uk/~thanaraj/altExtron/.

AT Rich Sequence↗

Distribution of infecting hepatitis C virus genotypes in end-stage liver disease patients at a large American transplantation center.

The distribution of hepatitis C virus (HCV) genotypes was studied in 202 anti-HCV-positive liver transplant candidates with end-stage liver disease. HCV sequences were successfully amplified from 185 patients: In the first 100, the genotype was determined by direct sequencing in the NS5 region, and in the remaining 85, type-specific primers were used for genotyping. Eighty-five patients (46.0%) were infected with type 1a HCV strains, 52 (28.1%) with type 1b, 14 (7.6%) with type 2b, 13 (7.0%) with type 4, 5 (2.7%) with type 3a, 2 (1.1%) with type 2a, and 1 (0.5%) with type 2c. Thirteen HCV-positive patients (7.0%) could not be genotyped. The relatively low prevalence of genotype 1b in this population of end-stage liver disease patients speaks against postulated higher pathogenicity of this genotype.

Adult↗

An evolutionary model for protein-coding regions with conserved RNA structure.

Here we present a model of nucleotide substitution in protein-coding regions that also encode the formation of conserved RNA structures. In such regions, apparent evolutionary context dependencies exist, both between nucleotides occupying the same codon and between nucleotides forming a base pair in the RNA structure. The overlap of these fundamental dependencies is sufficient to cause "contagious" context dependencies which cascade across many nucleotide sites. Such large-scale dependencies challenge the use of traditional phylogenetic models in evolutionary inference because they explicitly assume evolutionary independence between short nucleotide tuples. In our model we address this by replacing context dependencies within codons by annotation-specific heterogeneity in the substitution process. Through a general procedure, we fragment the alignment into sets of short nucleotide tuples based on both the protein coding and the structural annotation. These individual tuples are assumed to evolve independently, and the different tuple sets are assigned different annotation-specific substitution models shared between their members. This allows us to build a composite model of the substitution process from components of traditional phylogenetic models. We applied this to a data set of full-genome sequences from the hepatitis C virus where five RNA structures are mapped within the coding region. This allowed us to partition the effects of selection on different structural elements and to test various hypotheses concerning the relation of these effects. Of particular interest, we found evidence of a functional role of loop and bulge regions, as these were shown to evolve according to a different and more constrained selective regime than the nonpairing regions outside the RNA structures. Other potential applications of the model include comparative RNA structure prediction in coding regions and RNA virus phylogenetics.

Base Sequence↗

Identification and sequencing of two isopentenyladenosine-modified transfer RNAs from Chinese hamster ovary cells.

To determine the presence and identity of isopentenyladenosine-containing transfer RNAs (tRNAs) in a mammalian cell line, we adopted a novel method to isolate, clone and sequence these RNAs. This method was based on 3' polyadenylation of the tRNA prior to cDNA synthesis, PCR amplification, cloning and DNA sequencing. Using this unique procedure, we report the cloning and sequencing of the selenocysteine-tRNA and mitochondrial tryptophan-tRNA from Chinese hamster ovary cells which contain this specific tRNA modification. This new method will be useful in the identification of other tRNAs and other small RNAs where the primary sequence is unknown.

Animals↗

Silencer elements as possible inhibitors of pseudoexon splicing.

Human pre-mRNAs contain a definite number of exons and several pseudoexons which are located within intronic regions. We applied a computational approach to address the question of how pseudoexons are neglected in favor of exons and to possibly identify sequence elements preventing pseudoexon splicing. A search for possible splicing silencers was carried out on a pseudoexon selection that resembled exons in terms of splice site strength and exon splicing enhancer (ESE) representation; three motifs were retrieved through hexamer composition comparisons. One of these functions as a powerful silencer in transfection-based splicing assays and matches a previously identified silencer sequence with hnRNP H binding ability. The other two motifs are novel and failed to induce skipping of a constitutive exon, indicating that they might act as weak repressors or in synergy with other unidentified elements. All three motifs are enriched in pseudoexons compared with intronic regions and display higher frequencies in intronless gene-coding sequences compared with exons. We consider that a subpopulation of pseudoexons might rely on negative regulators for splicing repression; this hypothesis, if experimentally verified, might improve our understanding of exonic splicing regulatory sequences and provide the identification of a novel mutation target for human genetic diseases.

Animals↗

MicroInspector: a web tool for detection of miRNA binding sites in an RNA sequence.

Regulation of post-transcriptional gene expression by microRNAs (miRNA) has so far been validated for only a few mRNA targets. Based on the large number of miRNA genes and the possibility that one miRNA might influence gene expression of several targets simultaneously, the quantity of ribo-regulated genes is expected to be much higher. Here, we describe the web tool MicroInspector that will analyse a user-defined RNA sequence, which is typically an mRNA or a part of an mRNA, for the occurrence of binding sites for known and registered miRNAs. The program allows variation of temperature, the setting of energy values as well as the selection of different miRNA databases to identify miRNA-binding sites of different strength. MicroInspector could spot the correct sites for miRNA-interaction in known target mRNAs. Using other mRNAs, for which such an interaction has not yet been described, we discovered frequently potential miRNA binding sites of similar quality, which can now be analysed experimentally. The MicroInspector program is easy to use and does not require specific computer skills. The service can be accessed via the MicroInspector web server at http://www.imbb.forth.gr/microinspector.

Base Pairing↗

miRU: an automated plant miRNA target prediction server.

MicroRNAs (miRNAs) play important roles in gene expression regulation in animals and plants. Since plant miRNAs recognize their target mRNAs by near-perfect base pairing, computational sequence similarity search can be used to identify potential targets. A web-based integrated computing system, miRU, has been developed for plant miRNA target gene prediction in any plant, if a large number of sequences are available. Given a mature miRNA sequence from a plant species, the system thoroughly searches for potential complementary target sites with mismatches tolerable in miRNA-target recognition. True or false positives are estimated based on the number and type of mismatches in the target site, and on the evolutionary conservation of target complementarity in another genome which can be selected according to miRNA conservation. The output for predicted targets, ordered by mismatch scores, includes complementary sequences with mismatches highlighted in colors, original gene sequences and associated functional annotations. The miRU web server is available at http://bioinfo3.noble.org/miRU.htm.

Algorithms↗

The FOLDALIGN web server for pairwise structural RNA alignment and mutual motif search.

Foldalign is a Sankoff-based algorithm for making structural alignments of RNA sequences. Here, we present a web server for making pairwise alignments between two RNA sequences, using the recently updated version of foldalign. The server can be used to scan two sequences for a common structural RNA motif of limited size, or the entire sequences can be aligned locally or globally. The web server offers a graphical interface, which makes it simple to make alignments and manually browse the results. The web server can be accessed at http://foldalign.kvl.dk.

Algorithms↗

Evidence that public database records for many cancer-associated genes reflect a splice form found in tumors and lack normal splice forms.

Alternative splicing is widespread in the human genome, and it appears that many genes display different splice forms in cancerous tissue than in normal human tissues. However, since cDNAs for many cancer-associated genes were originally cloned from tumor samples, it is important to ask whether this repertoire of cDNAs provides a complete or representative picture of the transcript isoforms found in normal tissues. To answer this, we used bioinformatics and RT-PCR to identify novel splice forms, focusing on in-frame exonskips, for a panel of 50 cancer-associated genes in normal tissue samples. These data show that in nearly two-thirds of the genes, normal tissues expressed previously unknown splice forms, of which 40% were normally a dominant splice form. Surprisingly, the tumor-associated splice forms were twice as likely to be represented in GenBank than their normal tissue-associated splice forms, most likely because 70% of the mRNAs in GenBank for these genes were cloned from tumor samples. As an example, we describe a novel normal splice form of IKBbeta, an important regulator of the NFkappaB pathway. Our data suggest that systematic re-evaluation of cancer genes' splice forms in normal tissue will yield insights into their distinct functions in normal tissues and in cancer. Our database contains 1308 novel normal splice forms, including many known cancer genes.

Alternative Splicing↗

Abundance of correctly folded RNA motifs in sequence space, calculated on computational grids.

Although functional RNA molecules are known to be biased in overall composition, the effects of background composition on the probability of finding a particular active site by chance has received little attention. The probability of finding a particular motif has important implications both for understanding the distribution of functional RNAs in ancient and modern organisms with varying genome compositions and for tuning SELEX pools to optimize the chance of finding specific functions. Here we develop a new method for calculating the probability of finding a modular motif containing base-paired regions, and use a computational grid to fold several hundred million random RNA sequences containing the core elements of the isoleucine aptamer and the hammerhead ribozyme to estimate the probability that a sequence containing these structural elements will fold correctly when isolated from background sequences of different compositions. We find that the two motifs are most likely to be found in distinct regions of compositional space, and that the regions of greatest abundance are influenced by the probability of finding the conserved bases, finding the flanking helices, and folding, in that order of importance. Additionally, we can refine our estimates of the number of random sequences required for a 50% probability of finding an example of each site in unbiased random pools of length 100 to 4.1 x 10(9) for the isoleucine aptamer and 1.6 x 10(10) for the hammerhead ribozyme. These figures are consistent with the facile recovery of these motifs from SELEX experiments.

Base Composition↗

TACT: Transcriptome Auto-annotation Conducting Tool of H-InvDB.

Transcriptome Auto-annotation Conducting Tool (TACT) is a newly developed web-based automated tool for conducting functional annotation of transcripts by the integration of sequence similarity searches and functional motif predictions. We developed the TACT system by integrating two kinds of similarity searches, FASTY and BLASTX, against protein sequence databases, UniProtKB (Swiss-Prot/TrEMBL) and RefSeq, and a unified motif prediction program, InterProScan, into the ORF-prediction pipeline originally designed for the 'H-Invitational' human transcriptome annotation project. This system successively applies these constituent programs to an mRNA sequence in order to predict the most plausible ORF and the function of the protein encoded. In this study, we applied the TACT system to 19 574 non-redundant human transcripts registered in H-InvDB and evaluated its predictive power by the degree of agreement with human-curated functional annotation in H-InvDB. As a result, the TACT system could assign functional description to 12 559 transcripts (64.2%), the remainder being hypothetical proteins. Furthermore, the overall agreement of functional annotation with H-InvDB, including those transcripts annotated as hypothetical proteins, was 83.9% (16 432/19 574). These results show that the TACT system is useful for functional annotation and that the prediction of ORFs and protein functions is highly accurate and close to the results of human curation. TACT is freely available at http://www.jbirc.aist.go.jp/tact/.

Amino Acid Motifs↗

High sensitivity RNA pseudoknot prediction.

Most ab initio pseudoknot predicting methods provide very few folding scenarios for a given RNA sequence and have low sensitivities. RNA researchers, in many cases, would rather sacrifice the specificity for a much higher sensitivity for pseudoknot detection. In this study, we introduce the Pseudoknot Local Motif Model and Dynamic Partner Sequence Stacking (PLMM_DPSS) algorithm which predicts all PLM model pseudoknots within an RNA sequence in a neighboring-region-interference-free fashion. The PLM model is derived from the existing Pseudobase entries. The innovative DPSS approach calculates the optimally lowest stacking energy between two partner sequences. Combined with the Mfold, PLMM_DPSS can also be used in predicting complicated pseudoknots. The test results of PLMM_DPSS, PKNOTS, iterated loop matching, pknotsRG and HotKnots with Pseudobase sequences have shown that PLMM_DPSS is the most sensitive among the five methods. PLMM_DPSS also provides manageable pseudoknot folding scenarios for further structure determination.

Algorithms↗

NYD-SP16, a novel gene associated with spermatogenesis of human testis.

By hybridizing human adult testis cDNA microarrays with human adult and embryo testis cDNA probes, a novel human testis gene NYD-SP16 was identified. NYD-SP16 expression was 6.44-fold higher in adult testis than in fetal testis. NYD-SP16 contains 1595 base pairs (bp) and a 762-bp open reading frame encoding a 254-amino acid protein with 73% amino acid sequence identity with the mouse testis homologous protein. The NYD-SP16 gene was localized to human chromosome 5q14. The deduced structure of the NYD-SP16 protein contains one transmembrane domain, which was confirmed by GFP/NYD-SP16 fusion protein expression in the cytomembrane of the transfected human choriocarcinoma JAR cells, suggesting that it is a transmembrane protein. Multiple tissue distribution indicated that NYD-SP16 mRNA is highly expressed in the testes and pancreas, with little or no expression elsewhere. Further analysis of abnormal expression in infertile male patients revealed complete absence of NYD-SP16 in the testes of patients with Sertoli-cell-only syndrome and variable expression in patients with spermatogenic arrest. Homologous gene expression in mouse testis was confirmed in spermatogenic cells by in situ hybridization. The results of cDNA microarray, in situ hybridization, and semiquantitative polymerase chain reaction in mouse testis of different stages indicated that NYD-SP16 expression is developmentally regulated. These results suggest that the putative NYD-SP16 protein may play an important role in testicular development/spermatogenesis and may be an important factor in male infertility.

Adult↗

Kinetics of HIV-1 RNA and resistance-associated mutations after cessation of antiretroviral combination therapy.

OBJECTIVE: To study the kinetics of HIV-1 RNA and drug-induced mutations after cessation of antiretroviral therapy (ART). DESIGN AND METHODS: Successive plasma samples from 26 patients were tested for HIV-1 RNA by PCR and for mutations associated with drug resistance by sequencing of the pol gene. RESULTS: After cessation of ART the phase of undetectable virus (< 50 copies/ml), ranging from 6 to more than 29 days, was followed by a rapid viral increase, which slowed down before a plateau corresponding to pre-treatment levels or higher was reached in most cases (14/19 patients). In one patient virus was still undetectable at 4 weeks. Also, a significantly larger number of primary protease inhibitor (PI)-associated mutations reverted to wild-type, as compared with secondary PI-, and primary reverse transcriptase inhibitor (RTI)-associated mutations. During the rapid viral increase no mutations disappeared, which instead happened during the slower viral increase preceding the viral plateau level. CONCLUSION: After discontinuation of ART large individual variations were found for the time period until HIV-1 became detectable in plasma, possibly due to differences in the HIV-1 specific immunity. The more rapid loss of primary PI mutations suggests that they might cause a more impaired viral fitness than primary RTI mutations. However, the persistence of drug mutations during the initial viral load increase indicates that mutated strains may still replicate efficiently.

Adult↗