Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “long noncoding RNA”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

Nucleotide sequence of the Barmah Forest virus genome.

Barmah Forest virus (BFV) is an atypical alphavirus [Dalgarno, L., Short, N. J., Hardy, C. M., Bell, J. R., Strauss, J. H., and Marshall, I. D. (1984). Virology 133, 416-426] and has been classified as the sole known member of a seventh alphavirus serocomplex. The complete nucleotide sequence of BFV genomic RNA is 11,488 nucleotides in length excluding the poly(A) tail. Two long open reading frames in the RNA encode a nonstructural polyprotein of 2411 amino acids and a structural polyprotein of 1239 amino acids, respectively. The BFV envelope protein E2 is unique among sequenced alphaviruses in having no N-linked glycosylation sites; E1 carries two glycosylation sites. From amino acid sequence comparisons with sequenced alphaviruses BFV is most closely related to Ross River and Semliki Forest viruses. Sequence homology between BFV and other alphaviruses is relatively uniform along the length of the nonstructural and structural polyproteins, providing no evidence that BFV has arisen from recombination between ancestral alphaviruses in the coding region of the genome. The BFV 3' noncoding region of 445 nucleotides has unusual features. There are two unrelated sequence blocks of 48 nucleotides (sequence I) and 47 nucleotides (sequence II) both of which are repeated once. Sequence I is closely related to a repeat in the 3' noncoding region of Ross River and Getah viruses; sequence II is unrelated to repeat blocks in other sequenced alphaviruses. Thus, recombination between ancestral viruses may have played a role in the evolution of the BFV 3' noncoding region.

Alphavirus↗

The primary structure of human hemopexin deduced from cDNA sequence: evidence for internal, repeating homology.

We have cloned and analyzed a cDNA containing the coding sequence for human hemopexin. We have first identified, by immunological screening of 30.000 colonies of a liver cDNA library in the expression vector pEX1, a clone carrying an insert 1170 base pairs long that shows 100% homology with a known human hemopexin peptide. The complete sequence coding for hemopexin was isolated from a liver cDNA library in the vector pAT218. The DNA insert of 1523 base pairs shows an open reading frame coding for 439 amino acids, a 3' noncoding region of 159 nucleotides long, followed by a poly(A) tail. The insert spans the entire coding region and from which the primary structure of the protein was deduced. By computer assisted analysis of the amino acid sequence, it was possible to recognize a core unit, of about 45 amino acids, which is repeated 8 or possibly even 10 fold along the polypeptide chain. This feature suggests that the gene might have evolved through a series of duplications. This characteristic, together with prediction of secondary structure, suggest a rough model for the tridimensional folding that allows some speculations on the function of hemopexin. Blot hybridization of total RNA from human liver with nick translated hemopexin cDNA detected a message of about 1600 nucleotides. Southern blot experiments to identify the hemopexin gene (s) suggest that it is not a large multi-gene family, but that there is only one or at most a few genes in the human genome.

Amino Acid Sequence↗

RNA interference: the molecular immune system.

Introduction of double-stranded RNA (dsRNA) into cells expressing a homologous gene triggers RNA interference (RNAi), or RNA-based gene silencing (RBGS). The dsRNA degrades corresponding host mRNA into small interfering RNAs (siRNAs) by a protein complex containing Dicer. siRNAs in turn are incorporated into the RNA-induced silencing complex (RISC) that includes helicase, RecA, and exo- and endo-nucleases as well as other proteins. Following its assembly, the RISC guides the RNA degradation machinery to the target RNAs and cleaves the cognate target RNA in a sequence-specific, siRNA-dependent manner. RNAi has now been documented in a wide variety of organisms, including plants, fungi, flies, worms, and more recently, higher mammals. In eukaryotes, dsRNA directed against a range of viruses (i.e., HIV-1, RSV, HPV, poliovirus and others) and endogenous genes can induce sequence-specific inhibition of gene expression. In invertebrates, RNAi can be efficiently triggered by either long dsRNAs or 21- to 23-nt-long siRNAs. However, in jawed vertebrates, dsRNA longer than 30 bp can induce interferon and thus trigger undesirable side effects instead of initiating RNAi. siRNAs have been shown to act as potent inducers of RNAi in cultured mammalian cells. Many investigators have suggested that siRNAs may have evolved as a normal defense against endogenous and exogenous transposons and retroelements. Through a combination of genetic and biochemical approaches, some of the mechanisms underlying RNAi have been described. Recent data in C. elegans shows that two homologs of siRNAs, microRNAs (miRNAs) and tiny noncoding RNAs (tncRNAs) are endogenously expressed. However, many aspects of RNAi-induced gene silencing, including its origins and the selective pressures which maintain it, remain undefined. Its evolutionary history may pass through the more primitive immune functions of prokaryotes involving restriction enzymes that degrade plasmid DNA molecules that enter bacterial cells. RNAi has evolved further among eukaryotes, in which its wide distribution suggests early origins. RNAi seems to be involved in a variety of regulatory and immune functions that may differ among various kingdoms and phyla. We present here proposed mechanisms by which RBGS protects the host against endogenous and exogenous transposons and retroelements. The potential for therapeutic application of RBGS technology in treating viral infections such as HIV is also discussed.

AIDS Vaccines↗

Sequences of VP9 genes from short and supershort rotavirus strains.

Segment 10 genes from a short (RV-5, serotype G2) and a supershort (B37, a new G serotype) strain were cloned and their sequences compared to the (corresponding) segment 11 sequences of Wa, SA11, and UK rotaviruses. The determined nucleotide sequences were 817 (RV-5) and 947 (B37) bases in length and showed extensively conserved 5' noncoding and protein coding regions. The major open reading frame codes for a protein of 200 (RV-5) or 198 (B37) amino acids, and the newly proposed second open reading frame can code for a protein of 92 amino acids. Compared to long strain gene segments, the base sequences of the short and supershort strains were found to contain extended, AT-rich 3' noncoding regions which were not significantly homologous to each other, to other parts of the VP9 gene, or to other rotavirus genes that have been sequenced. The function(s) of these 3' regions is not apparent.

Amino Acid Sequence↗

Arabidopsis micro-RNA biogenesis through Dicer-like 1 protein functions.

Micro-RNAs (miRNAs) are small, noncoding RNAs of 18-25 nt in length that negatively regulate their complementary mRNAs at the posttranscriptional level. Previous work has shown that some RNase III-like enzymes such as Drosha and Dicer are known to be involved in miRNA biogenesis in animals. However, the mechanism of plant miRNA biogenesis still remains poorly understood. In this article, the process of Arabidopsis miR163 biogenesis was examined. The results revealed that two types of miR163 primary transcripts (pri-miR163s) are transcribed from a single gene by RNA polymerase II and that miR163 biogenesis requires at least three cleavage steps by RNase III-like enzymes at 21-nt-long intervals. The first step is from pri-miR163 to long miR163 precursor (premiR163), the second step is from long pre-miR163 to short premiR163, and the last step is from short pre-miR163 to mature miR163 and the remnant. It is interesting that, during the process, four small RNAs including miR163 are released. By using dcl1 mutants, it was demonstrated that Arabidopsis Dicer homologue Dicer-like 1 (DCL1) catalyzes at least the first and second cleavage steps and that double-stranded RNA-binding domains of DCL1 are involved in positioning of the cleavage sites. Our result is direct evidence that DCL1 is involved in processing of pri- and pre-miRNA.

Arabidopsis↗

Genomic organization and transcript analysis of ICAp69, a target antigen in diabetic autoimmunity.

Islet cell antigen p69 (ICAp69) is a target self-antigen in autoimmune (insulin-dependent) diabetes mellitus. Distributed over more than 100 kb on chromosome 6 (6{A1-A2}), the single murine genomic locus contains 14 coding exons, 39-271 bp in length. The identified human and mouse intron-exon junctions are identical, with intron sizes ranging from 94 bp to 24 kb and with conserved flanking region intron sequences. cDNA cloning identified alternatively spliced ICAp69 mRNA transcripts. The predominating alpha-transcripts lack exon 4, while beta-transcripts include this exon, which codes translation termination in all reading frames and a truncated molecule following in vitro expression. gamma-Transcripts show splice removal of exons 8-12, while delta-transcripts exclude exon 11. Transcripts use alternative polyadenylation signals including a less frequent ATTAAA sequence. 5'-Untranslated cDNA and genomic sequencing and long PCR analysis suggest the presence of more noncoding exons. All splice variants encode the conserved T-cell epitope (in exon 2) recognized by autoreactive T cells in diabetic children and diabetes-prone NOD mice.

Amino Acid Sequence↗

Structural analysis of an extremely long 5'-noncoding region of rat brain argininosuccinate lyase mRNA: presence of multiple B1 repeats and multiple upstream AUG codons, and a possibility of translational control.

The present detailed analysis of the sequence of the extremely long (967 bp) 5'-noncoding region of a rat brain argininosuccinate lyase cDNA clone, reveals several features of interest. Multiple copies of partial and inverted (antisense) B1 repeats and multiple upstream ATG codons are present in the region, which suggests a possibility of translational control of the argininosuccinate lyase gene expression in rat brain.

Amino Acid Sequence↗

Long, nearly identical untranslated sequences at the 3' terminal regions of the genomic RNAs of cherry leafroll virus (walnut strain).

Hybridization analyses of cDNA clones derived from the two genomic RNAs, RNA1 and RNA2, of the walnut strain of the nepovirus cherry leafroll nepovirus (wCLRV) demonstrated a long region of high homology between the two viral RNAs. Subsequent mapping and nucleotide sequencing revealed a long, noncoding, presumably untranslated, region (3' UTR) immediately 5' of the terminal polyadenylate, a region that is almost identical in the two RNAs. This 3' UTR is 1567 nucleotide residues long in RNA1. Homologies of about 80% were found with corresponding regions of genomic RNAs from other strains of CLRV, but not with the corresponding regions of other nepovirus genomic RNAs.

Base Sequence↗

Nucleotide sequence of the 26 S mRNA of the virulent Trinidad donkey strain of Venezuelan equine encephalitis virus and deduced sequence of the encoded structural proteins.

A cDNA clone containing all of the 26 S mRNA coding region of the RNA genome of Venezuelan equine encephalitis (VEE) virus, virulent strain Trinidad donkey (TRD), has been constructed and sequenced. The nucleotide and deduced amino acid sequences of the 26 S RNA of VEE virus conform to the general organization of the alphavirus subgenomic mRNA. Excluding the poly(A) tail, the VEE 26 S RNA is 3913 nucleotides long with a protein coding region of 3762 nucleotides. Codon usage in the translated region is nonrandom and correlates well with that reported for Sindbis (SIN), Semliki Forest (SF), and Ross River (RR) alphaviruses. Highly conserved sequences of 19 to 22 nucleotides representing putative replicase recognition sites occur at the 26 S RNA junction region of the 42 S genomic RNA and at the 3' terminus immediately preceding the poly(A) tail. The conserved sequence at the 26 S/42 S junction region of VEE virus differs from that of other alphaviruses in that an ochre termination codon (UAA) is substituted for a GGU (Gly) codon present in the other viruses. The 5' and 3' noncoding regions (30 and 121 nucleotides, respectively) of the VEE 26 S RNA are shorter than has been reported for several other alphaviruses. The approximate transmembrane domains of the VEE E1 and E2 envelope glycoproteins have been identified. VEE E1 contains a single asparagine-linked glycosylation site, whereas E2 has three such sites, all of which are apparently glycosylated. The deduced amino acid sequence of the VEE polyprotein shows an overall homology of 44 to 46% with the precursor polyproteins of SIN, SF, and RR viruses. VEE virus capsid, E1, and E2 structural proteins show 43 to 46%, 50 to 53%, and 36 to 41% homology, respectively, with the cognate proteins of SIN, SF, and RR viruses.

Amino Acid Sequence↗

Global amplification of cDNA from limiting amounts of tissue. An improved method for gene cloning and analysis.

In this study we present an improved polymerase chain reaction (PCR)-based methodology to generate large amounts of high-quality complementary DNA (cDNA) from small amounts of initial total RNA. Global amplification of cDNA makes it possible to simultaneously clone many cDNAs and to construct directional cDNA libraries from a sequence-abundance-normalized cDNA population, and also permits rapid amplification of cDNA ends (RACE), from a limited amount of starting material. The priming of cDNAs with an adapter oligo-deoxythymidine (oligo-dT) primer and the ligation of a modified oligonucleotide to the 3' end of single-stranded cDNAs, through the use of T4 RNA ligase, generates known sequences on either end of the cDNA population. This helps in the global amplification of cDNAs and in the sequence-abundance normalization of the cDNA population through the use of PCR. Utilization of a long-range PCR enzyme mix to amplify the cDNA population helps to reduce bias toward the preferential amplification of shorter molecules. Incorporation of restriction sites in the PCR primers allows the amplified cDNAs to be directionally cloned into appropriate cloning vectors to generate cDNA libraries. RACE-PCR done with biotinylated primers and streptavidin-coated para-magnetic particles are used for the efficient isolation of either full-length coding or noncoding strands.

Base Sequence↗

Sequence of chicken ovalbumin mRNA.

The complete sequence of chicken ovalbumin mRNA is presented; it is 1,859 residues long, excluding its terminal 'cap' and poly(A). The region coding for ovalbumin lies close to the 'cap' but is separated from the poly(A) by an extensive 3' noncoding region of 637 nucleotides which may have no function that is precisely dependent on its sequence.

Amino Acid Sequence↗

Differential expression of the human insulin-like growth factor II gene. Characterization of the IGF-II mRNAs and an mRNA encoding a putative IGF-II-associated protein.

Insulin-like growth factor II (IGF-II) is a polypeptide of 67 amino acids which is thought to play an important role in fetal growth and development. The human IGF-II gene is situated on chromosome 11, very close to the insulin gene. It extends over 30 kb of chromosomal DNA and consists of five noncoding exons (exons 1-4 and 4B) followed by three protein encoding exons (exons 5-7), one of which (exon 7) contains a long 3'-untranslated region. Here we show that differential initiation of transcription can occur at three distinct promoter sites, resulting in the appearance of mRNA species of different lengths. These promoters show a tissue-specific and a development-specific regulation of expression. Furthermore, we have determined the entire nucleotide sequence of the 3'-terminal exon, exon 7, which is about 4 kb long and contains 3.8 kb of 3'-untranslated sequences. This completes the elucidation of the human IGF-II gene structure. Surprisingly, Northern blot analysis of fetal and adult RNA with a probe derived from the 3'-nontranslated region of exon 7 detects a novel 1.8 kb mRNA which appears to be coordinately expressed with the IGF-II mRNAs. In vitro translation of this 1.8 kb mRNA results in the formation of a translation product of 8.3 kDa, which compares well with the size of a predicted translation product from a 252-nucleotides-long open reading frame.

Amino Acid Sequence↗

HCV-RNA assay in peripheral blood mononuclear cells in relation to IFN therapy.

Hepatitis C virus RNA (HCV-RNA) was serially assayed in the serum and peripheral blood mononuclear cells (PBMC) of 10 patients with chronic hepatitis C who underwent interferon (IFN) therapy, and whether detection of HCV-RNA from PBMC serves as an index of the response of chronic hepatitis C to IFN therapy was evaluated. HCV-RNA was assayed by reversed transcription and polymerase chain reaction using the 5'-noncoding region as a primer. IFN therapy was effective in 3 patients and ineffective in the other 7 patients. HCV-RNA disappeared from the serum during and immediately after the IFN therapy in all 3 patients in whom the therapy was effective and in 3 of 7 patients in whom the therapy was ineffective. HCV-RNA disappeared from PBMC in all 3 patients in whom the therapy was effective, but PBMC HCV-RNA remained positive in 6 of the 7 patients in whom the therapy was ineffective, and the serum HCV-RNA became positive again in 5 of these 6 patients after 6 months. The disappearance of HCV-RNA from PBMC was associated with long-term stabilization of the serum alanine aminotransferase value, so HCV-RNA assay in PBMC is considered to be useful as a prognostic marker of chronic hepatitis C after IFN therapy.

Adult↗

Identification of putative noncoding polyadenylated transcripts in Drosophila melanogaster.

Analysis of EST and cDNA collections from a number of metazoan species has identified genes encoding long polyadenylated transcripts that do not contain ORFs of lengths typical for protein-encoding mRNAs. Noncoding functions of such polyadenylated transcripts have been elucidated in only a few examples. The corresponding genes neither contain hallmark sequence motifs nor appear to have been conserved across phyla. Thus, it is impossible to systematically identify new members of this class of gene by using sequence homology and traditional gene-finding algorithms that depend on protein-coding potential. Consequently, even their approximate number has not been established for any metazoan genome. We curated polyadenylated transcripts with limited protein-coding capacity from intergenic regions of the Drosophila melanogaster genome. We used RT-PCR assays, hybridization to RNA blots and whole-mount embryos, and computational analyses to characterize candidate transcripts. We verify the structures and expression of 17 distinct, likely non-protein-coding polyadenylated transcripts. We show that the expression of many of these transcripts is conserved in other Drosophila species, indicating that they have important biological functions.

Animals↗

A population-based statistical approach identifies parameters characteristic of human microRNA-mRNA interactions.

BACKGROUND: MicroRNAs are approximately 17-24 nt. noncoding RNAs found in all eukaryotes that degrade messenger RNAs via RNA interference (if they bind in a perfect or near-perfect complementarity to the target mRNA), or arrest translation (if the binding is imperfect). Several microRNA targets have been identified in lower organisms, but only one mammalian microRNA target has yet been validated experimentally. RESULTS: We carried out a population-wide statistical analysis of how human microRNAs interact complementarily with human mRNAs, looking for characteristics that differ significantly as compared with scrambled control sequences. These characteristics were used to identify a set of 71 outlier mRNAs unlikely to have been hit by chance. Unlike the case in C. elegans and Drosophila, many human microRNAs exhibited long exact matches (10 or more bases in a row), up to and including perfect target complementarity. Human microRNAs hit outlier mRNAs within the protein coding region about 2/3 of the time. And, the stretches of perfect complementarity within microRNA hits onto outlier mRNAs were not biased near the 5'-end of the microRNA. In several cases, an individual microRNA hit multiple mRNAs that belonged to the same functional class. CONCLUSIONS: The analysis supports the notion that sequence complementarity is the basis by which microRNAs recognize their biological targets, but raises the possibility that human microRNA-mRNA target interactions follow different rules than have been previously characterized in Drosophila and C. elegans.

Computational Biology↗

Comparative sequence analysis of the 5' noncoding region of the enteroviruses and rhinoviruses.

A comparative sequence analysis of the 5' noncoding region of a subgroup of the picornaviruses, including the polioviruses, coxsackie B3, and the human rhinoviruses, reveals the conservation of certain features despite the divergence of sequence. In this subgroup, for which nine complete sequences are available, two long stretches of sequence, two pyrimidine-rich regions, and 22 hairpins are conserved. Based on these results, similar secondary structures encompassing the entire 5' noncoding regions of these viruses are predicted. The fact that sequence divergence occurred only in a manner that allowed conservation of these structures implicates a biologically functional role for this region. The possible roles it may have in the picornavirus life cycle are discussed.

Base Sequence↗

Cloning and nucleotide sequence of the simian rotavirus gene 6 that codes for the major inner capsid protein.

The nucleotide sequence of the gene that codes for the major inner capsid protein of the simian rotavirus SA11 has been determined. A DNA copy of mRNA from gene 6 was cloned in the E. coli plasmid pBR322. The full-length gene is 1357 nucleotides long with a 5'-noncoding region of 23 nucleotides and a 3'-noncoding region of 140 nucleotides. The gene contains a single, long, open reading-frame of 1194 nucleotides capable of coding for a protein of 397 amino acids with a molecular weight of 44,816. The predicted protein product is relatively proline-rich with a net charge at neutral pH of -3.5. One stretch of 53 amino acids (encoded by nucleotides 327-485) is basic.

Base Sequence↗

Nucleotide sequence and genetic map of cowpea severe mosaic virus RNA 2 and comparisons with RNA 2 of other comoviruses.

We report the nucleotide sequence of cowpea severe mosaic comovirus (CPSMV) genomic RNA 2. The molecule is composed of 3732 nucleotide (nt) residues, exclusive of the polyadenylate at the 3' end. Only one of the six reading frame registers has a long open reading frame, from nt 255 to nt 3260 in the polarity of encapsidated RNA and corresponding to a polyprotein of 1002 amino acid residues (aa). As has been reported for other comoviruses, a second in-frame AUG, at nt position 531, apparently also initiates translation, at least in vitro. Multiple alignments of the deduced CPSMV polyprotein aa sequence with those of bean pod mottle comovirus (BPMV), cowpea mosaic comovirus (CPMV), and red clover mottle comovirus (RCMV) were consistent with a similar size for each of the three genes: the putative movement protein, beginning at the second in-frame AUG, the large coat protein (L), and the small coat protein. Identical nucleotide sequences in the terminal noncoding regions of RNA 2 of the four viruses are limited to 9 nt at the 5' end and the 3' polyadenylate. However, extensive similarities in sequence and potential structure were found. For all three genes and the 5' untranslated region, CPSMV and BPMV are more similar to each other than either is to CPMV or RCMV, the last two being similar to each other. Observed similarities predict that both cleavage sites in the CPSMV RNA 2 polyprotein are at glutamine-serine dipeptides. A sequence of 16 aa at the amino terminus of L, determined by automated Edman degradation, matched a region of the deduced aa sequence in the polyprotein and is consistent with cleavage at the predicted glutamine-serine dipeptide.

Amino Acid Sequence↗