Search PubMedSearch

SEARCH · Search PubMed

Results for “Small non-coding RNA”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

83 records · Page 5Linked to original sources

Complete structure of the hamster alpha A crystallin gene. Reflection of an evolutionary history by means of exon shuffling.

The eye lens contains a structural protein, alpha crystallin, composed of two homologous primary gene products alpha A2 and alpha B2. In certain rodents, still another alpha crystallin polypeptide, alpha AIns, occurs, which is identical to alpha A2 except that it contains an insertion peptide between residues 63 and 64. In this paper we describe the complete alpha A crystallin gene that has been cloned from DNA isolated from Syrian golden hamster. Evidence is provided that the alpha A gene is present as a single copy in the hamster genome. The detailed organization of the gene has been established by means of DNA sequence analysis and S1 nuclease mapping, revealing that the gene consists of four exons. The first exon contains the information for the 68 base-pair long 5' non-coding region as well as the coding information for the first 63 amino acids. The second exon encodes the 23 amino acid insertion sequence, the third exon codes for amino acid 87 to 127 of the alpha AIns chain, whereas the last exon encodes the C-terminal 69 amino acids and contains the information for the 523 base-pair long 3' non-coding region. The second exon is bordered by a 3' splice junction (A X G/G X C), which deviates from the consensus for donor splice sites (A X G/G X T). This deviation is found in both hamster and mouse. An internal duplication was detected in the first exon by using a DIAGON-generated matrix for comparison. By means of similar DIAGON-generated matrices it was confirmed that the amino acids coded for by the third and fourth exons are homologous to the small heat-shock proteins of Drosophila, Caenorhabditis and soyabean. The implications of the differential splicing and the evolutionary aspects of the detected homologies are discussed.

Animals

Partial nucleotide sequence of the Murray Valley encephalitis virus genome. Comparison of the encoded polypeptides with yellow fever virus structural and non-structural proteins.

The sequence of 5400 bases corresponding to the 5'-terminal half of the Murray Valley encephalitis virus genome has been determined. The genome contains a 5' non-coding region of about 97 nucleotides, followed by a single continuous open reading frame that encodes the structural proteins followed by the non-structural proteins. Amino acid sequence homology between the Murray Valley encephalitis and yellow fever (Rice et al., 1985) polyproteins is 42% over the region sequenced. The start points of the various Murray Valley encephalitis virus-coded proteins have been assigned on the basis of this homology and a consistent set of potential proteolytic cleavage sites identified, the sequences of which are similar in Murray Valley encephalitis and yellow fever. The deduced Murray Valley encephalitis gene order is 5'-C-prM (M)-E-NS1-ns2a-ns2b-NS3-3'. The genome organization of Murray Valley encephalitis and yellow fever appears to be identical and the sizes of the predicted virus-coded proteins similar between the two viruses. Both viruses encode a basic capsid protein followed by three glycoproteins; the glycoproteins appear to have the conventional topology of N terminus outside with a C-terminal membrane-spanning domain. There are conserved glycosylation sites in prM, the precursor to the M protein of the virion, and in NS1, a non-structural protein of uncertain function. The glycosylation sites in E, the major envelope protein of the virion, are not conserved as to position. We predict the existence, in flavivirus-infected cells, of two small, hydrophobic peptides, ns2a and ns2b, which show only limited amino acid sequence homology. Finally, about half of the amino acid sequence of NS3 has been obtained; NS3 is a hydrophilic non-structural protein that shows 55% amino acid sequence similarity between Murray Valley encephalitis and yellow fever over the region sequenced and is probably involved in RNA replication.

Amino Acid Sequence

What constitutes the signal for the initiation of protein synthesis on Escherichia coli mRNAs?

Small DNA fragments (60 to 80 nucleotides), randomly obtained from a collection of 14 catabolic, biosynthetic or regulatory Escherichia coli genes, have been shot-gun cloned in place of the lacZ ribosome binding site. A total of 47 recombinants showing substantial beta-galactosidase synthesis (at least 1/30th of the wild-type) were isolated, and their newly acquired translational starts were characterized. Of these, 46 were found to carry a ribosome binding site from one of the original genes, and only one, a non-natural start. Moreover, 12 out of the 14 natural starts were found. The two that were not found are the only ones lacking a Shine-Dalgarno element. So, real starts are generally active in the lac mRNA, whereas the many sites (approx. 100 in this gene collection) that carry a Shine-Dalgarno element followed by AUG or GUG but are located in intra- or intergenic regions, or on non-transcribed strands, are inactive. I conclude that: (1) these "false" starts, being strongly discriminated against in the lac message, are presumably also inactive in their original mRNAs; (2) the discriminating information, being portable from one mRNA to another, must be contained within a small DNA region surrounding the starts. Indeed, I further show that it generally lies within a sequence of about 35 nucleotides bracketing real starts; and (3) this information must have a larger effect on initiation than the exact structure of the mRNA, because the discrimination persists despite a complete change of this structure. Previous statistical analysis has shown that real starts differ from false starts in having a non-random sequence composition from nucleotides -20 to +15 with respect to the start. To uncover whether these biases constitute the discriminating information or simply reflect coding constraints, translational starts were randomly searched in eukaryotic, largely non-coding, DNA. These "eukaryotic" starts all have an in-phase AUG or GUG, preceded by a typical Shine-Dalgarno sequence; outside these elements, the initiator region is strikingly rich in A, and poor in C. These biases match those found around real starts, demonstrating that they are indeed part of the initiation signal. Finally, I describe a simple procedure for introducing any DNA fragment in place of the lac operator site on the E. coli chromosome.

Base Sequence

Complete sequence of the S RNA of lymphocytic choriomeningitis virus (WE strain) compared to that of Pichinde arenavirus.

Previous studies have reported that the 3' half of the small, S, RNA species of the WE strain of lymphocytic choriomeningitis (LCM) virus codes for the viral nucleoprotein in a subgenomic, viral-complementary, mRNA species (Romanowski, V. and Bishop, D.H.L. (1985) Virus Res. 2, 35-51). The complete sequence of the LCM-WE S RNA has now been obtained, indicating that the 5' half of the RNA codes for the viral glycoprotein precursor in a viral-sense sequence that does not overlap the N gene. It is concluded that, like Pichinde virus (Auperin, D. et al. (1984) J. Virol. 52, 897-904), LCM has an ambisense S RNA coding strategy. The LCM-WE S RNA is 3375 nucleotides in length, has a size of 1.14 X 10(6) Da and base composition of 26.1% A, 23.2% C, 21.5% G, 29.2% U. The 3' and 5' end sequences of the S RNA are complementary for some 30 nucleotides, depending on the arrangement. The non-coding regions at the two ends are 77 (5') and 60 (3') nucleotides long. The glycoprotein precursor has a primary amino acid size of 56293 Da and is rich in potential glycosylation sites as well as histidine and cysteine residues. It has both amino and carboxy proximal hydrophobic regions. The LCM-WE S RNA and predicted protein sequence data have been compared to those of Pichinde arena-virus. Extensive RNA and protein sequence homology exists for the two S RNA species, although the homology for the glycoprotein sequences of the two viruses (39%) is less than the 50% observed for the two viral nucleoproteins.

Amino Acid Sequence

A novel peptide encoded by circTLL1 drives osimertinib resistance in lung cancer by modulating the NT5C2/Ras/PI3K axis.

BACKGROUND: Acquired resistance to osimertinib, a third-generation EGFR tyrosine kinase inhibitor, remains a major clinical challenge in the treatment of non-small cell lung cancer (NSCLC). Although circular RNAs (circRNAs) have been increasingly implicated in drug resistance, most studies have focused on their canonical role as microRNA sponges, while their capacity to encode functional micropeptides remains largely unexplored. This study aimed to identify novel circRNAs involved in osimertinib resistance and to characterize their regulatory functions at the protein level. METHODS: Osimertinib-resistant (OR) NSCLC cell lines were established and validated. High-throughput RNA sequencing was performed to compare the circRNA expression profiles between parental and OR cells. The function of the candidate circRNA was assessed through a series of in vitro and in vivo experiments, including cell viability assays, apoptosis analysis, and xenograft mouse models. Mechanistic investigations involved mass spectrometry, co-immunoprecipitation and western blotting to explore its protein-coding potential and downstream signaling pathways. RESULTS: We identified a novel circRNA, termed circTLL1, that was stably and significantly upregulated in OR-NSCLC cells. Functionally, overexpression of circTLL1 promoted osimertinib resistance, whereas its knockdown restored drug sensitivity both in vitro and in vivo. Mechanistically, we discovered that circTLL1 harbors an open reading frame (ORF) that is translated into a novel 90-amino-acid protein, which we designated circTLL1-90aa. Further investigation revealed that circTLL1-90aa directly interacts with and promotes the degradation of 5'-nucleotidase, cytosolic II (NT5C2), thereby uncoupling nucleotide metabolism from its normal regulatory constraints. The consequent downregulation of NT5C2 leads to elevated GTP levels and leading to the sustained activation of the downstream Ras/PI3K/AKT signaling pathway. CONCLUSION: Our findings unveil a previously unrecognized circRNA/micropeptide/metabolism cascade underlying osimertinib resistance. The identification of the circTLL1-90aa/NT5C2/Ras/PI3K axis not only expands the functional repertoire of the non-coding genome but also provides new insights into the complexity of drug resistance. Given its selective upregulation in resistant cells, circTLL1-90aa holds promise both as a predictive biomarker for treatment stratification and as an actionable therapeutic target, offering a novel strategy to overcome osimertinib resistance in NSCLC patients.

Pyrimidines

Cleavage region organizes the structural architecture of the SINE-derived B2 repressive ribozyme.

The SINE-encoded B2 retrotransposon is an RNA Polymerase III (POL-III)-derived transcript whose expression is substantially upregulated during various cellular stress responses. Beyond retrotransposition, the B2 non-coding RNA can directly bind and repress the activity of RNA Polymerase II (POL-II), leading to a significant downregulation of transcripts during stress. Notably, our recent findings have shown that B2 is a self-cleaving ribozyme whose activity can be induced by interactions with chromatin-modifying factors through non-canonical epigenetic mechanisms that co-regulate its function across distinct chromatin-binding target loci. Here, by integrating RNA chemical probing, small-angle X-ray scattering, and 3D motif modeling, we determine structural ensemble-to-function relations for the B2 SINE ribozyme RNA. Genetic perturbations of the RNA suggest that the B2 SINE ribozyme has a well-defined secondary and dynamic tertiary structure that depends on the integrity of the critical region, which confers ribozymatic activity and repressive extent by POL-II. Using an RNA engineering approach, we examine the effects of point mutations, deletions of the main cleavage site, and deletions of the cleavage domain on the structural ensemble of the RNA. Combining this approach with in vitro and in vivo functional perturbation methods highlights the relationships between structural ensembles and various biologically relevant functional outcomes.

RNA, Catalytic

Micropeptides encoded by lncRNAs associated with cancer progression reveal novel immunogenic epitopes.

MOTIVATION: Long non-coding RNAs (lncRNAs) regulate gene expression, chromatin organization, and cellular signaling. Recent studies indicate that ∼20% of the ∼36 000 human lncRNA genes harbor small open reading frames (sORFs) capable of producing micropeptides (MPs), whose functions remain largely unknown. Whether these peptides contribute to the cancer immunopeptidome is largely unexplored. RESULTS: We systematically analyzed lncRNAs with strong experimental and computational evidence of MP-encoding potential (∼13% of the initial MP collection). Using The Cancer Genome Atlas (TCGA), we identified 2606 high-confidence lncRNA-derived MPs encoded by 647 genes across 16 cancer types. We then focused on 501 MPs from 124 lncRNA genes whose expression changes significantly across tumor stages and metastatic transitions, representing cancer transitional lncRNAs (Tr-lncRNAs). Dipeptide composition and conservation analyses showed that these MPs differ from a size-matched human coding proteome, supporting their potential as neoantigens. All possible 9-mer peptides were evaluated for predicted binding to prevalent European HLA class I alleles. Approximately 60% of Tr-lncRNA genes and 184 (37%) of derived peptides exhibited strong predicted HLA binding. Peptides from XIST, PCAT7, PVT1, HAND2-AS1 showed broad HLA coverage. Notably, TTN-AS1, encoded an MP (79 aa) generated 33 predicted distinct epitopes spanning all 27 HLA alleles. Our analysis identifies lncRNA-derived MPs as a previously underexplored source of potential cancer neoantigens, highlighting their promise as biomarkers and targets for immunotherapy. AVAILABILITY: Data, code and supplementary materials are available in https://doi.org/10.5281/zenodo.20167452 and GitHub: https://github.com/stavzok1/lncrna_peptide_analysis.

Humans

Sequence of rat alpha- and gamma-casein mRNAs: evolutionary comparison of the calcium-dependent rat casein multigene family.

The complete sequences of rat alpha- and gamma-casein mRNAs have been determined. The 1402-nucleotide alpha- and 864-nucleotide gamma-casein mRNAs both encode 15 amino acid signal peptides and mature proteins of 269 and 164 residues, respectively. Considerable homology between the 5' non-coding regions, and the regions encoding the signal peptides and the phosphorylation sites, in these mRNAs as compared to several other rodent casein mRNAs, was observed. Significant homology was also detected between rat alpha- and bovine alpha s1-casein. Comparison of the rodent and bovine sequences suggests that the caseins evolved at about the time of the appearance of the primitive mammals. This may have occurred by intragenic duplication of a nucleotide sequence encoding a primitive phosphorylation site, -(Ser)n-Glu-Glu-, and intergenic duplication resulting in the small casein multigene family. A unique feature of the rat alpha-casein sequence is an insertion in the coding region containing 10 repeated elements of 18 nucleotides each. This insertion appears to have occurred 7-12 million years ago, just prior to the divergence of rat and mouse.

Amino Acid Sequence

The 5'-terminal sequence of potato leafroll virus RNA: evidence of recombination between virus and host RNA.

The discrepancy between published sequences of the 5' non-coding regions of RNA of a Scottish (S) and that of Dutch (D), Australian and Canadian isolates of potato leafroll virus (PLRV) was investigated. Reverse transcription followed by amplification by polymerase chain reaction showed that RNA from three distinct Scottish isolates of PLRV contained molecules with 5' ends like that of the original Scottish isolate. However, determination of the 5'-terminal sequences of RNA in two of these preparations showed that most RNA molecules had 5' termini like those of the Dutch and other non-Scottish sequences. Northern blot analysis confirmed that only a small fraction of PLRV RNA contained sequences homologous to the 5'-terminal 119 nucleotides of the PLRV-S sequence. Most PLRV-S RNA molecules therefore have termini like that reported for PLRV-D. The 5'-terminal 119 nucleotides of the minor species of PLRV-S RNA were very similar (109/119 nucleotides were identical) in sequence to an exon of tobacco chloroplast DNA open reading frame 196. The results therefore suggest that recombination has occurred between virus RNA and host RNA.

Base Sequence

Determination and comparative analysis of the small RNA genomic sequences of California encephalitis, Jamestown Canyon, Jerry Slough, Melao, Keystone and Trivittatus viruses (Bunyaviridae, genus Bunyavirus, California serogroup).

The nucleotide sequences of the small (S) genomic RNAs of six California (CAL) serogroup bunyaviruses (Bunyaviridae: genus Bunyavirus) were determined. The S RNAs of two California encephalitis virus strains, two Jamestown Canyon virus strains, Jerry Slough virus, Melao virus, Keystone virus and Trivittatus virus contained the overlapping nucleocapsid (N) and non-structural (NSs) protein open reading frames (ORFs) as described previously for the S RNAs of other CAL serogroup viruses. All N protein ORFs were 708 nucleotides in length and encoded a putative 235 amino acid gene product. The NSs ORFs were found to be of two lengths, 279 and 294 nucleotides, which potentially encode 92 and 97 amino acid proteins, respectively. The complementary termini and a purine-rich sequence in the 3' non-coding region (genome-complementary sense) were highly conserved amongst CAL serogroup bunyavirus S RNAs. Phylogenetic analyses of N ORF sequences indicate that the CAL serogroup bunyaviruses can be divided into three monophyletic lineages corresponding to three of the complexes previously derived by serological classification. The truncated version of the NSs protein, which is found in five CAL serogroup bunyaviruses, appears to have arisen twice during virus evolution.

Amino Acid Sequence

Cloned proteolipid protein and myelin basic protein cDNA. Transcription of the two genes during myelination.

cDNA clones of rat brain proteolipid protein (PLP), also named lipophilin, the major integral myelin membrane protein, and of myelin basic protein (MBP), the major extrinsic myelin protein, have been isolated from a rat brain cDNA library cloned into the PstI site of pBR322. Poly(A)+ RNA from actively myelinating 18-day-old rats has been reversely transcribed. Oligonucleotides synthesized according to the established amino-acid sequence of lipophilin and the nucleotide sequence of the small myelin basic protein of the N-terminal, the central and C-terminal region of their sequences were used as hybridization probes for screening. The largest insert in one of several lipophilin clones was 2,585 base pairs (bp) in length (pLp 1). It contained 521 bp of the C-terminal coding sequence and the complete 2,064 bp long non-coding 3' sequence. The myelin basic protein cDNA insert of clones pMBP5 and pMBP6 is 2,530 bp long and that of clones pMBP2 and pMBP3 640 bp. These clones were also characterized. pMBP2 was sequenced and used together with the lipophilin cDNA clones as hybridization probes to estimate the lipophilin and myelin basic protein mRNA levels of rat brain during the myelination period. The expression of the lipophilin and myelin basic protein genes during development of the myelin sheath appears to be strictly coordinated.

Animals