Search PubMedSearch

SEARCH · Search PubMed

Results for “Untranslated Regions”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Machine learning-based analysis of the impact of 5' untranslated region on protein expression.

The 5' untranslated region (5'UTR) plays a crucial regulatory role in messenger RNA (mRNA), with modified 5'UTRs extensively utilized in vaccine production, gene therapy, etc. Nevertheless, manually optimizing 5'UTRs may encounter difficulties in balancing the effects of various cis-elements. Consequently, multiple 5'UTR libraries have been created, and machine learning models have been employed to analyze and predict translation efficiency (TE) and protein expression, providing insights into critical regulatory features. On the one hand, these screening libraries, based on TE and mean ribosome load, struggle to accurately quantify protein expression; on the other hand, a precise method for quantifying 5'UTRs necessitates a significantly costlier library. To resolve this dilemma, we constructed a library utilizing firefly luciferase as the reporter to measure accurate protein expression. In addition, we optimized the library construction method by clustering mRNA sequences to reduce redundant data and minimize the size of the dataset. This dual strategy by increasing accuracy and reducing dataset size was found to be effective in predicting the 5'UTRs from the PC3 cell line.

5' Untranslated Regions

The nucleotide sequence of the 5' untranslated region of human gamma-globin mRNA.

The nucleotide sequence of the entire 5' untranslated region of human gamma-globin mRNA has been determined. This was accomplished by analyzing complementary DNA (cRNA) synthesized from the mRNA with reverse transcriptase. The CDNA was labeled at its 3' end with 32"p using terminal deoxynucleotidyl transferase, digested with the restriction endonuclease Hae III and the end-labeled fragment isolated ans sequenced by the method of Maxam and Gilbert. Including the initiation codon AUG, the 5' untranslated region of human gamma-globin mRNA contains 57 nucleotides, compared to 41 in alpha- and 54 in beta-globin mRNA. There is very little homology between alpha and gamma sequences in the 5' region. There is considerable homology between beta- and gamma-globin mRNAs in the regions proximal and distal to the initiation codon, but the entire sequence shows less homology than the human and rabbit beta-globin mRNAs. The hexanucleotide sequence CUUCUG is found near the 5' ends of all three human globin mRNAs, suggesting a possible role of this sequence or ribosomal binding. Both guanosine and cytidine were found at the 19th nucleotide position from the 5' end of the gamma mRNA. We believe this heterogeneity arises from the difference in nucleotide sequence between the A gamma and G gamma loci.

Base Sequence

Functional analysis of stem-loop structures within the SARS-CoV-2 5' untranslated region using a plasmid-based reporter system.

The 5' untranslated region (5'UTR) of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) contains highly conserved stem-loop structures that regulate viral gene expression. This study investigated the functional contributions of selected 5'UTR stem-loop elements to reporter gene expression using a plasmid-based mammalian expression system. Five constructs were tested using a non-integrating plasmid: the wild-type (WT) 5'UTR fused to GFP under the CMV promoter, and four deletion variants (&#x394;B, &#x394;C, &#x394;D, and &#x394;E) corresponding to deletions of stem-loop 4 (SL4), SL4.5, SL5, and SL5a, respectively. Following transfection into HEK293 cells, GFP fluorescence was quantified using a fluorescence microplate reader, and relative GFP transcript abundance was assessed by RT-qPCR. Deletion of SL4 (&#x394;B) resulted in marked reduction in both fluorescence and relative transcript abundance compared to WT construct, indicating substantially reduced reporter gene expression. In contrast, deletion of SL4.5, SL5, or SL5a did not produce the pronounced reduction observed for &#x394;B, although descriptive RT-qPCR analysis indicated differences in relative transcript abundance among these variants. Statistical analysis of fluorescence data demonstrated significant differences among constructs (one-way ANOVA, p&#x2009;<&#x2009;0.05). Because the reporter assay was based on plasmid expression, the observed differences likely reflect combined contributions from transcription, transcript abundance, RNA stability, and translation rather than translation alone. These findings demonstrate that the SL4 region contributes substantially to reporter gene expression in this experimental system, whereas the remaining stem-loop regions examined exert comparatively modest effects. This study provides additional insight into the functional organization of the SARS-CoV-2 5'UTR and establishes a framework for future investigations aimed at distinguished the transcriptional, post-transcriptional, and translational contributions of individual RNA structural elements.

5' Untranslated Regions

Pestivirus internal ribosome entry site (IRES) structure and function: elements in the 5' untranslated region important for IRES function.

The importance of certain structural features of the 5' untranslated region of classical swine fever virus (CSFV) RNA for the function of the internal ribosome entry site (IRES) was investigated by mutagenesis followed by in vitro transcription and translation. Deletions made from the 5' end of the CSFV genome sequence showed that the IRES boundary was close to nucleotide 65: thus, the IRES includes the whole of domain II but no sequences upstream of this domain. Deletions which invaded domain II even to a small extent reduced activity to about 20% that of the full-length structure, and this 20% residual activity persisted with more extensive deletions until the whole of domain II had been removed and the deletions invaded the pseudoknot, whereupon IRES activity fell to zero. The importance of both stems of the pseudoknot was verified by making mutations in both sides of each stem; this severely reduced IRES activity, but the compensating mutations which restored base pairing caused almost full IRES function to be regained. The importance of the length of the loop linking the two stems of the pseudoknot was demonstrated by the finding that a reduction in length from the wild-type AUAAAAUU to AUU almost completely abrogated IRES activity. Random A-->U substitutions in the wild-type sequence showed that IRES activity was fairly proportional to the number of A residues retained in this pseudoknot loop, with a preference for clustered neighboring A residues rather than dispersed As. Finally, it was found that the sequence of the highly conserved domain IIIa loop is, rather surprisingly, not important for the maintenance of full IRES activity, although amputation of the entire domain IIIa stem and loop was highly debilitating. These results are interpreted in the light of recent models, derived from cryo-electron microscopy, of the interaction of the closely related hepatitis C virus IRES with 40S ribosomal subunits.

5' Untranslated Regions

A trainable language model with potential to modulate translation rates in non-model organisms by generating upstream untranslated region sequence libraries.

Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5'-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanistic assumptions, prior knowledge that may not be available in non-model contexts, or the screening of sequence libraries. Here, we present a simple generative approach for creating synthetic 5'-UTR libraries based solely on the genomic sequence statistics of any desired organism. The method uses a sliding-window n-gram language model applied to native 5'-UTR sequences to produce novel sequences that preserve organism-specific base distributions and motifs without hard-coding specific motifs or mechanistic rules into inflexible statistical templates. We have applied this approach to the model bacterium Escherichia coli and the non-model probiotic Limosilactobacillus reuteri. Libraries of approximately 1,000 sequences were generated for each organism, from which about 100 unique sequences were experimentally tested for translation of a fluorescent reporter protein. In both organisms, the synthetic libraries yielded a broad range of translation levels from this relatively small number of tested variants. Sequences derived from an organism's own genomic statistics provided a more uniformly distributed range of translation rates in that organism than sequences derived from the other species. Correlations of individual sequence performance across the two species were weak, and thermodynamic predictions of ribosome binding strength showed very little predictive power, especially in the non-model L. reuteri. The results demonstrate that simple statistical language model approaches applied to genomic data can generate functional translational regulatory sequence libraries without detailed mechanistic knowledge or explicit reference to consensus motifs. The approach requires minimal computational resources, avoids reproducing native sequences, and can be readily applied to any organism with a sequenced genome. This strategy may lower technical barriers to expression tuning in non-model organisms.

5' Untranslated Regions

Structure of the constant and 3' untranslated regions of the murine gamma 2b heavy chain messenger RNA.

The complete coding sequence for the constant region of the mouse gamma 2b immunoglobulin heavy chain and the 3' untranslated region has been determined. The coding portion of the sequence is 1008 nucleotides long (amino acid residues 114 to 449), and the 3' noncoding region contains 102 nucleotides preceeding the polyadenylate. An extra carboxyl-terminal lysine residue which had not been observed in the gamma 2b or other gamma subclass protein sequences occurs in the nucleotide sequence and is probably processed posttranslationally. A 17-nucleotide sequence occurs with slight variation twice in CH1 and once in CH2 domains in the same relative location but with different translational phase. This sequence may be the site of crossover in a gamma 2b . gamma 2a heavy chain variant, an indication of possible recombinational activity of some kind.

Animals

Occurrence of reiterated sequences in an untranslated region of Simian virus 40 DNA determined by nucleotide sequence analysis.

An earlier report (Subramanian, Dhar, and Weissman, 1977c) presented the nucleotide sequence of Eco RII-G fragment of SV40 DNA, which contains the origin of DNA replication. The nucleotide sequence of Eco RII-N fragment located next to Eco RII-G on the physical map of SV40 DNA is presented in this report. Eco RII-N is found to be a tandem duplication of the last 55 nucleotides of Eco RII-G. This tandem repeat is immediately preceded by two other reiterated sequences occurring within Eco RII-G, one of them being a tandem repeat of 21 nucleotides and the other a nontandem repeat of 10 nucleotides. These repetitive sequences occur in close proximity to the origin of DNA replication which is known to contain other specialized sequences such as a few palindromes (one of which is 27 long and possesses a perfect 2-fold axis of symmetry), one "true" palindrome, and a long A/T-rich cluster. The repeats (and the replication origin) occur within an untranslated region of SV40 DNA flanked by (the few) structural genes coding for the "late" proteins on the one side and that (those) coding for the "early" protein(s) on the other side. The reiterated sequences are comparable in some respects to repetitive sequences occurring in eucaryotic DNAs. Possible biological functions of the repeats are discussed.

Base Sequence

Protamine messenger RNA from rainbow trout testis contains the nucleotide sequence A-A-U-A-A-A in an untranslated region.

Full-length, complementary DNAs were prepared to rainbow trout protamine mRNA using reverse transcriptase and were labelled during synthesis by the replacement of dATP by [alpha32P]dATP or dTTP by [alpha32P]dTTP. The 32P-labelled protamine complementary DNAs were digested with T4 endonuclease IV. Fragments from the digests were separated in two dimensions, and those discrete fragments which could be identified from both the A-labelled and T-labelled complementary DNAs were subjected to sequence analysis. The sequences described here all arise from the non-coding region. One pentadecanucleotide contained the sequence A-A-U-A-A-A which has been reported by Proudfoot and Brownlee ((1976) Nature 263, 211-214) to occur in the noncoding regions of six other eukaryotic mRNAs.

Animals

Nucleotide sequence of the Hind-C fragment of simian virus 40 DNA. Comparison of the 5'-untranslated region of wild-type virus and of some deletion Mutants.

We report here the nucleotide sequence of the wild-type simian virus 40 (strain 776) restriction fragment Hind-C-P1 DNA and of the homologous region of various mutant DNAs which lack part of this fragment. During this work, we detected between EcoRII fragments N and G an additional, 17-base-pair EcoRII fragment, fragment P, which had previously been overlooked. Also, an additional dTpdG dinucleotide at residues L 339--340 was observed by sequence analysis of the DNA minus (E) strand; the presence of this dinucleotide was masked on sequencing patterns of the plus strand due to the persistence (during gel electrophoresis) of some secondary structures in the strand's 5'-terminal region. These nucleotide additions raise the total length of SV40 DNA to 5243 base pairs. The longest tandemly repeated segment in SV40 DNA now extends over 72 base pairs. SV40 deletion mutants dl 893 and dl 894 and SV40 strains Rh 911 and 1801 all lack an identical 72-base-pair-long DNA segment in the Hind-C region. This deletion corresponds precisely to one of the two aforementioned large tandemly repeated sequences. Mutant dl 895 lacks 66 base pairs, 63 of which are part of the former repetition. All these mutants, except dl 895, very probably were generated by an intramolecular, homologous recombination event. The 40-base-pair deletion in mutant dl 1811 includes the major capping site of SV40 late RNA. dl 1812 lacks only three base pairs, which are part of the overlapping HhaI and HpaII restriction sites at position 0.725--0.726.

Base Sequence

The nucleotide sequences of the untranslated 5' regions of human alpha- and beta-globin mRNAs.

The complete sequences of the untranslated 5' regions of human alpha- and beta-globin mRNAs were determined by sequence analysis of full-length cDNAs. The single-stranded cDNAs were digested with the restriction endonuclease Hae III, and the two 3'-terminal fragments of 75 and 132 nucleotides, complementary to the 5' termini of the alpha- and beta-globin mRNAs, respectively, were isolated and sequenced. Including the initiation codon AUG, the untranslated 5' regions of human alpha- and beta-globin mRNAs contain 41 and 54 nucleotides, respectively, and exhibit striking homologies with the corresponding sequences in the rabbit. Human alpha- and beta-globin mRNAs have five bases in the region of the initiation codon that may form base pairs with the 3' terminus of 18S rRNA. Stable secondary structures with hairpin loops can be constructed in the untranslated 5' regions.

Base Sequence

Non-coding regions of nuclear-DNA-encoded mitochondrial genes and intergenic sequences are targeted by autoantibodies in breast cancer.

Autoantibodies against mitochondrial-derived antigens play a key role in chronic tissue inflammation in autoimmune disorders and cancers. Here, we identify autoreactive nuclear genomic DNA (nDNA)-encoded mitochondrial gene products (GAPDH, PKM2, GSTP1, SPATA5, MFF, TSPOAP1, PHB2, COA4, and HAGH) recognized by breast cancer (BC) patients' sera as nonself, supporting a direct relationship of mitochondrial autoimmunity to breast carcinogenesis. Autoreactivity of multiple nDNA-encoded mitochondrial gene products was mapped to protein-coding regions, 3' untranslated regions (UTRs), as well as introns. In addition, autoantibodies in BC sera targeted intergenic sequences that may be parts of long non-coding RNA (lncRNA) genes, including LINC02381 and other putative lncRNA neighbors of the protein-coding genes ERCC4, CXCL13, SOX3, PCDH1, EDDM3B, and GRB2. Increasing evidence indicates that lncRNAs play a key role in carcinogenesis. Consistent with this, our findings suggest that lncRNAs, as well as mRNAs of nDNA-encoded mitochondrial genes, mechanistically contribute to BC progression. This work supports a new paradigm of breast carcinogenesis based on a globally dysfunctional genome with altered function of multiple mitochondrial and non-mitochondrial oncogenic pathways caused by the effects of autoreactivity-induced dysregulation of multiple genes and their products. This autoimmunity-based model of carcinogenesis will open novel avenues for BC treatment.

autoimmunity

Sequences of mouse immunoglobulin light chain genes before and after somatic changes.

We have determined the nucleotide sequences of the germ line gene as well as a corresponding somatically mutated and rearranged gene coding for a mouse immunoglobulin lambdaI type light chain. These sequencing studies were carried out on three Eco RI-DNA fragments which had been cloned from BALB/c mouse embryos or a lambdaI chainsecreting myeloma, H2020. The embryonic DNA clone Ig 99lambda contains two protein-encoding segments, one for the majority of the hydrophobic leader (L) and the other for the rest of the leader and the variable (V) region of the lambda0 chain (Cohn et al., 1974); these segments are separated by a 93 base pair (bp) intervening sequence (I-small). The coding of the V region ends with His at residue 97. The second embryonic DNA clone Ig 25lambda includes a 39 bp DNA segment (J) coding for the rest of the conventionally defined V region (that is, up to residue 110), and also contains the sequence coding for the constant (C) region approximately 1250 untranslated bp (I-large) away from the J sequence. The J sequence is directly linked with the V-coding sequence in the myeloma DNA clone, Ig 303lambda, which has the various DNA segments arranged in the following order: 5' untranslated region, L, l-small, V linked with J, l-large, C, 3' untranslated sequence. The lg 303lambda V DNA sequence codes for the V region synthesized by the H2020 myeloma and is different from the lg 99lambda V DNA sequence by only two bases. No silent base change was observed between the two DNA clones for the entire sequence spanning the 5' untranslated regions and the V-coding segments. These results confirm the previously drawn conclusion that an active complete lambdaI gene arises by somatic recombination that takes place at the ends of the V-coding DNA segment and the J sequence. No sequence homology was observed at or near the sites of the recombination.

Animals

Molecular cloning and sequence analysis of adult chicken betal globin cDNA.

The molecular cloning and nucleotide sequence analysis of adult chicken beta globin mRNA is reported. DNA sequences derived from in vitro transcrption of globin mRNA were purified and amplified as recombinant DNA using the plasmid pBR322. Sequence analysis of several clones coding for beta globin strongly suggests that transcription errors may be generated near the 5' end of transcripts in vitro by reverse transcription. The complete sequence of the longest beta globin insert containing 51 bases of the 5' untranslated region as well as the complete coding and 3' untranslated regions has been determined.

Amino Acid Sequence

Human growth hormone: complementary DNA cloning and expression in bacteria.

The nucleotide sequence of a DNA complementary to human growth hormone messenger RNA was cloned; it contains 29 nucleotides in its 5' untranslated region, the 651 nucleotides coding for the prehormone, and the entire 3' untranslated region (108 nucleotides). The data reported predict the previously unknown sequence of the signal peptide of human growth hormone and, by comparison with the previously determined sequences of rat growth hormone and human chorionic somatomammotropin, strengthens the hypothesis that these genes evolved by gene duplication from a common ancestral sequence. The human growth hormone gene sequences have been linked in phase to a fragment of the trp D gene of Escherichia coli in a plasmid vehicle, and a fusion protein is synthesized at high level (approximately 3 percent of bacterial protein) under the control of the regulatory region of the trp operon. This fusion protein (70 percent of whose amino acids are coded for by the human growth hormone gene) reacts specifically with antibodies to human growth hormone and is stable in E. coli.

Amino Acid Sequence

N6-methyladenosine modification of the subgroup J avian leukosis viral RNAs attenuates host innate immunity via MDA5 signaling.

Subgroup J avian leukosis virus (ALV-J), a retrovirus, elicits immunosuppression and persistent infections in chickens. Although it is widely acknowledged that ALV-J can evade the host's innate immune defenses, the mechanisms behind this immune evasion remain elusive. N6-methyladenosine (m6A), the most prevalent internal RNA modification, plays a role in innate immune evasion. Our research identified ALV-J as an inefficient stimulator of innate immunity in vitro and in vivo, with its genomic RNA featuring m6A modifications predominantly in the envelope protein (Env) region and 3' untranslated region (3'UTR). To elucidate the functional consequences of m6A modification, we subsequently generated m6A-deficient ALV-J through its culturing in the DF-1 overexpressing fat mass and obesity-associated protein (FTO) cells. The m6A-deficient ALV-J virus, or its RNAs significantly enhanced IFN-&#x3b2; production compared to the wild-type (wt) ALV-J, suggesting a pivotal regulatory function of m6A modifications in modulating innate immune response. Mechanistically, the m6A modification of the ALV-J genomic RNA directly impacted its recognition by MDA5, weakening its binding and ubiquitination and attenuating IFN-&#x3b2; activation. Moreover, m6A-deficient ALV-J, created by inducing mutations in m6A sites within Env and 3'UTR, exhibited reduced replication capacity and elevated IFN-&#x3b2; expression in host cells. Importantly, this phenomenon was abolished in MDA5-knockout DF-1 cells, further demonstrating the core role of MDA5. These data demonstrate that m6A modification of ALV-J genomic RNA dampens the host's innate immune response through MDA5 signaling pathway.

Animals

Genome-wide detection of human 5' UTR variants that impact protein translation.

The 5' untranslated region (5' UTR) of messenger RNAs (mRNAs) plays a central role in regulating protein synthesis initiation, particularly through the Kozak sequence and upstream open reading frames (uORFs). Genetic variants within these regulatory elements could affect translation, altering gene expression and contributing to clinical phenotypes in humans. We developed a computational method called 5ULTRA (5' Untranslated Region Annotation) for analysis of whole-exome sequencing and whole-genome sequencing data to detect, annotate, and prioritize 5' UTR variants with potential translation impact. 5ULTRA identifies single-nucleotide variants, indels, and splicing variants that affect uORFs by creating or disrupting start/stop codons and that alter Kozak sequence strength of either the uORFs or the main coding sequence. 5ULTRA incorporates recent uORF databases and provides comprehensive annotations. 5ULTRA implements a machine-learning score to prioritize candidate variants with predicted effects on translation and also provides specific mechanistic predictions. The score correlates strongly with experimentally measured protein-level effects of 5' UTR variants. We applied 5ULTRA to multiple genetics datasets across diverse disease contexts, identifying candidate variants including potential cancer-driving somatic mutations predicted to decrease ABI1 level or increase NRAS abundance; common variants associated with traits such as multiple sclerosis, lung function, and cardiovascular function, by altering protein levels of TAGAP, VRTN, and SPAAR, respectively; and rare germline variants in our cohort, including a splicing variant of RPSA leading to 5' UTR sequence alteration that causes congenital asplenia and a variant of TNF that could predispose to tuberculosis.

Humans

Abnormal levels of miRNA in pancreatic cancer are linked to tumor progression by regulating the translation of tumor-associated mRNA.

BACKGROUND: Pancreatic cancer remains one of the most malignant tumors, characterized by limited treatment efficacy. MAIN FINDINGS: microRNAs (miRNAs) play a crucial role in regulating the proliferation, invasion, migration, drug resistance, apoptosis, and cell cycle progression of pancreatic cancer cells by inhibiting tumor-associated proteins. Metscape analysis revealed that miRNA-targeted proteins associated with pancreatic cancer are enriched in processes such as cell proliferation, mitosis, and cell migration, and participate in multiple signaling pathways. These proteins primarily localize to classical pathways, including JAK/STAT, PI3K/AKT, and Wnt/&#x3b2;-catenin. Furthermore, gene mutations or abnormal alternative poly(A)denylation (APA) within miRNA-targeted regions can disrupt base pairing to the 3'-Untranslated Region (3'-UTR), thereby enhancing the translation of oncogenic mRNA translation. FUTURE DIRECTIONS: Collectively, these findings indicate that multiple miRNAs act cooperatively to influence pancreatic cancer progression. Consequently, therapeutic strategies aimed at restoring the balance of the miRNA system are essential to disrupt the 'mRNA-oncogene' vicious cycle.

Humans

Aldosterone suppresses Na+/H+ exchanger-3 expression through miR-204-5P-mediated posttranscriptional regulation in distal colon.

Na+/H+ exchanger-3 (NHE3) is a major mediator of electroneutral NaCl absorption in the intestine and colon. In the distal colon, chronic aldosterone exposure suppresses NHE3 expression, but the molecular mechanism responsible for this regulation remains unclear. Here, we tested whether aldosterone represses NHE3 through microRNA-dependent posttranscriptional regulation. Transcriptomic analysis of distal colon from dietary Na+-depleted rats identified miR-204-5P (miR-204-5P) as markedly upregulated. Aldosterone increased miR-204-5P abundance and concomitantly reduced NHE3 mRNA, protein expression, and transport activity in rat and human distal colonic epithelium and in SK-CO15 cells. Bioinformatic and reporter analyses identified a conserved miR-204-5P binding site within the NHE3 3'-untranslated region, and miR-204-5P mimic transfection markedly suppressed NHE3 expression and transport activity without affecting other Na+/H+ exchanger isoforms. These findings identify a previously unrecognized aldosterone-microRNA signaling pathway that mediates chronic repression of NHE3 and provide new insight into hormonal regulation of colonic Na+ absorption.NEW & NOTEWORTHY This study identifies a previously unrecognized aldosterone-microRNA signaling mechanism regulating colonic Na+ absorption. We demonstrate that aldosterone induces miR-204-5P, which directly targets the NHE3 3'-untranslated region and suppresses NHE3 expression and transport activity in distal colonic epithelium. These findings reveal a microRNA-mediated pathway linking mineralocorticoid signaling to long-term inhibition of electroneutral NaCl absorption, providing new insight into hormonal regulation of intestinal electrolyte transport.

Animals