Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Noncoding RNAs”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Deletion or substitution of the aphthovirus 3' NCR abrogates infectivity and virus replication.

The 3' noncoding region (NCR) of the genomic picornaviral RNA is believed to contain major cis-acting signals required for negative-strand RNA synthesis. The 3' NCR of foot-and-mouth disease virus (FMDV) was studied in the context of a full-length infectious clone in which the genetic element was deleted or exchanged for the equivalent region of a distantly related swine picornavirus, swine vesicular disease virus (SVDV). Deletion of the 3' NCR, while maintaining the intact poly(A) tail as well as its replacement for the SVDV counterpart, abrogated virus replication in susceptible cells as determined by infectivity and Northern blot assays. Nevertheless, the presence of the SVDV sequence allowed the synthesis of low amounts of chimeric viral RNA at extended times post-transfection as compared to RNAs harbouring the 3' NCR deletion. The failure to recover viable viruses or revertants after several passages on susceptible cells suggests that the presence of specific sequences contained within the FMDV 3' NCR is essential to complete a full replication cycle and that FMDV and SVDV 3' NCRs are not functionally interchangeable.

3' Untranslated Regions↗

Yellow fever 5' noncoding region as a potential element to improve hepatitis C virus production through modification of translational control.

The lengthy 5' noncoding region (5' NCR) of hepatitis C virus (HCV) RNA forms a highly ordered secondary structure, very conserved among different strains. It includes an internal ribosome entry site (IRES) element, responsible for the cap-independent translation initiation of HCV RNA. Similarly to the IRES of hepatitis A virus (HAV), another human hepatitis virus, HCV IRES, activity in internal initiation of translation is weak. Furthermore, both viruses exhibit a poor growth phenotype that may result at least partially from an inhibitory control of translation. To enhance HCV translation, as a preliminary step in designing constructs for improvement in viral production, we sought to evaluate a chimeric construct containing the yellow fever virus (YFV) 5' NCR fused to the initiation codon of the HCV coding sequence. YF viral RNA, as the majority of eukaryotic messenger RNAs, is translated by a ribosome scanning mechanism in a cap-dependent manner. The efficiency of translation initiation of the parental HCV construct was compared in vitro in rabbit reticulocyte lysates with that of the chimeric construct containing YFV 5' NCR. Surprisingly, the related distanced YFV 5' NCR was fivefold more active than was the wild-type HCV IRES in directing that function. Furthermore, chimeric transcripts were shown to be effective in vivo after transfection of eukaryotic cells. Taken together, these results raise the following question: why has the HCV genus evolved to the acquisition of an IRES element within its 5' NCR among the Flaviviridae family?

5' Untranslated Regions↗

Nucleotide sequence shows that Bean leafroll virus has a Luteovirus-like genome organization.

The complete nucleotide sequence of the Bean leafroll virus (BLRV) genomic RNA and the termini of its smallest subgenomic RNAs were determined to better understand its mechanisms of gene expression and replication and its phylogenetic position within the Luteoviridae: The number and placement of open reading frames (ORFs) within the BLRV genome was Luteovirus-like. The nucleotide and predicted amino acid sequences of BLRV were most similar to those of Soybean dwarf virus (SbDV). Phylogenetic analyses employing the neighbour-joining method and sister-scanning analysis indicated that the BLRV nonstructural proteins were closely related to those of Barley yellow dwarf virus-PAV (BYDV-PAV), a luteovirus: The region surrounding the frameshift at the junction between ORFs 1 and 2 also contained sequences very similar to those of BYDV-PAV and a Dianthovirus, Red clover necrotic mosaic virus. Similar analyses showed that the structural proteins were most similar to those of the Polerovirus genus. The 3'-noncoding regions downstream of ORF5 contained sequences similar to translational control elements identified in the BYDV-PAV genome. These data suggest that BLRV, like SbDV, is derived either through selection from a common ancestor with BYDV-PAV or that BLRV is the product of two recombination events between luteovirus-like and polerovirus-like ancestors where the 5' 2900 nt and 3' 700 nt of the BLRV genome are from a Luteovirus and the intervening sequences are derived from a Polerovirus:

3' Untranslated Regions↗

U6 small nuclear RNA is transcribed by RNA polymerase III.

A DNA fragment homologous to U6 small nuclear RNA was isolated from a human genomic library and sequenced. The immediate 5'-flanking region of the U6 DNA clone had significant homology with a potential mouse U6 gene, including a "TATA box" at a position 26-29 nucleotides upstream from the transcription start site. Although this sequence element is characteristic of RNA polymerase II promoters, the U6 gene also contained a polymerase III "box A" intragenic control region and a typical run of five thymines at the 3' terminus (noncoding strand). The human U6 DNA clone was accurately transcribed in a HeLa cell S100 extract lacking polymerase II activity. U6 RNA transcription in the S100 extract was resistant to alpha-amanitin at 1 microgram/ml but was completely inhibited at 200 micrograms/ml. A comparison of fingerprints of the in vitro transcript and of U6 RNA synthesized in vivo revealed sequence congruence. U6 RNA synthesis in isolated HeLa cell nuclei also displayed low sensitivity to alpha-amanitin, in contrast to U1 and U2 RNA transcription, which was inhibited greater than 90% at 1 microgram/ml. In addition, U6 RNA synthesized in isolated nuclei was efficiently immunoprecipitated by an antibody against the La antigen, a protein known to bind most other RNA polymerase III transcripts. These results establish that, in contrast to the polymerase II-directed transcription of mammalian genes for U1-U5 small nuclear RNAs, human U6 RNA is transcribed by RNA polymerase III.

Amanitins↗

Identification of specific nucleotide sequences within the conserved 3'-SL in the dengue type 2 virus genome required for replication.

The flavivirus genome is a positive-stranded approximately 11-kb RNA including 5' and 3' noncoding regions (NCR) of approximately 100 and 400 to 600 nucleotides (nt), respectively. The 3' NCR contains adjacent, thermodynamically stable, conserved short and long stem-and-loop structures (the 3'-SL), formed by the 3'-terminal approximately 100 nt. The nucleotide sequences within the 3'-SL are not well conserved among species. We examined the requirement for the 3'-SL in the context of dengue virus type 2 (DEN2) replication by mutagenesis of an infectious cDNA copy of a DEN2 genome. Genomic full-length RNA was transcribed in vitro and used to transfect monkey kidney cells. A substitution mutation, in which the 3'-terminal 93 nt constituting the wild-type (wt) DEN2 3'-SL sequence were replaced by the 96-nt sequence of the West Nile virus (WN) 3'-SL, was sublethal for virus replication. An analysis of the growth phenotypes of additional mutant viruses derived from RNAs containing DEN2-WN chimeric 3'-SL structures suggested that the wt DEN2 nucleotide sequence forming the bottom half of the long stem and loop in the 3'-SL was required for viability. One 7-bp substitution mutation in this domain resulted in a mutant virus that grew well in monkey kidney cells but was severely restricted in cultured mosquito cells. In contrast, transpositions of and/or substitutions in the wt DEN2 nucleotide sequence in the top half of the long stem and in the short stem and loop were relatively well tolerated, provided the stem-loop secondary structure was conserved.

Animals↗

Identification of eukaryotic mRNAs that are translated at reduced cap binding complex eIF4F concentrations using a cDNA microarray.

Although most eukaryotic mRNAs need a functional cap binding complex eIF4F for efficient 5' end- dependent scanning to initiate translation, picornaviral, hepatitis C viral, and a few cellular RNAs have been shown to be translated by internal ribosome entry, a mechanism that can operate in the presence of low levels of functional eIF4F. To identify cellular mRNAs that can be translated when eIF4F is depleted or in low abundance and that, therefore, may contain internal ribosome entry sites, mRNAs that remained associated with polysomes were isolated from human cells after infection with poliovirus and were identified by using a cDNA microarray. Approximately 200 of the 7000 mRNAs analyzed remained associated with polysomes under these conditions. Among the gene products encoded by these polysome-associated mRNAs were immediate-early transcription factors, kinases, and phosphatases of the mitogen-activated protein kinase pathways and several protooncogenes, including c-myc and Pim-1. In addition, the mRNA encoding Cyr61, a secreted factor that can promote angiogenesis and tumor growth, was selectively mobilized into polysomes when eIF4F concentrations were reduced, although its overall abundance changed only slightly. Subsequent tests confirmed the presence of internal ribosome entry sites in the 5' noncoding regions of both Cyr61 and Pim-1 mRNAs. Overall, this study suggests that diverse mRNAs whose gene products have been implicated in a variety of stress responses, including inflammation, angiogenesis, and the response to serum, can use translational initiation mechanisms that require little or no intact cap binding protein complex eIF4F.

Cysteine-Rich Protein 61↗

bic, a novel gene activated by proviral insertions in avian leukosis virus-induced lymphomas, is likely to function through its noncoding RNA.

The bic locus is a common retroviral integration site in avian leukosis virus (ALV)-induced B-cell lymphomas originally identified by infection of chickens with ALVs of two different subgroups (Clurman and Hayward, Mol. Cell. Biol. 9:2657-2664, 1989). Based on its frequent association with c-myc activation and its preferential activation in metastatic tumors, the bic locus is thought to harbor a gene that can collaborate with c-myc in lymphomagenesis and presumably plays a role in late stages of tumor progression. In the present study, we have cloned and characterized two novel genes, bdw and bic, at the bic locus. bdw encoded a putative novel protein of 345 amino acids. However, its expression did not appear to be altered in tumor tissues, suggesting that it is not involved in oncogenesis. The bic gene consisted of two exons and was expressed as two spliced and alternatively polyadenylated transcripts at low levels in lymphoid/hematopoietic tissues. In tumors harboring bic integrations, proviruses drove bic gene expression by promoter insertion, resulting in high levels of expression of a chimeric RNA containing bic exon 2. Interestingly, bic lacked an extensive open reading frame, implying that it may function through its RNA. Computer analysis of RNA from small exon 2 of bic predicted extensive double-stranded structures, including a highly ordered RNA duplex between nucleotides 316 and 461. The possible role of bic in cell growth and differentiation is discussed in view of the emerging evidence that untranslated RNAs play a role in growth control.

Amino Acid Sequence↗

Transport and localization elements in myelin basic protein mRNA.

Myelin basic protein (MBP) mRNA is localized to myelin produced by oligodendrocytes of the central nervous system. MBP mRNA microinjected into oligodendrocytes in primary culture is assembled into granules in the perikaryon, transported along the processes, and localized to the myelin compartment. In this work, microinjection of various deleted and chimeric RNAs was used to delineate regions in MBP mRNA that are required for transport and localization in oligodendrocytes. The results indicate that transport requires a 21-nucleotide sequence, termed the RNA transport signal (RTS), in the 3' UTR of MBP mRNA. Homologous sequences are present in several other localized mRNAs, suggesting that the RTS represents a general transport signal in a variety of different cell types. Insertion of the RTS from MBP mRNA into nontransported mRNAs, causes the RNA to be transported to the oligodendrocyte processes. Localization of mRNA to the myelin compartment requires an additional element, termed the RNA localization region (RLR), contained between nucleotide 1,130 and 1, 473 in the 3' UTR of MBP mRNA. Computer analysis predicts that this region contains a stable secondary structure. If the coding region of the mRNA is deleted, the RLR is no longer required for localization, and the region between nucleotide 667 and 953, containing the RTS, is sufficient for both RNA transport and localization. Thus, localization of coding RNA is RLR dependent, and localization of noncoding RNA is RLR independent, suggesting that they are localized by different pathways.

Actins↗

Characterization of the Hantaan nucleocapsid protein-ribonucleic acid interaction.

The nucleocapsid (N) protein functions in hantavirus replication through its interactions with the viral genomic and antigenomic RNAs. To address the biological functions of the N protein, it was critical to first define this binding interaction. The dissociation constant, K(d), for the interaction of the Hantaan virus (HTNV) N protein and its genomic S segment (vRNA) was measured under several solution conditions. Overall, increasing the NaCl and Mg(2+) in these binding reactions had little impact on the K(d). However, the HTNV N protein showed an enhanced specificity for HTNV vRNA as compared with the S segment open reading frame RNA or a nonviral RNA with increasing ionic strength and the presence of Mg(2+). In contrast, the assembly of Sin Nombre virus N protein-HTNV vRNA complexes was inhibited by the presence of Mg(2+) or an increase in the ionic strength. The K(d) values for HTNV and Sin Nombre virus N proteins were nearly identical for the S segment open reading frame RNA, showing weak affinity over several binding reaction conditions. Our data suggest a model in which specific recognition of the HTNV vRNA by the HTNV N protein resides in the noncoding regions of the HTNV vRNA.

Capsid↗

Mechanism of attenuation of a chimeric influenza A/B transfectant virus.

The ribonucleoprotein transfection system for influenza virus allowed us to construct an influenza A virus containing a chimeric neuraminidase (NA) gene in which the noncoding sequence is derived from the NS gene of influenza B virus (T. Muster, E. K. Subbarao, M. Enami, B. P. Murphy, and P. Palese, Proc. Natl. Acad. Sci. USA 88:5177-5181, 1991). This transfectant virus is attenuated in mice and grows to lower titers in tissue culture than wild-type virus. Since such a virus has characteristics desirable for a live attenuated vaccine strain, attempts were made to characterize this virus at the molecular level. Our analysis suggests that the attenuation of the virus is due to changes in the cis signal sequences, which resulted in a reduction of transcription and replication of the chimeric NA gene. The major finding concerns a sixfold reduction in NA-specific viral RNA in the virion, causing a reduction in the ratio of infectious particles to physical particles compared with the ratio in wild-type virus. Although the NA-specific mRNA level is also reduced in transfectant virus-infected cells, it does not appear to contribute to the attenuation characteristics of the virus. The levels of the other RNAs and their expression appear to be unchanged for the transfectant virus. It is suggested that downregulation of the synthesis of one viral RNA segment leads to the generation of defective viruses during each replication cycle. We believe that this represents a general principle for attenuation which may be applied to other segmented viruses containing either single-stranded or double-stranded RNA.

Animals↗

Replication of subgenomic hepatitis A virus RNAs expressing firefly luciferase is enhanced by mutations associated with adaptation of virus to growth in cultured cells.

Replication of hepatitis A virus (HAV) in cultured cells is inefficient and difficult to study due to its protracted and generally noncytopathic cycle. To gain a better understanding of the mechanisms involved, we constructed a subgenomic HAV replicon by replacing most of the P1 capsid-coding sequence from an infectious cDNA copy of the cell culture-adapted HM175/18f virus genome with sequence encoding firefly luciferase. Replication of this RNA in transfected Huh-7 cells (derived from a human hepatocellular carcinoma) led to increased expression of luciferase relative to that in cells transfected with similar RNA transcripts containing a lethal premature termination mutation in 3D(pol) (RNA polymerase). However, replication could not be confirmed in either FrhK4 cells or BSC-1 cells, cells that are typically used for propagation of HAV. Replication was substantially slower than that observed with replicons derived from other picornaviruses, as the basal luciferase activity produced by translation of input RNA did not begin to increase until 24 to 48 h after transfection. Replication of the RNA was reversibly inhibited by guanidine. The inclusion of VP4 sequence downstream of the viral internal ribosomal entry site had no effect on the basal level of luciferase or subsequent increases in luciferase related to its amplification. Thus, in this system this sequence does not contribute to viral translation or replication, as suggested previously. Amplification of the replicon RNA was profoundly enhanced by the inclusion of P2 (but not 5' noncoding sequence or P3) segment mutations associated with adaptation of wild-type virus to growth in cell culture. These results provide a simple reporter system for monitoring the translation and replication of HAV RNA and show that critical mutations that enhance the growth of virus in cultured cells do so by promoting replication of viral RNA in the absence of encapsidation, packaging, and cellular export of the viral genome.

5' Untranslated Regions↗

The importance of a single G in the hairpin loop of the iron responsive element (IRE) in ferritin mRNA for structure: an NMR spectroscopy study.

Noncoding sequences regulate the function of mRNA and DNA. In animal mRNAs, iron responsive elements (IREs) regulate the synthesis of proteins for iron storage, uptake and red cell heme formation. Folding of the IRE was indicated previously by reactivity with chemical and enzymatic probes. 1H- and 31P-NMR spectra now confirm the IRE folding; an atypical 31P-spectrum, differential accessibility of imino protons to solvents, multiple long-range NOEs and heat stable subdomains were observed. Biphasic hyperchromic transitions occurred (52 and 73 degrees C). A G-C base pair occurs in the hairpin loop (HL) (based on dimethylsulfate, RNAse T1 previously used, and changes in NMR imino proton resonances typical of G-C base pairs after G/A substitution). Mutation of the hairpin loop also decreased temperature stability and changed the 31P-NMR spectrum; regulation and protein (IRP) binding were previously shown to change. Alteration of IRE structure shown by NMR spectroscopy, occurred at temperatures used in studies of IRE function, explaining loss of IRP binding. The effect of the HL mutation on the IRE emphasizes the importance of HL structure in other mRNAs, viral RNAs (e.g. HIV-TAR), and ribozymes.

Animals↗

Attenuation stem-loop lesions in the 5' noncoding region of poliovirus RNA: neuronal cell-specific translation defects.

The nucleotide at position 480 in the 5' noncoding region of the viral RNA genome plays an important role in directing the attenuation phenotype of the Sabin vaccine strain of poliovirus type 1. In vitro translation studies have shown that the attenuated viral genomes of the Sabin strains direct levels of viral protein synthesis lower than those of their neurovirulent counterparts. We previously described the isolation of pseudorevertant polioviruses derived from transfections of HeLa cells with genome-length RNA harboring an eight-nucleotide lesion in a stem-loop structure (stem-loop V) that contains the attenuation determinant at position 480 (A. A. Haller and B. L. Semler, J. Virol. 66:5075-5086, 1992). This stem-loop structure is a major component of the poliovirus internal ribosome entry site required for initiation of viral protein synthesis. The eight-nucleotide lesion (X472) was lethal for virus growth and gave rise only to viruses which had partially reverted nucleotides within the original substituted sequences. In this study, we analyzed two of the poliovirus revertants (X472RI and X472R2) for cell-type-specific growth properties. The X472RI and X472R2 RNA templates directed protein synthesis to wild-type levels in in vitro translation reaction mixtures supplemented with crude cytoplasmic HeLa cell extracts. In contrast, the same X472 revertant RNAs displayed a decreased translation initiation efficiency when translated in a cell-free system supplemented with extracts from neuronal cells. This translation initiation defect of the X472R templates correlated with reduced yields of infectious virus particles in neuronal cells compared with those obtained from HeLa cells infected with the X472 poliovirus revertants. Our results underscore the important of RNA secondary structures within the poliovirus internal ribosome entry site in directing translation initiation and suggest that such structures interact with neuronal cell factors in a specific manner.

Base Sequence↗

Expression regulation network in papillae of sea cucumbers: Whole-transcriptome and DNA methylation datasets.

To elucidate the expression regulation network of papilla size of sea cucumbers (Apostichopus japonicus), the whole-transcriptome and DNA methylome datasets of different sizes of papillae in sea cucumbers were generated. Average clean bases of whole-transcriptome (16.35 G) and DNA methylome (28.92 G) were obtained using RNA sequencing and whole-genome bisulfite sequencing techniques. A total of 3,188 ceRNA networks were also identified including 3,081 long non-coding RNAs (lncRNA)/microRNAs (miRNA)/mRNA networks and 107 circular RNA (circRNA)/miRNA/mRNA networks. Methylome data indicate that there were 3,307 and 3,776 differentially methylated regions (DMRs) with high-level methylation as well as 3,125 and 3,016 DMRs with low-level methylation in big papillae compared to small papillae. The identified DMRs were mainly distributed in introns, promotors, or exons. The whole-transcriptome and DNA methylome datasets generated from this study not only established a robust theoretical foundation (especially from the epigenetic aspect) for elucidating expression regulation network determining papilla size in sea cucumbers but also can be a valuable resource of biomarker mining for papilla appearance-based selective breeding in sea cucumbers.

DNA Methylation↗

Association of H19 promoter methylation with the expression of H19 and IGF-II genes in adrenocortical tumors.

Low H19 and abundant IGF-II expression may have a role in the development of adrenocortical carcinomas. In the mouse, the H19 promoter area has been found to be methylated when transcription of the H19 gene is silent and unmethylated when it is active. We used PCR-based methylation analysis and bisulfite genomic sequencing to study the cytosine methylation status of the H19 promoter region in 16 normal adrenals and 30 pathological adrenocortical samples. PCR-based analysis showed higher methylation status at three HpaII-cutting CpG sites of the H19 promoter in adrenocortical carcinomas and in a virilizing adenoma than in their adjacent normal adrenal tissues. Bisulfite genomic sequencing revealed a significantly higher mean degree of methylation at each of 12 CpG sites of the H19 promoter in adrenocortical carcinomas than in normal adrenals (P < 0.01 for all sites) or adrenocortical adenomas (P < 0.01, except P < 0.05 for site 12 and P > 0.05 for site 11). The mean methylation degree of the 12 CpG sites was significantly higher in the adrenocortical carcinomas (mean +/- SE, 76 +/- 7%) than in normal adrenals (41 +/- 2%) or adrenocortical adenomas (45 +/- 3%; both P < 0.005). RNA analysis indicated that the adrenocortical carcinomas expressed less H19 but more IGF-II RNAs than normal adrenal tissues did. The mean methylation degree of the 12 H19 promoter CpG sites correlated negatively with H19 RNA levels (r = -0.550; P < 0.01), but positively with IGF-II mRNA levels (r = 0.805; P < 0.001). In the adrenocortical carcinoma cell line NCI-H295R, abundant IGF-II, but minimal H19, RNA expression was detected by Northern blotting. Treatment with a cytosine methylation inhibitor, 5-aza-2'-deoxycytidine, increased H19 RNA expression, whereas it decreased IGF-II mRNA accumulation dose- and time-dependently (both P < 0.005) and reduced cell proliferation to 10% in 7 d. Our results suggest that altered DNA methylation of the H19 promoter is involved in the abnormal expression of both H19 and IGF-II genes in human adrenocortical carcinomas.

Adenoma↗

Identification and characterization of BC1 RNP particles.

Rodent brain-specific small cytoplasmic BC1 RNA is an unusual RNA in several respects. It is an RNA polymerase III transcript expressed specifically in neurons, with regional and developmental regulation. Moreover, it is one of a few RNAs actively transported into dendrites. Three findings indicate that BC1 RNA exists as a ribonucleoprotein complex in vivo. First, the buoyant density of fractions containing BC1 RNA from brain extract on CsCI and Cs2SO4 gradients is 1.45 g/ml and 1.55 g/ml, respectively; this is consistent with the density of RNA-protein complexes. Second, in sucrose gradients, the BC1 particle has a larger S value (8.7S) than naked RNA (6.1S). Third, BC1 RNA from brain extracts migrates with retarded mobility compared to naked BC1 RNA during agarose gel electrophoresis. Additionally, in comparison to the signal recognition particle (SRP), the BC1 RNP is more heat resistant and less Mg(2+)-dependent. The buoyant density of the BC1 RNP suggests the presence of protein(s) with a total mass of about 138kD.

Animals↗

Effect of TSIX disruption on XIST expression in male ES cells.

XIST and its antisense partner, TSIX, encode non-coding RNAs and play key roles in X chromosome inactivation. Targeted disruption of TSIX causes ectopic expression of XIST in the extraembryonic tissues upon maternal transmission, which subsequently results in embryonic lethality due to inactivation of both X chromosomes in females and a single X chromosome in males. TSIX, therefore, plays a crucial role in maintaining the silenced state of XIST in CIS and regulates the imprinted X inactivation in the extraembryonic tissues. In this study, we examined the effect of TSIX disruption on XIST expression in the embryonic lineage using embryonic stem (ES) cells as a model system. Upon differentiation, XIST is ectopically activated in a subset of the nuclei of male ES cells harboring the TSIX-deficient X chromosome. Such ectopic expression, however, eventually ceased during prolonged culture. It is likely that surveillance by the X chromosome counting mechanism somehow shuts off the ectopic expression of XIST before inactivation of the X chromosome.

Animals↗

Contrasting regulation of protein-coding genes and lncRNA homeologs in allotetraploid Coffea arabica.

A chromosome-level Bourbon assembly revealed that protein-coding homeologs are predominantly co-regulated between subgenomes. In contrast, intergenic lncRNAs display a modest, but statistically consistent bias toward subgenome E across diverse developmental and stress contexts. Coffea arabica is an allotetraploid species derived from natural hybridization between C. canephora and C. eugenioides, which contributed the C and E subgenomes, respectively. This genomic origin poses major challenges for genome assembly, annotation, and the interpretation of gene regulation. In this study, a high-quality genome assembly of C. arabica was generated and annotated, with particular emphasis on identifying protein-coding genes and intergenic long non-coding RNAs (lincRNAs). Homeologous relationships between genes from the C and E subgenomes were established, providing a robust framework to investigate subgenomic conservation and regulatory divergence. Using an extensive collection of publicly available RNA-seq libraries spanning multiple developmental stages, tissues, and environmental conditions, the relative transcriptional contribution of each subgenome was evaluated. On a global scale, gene expression was largely balanced between subgenomes, with no consistent evidence of subgenome dominance. While protein-coding genes showed comparable regulatory behavior across subgenomes, lincRNAs exhibited a more asymmetric expression pattern, suggesting higher subgenome-specific expression that is interpreted here as a consistent directional tendency rather than as evidence of subgenome dominance. Together, these results provide new insights into the regulatory architecture of the C. arabica genome and establish a foundational genomic and transcriptomic resource for future functional studies and crop improvement efforts.

Coffea↗