Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Cell proteins bind to sites within the 3' noncoding region and the positive-strand leader sequence of measles virus RNA.

The genomic 3' noncoding region (NCR) of nonsegmented negative-strand RNA viruses contains recognition site(s) for the polymerase complex, while the RNA plus-strand leader sequence (LS) is probably involved in RNA encapsidation. It is known that host-encoded factors play a role in transcription and replication of some of this group of viruses. Here we report that cellular proteins interact with the genomic 3' NCR and with the plus-strand LS RNA of an important human pathogen, measles virus (MV), a member of the family Paramyxoviridae. Using gel retardation assay and RNA footprinting analysis, we demonstrated that in Vero cells, host-encoded proteins bind specifically to domains within these two sequences. A polypeptide of about 20 kDa binding to the 3' NCR and two polypeptides of about 22 and 30 kDa interacting with plus-strand LS were detected by RNA-protein UV cross-linking. Different RNA-binding activities were found in cells differing in permissiveness to MV replication. The results suggest a role for host-encoded proteins in MV replication.

Animals↗

Cell-dependent role for the poliovirus 3' noncoding region in positive-strand RNA synthesis.

We previously reported the isolation of a mutant poliovirus lacking the entire genomic RNA 3' noncoding region. Infection of HeLa cell monolayers with this deletion mutant revealed only a minor defect in the levels of viral RNA replication. To further analyze the consequences of the genomic 3' noncoding region deletion, we examined viral RNA replication in a neuroblastoma cell line, SK-N-SH cells. The minor genomic RNA replication defect in HeLa cells was significantly exacerbated in the SK-N-SH cells, resulting in a decreased capacity for mutant virus growth. Analysis of the nature of the RNA replication deficiency revealed that deleting the poliovirus genomic 3' noncoding region resulted in a positive-strand RNA synthesis defect. The RNA replication deficiency in SK-N-SH cells was not due to a major defect in viral translation or viral protein processing. Neurovirulence of the mutant virus was determined in a transgenic mouse line expressing the human poliovirus receptor. Greater than 1,000 times more mutant virus was required to paralyze 50% of inoculated mice, compared to that with wild-type virus. These data suggest that, together with a cellular factor(s) that is limiting in neuronal cells, the poliovirus 3' noncoding region is involved in positive-strand synthesis during genome replication.

3' Untranslated Regions↗

Segment-specific and common nucleotide sequences in the noncoding regions of influenza B virus genome RNAs.

The nucleotide sequences of the 3' noncoding regions of all eight segments of influenza B virus RNA and the sequences of the 5' noncoding regions of segments 4-8 were determined in virus strains isolated over a period of 40 years. Nearly complete conservation of the noncoding sequences was found. Nine nucleotides at the 3' termini and 11 nucleotides at the 5' termini were common to all segments examined. In the region immediately adjacent to the common 3' terminal region, the nucleotides were specific for each segment and these segment-specific sequences were conserved in all strains examined. In each of the five segments in which both termini were examined, the segment-specific 3' sequences exhibited perfect inverted complementarity to a segment-specific sequence adjacent to the common 5' terminus. In addition, in the 3' noncoding region of RNA segments 1-3, which encode proteins involved in RNA synthesis, a single nucleotide substitution at position 10 was found that distinguishes these segments from segments 4-8. Comparison of these data with published reports has revealed that some of the features found in the noncoding regions of influenza B virus are also present in influenza A and C virus RNAs. In the RNAs of all three virus types, there is a segment-specific sequence of nucleotides near the 3' terminus that shows inverted complementarity to a sequence near the 5' terminus. This segment-specific sequence may play a role in the transcription of individual segments or in sorting of segments during virion assembly.

Animals↗

Molecular evolution of noncoding regions of the chloroplast genome in the Crassulaceae and related species.

Universal primers were used for PCR amplification of three noncoding regions of chloroplast DNA (cpDNA) in order to study sequence-length variation in the Crassulaceae and in related species. Several length mutations were observed that are of diagnostic value for evolutionary relationships in the Crassulaceae and the Saxifragaceae. Length variation and sequence divergence in the intergenic spacer between the trnL (UAA) 3' exon and the trnF (GAA) gene among 15 species were studied in detail by nucleotide-sequence analysis. A total of 50 insertion/deletion mutations were observed, accounting for a spacer-length variation in the range of 228-360 bp. Eighteen short direct repeat motifs (4-11 bp) and two inverted repeat motifs (7-11 bp) were found to be associated with length variation. Phylogenetic analysis of the sequence data indicated a pattern of relationships that was largely consistent with a previous analysis of cpDNA restriction-site variation. Evaluation of the level of homoplasy in insertion/deletion mutations within a phylogenetic framework revealed that only 1 out of 34 length mutations longer than 2 bp must have had multiple origins. The feasibility of the noncoding chloroplast DNA regions for molecular evolutionary studies is discussed.

Base Sequence↗

Compositional evolution of noncoding DNA in the human and chimpanzee genomes.

We have examined the compositional evolution of noncoding DNA in the primate genome by comparison of lineage-specific substitutions observed in 1.8 Mb of genomic alignments of human, chimpanzee, and baboon with 6542 human single-nucleotide polymorphisms (SNPs) rooted using chimpanzee sequence. The pattern of compositional evolution, measured in terms of the numbers of GC-->AT and AT-->GC changes, differs significantly between fixed and polymorphic sites, and indicates that there is a bias toward fixation of AT-->GC mutations, which could result from weak directional selection or biased gene conversion in favor of high GC content. Comparison of the frequency distributions of a subset of the SNPs revealed no significant difference between GC-->AT and AT-->GC polymorphisms, although AT-->GC polymorphisms in regions of high GC segregate at slightly higher frequencies on average than GC-->AT polymorphisms, which is consistent with a fixation bias favoring high GC in these regions. However, the substitution data suggest that this fixation bias is relatively weak, because the compositional structure of the human and chimpanzee genomes is becoming homogenized, with regions of high GC decreasing in GC content and regions of low GC increasing in GC content. The rate and pattern of nucleotide substitution in 333 Alu repeats within the human-chimpanzee-baboon alignments are not significantly affected by the GC content of the region in which they are inserted, providing further evidence that, since the time of the human-chimpanzee ancestor, there has been little or no regional variation in mutation bias.

Alleles↗

DNA sequence alterations affect nucleosome array formation of the chicken ovalbumin gene.

The role of the large amount (more than half of the genome) of noncoding DNA in higher organisms is not well understood. DNA evolved to function in the context of chromatin, and the possibility exists that some of the noncoding DNA serves to influence chromatin structure and function. In this age of genomics and bioinformatics, genomic DNA sequences are being searched for informational content beyond the known genetic code. The discovery that period-10 non-T, A/T, G (VWG) triplets are among the most abundant motifs in human genomic DNA suggests that they may serve some function in higher organisms. In this paper, we provide direct evidence that the regular oscillation of period-10 VWG that occurs in the chicken ovalbumin gene sequence with a dinucleosome-like period facilitates nucleosome array formation. Using a linker histone-dependent in vitro chromatin assembly system that spontaneously aligns nucleosomes into a physiological array, we show that nucleosomes tend to avoid DNA regions with low period-10 VWG counts. This avoidance leads to the formation of an array with a nucleosome repeat equal to half the period value of the oscillation in period-10 VWG, as determined by Fourier analysis. Two different half-period deletions in the wild-type DNA sequence altered the nucleosome array, as predicted computationally. In contrast, a full-period deletion had an insignificant effect on the nucleosome array formed, also consistent with the prediction. An inversion mutation, with no DNA sequences deleted, again altered the nucleosome array formed, as predicted computationally. Hence, a VWG dinucleosome signal is plausible.

Animals↗

SINEs and LINEs: the art of biting the hand that feeds you.

SINEs and LINEs are short and long interspersed retrotransposable elements, respectively, that invade new genomic sites using RNA intermediates. SINEs and LINEs are found in almost all eukaryotes (although not in Saccharomyces cerevisiae) and together account for at least 34% of the human genome. The noncoding SINEs depend on reverse transcriptase and endonuclease functions encoded by partner LINEs. With the completion of many genome sequences, including our own, the database of SINEs and LINEs has taken a great leap forward. The new data pose new questions that can only be answered by detailed studies of the mechanism of retroposition. Current work ranges from the biochemistry of reverse transcription and integration invitro, target site selection in vivo, nucleocytoplasmic transport of the RNA and ribonucleoprotein intermediates, and mechanisms of genomic turnover. Two particularly exciting new ideas are that SINEs may help cells survive physiological stress, and that the evolution of SINEs and LINEs has been shaped by the forces of RNA interference. Taken together, these studies promise to explain the birth and death of SINEs and LINEs, and the contribution of these repetitive sequence families to the evolution of genomes.

Animals↗

Antisense RNA in imprinting: spreading silence through Air.

In some animals, including mammals, a number of genes are expressed differently according to whether they have been inherited from the mother or from the father, through a process known as genomic imprinting. Noncoding RNAs have increasingly been found associated with imprinted genes, but their role, if any, has remained enigmatic. A recent study provides the first evidence that, at least in one case, a noncoding RNA has a direct role in regulating imprinted gene expression in cis.

Animals↗

Prospects for identifying functional variation across the genome.

The genetic factors contributing to complex trait variation may reside in regulatory, rather than protein-coding portions of the genome. Within noncoding regions, SNPs in regulatory elements are more likely to contribute to phenotypic variation than those in nonregulatory regions. Thus, it is important to be able to identify and annotate noncoding regulatory elements. DNA conservation among diverged species successfully identifies noncoding regulatory regions. However, because rapidly evolving regulatory regions will not generally be conserved across species, these will not detected by using purely conservation-based methods. Here we describe additional approaches that can be used to identify putative regulatory elements via signatures of nonneutral evolution. An examination of the pattern of polymorphism both within and between populations of Drosophila melanogaster, as well as divergence with its sibling species Drosophila simulans, across 24.2 kb of noncoding DNA identifies several nonneutrally evolving regions not identified by conservation. Because different methods tag different regions, it appears that the methods are complementary. Patterns of variation at different elements are consistent with the action of selective sweeps, balancing selection, or population differentiation. Together with regions conserved between D. melanogaster and Drosophila pseudoobscura, we tag 5.3 kb of noncoding DNA as potentially regulatory. Ninety-seven of the 408 common noncoding SNPs surveyed are within putatively regulatory regions. If these methods collectively identify the majority of functional noncoding polymorphisms, genotyping only these SNPs in an association mapping framework would reduce genotyping effort for noncoding regions 4-fold.

Animals↗

Pseudogenes, junk DNA, and the dynamics of Rickettsia genomes.

Studies of neutrally evolving sequences suggest that differences in eukaryotic genome sizes result from different rates of DNA loss. However, very few pseudogenes have been identified in microbial species, and the processes whereby genes and genomes deteriorate in bacteria remain largely unresolved. The typhus-causing agent, Rickettsia prowazekii, is exceptional in that as much as 24% of its 1.1-Mb genome consists of noncoding DNA and pseudogenes. To test the hypothesis that the noncoding DNA in the R. prowazekii genome represents degraded remnants of ancestral genes, we systematically examined all of the identified pseudogenes and their flanking sequences in three additional Rickettsia species. Consistent with the hypothesis, we observe sequence similarities between genes and pseudogenes in one species and intergenic DNA in another species. We show that the frequencies and average sizes of deletions are larger than insertions in neutrally evolving pseudogene sequences. Our results suggest that inactivated genetic material in the Rickettsia genomes deteriorates spontaneously due to a mutation bias for deletions and that the noncoding sequences represent DNA in the final stages of this degenerative process.

Base Sequence↗

How does replication-associated mutational pressure influence amino acid composition of proteins?

We have performed detrended DNA walks on whole prokaryotic genomes, on noncoding sequences and, separately, on each position in codons of coding sequences. Our method enables us to distinguish between the mutational pressure associated with replication and the mutational pressure associated with transcription and other mechanisms that introduce asymmetry into prokaryotic chromosomes. In many prokaryotic genomes, each component of mutational pressure affects coding sequences not only in silent positions but also in positions in which changes cause amino acid substitutions in coded proteins. Asymmetry in the silent positions of codons differentiates the rate of translation of mRNA produced from leading and lagging strands. Asymmetry in the amino acid composition of proteins resulting from replication-associated mutational pressure also corresponds to leading and lagging roles of DNA strands, whereas asymmetry connected with transcription and coding function corresponds to the distance of genes from the origin or terminus of chromosome replication.

Amino Acid Sequence↗

The complete mitochondrial DNA sequence of the horseshoe crab Limulus polyphemus.

We determined the complete 14,985-nt sequence of the mitochondrial DNA of the horseshoe crab Limulus polyphemus (Arthropoda: Xiphosura). This mtDNA encodes the 13 protein, 2 rRNA, and 22 tRNA genes typical for metazoans. The arrangement of these genes and about half of the sequence was reported previously; however, the sequence contained a large number of errors, which are corrected here. The two strands of Limulus mtDNA have significantly different nucleotide compositions. The strand encoding most mitochondrial proteins has 1. 25 times as many A's as T's and 2.33 times as many C's as G's. This nucleotide bias correlates with the biases in amino acid content and synonymous codon usage in proteins encoded by different strands and with the number of non-Watson-Crick base pairs in the stem regions of encoded tRNAs. The sizes of most mitochondrial protein genes in Limulus are either identical to or slightly smaller than those of their Drosophila counterparts. The usage of the initiation and termination codons in these genes seems to follow patterns that are conserved among most arthropod and some other metazoan mitochondrial genomes. The noncoding region of Limulus mtDNA contains a potential stem-loop structure, and we found a similar structure in the noncoding region of the published mtDNA of the prostriate tick Ixodes hexagonus. A simulation study was designed to evaluate the significance of these secondary structures; it revealed that they are statistically significant. No significant, comparable structure can be identified for the metastriate ticks Rhipicephalus sanguineus and Boophilus microplus. The latter two animals also share a mitochondrial gene rearrangement and an unusual structure of mt-tRNA(C) that is exactly the same association of changes as previously reported for a group of lizards. This suggests that the changes observed are not independent and that the stem-loop structure found in the noncoding regions of Limulus and Ixodes mtDNA may play the same role as that between trnN and trnC in vertebrates, i.e., the role of lagging strand origin of replication.

Animals↗

Contrasting regulation of protein-coding genes and lncRNA homeologs in allotetraploid Coffea arabica.

A chromosome-level Bourbon assembly revealed that protein-coding homeologs are predominantly co-regulated between subgenomes. In contrast, intergenic lncRNAs display a modest, but statistically consistent bias toward subgenome E across diverse developmental and stress contexts. Coffea arabica is an allotetraploid species derived from natural hybridization between C. canephora and C. eugenioides, which contributed the C and E subgenomes, respectively. This genomic origin poses major challenges for genome assembly, annotation, and the interpretation of gene regulation. In this study, a high-quality genome assembly of C. arabica was generated and annotated, with particular emphasis on identifying protein-coding genes and intergenic long non-coding RNAs (lincRNAs). Homeologous relationships between genes from the C and E subgenomes were established, providing a robust framework to investigate subgenomic conservation and regulatory divergence. Using an extensive collection of publicly available RNA-seq libraries spanning multiple developmental stages, tissues, and environmental conditions, the relative transcriptional contribution of each subgenome was evaluated. On a global scale, gene expression was largely balanced between subgenomes, with no consistent evidence of subgenome dominance. While protein-coding genes showed comparable regulatory behavior across subgenomes, lincRNAs exhibited a more asymmetric expression pattern, suggesting higher subgenome-specific expression that is interpreted here as a consistent directional tendency rather than as evidence of subgenome dominance. Together, these results provide new insights into the regulatory architecture of the C. arabica genome and establish a foundational genomic and transcriptomic resource for future functional studies and crop improvement efforts.

Coffea↗

Genome imbalance modulates the expression of long non-coding RNAs in maize.

Genome imbalance, resulting from varying the dosage of individual chromosomes (aneuploidy), has a more detrimental effect than changes in complete sets of chromosomes (haploidy/polyploidy). This imbalance is likely due to disruptions in stoichiometry and interactions among macromolecular assemblies. Previous research has shown that aneuploidy causes global modulation of protein-coding genes (PCGs), microRNAs, and transposable elements (TEs), affecting both the varied chromosome (cis-located) and unvaried genome regions (trans-located) across various taxa. While long non-coding RNAs (lncRNAs) are important gene expression regulators, their roles in the context of genomic imbalance remain largely unexplored. In this study, we analyzed and compared the impact of aneuploidy and haploidy/polyploidy on lncRNA expression using RNA-seq data from maize mature leaf tissue. Our results indicate that cis-located lncRNAs are modulated from dosage compensation to a gene dosage effect, while trans-located lncRNAs exhibit trends ranging from an inverse effect to a positive correlation with chromosomal dosage. Remarkably, the ploidy series showed a lesser degree of lncRNA modulation. LncRNAs and TEs display a similar trend of inverse modulation but exhibit greater sensitivity to dosage changes compared to PCGs. The construction of cis-acting and trans-acting lncRNA co-expression networks indicates that lncRNAs likely function as dosage-sensitive regulators of gene expression under conditions of genomic imbalance. Overall, this study not only elucidates the dosage effect of plant lncRNAs but also serves as a valuable resource for exploring potential regulators of PCGs that play significant biological functions.

Zea mays↗

Filaggrin, an intermediate filament-associated protein: structural and functional implications from the sequence of a cDNA from rat.

Filaggrin is an intermediate filament-associated protein that is involved in aggregation of keratin filaments in fully cornified cells of the mammalian epidermis, and is an important marker for epidermal differentiation. In this report, the sequence of a rat cDNA clone coding for a portion of the polymeric precursor, profilaggrin, is presented. The cDNA is 2,314 bp long with 1,875 bp of coding region ending with an A-T-rich 3' noncoding region. Genomic analysis indicates that the profilaggrin gene consists of 20 +/- 2 repeats of 1,218 bp of sequence coding for 406 amino acids, making the mRNA at least 25-27 kb in length. Each repeat consists of a filaggrin domain and a linker sequence with an estimated size of 380 and 26 amino acids, respectively. High levels of profilaggrin mRNA are found only in keratinizing epithelia. Comparison of the rat filaggrin sequence with that of mouse and human filaggrin and with the sequence of phosphorylated peptides from mouse profilaggrin indicates that the proteins share extensive amino acid sequence similarities, especially in the two phosphorylated regions. Proteolytic processing sites are also quite similar in rat and mouse. The three species show blocks of sequence that are similar in length and composition which alternate with sequences that are variable in length. This analysis suggests that the evolution of the present-day filaggrins has been constrained by maintenance of phosphorylation sites and overall amino acid composition. The cDNAs for the profilaggrins are similar in structure, reflecting genes that have simple repeating structures and lack introns within their coding regions. Mouse and rat profilaggrin terminate with a nonpolar sequence atypical of the rest of the coding region, and have similar 3' noncoding regions. To explain these observations, a novel evolutionary model is proposed.

Amino Acid Sequence↗

Asymmetry of coding versus noncoding strand in coding sequences of different genomes.

We have used the asymmetry between the coding and noncoding strands in different codon positions of coding sequences of DNA as a parameter to evaluate the coding probability for open reading frames (ORFs). The method enables an approximation of the total number of coding ORFs in the set of analyzed sequences as well as an estimation of the coding probability for the ORFs. The asymmetry observed in the nucleotide composition of codons in coding sequences has been used successfully for analysis of the genomes completed at the time of this analysis.

Codon↗

A Programmable Nanovaccine Platform Based on M13 Bacteriophage for Personalized Cancer Vaccine and Therapy.

Nanovaccines co-assemble antigens and adjuvants to elicit robust immune responses but often require complex synthesis and post-modification procedures. Here, a programmable nanovaccine platform based on the M13 bacteriophage is developed for the scalable production of vaccines and single-step modular engineering of adjuvanticity, length, and antigen density. By reprogramming the sequence and size of the noncoding phage genome, the Toll-like receptor 9 activation and the length of the phage are precisely controlled. With a novel molecular engineering approach, the antigen density is tuned from 13.6% to 70.3%. A systematic modulation reveals an optimal adjuvanticity at a constant antigen density for maximum anti-tumor CD8+ T cell response, and vice versa, using the model antigen SIINFEKL. The M13 phage-based nanovaccine induces durable memory immunity lasting over a year. In addition, a 24-fold increase in neoantigen-specific CD8+ T cell frequency is achieved when increasing both the adjuvanticity and antigen density. Furthermore, when combined with anti-PD-1 therapy, the M13 phage-based personalized vaccine eradicates established MC-38 tumors in 75% of treated animals and they develop 100% resistance against tumor invasion when challenged 5 months after treatment. These findings establish M13 phage as a powerful and versatile nanovaccine platform with transformative potential for personalized cancer immunotherapy.

Cancer Vaccines↗

Organization of the 5'-End of the porcine gamma-glutamyl transpeptidase gene and identification of three different mRNAs in the kidney.

Three different mRNAs coding for the porcine gamma-glutamyl transpeptidase (GGT) in the kidney were identified by 5'-RACE-PCR. These differ in their 5'-noncoding region. Genomic Southern blot analysis has demonstrated the existence of a single GGT gene in the porcine genome. Thus, the existence of multiple mRNAs can only be explained by the use of different promoters or alternative splicing. Four GGT-specific genomic clones containing the complete 5'-end of the gene were isolated and characterized, revealing six exons common to all three mRNAs. Four of these exons were located in the coding region comprising the codons for amino acids 1 to 138. Two exons and an intervening sequence were identified upstream from these six common exons representing the unique 5'-ends of the three mRNAs. The coding exons show a significant sequence homology to mouse, rat, and human GGT cDNA, whereas exons 1 and 3 display no homology.

Animals↗