Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Convergence of natural and artificial evolution on an RNA loop-loop interaction: the HIV-1 dimerization initiation site.

Loop-loop interactions among nucleic acids constitute an important form of molecular recognition in a variety of biological systems. In HIV-1, genomic dimerization involves an intermolecular RNA loop-loop interaction at the dimerization initiation site (DIS), a hairpin located in the 5' noncoding region that contains an autocomplementary sequence in the loop. Only two major DIS loop sequence variants are observed among natural viral isolates. To investigate sequence and structural constraints on genomic RNA dimerization as well as loop-loop interactions in general, we randomized several or all of the nucleotides in the DIS loop and selected in vitro for dimerization-competent sequences. Surprisingly, increasing interloop complementarity above a threshold of 6 bp did not enhance dimerization, although the combinations of nucleotides forming the theoretically most stable hexanucleotide duplexes were selected. Noncanonical interactions contributed significantly to the stability and/or specificity of the dimeric complexes as demonstrated by the overwhelming bias for noncanonical base pairs closing the loop and covariations between flanking and central loop nucleotides. Degeneration of the entire loop yielded a complex population of dimerization-competent sequences whose consensus sequence resembles that of wild-type HIV-1. We conclude from these findings that the DIS has evolved to satisfy simultaneous constraints for optimal dimerization affinity and the capacity for homodimerization. Furthermore, the most constrained features of the DIS identified by our experiments could be the basis for the rational design of DIS-targeted antiviral compounds.

Codon, Initiator↗

The 5'-terminal 32 basepairs conserved between genome segments A and B contain a major promoter element of infectious bursal disease virus.

The regions of the infectious bursal disease virus (IBDV) genome with regulatory function are not known. In the present study, progressively deleted lengths of the 5' noncoding region of segment A were constructed in pGL3 vectors having SV40 enhancer or promoter, and a luciferase (LUC) reporter gene. Transient transfections of the constructs made in a promoter-less pGL3-Enhancer vector when transfected in Vero cells and the lysates assayed for LUC expression, allowed the localization of maximal activity to the 32-nucleotide stretch (precursor polyprotein ORF positions -131 to -100), which is highly conserved at the 5' end of both genome segments. This fragment, when evaluated in parallel in an enhancer-less pGL3-Promoter vector demonstrated no activity. To determine if this region is recognized by IBDV replicative proteins, we engineered modifications in an enhancer-less pGL3-Promoter vector where the terminal 32-bp fragment, the full-length noncoding region, or the noncoding region with the 32-bp fragment deleted was positioned in either the plus-sense or the minus-sense orientation immediately downstream of the SV40 promoter and upstream of the LUC gene. Transfections of these constructs in IBDV-infected and uninfected Vero cells resulted in the endogenous generation of recombinant viral-LUC RNAs containing the 5' terminal viral RNA sequences in either the plus-sense or the minus-sense orientation. LUC assays of the infected cell lysates showed up-regulated expression of LUC only with constructs containing the 32-bp fragment in the minus-sense orientation. Deletion of this 32-bp fragment abolished such LUC expression. We therefore conclude that the 5'-terminal 32 base pairs of genomic segment A contain a major promoter element in IBDV. In addition, our results show that IBDV replicative proteins recognize and transcribe single-stranded RNA in vivo.

Animals↗

Canine oral papillomavirus genomic sequence: a unique 1.5-kb intervening sequence between the E2 and L2 open reading frames.

The canine oral papillomavirus (COPV) is associated with oropharyngeal papillomatosis in dogs, coyotes, and wolves. We have determined the complete nucleotide sequence of COPV, the largest of all known PV genomes (8607 bp). The genomic architecture of the COPV genome is similar to that of other PVs except for a unique and large noncoding region of 1.5 kb between the end of the early region (E2) and the beginning of the late region (L2) and a small (345 bp) upstream regulatory region between the end of L1 and the beginning of E6. Although COPV displays a primarily mucosal tropism, the COPV nucleotide sequence showed the highest overall similarity to cutaneous papillomaviruses such as HPV-1, HPV-63, CRPV (cottontail rabbit PV), FdPV (Felis domesticus PV), and MnPV (Mastomys natalensis PV).

Amino Acid Sequence↗

A source of small repeats in genomic DNA.

The processes of spontaneous mutation are known to be influenced by neighboring DNA. Imperfect nearby repeats in the neighboring DNA have been observed to mutate to form perfect repeats. The repeats may be either direct or inverted. Such a mutational process should create perfect direct and inverted repeats in intergenic DNA. A larger than expected number of direct repeats has generally been observed in a wide range of species in both coding and noncoding DNA. Simulations are carried out to determine how this process might influence the repetitive structure of genomic DNA. These simulations show that small repeats created by this kind of a mutational process can explain the excess number of repeats in intergenic DNA. The simulations suggest that this mechanism may be a common cause of mutations, including single-base changes. The influences of the distance between imperfect repeats and of their degree of similarity are investigated.

Animals↗

Genome characterization of a Korean isolate of cymbidium mosaic virus.

The complete nucleotide sequence of the genomic RNA of a Korean isolate of cymbidium mosaic virus (CymMV-K2) was determined. The genomic RNA is 6227 nucleotides in length, excluding the poly(A) tail. It contains a 5'-noncoding region (NCR) of 73 nucleotides, five open reading frames (ORFs 1 to 5) which encode proteins with M(r) 160 kDa RNA-dependent RNA polymerase (ORF1), 26 kDa movement protein 1 (ORF2), 13 kDa movement protein 2 (ORF3), 10 kDa movement protein 3 (ORF4), 24 kDa coat protein (OFR5), and a 3' NCR of 76 nucleotides. The 5'-end of the CymMV-K2 genome initiates with GGAAAA which contrasts to GAAAA at the 5'-ends of other potexviruses, including a Singapore isolate of CymMV (CymMV-S2). When compared with CymMV-S2, 171 base substitutions were observed in the CymMV-K2 genome. Substitutions in the overlapping ORFs (ORFs 2 to 4) occurred more frequently than those in 5' NCR, ORF1, and 3' NCR. In addition to substitutions, two single-base deletions, one in the intercistronic region between ORF1 and ORF2 and the other in the ORF2, were found on the CymMV-K2 genome. The deletion in the ORF2 induced a frameshift which altered the C-terminal domain of movement protein 1. ORF3 and ORF4 of the CymMV-K2 genome are partially different from those of another Singapore CymMV genome (CymMV-S1) which has four frameshifts due to nucleotide deletions within these ORFs. Interestingly, the frameshifts resulted in no change in the conserved sequences of the movement proteins but reconstructed their transmembrane domains.

Amino Acid Sequence↗

Structure and functional properties of prokaryotic small noncoding RNAs.

Most biochemical, computational and genetic approaches to gene finding assume the Central Dogma and look for genes that make mRNA and have ORFs. These approaches essentially do not work for one class of genes--the noncoding RNA. In all living organisms RNA is involved in a number of essential cell processes. Functional analysis of genome sequences has largely ignored RNA genes and their structures. Different RNA species including rRNA, tRNA, mRNA and sRNA (small RNA) are important structural, transfer, informational, and regulatory molecules containing complex folded conformations that participate in recognition and catalytic processes. Noncoding RNAs play an number of important structural, catalytic and regulatory roles in the cell. The size of the sRNA genes ranges from 70 to 500 nucleotides. Several transcripts of these genes are processed by RNAases and their final products are smaller. The encoding genes are localized between two ORFs and do not overlap with ORFs on the complementary DNA strand. As aptamers, some sRNA bind small molecular components (metal ions, peptides and nucleotides). This review summarizes recent data on the functions of prokaryotic sRNAs and approaches to their identification.

Bacteria↗

Eukaryotic regulatory element conservation analysis and identification using comparative genomics.

Comparative genomics is a promising approach to the challenging problem of eukaryotic regulatory element identification, because functional noncoding sequences may be conserved across species from evolutionary constraints. We systematically analyzed known human and Saccharomyces cerevisiae regulatory elements and discovered that human regulatory elements are more conserved between human and mouse than are background sequences. Although S. cerevisiae regulatory elements do not appear to be more conserved by comparison of S. cerevisiae to Schizosaccharomyces pombe, they are more conserved when compared with multiple other yeast genomes (Saccharomyces paradoxus, Saccharomyces mikatae, and Saccharomyces bayanus). Based on these analyses, we developed a sequence-motif-finding algorithm called CompareProspector, which extends Gibbs sampling by biasing the search in regions conserved across species. Using human-mouse comparison, CompareProspector identified known motifs for transcription factors Mef2, Myf, Srf, and Sp1 from a set of human-muscle-specific genes. It also discovered the NFAT motif from genes up-regulated by CD28 stimulation in T-cells, which implies the direct involvement of NFAT in mediating the CD28 stimulatory signal. Using Caenorhabditis elegans-Caenorhabditis briggsae comparison, CompareProspector found the PHA-4 motif and the UNC-86 motif. CompareProspector outperformed many other computational motif-finding programs, demonstrating the power of comparative genomics-based biased sampling in eukaryotic regulatory element identification.

Algorithms↗

Wnt/beta-catenin/Tcf signaling induces the transcription of Axin2, a negative regulator of the signaling pathway.

Axin2/Conductin/Axil and its ortholog Axin are negative regulators of the Wnt signaling pathway, which promote the phosphorylation and degradation of beta-catenin. While Axin is expressed ubiquitously, Axin2 mRNA was seen in a restricted pattern during mouse embryogenesis and organogenesis. Because many sites of Axin2 expression overlapped with those of several Wnt genes, we tested whether Axin2 was induced by Wnt signaling. Endogenous Axin2 mRNA and protein expression could be rapidly induced by activation of the Wnt pathway, and Axin2 reporter constructs, containing a 5.6-kb DNA fragment including the promoter and first intron, were also induced. This genomic region contains eight Tcf/LEF consensus binding sites, five of which are located within longer, highly conserved noncoding sequences. The mutation or deletion of these Tcf/LEF sites greatly diminished induction by beta-catenin, and mutation of the Tcf/LEF site T2 abolished protein binding in an electrophoretic mobility shift assay. These results strongly suggest that Axin2 is a direct target of the Wnt pathway, mediated through Tcf/LEF factors. The 5.6-kb genomic sequence was sufficient to direct the tissue-specific expression of d2EGFP in transgenic embryos, consistent with a role for the Tcf/LEF sites and surrounding conserved sequences in the in vivo expression pattern of Axin2. Our results suggest that Axin2 participates in a negative feedback loop, which could serve to limit the duration or intensity of a Wnt-initiated signal.

Animals↗

Cloning and analysis of the rat gamma-glutamyltransferase gene.

We have isolated and characterized a complete structural gene encoding the enzyme gamma-glutamyltransferase ((5-glutamyl)-peptide:amino acid 5-glutamyltransferase; EC 2.3.2.2). The gene contains 8 exons and spans approximately 12 kilobases. Ras-transformed rat liver epithelial cells and rat kidney express RNAs which differ in length by approximately 0.3 kilobase pair. Comparison of the genomic sequence with kidney gamma GT cDNA sequence indicates that the first exon is noncoding, and nuclease protection and primer extension data have identified a potential kidney transcription start site (defined as +1) for this exon. The site is not associated with a TATA box, but there are two CCAAT boxes (-136 and -599) and two sites (-101 and -746) containing the consensus sequences to which the transcription factor SP1 is known to bind. There is also a sequence at -453 (TGTGGTTG) that is highly homologous to the core sequence (TGTGG(T)3-5G) of SV40 and polyoma viral enhancers.

Base Sequence↗

Bovine and rodent tamm-horsfall protein (THP) genes: cloning, structural analysis, and promoter identification.

We have isolated bovine and rodent cDNA and genomic clones encoding the kidney-specific Tamm-Horsfall protein (THP). In both species the gene contains 11 exons, the first of which is noncoding. Exon/intron junctions were analyzed and all were shown to follow the AG/GT rule. A kidney-specific DNase I hypersensitive site was mapped onto a rodent genomic fragment for which the sequence is highly conserved in three species (rat, cow, and human) over a stretch of 350 base pairs. Primer extension and RNase protection analysis identified a transcription start site at the 3' end of this conserved region. A TATA box is located at 32 nucleotides upstream of the start site in the bovine gene and 34 nucleotides upstream in the rodent gene. An inverted CCAAT motif occurs at 65 and 66 nucleotides upstream of the start site in the bovine and rodent genes, respectively. Other highly conserved regions were noted in this 350 bp region and these are likely to be binding sites for transcription factors. A functional assay based on an in vitro transcription system confirmed that the conserved region is an RNA Pol II promoter. The in vitro system accurately initiated transcription from the in vivo start site and was highly sensitive to inhibition by alpha-amanitin at a concentration of 2.5 micrograms/ml. These studies set the stage for the further definition of cis-acting sequences and trans-factors regulating expression of the THP gene, a model for kidney-specific gene expression.

Amino Acid Sequence↗

Nucleotide sequence of the virulent SA-14 strain of Japanese encephalitis virus and its attenuated vaccine derivative, SA-14-14-2.

The attenuated SA-14-14-2 strain of Japanese encephalitis (JE) virus has been used to immunize people in the People's Republic of China. Oligonucleotide fingerprints of the parent SA-14 and vaccine strain indicate that multiple genetic changes occurred during attenuation of the virus. We have cloned and sequenced the genomes of both the virulent SA-14 and attenuated SA-14-14-2 viruses to define molecular differences in the genomes. Forty-five nucleotide differences, resulting in 15 amino acid substitutions, were found by comparing sequences of the SA-14 and SA-14-14-2 genomes. Transversion of U to A occurred at position 39 in the 5'-noncoding region of SA-14-14-2 and another SA-14 vaccine derivative SA-14-5-3. A single nucleotide change in the capsid gene of SA-14-14-2 altered a single amino acid which changed its predicted secondary structure. A silent nucleotide change was found in the prM gene sequence and the M-protein was unchanged. There are seven nucleotide differences, resulting in five amino acid changes, in the E glycoprotein sequence of the two viruses. Nine amino acid differences were found in the nonstructural proteins of SA-14 and SA-14-14-2: one in NS2A, two in NS2B, three in NS3, one in ns4a, and two in NS5. A single nucleotide change at position 10,428 in the 3'-noncoding region is vaccine virus-specific. The nucleotide and deduced amino acid sequences of the vaccine strain SA-14-14-2, the parent virus SA-14, and virulent strains JaOArS982 and Beijing-1 have been compared and are highly conserved.

Aedes↗

Role of poly(A) tail length in Alu retrotransposition.

Alu are mobile noncoding Short INterspersed Elements (SINEs) present at a million copies in the human genome. Using marked Alu sequences in an ex vivo assay, we previously showed that they are mobilized through diversion of the LINE (Long INterspersed Elements) retrotransposition machinery, with the poly(A) tail of the Alu being required for their mobility. Here we show that other homopolymeric tracts cannot functionally replace the Alu poly(A) tail, and that the Alu transposition rate varies over a two-log range depending on the poly(A) tail length. Variation is according to a sigmoid-shaped curve with a lag observed for tails shorter than 15 nt and a plateau reached for tails longer than 50 nt, consistent with the binding of a limited number of a protein component requiring multiple contacts for a productive interaction with the poly(A) stretch. This analysis indicates that most of the naturally occurring genomic Alu, owing to their pA tail length, should be poor substrates for the LINE machinery, a feature possibly "selected" for the host sake.

Alu Elements↗

High-resolution whole-organ mapping with SNPs and its significance to early events of carcinogenesis.

We attempted to identify deleted segments in two model tumor suppressor gene loci on chromosomes 13q14 and 17p13 that were associated with clonal expansion of in situ bladder preneoplasia using single nucleotide polymorphisms (SNPs)-based whole-organ histologic and genetic mapping. For mapping with SNPs, the sequence-based maps spanning approximately 27 and 5 Mb centered around RB1 and p53, respectively, were assembled. The integrated gene and SNP maps of the regions were used to select 661 and 960 SNPs, which were genotyped by pyrosequencing. Genotyping of SNPs was performed on DNA samples corresponding to histologic maps of the entire bladder mucosa in human cystectomy specimens with invasive urothelial carcinoma. By using this approach, we have identified deleted regions associated with clonal expansion of intraurothelial neoplasia; which ranged from 0.001 to 4.3 Mb (average 0.67 Mb) and formed clusters of discontinuous deleted segments. The high resolution of such maps is a prerequisite for future positional targeting of genes involved in early phases of bladder neoplasia. This approach also permits analysis of the overall genomic landscape of the involved region and discloses that a unique composition of noncoding DNA characterized by a high concentration of repetitive sequences may predispose to deletions.

Carcinoma in Situ↗

A small bacterial RNA regulates a putative ABC transporter.

A small noncoding bacterial ribonucleic acid of 62-64 nucleotides, RydC, was identified in the genomes of Escherichia coli, Salmonella, and Shigella. In vivo, RydC binds to the RNA-binding protein Hfq, and it is unstable when Hfq is absent. Mobility assays reveal that complex formation between RydC and Hfq is specific, with an apparent binding constant of approximately 300 nm. Sequence alignments combined with structural probing demonstrate that RydC folds as a pseudoknot. Hfq binds the loops crossing the deep and shallow grooves of the pseudoknotted RNA and reorganizes its overall conformation. An interaction with a polycistronic mRNA, yejABEF, which encodes a putative ABC transporter, was detected by affinity purification of immobilized RNA-Hfq complexes. In vivo, the yejABEF operon is expressed on minimal medium. Remarkably, its expression is reduced when RydC is absent, and the operon is degraded when RydC expression is stimulated. This observation correlates with the growth defects associated with a stimulation of its expression in vivo, generating a thermosensitive phenotype that affects growth on minimal media supplemented with glycerol, maltose, or ribose. We conclude that RydC regulates the yejABEF-encoded ABC permease at the mRNA level. This small RNA may contribute to optimal adaptation of some Enterobacteria to environmental conditions.

ATP-Binding Cassette Transporters↗

Isolation and analysis of inducibility of the rat N-methylpurine-DNA glycosylase promoter.

Alkylations at base nitrogens in DNA are removed by excision repair, the first step of which is catalyzed by the repair enzyme N-methylpurine-DNA glycosylase (MPG). To study regulation of MPG expression, we have cloned the rat MPG promoter. A cosmid clone containing the rat MPG gene was isolated from a library using rat MPG cDNA as a probe. The 5' part of the MPG gene and the nontranscribed 5'-flanking region were isolated and characterized. Transcription start sites of the rat MPG gene were identified by primer extension and S1 nuclease protection analysis of RNA from primary rat hepatocytes. Promoter activity of the 5'-flanking noncoding region was shown by transfection in H4IIE rat hepatoma cells of various genomic MPG fragments cloned in front of the reporter gene chloramphenicol acetyltransferase. The rat MPG promoter does not contain a TATA box, but has a CCAAT sequence element and putative binding sites for the transcription factors Sp1, AP-2, AP-3, Ets-1, PEA3, NF-1, p53, c-Myc, NF-kappa B, and the glucocorticoid receptor. The activity of the rat MPG promoter was found to be inducible by the tumor promoter TPA and UV light, but not to a significant extent by methylating agents and ionizing radiation.

Animals↗

Complete nucleotide sequence of alfalfa mosaic virus RNA 1.

Double-stranded cDNA of alfalfa mosaic virus (AlMV) RNA 1 has been cloned and sequenced. From clones with overlapping inserts, and other sequence data, the complete primary sequence of the 3644 nucleotides of RNA 1 was deduced: a long open reading frame for a protein of Mr 125,685 is flanked by a 5'-terminal sequence of 100 nucleotides and a 3' noncoding region of 163 nucleotides, including the sequence of 145 nucleotides the three genomic RNAs of AlMV have in common. The two UGA-termination codons halfway RNA 1, that were postulated by Van Tol et al. (FEBS Lett. 118, 67-71, 1980) to account for partial translation of RNA 1 in vitro into Mr 58,000 and Mr 62,000 proteins, were not found in the reading frame of the Mr 125,685 protein.

Amino Acid Sequence↗

Characterization of a transcriptional promoter of human papillomavirus 18 and modulation of its expression by simian virus 40 and adenovirus early antigens.

RNA present in cells derived from cervical carcinoma that contained human papillomavirus 18 genomes was initiated in the 1.053-kilobase BamHI fragment that covered the complete noncoding region of this virus. When cloned upstream of the chloramphenicol acetyltransferase gene, this viral fragment directed the expression of the bacterial enzyme only in the sense orientation. Initiation sites were mapped around the ATG of open reading frame E6. This promoter was active in some human and simian cell lines, and its expression was modulated positively by simian virus 40 large T antigen and negatively by adenovirus type 5 E1a antigen.

Acetyltransferases↗

Complete nucleotide sequence of wild-type hepatitis A virus: comparison with different strains of hepatitis A virus and other picornaviruses.

The complete nucleotide sequence of wild-type hepatitis A virus (HAV) HM-175 was determined. The sequence was compared with that of a cell culture-adapted HAV strain (R. Najarian, D. Caput, W. Gee, S.J. Potter, A. Renard, J. Merryweather, G.V. Nest, and D. Dina, Proc. Natl. Acad. Sci. USA 82:2627-2631, 1985). Both strains have a genome length of 7,478 nucleotides followed by a poly(A) tail, and both encode a polyprotein of 2,227 amino acids. Sequence comparison showed 624 nucleotide differences (91.7% identity) but only 34 amino acid differences (98.5% identity). All of the dipeptide cleavage sites mapped in this study were conserved between the two strains. The sequences of these two HAV strains were compared with the partial sequences of three other HAV strains. Most amino acid differences were located in the capsid region, especially in VP1. Whereas changes in amino acids were localized to certain portions of the genome, nucleotide differences occurred randomly throughout the genome. The most extensive nucleotide homology between the strains was in the 5' noncoding region (96% identity for cell culture-adapted strains versus wild type; greater than 99% identity among cell culture-adapted strains). HAV proteins are less homologous with those of any other picornavirus than the latter proteins are when compared with each other. When the sequences of wild-type and cell culture-adapted HAV strains are compared, the nucleotide differences in the 5' noncoding region and the amino acid differences in the capsid region suggest areas that may contain markers for cell culture adaptation and for attenuation.

Base Sequence↗