Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

The evolution of noncoding DNA: how much junk, how much func?

Comparative sequence analysis on a genomic scale has opened the door for the systematic analysis of cis-acting regulatory DNA. It is now possible to begin to answer basic questions such as, how much meaningful noncoding sequence is in the genome? How strong is natural selection on functional noncoding sequences in different species? Two recent articles have capitalized on the comparative genomic approach in an attempt to answer these questions with surprising results.

DNA, Intergenic↗

Selection at linked sites in the partial selfer Caenorhabditis elegans.

Natural selection can produce a correlation between local recombination rates and levels of neutral DNA polymorphism as a consequence of genetic hitchhiking and background selection. Theory suggests that selection at linked sites should affect patterns of neutral variation in partially selfing populations more dramatically than in outcrossing populations. However, empirical investigations of selection at linked sites have focused primarily on outcrossing species. To assess the potential role of selection as a determinant of neutral polymorphism in the context of partial self-fertilization, we conducted a multivariate analysis of single-nucleotide polymorphism (SNP) density throughout the genome of the nematode Caenorhabditis elegans. We based the analysis on a published SNP data set and partitioned the genome into windows to calculate SNP densities, recombination rates, and gene densities across all six chromosomes. Our analyses identify a strong, positive correlation between recombination rate and neutral polymorphism (as estimated by noncoding SNP density) across the genome of C. elegans. Furthermore, we find that levels of neutral polymorphism are lower in gene-dense regions than in gene-poor regions in some analyses. Analyses incorporating local estimates of divergence between C. elegans and C. briggsae indicate that a mutational explanation alone is unlikely to explain the observed patterns. Consequently, we interpret these findings as evidence that natural selection shapes genome-wide patterns of neutral polymorphism in C. elegans. Our study provides the first demonstration of such an effect in a partially selfing animal. Explicit models of genetic hitchhiking and background selection can each adequately describe the relationship between recombination rate and SNP density, but only when they incorporate selfing rate. Clarification of the relative roles of genetic hitchhiking and background selection in C. elegans awaits the development of specific theoretical predictions that account for partial self-fertilization and biased sex ratios.

Animals↗

Tissue specificity of geminivirus infection is genetically determined.

The types of cells and tissues infected by a virus define its tissue tropism. Determinants of tissue tropism in animal-infecting viruses have been extensively investigated, but little is known about plant viruses in this regard. Some geminiviruses in the genus Begomovirus exhibit phloem limitation and are restricted to cells of the vascular system, whereas others can invade mesophyll tissue. To identify viral genetic determinants of tissue tropism, we established a model system using two begomoviruses and their common host plant, Nicotiana benthamiana. Analysis by DNA in situ hybridization confirmed that tomato golden mosaic virus invades mesophyll tissues in systemically infected leaves, whereas bean golden mosaic virus remains phloem limited. Through genetic complementation and analysis of recombinant hybrid viruses, we demonstrated that three genetic elements of tomato golden mosaic virus determine its mesophyll tissue tropism. A noncoding region of the viral genome is essential for the phenotype, but it must be accompanied by one of two different coding regions. To our knowledge, this is the first example documented in a plant virus of noncoding DNA sequences that determine tissue tropism.

DNA, Viral↗

Hepatitis C virus replication during acute infection in the chimpanzee.

The events following experimental infection of 2 chimpanzees with the H strain of hepatitis C virus (HCV) were studied by quantitating the levels of HCV RNA in liver and serum. Serum and liver samples were tested every 1-3 weeks for up to 32 weeks. The genomic and antigenomic strands of HCV RNA were individually detected in liver and serum by strand-specific reverse transcription followed by polymerase chain reaction (PCR) using nested primers specific for the 5' noncoding region of the HCV genome and were quantitated by end-point dilution of the nucleic acid extract. Both genomic and antigenomic strands were detected in liver and serum within 1 week after inoculation and approximately 1 week before the development of elevated levels of alanine aminotransferase (ALT) activity. Changes in levels of antigenomic strand paralleled those of the genomic strand in both serum and liver. Titers of HCV RNA in serum and liver generally correlated with changes in ALT levels.

Acute Disease↗

Ongoing activity of RNA polymerase II confers preferential repair of nitrogen mustard-induced N-alkylpurines in the hamster dihydrofolate reductase gene.

Recently, it has been demonstrated that nitrogen mustard-induced N-alkylpurines are excised rapidly from actively transcribing genes, while they persist longer in noncoding regions and in the genome overall. It was suggested that transcriptional activity is implicated as a regulatory element in the efficient removal of lesions. By treating cells or not with the transcription inhibitor alpha-amanitin, we have explored whether ongoing activity of RNA polymerase II was coordinately related to proficient repair of nitrogen mustard-induced alkylation products in the actively transcribed dihydrofolate reductase gene in the Chinese hamster ovary B11 cells. Nuclear run-off transcription analysis verified that alpha-amanitin completely and selectively inhibited transcription by RNA polymerase II. At the drug exposure examined, nitrogen mustard induced DNA damage capable of a complete transcription termination in the RNA polymerase II-transcribed dihydrofolate reductase gene and reduced 28S rDNA transcription by a factor of 7.9. The transcription activity did partially recover following reincubation in drug-free medium; this recovery was about 34 and 76% of ribosomal 28S gene transcripts and dihydrofolate reductase gene transcripts, respectively, after 6 h of repair incubation. alpha-Amanitin significantly inhibited the removal of nitrogen mustard-induced N-alkylpurines in the 5'-half of the essential, constitutively active dihydrofolate reductase gene, while no effect of alpha-amanitin was observed on the lesion removal from a noncoding region 3'-flanking to the gene and from the genome overall. In the actively transcribed gene region, about 77% of N-alkylpurines were removed 21 h following drug exposure of cells not treated with alpha-amanitin and about 47% in 21 h in alpha-amanitin treated cells. The global semiconservative replication seemed unaffected by the alpha-amanitin treatment. From these results we suggest that gene-specific repair of nitrogen mustard-induced N-alkylpurines is dependent on ongoing activity of the transcribing RNA polymerase II. The findings are discussed in terms of the current ideas about the mechanism of preferential DNA repair.

Amanitins↗

Nucleotide sequences important for translation initiation of enterovirus RNA.

An infectious cDNA clone was constructed from the genome of coxsackievirus B1 strain. A number of RNA transcripts that have mutations in the 5' noncoding region were synthesized in vitro from the modified cDNA clones and examined for their abilities to act as mRNAs in a cell-free translation system prepared from HeLa S3 cells. RNAs that lack nucleotide sequences at positions 568 to 726 and 565 to 726 were found to be less efficient and inactive mRNAs, respectively. To understand the biological significance of this region of RNA, small deletions and point mutations were introduced in the nucleotide sequence between positions 538 and 601. Except for a nucleotide substitution at 592 (U----C) within the 7-base conserved sequence, mutations introduced in the sequence downstream of position 568 did not affect much, if any, of the ability of RNA to act as mRNA. Except for a point mutation at 558 (C----U), mutations upstream of position 567 appeared to inactivate the mRNA. In the upstream region, a sequence consisting of 21 nucleotides at positions 546 to 566 is perfectly conserved in the 5' noncoding regions of enterovirus and rhinovirus genomes. These results suggest that the 7-base conserved sequence functions to maintain the efficiency of translation initiation and that the nucleotide sequence upstream of position 567, including the 21-base conserved sequence, plays essential roles in translation initiation. A deletion mutant whose genome lacks the nucleotide sequence at positions 568 to 726 showed a small-plaque phenotype and less virulence against suckling mice than the wild-type virus. Thus, reduction of the efficiency of translation initiation may result in the construction of enteroviruses with the lower-virulence phenotype.

Animals↗

Meta-analysis discovery of tissue-specific DNA sequence motifs from mammalian gene expression data.

BACKGROUND: A key step in the regulation of gene expression is the sequence-specific binding of transcription factors (TFs) to their DNA recognition sites. However, elucidating TF binding site (TFBS) motifs in higher eukaryotes has been challenging, even when employing cross-species sequence conservation. We hypothesized that for human and mouse, many orthologous genes expressed in a similarly tissue-specific manner in both human and mouse gene expression data, are likely to be co-regulated by orthologous TFs that bind to DNA sequence motifs present within noncoding sequence conserved between these genomes. RESULTS: We performed automated motif searching and merging across four different motif finding algorithms, followed by filtering of the resulting motifs for those that contain blocks of information content. Applying this motif finding strategy to conserved noncoding regions surrounding co-expressed tissue-specific human genes allowed us to discover both previously known, and many novel candidate, regulatory DNA motifs in all 18 tissue-specific expression clusters that we examined. For previously known TFBS motifs, we observed that if a TF was expressed in the specified tissue of interest, then in most cases we identified a motif that matched its TRANSFAC motif; conversely, of all those discovered motifs that matched TRANSFAC motifs, most of the corresponding TF transcripts were expressed in the tissue(s) corresponding to the expression cluster for which the motif was found. CONCLUSION: Our results indicate that the integration of the results from multiple motif finding tools identifies and ranks highly more known and novel motifs than does the use of just one of these tools. In addition, we believe that our simultaneous enrichment strategies helped to identify likely human cis regulatory elements. A number of the discovered motifs may correspond to novel binding site motifs for as yet uncharacterized tissue-specific TFs. We expect this strategy to be useful for identifying motifs in other metazoan genomes.

Algorithms↗

Expression of enhancers is altered in Drosophila melanogaster hybrids.

The molecular foundations of evolution are difficult to trace because most protein sequences are virtually identical in closely related species. The largest fraction of sequence within the genome, however, is composed of noncoding sequences where regulatory elements locate to various sites. It has been suggested that changes in the activity of these elements may trigger evolutionary change. In Drosophila, the enhancer trap procedure identifies regulatory sequences in the genome after the insertion of a P-element-based construct. We generated new insertions and characterized their expression domains in the adult eye and larval imaginal disks using the white and LacZ reporter genes. Lines with robust expression patterns in D. melanogaster were analyzed in hybrids to test the conservation of regulatory mechanisms between species. Most of the enhancers used in this study modified their expression in hybrids with the mating species D. mauritiana and D. simulans. Expression changes resulted either in gain or loss of expression and were cell-type or hybrid-genome specific. Further characterization of a limited number of enhancers in D. melanogaster showed that expression domains could adapt to changes in cell number during development but not after the completion of cell proliferation. Also, expression of some enhancers appeared to be sensitive to heterochromatin from the Y but not the X chromosome. Taken together, these results demonstrate the high sensitivity of regulatory mechanisms of gene expression as a prime source of evolutionary change and suggest quantitative changes in available transcription factors as one of the mechanisms involved.

Animals↗

Genome sequence of bovine herpesvirus 4, a bovine Rhadinovirus, and identification of an origin of DNA replication.

Bovine herpesvirus 4 (BoHV-4) is a gammaherpesvirus of cattle. The complete long unique coding region (LUR) of BoHV-4 strain 66-p-347 was determined by a shotgun approach. Together with the previously published noncoding terminal repeats, the entire genome sequence of BoHV-4 is now available. The LUR consists of 108,873 bp with an overall G+C content of 41.4%. At least 79 open reading frames (ORFs) are present in this coding region, 17 of them unique to BoHV-4. In contrast to herpesvirus saimiri and human herpesvirus 8, BoHV-4 has a reduced set of ORFs homologous to cellular genes. Gene arrangement as well as phylogenetic analysis confirmed that BoHV-4 is a member of the genus Rhadinovirus. In addition, an origin of replication (ori) in the genome of BoHV-4 was identified by DpnI assays. A minimum of 1.69 kbp located between ORFs 69 and 71 was sufficient to act as a cis signal for replication.

Animals↗

Splice site requirement for the efficient accumulation of polyoma virus late mRNAs.

Polyoma virus late nuclear primary transcripts are giant and heterogeneous, containing tandem repeats of the late strand of the circular viral genome. Late pre-mRNA processing involves the splicing of noncoding 'leader' exons to each other (removing genome-length introns), with the joining of the last leader to a coding 'body' exon. We have constructed a number of mutants blocked only in leader-leader splicing, or blocked in both leader-leader and leader-body splicing. We examined the accumulation of both nuclear and cytoplasmic late-strand RNAs in NIH3T3 cells. Consistent with our previous results, mutants lacking the 3' splice site of the late leader (leader-leader splicing blocked) showed a 10-20 fold defect in late RNA accumulation. Mutants which lacked the leader 5' splice site (leader-body splicing blocked) had a more profound defect, exhibiting virtually no late-strand cytoplasmic or nuclear RNA. This result was unexpected as a substantial proportion of wild type late cytoplasmic messages are unspliced. A mutant with no intron, but having functional 3' and 5' splice sites bordering the leader exon, is capable of producing large amounts of unspliced late mRNA. This demonstrates that an excisable intron is not a requirement for late mRNA accumulation. The accumulation of polyoma late mRNAs requires the presence of leader exons bordered by functional 3' and 5' splice sites, whether or not these sites are used during pre-mRNA processing.

Base Sequence↗

The evolution of isochores: evidence from SNP frequency distributions.

The large-scale systematic variation in nucleotide composition along mammalian and avian genomes has been a focus of the debate between neutralist and selectionist views of molecular evolution. Here we test whether the compositional variation is due to mutation bias using two new tests, which do not assume compositional equilibrium. In the first test we assume a standard population genetics model, but in the second we make no assumptions about the underlying population genetics. We apply the tests to single-nucleotide polymorphism data from noncoding regions of the human genome. Both models of neutral mutation bias fit the frequency distributions of SNPs segregating in low- and medium-GC-content regions of the genome adequately, although both suggest compositional nonequilibrium. However, neither model fits the frequency distribution of SNPs from the high-GC-content regions. In contrast, a simple population genetics model that incorporates selection or biased gene conversion cannot be rejected. The results suggest that mutation biases are not solely responsible for the compositional biases found in noncoding regions.

Alleles↗

Selection at the wobble position of codons read by the same tRNA in Saccharomyces cerevisiae.

The transfer RNA gene complement of Saccharomyces cerevisiae was utilized for a whole-genome analysis of the deviation from a neutral usage of pyrimidine-ending cognate codons, that is, codons read by a single tRNA species having either inosine or guanosine as the first anticodon base. Mutational pressure at the wobble position was estimated from the base composition of the noncoding portion of the yeast genome. The selective pressure for translational efficiency was inferred from the degree of codon adaptation to tRNA gene redundancy and from mRNA abundance data derived from yeast transcriptome analysis. Amino acid conservation in orthologous comparisons with wholly sequenced microbial genomes was used to estimate translational accuracy requirements. A close correspondence was observed between the usage of wobble position pyrimidines and the frequency predicted by mutational bias. However, in the case of four cognate pairs (Gly: ggu/ggc; Asn: aau/aac; Phe: uuu/uuc; Tyr: uau/ uac) all read by guanosine-starting anticodons, we found evidence for a strong selective pressure driven by translational efficiency. Only for the glycine pair, wobble pyrimidine choice also appears to fulfill a translational accuracy requirement. Wobble pyrimidine selection is strictly related to the number of hydrogen bonds formed by alternative cognate codons: whenever a different number of hydrogen bonds can be formed at the wobble position, there is selection against six- or nine-hydrogen-bonded codon-anticodon pairs. Our results indicate that an intrinsic codon preference, critically dependent on the stability of codon-anticodon interaction and mainly reflecting selection for the optimization of translational efficiency, is built into the translational apparatus.

Codon↗

Sequence analysis of cloned dengue virus type 2 genome (New Guinea-C strain).

Sequences totalling 5472 nucleotides (nt) from four complementary DNA (cDNA) clones of the dengue virus type 2 (DEN-2) RNA (New Guinea strain, NGS-C) have been reported previously [Yaegashi et al., Gene 46 (1986) 257-267; Putnak et al., Virology 163 (1988) 93-103]. This report describes the complete nucleotide sequence, with the exception of about 7 nt at the 5'-noncoding region, of this RNA genome derived from several cDNA clones. It is 10,723 nt in length and contains a single long open reading frame of 10,173 nt, encoding a polyprotein of 3391 amino acids. The genomic organization is similar to that of other flaviviruses that have recently been reported. Among the three DEN-2 strains - the Jamaica genotype (DEN-2JAM), the DEN-2NGS-C, and the S1 candidate vaccine strain derived from Puerto Rico (PR)-159 isolate (DEN-2S1) - which have been sequenced to date, the amino acid sequences of the polyproteins bear 94%-99% similarity. When the amino acid sequences of DEN-2NGS-C are compared with those of the other two strains, the variations are greater in the DEN-2S1 than in the DEN-2JAM. When DEN-2 and DEN-4 are compared, the overall amino acid identities range from 30% to 80% in both the structural and nonstructural proteins; whereas between DEN-2 and DEN-1, they range from 68% to 79% in the region encoding the structural proteins and the nonstructural protein NS1.

Amino Acid Sequence↗

Analysis of 94 kb of the chlorella virus PBCV-1 330-kb genome: map positions 88 to 182.

Analysis of 94 kb of DNA, located between map positions 88 and 182 kb in the 330-kb chlorella virus PBCV-1 genome, revealed 195 open reading frames (ORFs) 65 codons or longer. One hundred and five of the 195 ORFs were considered major ORFs. Twenty-six of the 105 major ORFs resembled genes in the databases including three chitinases, a chitosanase, three serine/threonine protein kinases, two additional protein kinases, a tyrosine protein phosphatase, two ankyrins, an ornithine decarboxylase, a copper/zinc-superoxide dismutase, a proliferating cell nuclear antigen, a DNA polymerase, a fibronectin-binding protein, the yeast Ski2 protein, an adenine DNA methyltransferase and its corresponding DNA site-specific endonuclease, and an amidase. The genes for the 105 major ORFs were evenly distributed along the genome and, except for one noncoding 1788-nucleotide stretch, the genes were close together. Unexpectedly, a 900-bp region in the 1788-bp noncoding sequence resembled a CpG island.

Amino Acid Sequence↗

Biological consequences of deletions within the 3'-untranslated region of flaviviruses may be due to rearrangements of RNA secondary structure.

It was previously reported that deletions introduced into the 3'-untranslated region (3'-UTR) of dengue type 4 (DEN 4) virus (Men, R., Bray, M., Clark, D., Chanock, R.M., Lai, C.J., 1996. DEN 4 virus mutants containing deletions in the 3'-noncoding region of the RNA genome: analysis of growth restriction in cell culture and altered viremia pattern and immunogenicity in Rhesus monkeys. J. Virol. 70, 3930-3937), tick-borne encephalitis (TBE) virus (Mandl, C.W., Holzmann, H., Meixner, T., Rauscher, S., Stadler, P.F., Allison, S.L. , Heinz, F.X., 1998. Spontaneous and engineered deletions in the 3'-noncoding region of TBE virus: construction of highly attenuated mutants of a flavivirus. J. Virol. 72, 2132-2140) and subgenomic replicons of Kunjin virus (Khromykh, A.A., Westaway, E.G., 1997. Subgenomic replicons of the flavivirus Kunjin: construction and applications. J. Virol. 71, 1497-1505) altered the infectivity of the mutants and reduced the efficiency of RNA replication. Here, these deletions were superimposed onto the models of secondary structure we constructed previously and the folding of the modified 3'-UTR sequences was simulated. The analysis showed that most of the deletions disrupted or reshaped conserved elements of secondary structure and that the biological effects of these deletions are likely to represent structural rearrangements in the 3'-UTR, rather than the loss of sequence motifs. The analysis also suggested that the overall structural integrity of the flaviviral 3'-UTR is essential for optimal performance of its promotor function, although two distinct parts can be defined: the most 3'-terminal structures and sequences which may be critical for the initiation of minus-strand RNA synthesis, and more proximal structures and sequences that possibly function as enhancers of viral RNA replication. The functional significance of certain structural elements and their possible effect on the efficiency of viral replication in different cells are also discussed.

3' Untranslated Regions↗

The regulatory content of intergenic DNA shapes genome architecture.

BACKGROUND: Factors affecting the organization and spacing of functionally unrelated genes in metazoan genomes are not well understood. Because of the vast size of a typical metazoan genome compared to known regulatory and protein-coding regions, functional DNA is generally considered to have a negligible impact on gene spacing and genome organization. In particular, it has been impossible to estimate the global impact, if any, of regulatory elements on genome architecture. RESULTS: To investigate this, we examined the relationship between regulatory complexity and gene spacing in Caenorhabditis elegans and Drosophila melanogaster. We found that gene density directly reflects local regulatory complexity, such that the amount of noncoding DNA between a gene and its nearest neighbors correlates positively with that gene's regulatory complexity. Genes with complex functions are flanked by significantly more noncoding DNA than genes with simple or housekeeping functions. Genes of low regulatory complexity are associated with approximately the same amount of noncoding DNA in D. melanogaster and C. elegans, while loci of high regulatory complexity are significantly larger in the more complex animal. Complex genes in C. elegans have larger 5' than 3' noncoding intervals, whereas those in D. melanogaster have roughly equivalent 5' and 3' noncoding intervals. CONCLUSIONS: Intergenic distance, and hence genome architecture, is highly nonrandom. Rather, it is shaped by regulatory information contained in noncoding DNA. Our findings suggest that in compact genomes, the species-specific loss of nonfunctional DNA reveals a landscape of regulatory information by leaving a profile of functional DNA in its wake.

Animals↗

The mosaic organization of the mitochondrial introns of Saccharomyces cerevisiae: features and evolutionary origins.

The introns of three genes (oxi3, cob and 21S) from the mitochondrial (mt) genome of Saccharomyces cerevisiae contain closed reading frames (CRFs). In the present work, we have analyzed these sequences in their oligodeoxyribonucleotide (oligo; isostich) patterns. We have shown that the relative amounts of di- to hexanucleotides, when compared to random sequences having the same sizes and compositions, exhibit the same deviations as the intergenic noncoding sequences of the mt genome (except for the CRFs from 21S intron). In contrast, intronic open reading frames (ORFs) showed oligo patterns which were generally quite distinct from those of CRFs, although some similarities could be detected in some cases (especially for aI5 alpha). The mt introns of yeast, therefore, are endowed with a mosaic structure, in which CRFs derive from mt intergenic sequences, whereas ORFs have a different origin (indicated as exogenous by other evidences) yet show, in some cases, the effects of 'sequence assimilation' with CRFs.

Base Composition↗

Nucleotide sequence of the 3'-noncoding region of alfalfa mosaic virus RNA 4 and its homology with the genomic RNAs.

A 226-nucleotide fragment was derived from alfalfa mosaic virus RNA 4 (ALMV RNA 4), the subgenomic messenger for viral coat protein, and its sequence was deduced by in vitro labeling with polynucleotide kinase and application of RNA sequencing techniques. The fragment contains the 3'-terminal 45 nucleotides of the coat protein cistron and the complete 3'-noncoding region of 182 nucleotides. The total length of RNA 4 was calculated to be 881 nucleotides. AlMV RNAs 1, 2 and 3 were elongated with a 3'-terminal poly(A) stretch and subjected to sequence analysis by using a specific primer, reverse transcriptase and chain terminators. This revealed and extensive homology between the 3'-terminal 140 to 150 nucleotides of all four ALMV RNAs. Despite a number of base substitutions, the secondary structure of the homologous region is highly conserved. The observed homology indicates that, as with RNA 4, the sites with a high affinity for the viral coat protein are located at the 3'-termini of the genomic RNAs.

Base Sequence↗