Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Polyomavirus late pre-mRNA processing: DNA replication-associated changes in leader exon multiplicity suggest a role for leader-to-leader splicing in the early-late switch.

Polyomavirus late mRNAs contain at their 5' ends multiple, tandem repeats of a 57-base noncoding sequence, the late leader, whose sequence appears only once in the viral genome. Pre-mRNA molecules are processed by a pathway that includes the splicing of late leader exons to each other in giant, multigenome-length precursors which are the result of inefficient transcription termination. We have devised a method involving reverse transcription and the polymerase chain reaction to determine the number of tandem late leader units on polyomavirus late RNA molecules. Using this technique, we have shown that each class of late viral mRNA (mVP1, mVP2, and mVP3) consists of molecules with between 1 and 12 tandem leader units at their 5' ends. Importantly, single-leader RNAs are underrepresented in both the cytoplasm and the nucleus, suggesting that single-leader primary transcripts are preferentially degraded in the nucleus. In addition, the average number of leaders on late RNAs increases in the presence of DNA replication. Taken together with previous work from our laboratory, the results presented here are consistent with a model for the control of late gene expression at the level of RNA splicing and stability which is in turn controlled by the efficiency of transcription termination.

Animals↗

Genetic analysis of porcine respiratory coronavirus, an attenuated variant of transmissible gastroenteritis virus.

The genome and transcriptional pattern of a newly identified respiratory variant of transmissible gastroenteritis virus were analyzed and compared with those of classical enterotropic transmissible gastroenteritis virus. The transcriptional patterns of the two viruses indicated that differences occurred in RNAs 1 and 2(S) and that RNA 3 was absent in the porcine respiratory coronavirus (PRCV) variant. The smaller RNA 2(S) of PRCV was due to a 681-nucleotide (nt) deletion after base 62 of the PRCV peplomer or spike (S) gene. The PRCV S gene still retained information for the 16-amino-acid signal peptide and the first 6 amino acid residues at the N terminus of the mature S protein, but the adjacent 227 residues were deleted. Two additional deletions (3 and 5 nt) were detected in the PRCV genome downstream of the S gene. The 3-nt deletion occurred in a noncoding region; however, the 5-nt deletion shortened the potential open reading frame A polypeptide from 72 to 53 amino acid residues. Significantly, a C-to-T substitution was detected in the last base position of the transcription recognition sequence upstream of open reading frame A, which rendered RNA 3 nondetectable in PRCV-infected cell cultures.

Base Sequence↗

Unit-length line-1 transcripts in human teratocarcinoma cells.

We have characterized the approximately 6.5-kilobase cytoplasmic poly(A)+ Line-1 (L1) RNA present in a human teratocarcinoma cell line, NTera2D1, by primer extension and by analysis of cloned cDNAs. The bulk of the RNA begins (5' end) at the residue previously identified as the 5' terminus of the longest known primate genomic L1 elements, presumed to represent "unit" length. Several of the cDNA clones are close to 6 kilobase pairs, that is, close to full length. The partial sequences of 18 cDNA clones and full sequence of one (5,975 base pairs) indicate that many different genomic L1 elements contribute transcripts to the 6.5-kilobase cytoplasmic poly(A)+ RNA in NTera2D1 cells because no 2 of the 19 cDNAs analyzed had identical sequences. The transcribed elements appear to represent a subset of the total genomic L1s, a subset that has a characteristic consensus sequence in the 3' noncoding region and a high degree of sequence conservation throughout. Two open reading frames (ORFs) of 1,122 (ORF1) and 3,852 (ORF2) bases, flanked by about 800 and 200 bases of sequence at the 5' and 3' ends, respectively, can be identified in the cDNAs. Both ORFs are in the same frame, and they are separated by 33 bases bracketed by two conserved in-frame stop codons. ORF 2 is interrupted by at least one randomly positioned stop codon in the majority of the cDNAs. The data support proposals suggesting that the human L1 family includes one or more functional genes as well as an extraordinarily large number of pseudogenes whose ORFs are broken by stop codons. The cDNA structures suggest that both genes and pseudogenes are transcribed. At least one of the cDNAs (cD11), which was sequenced in its entirety, could, in principle, represent an mRNA for production of the ORF1 polypeptide. The similarity of mammalian L1s to several recently described invertebrate movable elements defines a new widely distributed class of elements which we term class II retrotransposons.

Amino Acid Sequence↗

[Virology of hepatitis C virus].

Hepatitis C virus (HCV) along with hepatitis G virus are members of the hepacivirus genus of the flavi-viridae family to which the flavi- and pestis viruses also belong. The HCV genome has only one ORF which is flanked by a 5' and 3' noncoding region. The ORF encodes for a single polyprotein, which is stepwise cleaved into the 3 structural proteins, core (C), envelope 1 and 2 (E1,2) as well as into 7 non-structural proteins (p7, NS2, NS3, NS4A, NS4B, NS5A, NS5B). The proteolytic active NS3 plays a central role in this processing. Whereas expectations for development of an HCV vaccine are not very optimistic today, there is great hope in new therapeutic possibilities by inhibition of NS3 function. HCV is highly variable. In the liver and serum of a single patient, genetically slightly different virus particles (quasispecies) can be found. Worldwide, hepatitis C virus has been classified into genotypes and subtypes. This differentiation is not only important epidemiologically, but also has biological and therapeutical implications.

Genotype↗

An encyclopedia of human enhancer-gene regulatory interactions.

Identifying transcriptional enhancers and their target genes is essential for understanding gene regulation and the effect of human genetic variation on disease1-6. Here we create and evaluate a resource of more than 92 million enhancer-gene regulatory interactions across 1,458 biosamples covering 369 cell types and tissues, by integrating predictive models, chromatin states, three-dimensional contacts and large-scale genetic perturbations generated by the ENCODE Consortium7. We first create a systematic benchmarking pipeline to compare predictive models, assembling a dataset of 10,356 element-gene pairs measured in CRISPR perturbation experiments, more than 30,000 fine-mapped expression quantitative trait loci and 569 fine-mapped genome-wide association study (GWAS) variants linked to a probable causal gene. Using this framework, we develop ENCODE-rE2G, a predictive model achieving state-of-the-art performance across several prediction tasks, demonstrating that iterative perturbations and supervised machine learning can build increasingly accurate predictive models of enhancer regulation. Using ENCODE-rE2G, we build an encyclopedia of enhancer-gene regulatory interactions in the human genome, revealing global properties of enhancer networks, identifying differences in regulatory complexity across genes and improving analyses linking noncoding variants to target genes and cell types for common complex diseases. By interpreting the model, we find that beyond enhancer activity and three-dimensional enhancer-promoter contacts, additional features that guide enhancer-promoter communication include promoter class and enhancer-enhancer synergy. These genome-wide maps of enhancer-gene regulatory interactions, benchmarking software, predictive models and insights about enhancer function provide a valuable resource for future studies of gene regulation and human genetics.

Humans↗

Substantial portions of the 5' and intercistronic noncoding regions of cowpea chlorotic mottle virus RNA3 are dispensable for systemic infection but influence viral competitiveness and infection pathology.

Cowpea chlorotic mottle virus (CCMV) has a tripartite, positive strand RNA genome. Genomic RNA3 (2.2 kb) encodes the 3a nonstructural protein and the coat protein, which are dispensable for viral RNA synthesis in protoplasts, but required for systemic infection of whole plants. In protoplasts, portions of the 5' and intercistronic noncoding regions of CCMV RNA3 are also dispensable for RNA3 replication and for transcription of the subgenomic coat protein mRNA. To determine whether these noncoding sequences are required for systemic infection, a series of 5' and intercistronic deletions in RNA3 were tested for their effects on the infection of cowpea plants, a natural host of CCMV. The results refine the mapping of the subgenomic mRNA promoter and show that at least 144 bases in the 5' noncoding region and at least 125 bases in the intercistronic noncoding region of CCMV RNA3 are dispensable for systemic infection. For mutants with deletions within these regions, no differences were noted in the rate of infection spread, and the level of virus accumulation in systemically infected tissue 10-14 days postinoculation was 60-100% of wild type (wt). However, the largest viable intercistronic deletion transformed the nearly symptomless appearance of wt CCMV infections to an extensive, bright yellow chlorosis, showing that infection pathology can be altered by mutations with a regulatory rather than a protein-coding character. In addition, neither 5' nor intercistronic deletion mutants competed effectively with wt CCMV in whole plant co-infection experiments; i.e., such mutants were not detectable in systemically infected tissue after co-inoculation with wt CCMV. Thus, although substantial portions of both the 5' and the intercistronic noncoding regions of CCMV RNA3 are dispensable for individual systemic infection, these segments contribute to the competitive fitness of the virus and influence interaction with the host, as evidenced by symptom response.

Base Sequence↗

The complete mitochondrial genome sequence and characterization of single-nucleotide polymorphisms in the control region of the Asian seabass (Lates calcarifer).

We determined the complete mtDNA nucleotide sequence of Lates calcarifer using the shotgun sequencing method. The mitochondrial DNA (mtDNA) was 16,535 base pairs (bp) in length, and contained 13 protein coding genes, 22 transfer RNAs, 2 ribosomal RNAs, and one major noncoding control region (CR). The CR was unusually short at only 768 bp. A striking feature of the mitochondrial genome was the high G+C content (46.1%), which is among the highest in fish. The gene order was identical to that of a typical vertebrate. Phylogenetic analyses using concatenated amino acid sequences of 12 protein-coding genes of 30 fish species representing 14 suborders clearly showed Lates calcarifer was located in the cluster of fish species from the order Perciformes, supporting the traditional systematic classification. We characterized single-nucleotide polymorphisms (SNPs) in the CR by sequencing the complete CR of 25 individuals obtained from Australia and Singapore. A total of 68 SNPs were detected. Eighteen SNPs were fixed with alternative nucleotides in Australian and Singapore seabass, and these SNPs could be used for differentiating fish from the two countries.

Animals↗

The complete genome sequence of Perina nuda picorna-like virus, an insect-infecting RNA virus with a genome organization similar to that of the mammalian picornaviruses.

Perina nuda picorna-like virus (PnPV) is an insect-infecting RNA virus with morphological and physicochemical characters similar to the Picornaviridae. In this article, we determine the complete genome sequence and analyze the gene organization of PnPV. The genome of PnPV consists of 9476 nucleotides (nts) excluding the poly(A) tail and contains a single large open reading frame (ORF) of 8958 nts (2986 codons) flanked by 473 and 45 nt noncoding regions on the 5' and 3' ends, respectively. Northern blotting did not detect the presence of any subgenomic RNA. The PnPV genome codes for four structural proteins (CP1-4), and determination of their N-terminal sequences by Edman degradation, showed that all four are located in the 5' region of the genome. The 3' part of the PnPV genome contains the consensus sequence motifs for picornavirus RNA helicase, cysteine protease, and RNA-dependent RNA polymerase (RdRp) in that order from the 5' to the 3' end. In all of these characters, the genome organization of PnPV resembles the mammalian picornaviruses and two other insect picorna-like viruses, infectious flacherie virus (IFV) of the silkworm and Sacbrood virus (SBV) of the honeybee. In a phylogenetic tree based on the eight conserved domains in the RdRp sequence, PnPV formed a separate cluster with IFV and SBV, which suggests that these three insect picorna-like viruses might constitute a novel group of insect-infecting RNA viruses.

Amino Acid Sequence↗

A comparative genomic analysis of the cow, pig, and human CFTR genes identifies potential intronic regulatory elements.

The identification of sequences within noncoding regions of genes that are conserved between several species may indicate potential regulatory elements. This is important for genes with complex control mechanisms such as the cystic fibrosis transmembrane conductance regulator (CFTR). CFTR demonstrates similar patterns of temporal and spatial expression in human and sheep, but these differ significantly in mouse cftr. The complete sheep CFTR sequence is unavailable so we annotated BAC clones encompassing the CFTR gene from two other artiodactyl species (cow and pig) for comparative sequence analysis. Regions of introns 2, 3, 10, 17a, 18, and 21 and 3' flanking sequence corresponding to human CFTR DNase I hypersensitive sites (DHS) showed high homology in the cow and pig. Cross-species sequence conservation also enabled finer mapping of other human DHS, including those in introns 1, 16, and 20. Additional potential regulatory elements not associated with human DHS were also identified.

Animals↗

Sequence of the cDNA and gene for angiogenin, a human angiogenesis factor.

Human cDNAs coding for angiogenin, a human tumor derived angiogenesis factor, were isolated from a cDNA library prepared from human liver poly(A) mRNA employing a synthetic oligonucleotide as a hybridization probe. The largest cDNA insert (697 base pairs) contained a short 5'-noncoding sequence followed by a sequence coding for a signal peptide of 24 (or 22) amino acids, 369 nucleotides coding for the mature protein of 123 amino acids, a stop codon, a 3'-noncoding sequence of 175 nucleotides, and a poly(A) tail. The gene coding for human angiogenin was then isolated from a genomic lambda Charon 4A bacteriophage library employing the cDNA as a probe. The nucleotide sequence of the gene and the adjacent 5'- and 3'-flanking regions (4688 base pairs) was then determined. The coding and 3'-noncoding regions of the gene for human angiogenin were found to be free of introns, and the DNA sequence for the gene agreed well with that of the cDNA. The gene contained a potential TATA box in the 5' end in addition to two Alu repetitive sequences immediately flanking the 5' and 3' ends of the gene. The third Alu sequence was also found about 500 nucleotides downstream from the Alu sequence at the 3' end of the gene. The amino acid sequence of human angiogenin as predicted from the gene sequence was in complete agreement with that determined by amino acid sequence analysis. It is about 35% homologous with human pancreatic ribonuclease, and the amino acid residues that are essential for the activity of ribonuclease are also conserved in angiogenin. This provocative finding is thought to have important physiological implications.

Amino Acid Sequence↗

Poliovirus temperature-sensitive mutant containing a single nucleotide deletion in the 5'-noncoding region of the viral RNA.

The effect on viral replication of deleting nucleotide 10 of the poliovirus RNA genome was determined. This deletion, which removes a base pair from a predicted hairpin structure in the viral RNA, was introduced into full-length cDNA. Virus recovered after transfection of HeLa cells with the mutated cDNA contained the expected deletion and was temperature sensitive for plaque formation. Analysis of viral replication by one-step growth experiments indicated that mutant virus production at the nonpermissive temperature was at least 100 times less than that of wild type virus, and release of virus from mutant-infected cells was delayed. The synthesis of positive- and negative-strand viral RNA in mutant virus-infected cells was temperature sensitive. Virus-specific protein synthesis in mutant virus-infected cells was not temperature sensitive but occurred at a slower rate than that of wild type virus at permissive and nonpermissive temperatures. Replication of the mutant virus was sensitive to actinomycin D, in contrast to the wild type parent virus, which was resistant to the drug. Mutant virus stocks contained a small percentage of ts+ viruses that were able to form plaques at the nonpermissive temperature. Nucleotide sequence analysis of genomic RNA from these ts+ viruses revealed a single base change at position 34 from a G to U. In the positive RNA strand, the effect of this mutation is to restore to the hairpin structure the single base pair whose formation was prevented by the original deletion. The ts+ pseudorevertants replicated to similar titers as wild type virus at 33 and 38.5 degrees and were partially sensitive to actinomycin D.

Base Sequence↗

[Genetic differentiation of the inhabitants of Mongolia. Geographic distribution of mitochondrial DNA RFLPs and mitotypes in the inhabitants of Mongolia and a population assessment of the mutation rate in the mitochondrial genome].

The geographical distribution of the Asian specific deletion--insertion polymorphisms and or the RFLP's in the V noncoding region and the D-loop and of the mitotypes was analysed in Mongolia. The frequencies of the mtDNA markers demonstrated homogeneity of 18 local groups in Mongolia. The geographical distribution of the mitotypes showed the existence of two ancestral maternal lineages in mongols. There was no significant difference in the average FST values between mitochondrial gene flow and the nuclear gene flow of the Mongolian population. The equality of FST values permit to calculate the mutation rate for the human mtDNA--6.10(-9) per nucleotide per year. The data reveals the Mongolian population is in the equilibrium.

DNA, Mitochondrial↗

An RNA stem-loop structure involved in the packaging of bovine leukemia virus genomic RNA in vivo.

An RNA secondary structure of the bovine leukemia virus (BLV) 5'-terminal RNA sequence was constructed by computer-assisted RNA secondary structure analysis. Mutations were created in the noncoding region (NCR) of BLV, which contains a conserved consensus sequence, to disrupt predicted secondary structure of this region. After transfection of these constructs into FLK-BLV cells and analysis of viral particles a reduction in mutant RNA content was observed relative to that of unmutated vector RNA. The packaging efficiency of the mutant with a substitution in the consensus sequence was reduced threefold and that of the mutant with a deleted 5' NCR was reduced fivefold. We conclude that predicted RNA secondary structure and/or nucleotide sequence of the BLV noncoding region is essential for BLV RNA packaging in vivo.

Animals↗

Complete sequence of a sea lamprey (Petromyzon marinus) mitochondrial genome: early establishment of the vertebrate genome organization.

The complete nucleotide sequence of a sea lamprey (Petromyzon marinus) mitochondrial genome has been determined. The lamprey genome is 16,201 bp in length and contains genes for 13 proteins, two rRNAs, 22 tRNAs and two major noncoding regions. The order and transcriptional polarities of protein-coding genes are basically identical to those of other chordate mtDNAs, demonstrating that the common mitochondrial gene organization of vertebrates was established at an early stage of vertebrate evolution. The two major noncoding regions are separated by two tRNA genes. The first region probably functions as the control region because it contains distinctive conserved sequence blocks (CSB-II and III) common to other vertebrate control regions. The central conserved domain observed in other vertebrate control regions is not found in the lamprey, suggesting that it is a recently evolved functional domain in vertebrates. Noncoding segments are not found in the expected position of the origin of replication for the second strand, suggesting either that one of the tRNA genes has a dual function or that the second noncoding region may function as the second-strand origin. The base composition at the wobble positions of fourfold degenerate codon families is highly biased toward thymine (32.7%). Values of GC- and AT-skew are typical of vertebrate mitochondrial genomes.

Amino Acid Sequence↗

MicroRNAs: a new insight into cancer genome.

Cancer is a disease involving multi-step dynamic changes in the genome. However, studies on cancer genome so far have focused most heavily on protein-coding genes, and our knowledge on alterations of the functional noncoding sequences in cancer is largely absent. MicroRNAs (miRNA) are approximately 22 nt noncoding RNAs, which regulate gene expression in a sequence-specific manner via translational inhibition or mRNA degradation. Mounting evidence is showing that miRNAs may play important roles in tumor development, and a better understanding of their alteration in cancer genome and oncogenic property should contribute to the diagnosis and treatment of cancer.

Epigenesis, Genetic↗

Determinants in the 5' noncoding region of poliovirus Sabin 1 RNA that influence the attenuation phenotype.

A number of recombinants between the virulent Mahoney and attenuated Sabin strains of type 1 poliovirus were constructed by using infectious cDNA clones of the two strains. To identify a strong neurovirulence determinant(s) residing in the genome region upstream of nucleotide position 1122, these recombinant viruses were subjected to biological tests, including monkey neurovirulence tests. The results of the monkey neurovirulence tests suggested the important contribution of an adenine residue (Mahoney type) at position 480 to the expression of the neurovirulence phenotype of type 1 poliovirus. This nucleotide, however, had only a minor effect, if any, on viral temperature sensitivity. Monkey neurovirulence tests on the recombinant virus whose genome had a guanine residue (Sabin type) at position 480 and variants generated from this recombinant virus in the central nervous system of monkeys strongly suggested that only one nucleotide change, from adenine to guanine, was not sufficient for full expression of the attenuation phenotype encoded by this genome region. These results suggest that the expression of the attenuation phenotype depends on the highly ordered structure formed in the 5' noncoding sequence and that the formation of such a structure is possibly influenced by the nucleotide at position 480. Furthermore, in vitro biological tests performed on viruses recovered from the central nervous system of monkeys injected with a temperature-sensitive recombinant virus showing the small-plaque and d phenotypes revealed that most of the recovered viruses had even higher temperature sensitivities and that all of the recovered viruses that had acquired the large-plaque phenotype had lost the d phenotype to some extent. These results indicate that there may be an unknown selection pressure(s) in the central nervous system and that common determinants might be involved in the expression of the small-plaque and d phenotypes.

Animals↗

Dimethylation of histone H3 at lysine 36 demarcates regulatory and nonregulatory chromatin genome-wide.

Set2p, which mediates histone H3 lysine 36 dimethylation (H3K36me2) in Saccharomyces cerevisiae, has been shown to associate with RNA polymerase II (RNAP II) at individual loci. Here, chromatin immunoprecipitation-microarray experiments normalized to general nucleosome occupancy reveal that nucleosomes within open reading frames (ORFs) and downstream noncoding chromatin were highly dimethylated at H3K36 and that Set2p activity begins at a stereotypic distance from the initiation of transcription genome-wide. H3K36me2 is scarce in regions upstream of divergently transcribed genes, telomeres, silenced mating loci, and regions transcribed by RNA polymerase III, providing evidence that the enzymatic activity of Set2p is restricted to its association with RNAP II. The presence of H3K36me2 within ORFs correlated with the "on" or "off" state of transcription, but the degree of H3K36 dimethylation within ORFs did not correlate with transcription frequency. This provides evidence that H3K36me2 is established during the initial instances of gene transcription, with subsequent transcription having at most a maintenance role. Accordingly, newly activated genes acquire H3K36me2 in a manner that does not correlate with gene transcript levels. Finally, nucleosomes dimethylated at H3K36 appear to be refractory to loss from highly transcribed chromatin. Thus, H3K36me2, which is highly conserved throughout eukaryotic evolution, provides a stable molecular mechanism for establishing chromatin context throughout the genome by distinguishing potential regulatory regions from transcribed chromatin.

Chromatin↗

Genomewide demarcation of RNA polymerase II transcription units revealed by physical fractionation of chromatin.

Epigenetic modifications of chromatin serve an important role in regulating the expression and accessibility of genomic DNA. We report here a genomewide approach for fractionating yeast chromatin into two functionally distinct parts, one containing RNA polymerase II transcribed sequences, and the other comprising noncoding sequences and genes transcribed by RNA polymerases I and III. Noncoding regions could be further fractionated into promoters and segments lacking promoters. The observed separations were apparently based on differential crosslinking efficiency of chromatin in different genomic regions. The results reveal a genomewide molecular mechanism for marking promoters and genomic regions that have a license to be transcribed by RNA polymerase II, a previously unrecognized level of genomic complexity that may exist in all eukaryotes. Our approach has broad potential use as a tool for genome annotation and for the characterization of global changes in chromatin structure that accompany different genetic, environmental, and disease states.

DNA Primers↗