Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Comparative genomic analysis reveals a distant liver enhancer upstream of the COUP-TFII gene.

COUP-TFII is a central nuclear hormone receptor that tightly regulates the expression of numerous target lipid metabolism genes in vertebrates. However, it remains unclear how COUP-TFII itself is transcriptionally controlled since studies with its promoter and upstream region fail to recapitulate the gene's liver expression. In an attempt to identify liver enhancers in the vicinity of COUP-TFII, we employed a comparative genomic approach. Initial comparisons between humans and mice of the 3470-kb gene-poor region surrounding COUP-TFII revealed 2023 conserved noncoding elements. To prioritize a subset of these elements for functional studies, we performed further genomic comparisons with the orthologous pufferfish (Fugu rubripes) locus and uncovered two anciently conserved noncoding sequences (CNS) upstream of COUP-TFII (CNS-62kb and CNS-66kb). Testing these two elements using reporter constructs in liver cells (HepG2) revealed that CNS-66kb, but not CNS-62kb, yielded robust in vitro enhancer activity. In addition, an in vivo reporter assay using naked DNA transfer with CNS-66kb linked to luciferase displayed strong reproducible liver expression in adult mice, further supporting its role as a liver enhancer. Together, these studies further support the utility of comparative genomics to uncover gene regulatory sequences based on evolutionary conservation and provide the substrates to better understand the regulation and expression of COUP-TFII.

Animals↗

Noncoding sequences conserved in a limited number of mammals in the SIM2 interval are frequently functional.

Cross-species DNA sequence comparison is a fundamental method for identifying biologically important elements, because functional sequences are evolutionarily conserved, wheres nonfunctional sequences drift. A recent genome-wide comparison of human and mouse DNA discovered over 200,000 conserved noncoding sequences with unknown function. Multispecies DNA comparison has been proposed as a method to prioritize these conserved noncoding sequences for functional analysis based on the hypothesis that elements present in many species are more likely to be functional than elements present in limited numbers of species. Here, we perform a comparative analysis of the single-minded 2 (SIM2) gene interval on human chromosome 21 with horse, cow, pig, dog, cat, and mouse DNA. We classify conserved sequences based on the number of mammals in which they are present, and experimentally test sequences in each class for function. As hypothesized, conserved sequences present in many mammals are frequently functional. Additionally, we demonstrate that sequences conserved in a limited number of mammals are also frequently functional. Examination of genomic deletions in chimpanzee and rhesus macaque DNA showed that several putatively functional conserved noncoding human sequences were absent in these primates. These findings suggest that functional conserved noncoding human sequences can be missing in other mammals, even closely related primate species.

Animals↗

DNA microarrays of the complex human cytomegalovirus genome: profiling kinetic class with drug sensitivity of viral gene expression.

We describe, for the first time, the generation of a viral DNA chip for simultaneous expression measurements of nearly all known open reading frames (ORFs) in the largest member of the herpesvirus family, human cytomegalovirus (HCMV). In this study, an HCMV chip was fabricated and used to characterize the temporal class of viral gene expression. The viral chip is composed of microarrays of viral DNA prepared by robotic deposition of oligonucleotides on glass for ORFs in the HCMV genome. Viral gene expression was monitored by hybridization to the oligonucleotide microarrays with fluorescently labelled cDNAs prepared from mock-infected or infected human foreskin fibroblast cells. By using cycloheximide and ganciclovir to block de novo viral protein synthesis and viral DNA replication, respectively, the kinetic classes of array elements were classified. The expression profiles of known ORFs and many previously uncharacterized ORFs provided a temporal map of immediate-early (alpha), early (beta), early-late (gamma1), and late (gamma2) genes in the entire genome of HCMV. Sequence compositional analysis of the 5' noncoding DNA sequences of the temporal classes, performed by using algorithms that automatically search for defined and recurring motifs in unaligned sequences, indicated the presence of potential regulatory motifs for beta, gamma1, and gamma2 genes. In summary, these fabricated microarrays of viral DNA allow rapid and parallel analysis of gene expression at the whole viral genome level. The viral chip approach coupled with global biochemical and genetic strategies should greatly speed the functional analysis of established as well as newly discovered large viral genomes.

Base Sequence↗

An RNA hairpin at the extreme 5' end of the poliovirus RNA genome modulates viral translation in human cells.

Several mutations were introduced into an infectious poliovirus cDNA clone by inserting different oligodeoxynucleotide linkers into preexisting DNA restriction endonuclease sites in the viral cDNA. Ten mutated DNAs were constructed whose lesions mapped in the 5' noncoding region or in the capsid coding region of the viral genome. Eight of these mutated cDNAs did not give rise to infectious virus upon transfection into human cells, one yielded virus with a wild-type phenotype, and one gave rise to a viral mutant with a small-plaque phenotype. This last mutant, designated 1-5NC-S21, bears a 6-nucleotide insertion in the loop of a stable RNA hairpin at the very 5' end of the viral genome. Detailed analysis of the biological properties of 1-5NC-S21 showed that the primary defect in mutant-infected cells is a fivefold decrease in translation relative to wild-type-infected cells. Transfection into HeLa cells of in vitro-synthesized RNA molecules bearing either the 5' noncoding region of 1-5NC-S21 or wild-type poliovirus upstream of a luciferase reporter gene showed that the mutated RNA hairpin was responsible for the observed decrease in viral translation in mutant-infected cells and conferred this defect to heterologous RNAs. These findings indicate that an RNA hairpin located at the extreme 5' end of the viral RNA and highly conserved among enteroviruses and rhinoviruses profoundly affects the translation efficiency of poliovirus RNA in infected cells.

Base Sequence↗

Genomic organization of a 225-kb region in Xq28 containing the gene for X-linked myotubular myopathy (MTM1) and a related gene (MTMR1).

MTM1 is responsible for X-linked recessive myotubular myopathy, which is a congenital muscle disorder linked to Xq28. MTM1 is highly conserved from yeast to humans. A number of related genes also exist. The MTM1 gene family contains a consensus sequence consisting of the active enzyme site of protein tyrosine phosphatases (PTPs), suggesting that they belong to a new family of PTPs. Database searches revealed homology of myotubularin and all related peptides to the cisplatin resistance-associated alpha protein, which implicates an as yet unknown function. In addition, homology to the Sbf1 protein (SET binding factor 1), involved in the oncogenic transformation of fibroblasts and differentiation of myoblasts, was also evident. We describe 225 kb of genomic sequence containing MTM1 and the related gene, MTMR1, which lies 20 kb distal to MTM1. Although there is only moderate conservation of the exons, the striking similarity in the gene structures indicates that these two genes arose by duplication. Calculations suggest that this event occurred early in evolution long before separation of the human and mouse lineages. So far, mutations have been identified in the coding sequence of only 65% of the patients analyzed, indicating that the remaining mutations may lie in noncoding regions of MTM1 or possibly in MTMR1. Knowledge of the genomic sequence will facilitate mutation analyses of the coding and noncoding sequences of MTM1 and MTMR1.

Amino Acid Sequence↗

Mutation pattern variation among regions of the primate genome.

We sequenced three argininosuccinate-synthetase-processed pseudogenes (PsiAS-A1, PsiAS-A3, PsiAS-3) and their noncoding flanking sequences in human, orangutan, baboon, and colobus. Our data showed that these pseudogenes were incorporated into the genome of the Old World monkeys after the divergence of the Old World and New World monkey lineages. These pseudogene flanking regions show variable mutation rates and patterns. The variation in the G/C to A/T mutation rate (u) can account for the unequal GC contents at equilibrium: 34.9, 36.9, and 41.7% in the pseudogene PsiAS-A1, PsiAS-A3, and PsiAS-3 flanking regions, respectively. The A/T to G/C mutation rate (v) seems stable and the u/v ratios equal 1.9, 1.7, and 1.4 in the flanking regions of PsiAS-A1, PsiAS-A3, and PsiAS-3, respectively. These "regional" variations of the mutation rate affect the evolution of the pseudogenes, too. The ratio u/v being greater than 1.0 in each case, the overall mutation rate in the GC-rich pseudogenes is, as expected, higher than in their GC-poor flanking regions. Moreover, a "sequence effect" has been found. In the three cases examined u and v are higher (at least 20%) in the pseudogene than in its flanking region-i.e., the pseudogene appears as mutation "hot" spots embedded in "cold" regions. This observation could be partly linked to the fact that the pseudogene flanking regions are long-standing unconstrained DNA sequences, whereas the pseudogenes were relieved of selection on their coding functions only around 30-40 million years ago. We suspect that relatively more mutable sites maintained unchanged during the evolution of the argininosuccinate gene are able to change in the pseudogenes, such sites being eliminated or rare in the flanking regions which have been void of strong selective constraints over a much longer period. Our results shed light on (1) the multiplicity of factors that tune the spontaneous mutation rate and (2) the impact of the genomic position of a sequence on its evolution.

Animals↗

Cloning and characterization of the extreme 5'-terminal sequences of the RNA genomes of GB virus C/hepatitis G virus.

The extreme 5'-terminal sequences of the GB virus C/hepatitis G virus (GBV-C/HGV), containing elements essential for regulation of viral gene expression and replication, have not been determined. By using a RNA-ligase-mediated RACE (rapid amplification of the cDNA ends) procedure, we have cloned the extreme 5'-terminal sequences of the viral genome from the serum of three Taiwanese patients. Sequence analysis of the 5' noncoding region in alignment with one West African and two American isolates showed that (i) a consensus 5'-end sequence was cloned; (ii) about 97% of sequences were homologous among the three Taiwan isolates and also between the two American isolates, whereas about 90% of sequences were homologous among the isolates from the three different geographic areas; (iii) the sequence heterogeneity related to geographic separation is confined mainly to three domains; and (iv) a potential hairpin structure, resembling the hairpin structure found in the 5' end of hepatitis C virus genome, was detected in the 5' end of the noncoding region. Our data support the hypotheses that (i) the extreme 5' end of the hepatitis GBV-C/HGV viral genome has been cloned, (ii) there are different genotypes correlated with geographic separation, and (iii) the viral translation and replication mechanisms may be similar to that of hepatitis C virus and pestiviruses. Our data have not only shed light on the viral replication mechanism but also offer information for selection of optimal primer sequences for the detection and genotyping of the hepatitis GBV-C/HGV virus by PCR assays.

Base Sequence↗

The genome of Nanoarchaeum equitans: insights into early archaeal evolution and derived parasitism.

The hyperthermophile Nanoarchaeum equitans is an obligate symbiont growing in coculture with the crenarchaeon Ignicoccus. Ribosomal protein and rRNA-based phylogenies place its branching point early in the archaeal lineage, representing the new archaeal kingdom Nanoarchaeota. The N. equitans genome (490,885 base pairs) encodes the machinery for information processing and repair, but lacks genes for lipid, cofactor, amino acid, or nucleotide biosyntheses. It is the smallest microbial genome sequenced to date, and also one of the most compact, with 95% of the DNA predicted to encode proteins or stable RNAs. Its limited biosynthetic and catabolic capacity indicates that N. equitans' symbiotic relationship to Ignicoccus is parasitic, making it the only known archaeal parasite. Unlike the small genomes of bacterial parasites that are undergoing reductive evolution, N. equitans has few pseudogenes or extensive regions of noncoding DNA. This organism represents a basal archaeal lineage and has a highly reduced genome.

Arabidopsis↗

DIS3 licenses B cells for plasma cell differentiation in humans.

DIS3 is the main catalytic subunit of the nuclear RNA exosome, a complex playing a crucial role in RNA processing and the degradation of various noncoding RNA substrates. In mice, DIS3 is essential for genomic rearrangements during B cell development, but its role in terminal plasma cell (PC) differentiation has not been explored. Although DIS3 gene alterations are frequent in multiple myeloma (MM), a PC malignancy, their molecular impact remains poorly understood. In this study, we developed an antisense oligonucleotide strategy to knock down DIS3 expression in a well-characterized model of human PC differentiation. Reducing DIS3 expression systematically led to decreased B cell proliferation and impaired PC differentiation with lower levels of switched immunoglobulin secretion. Transcriptome analyses confirmed alterations in the proliferation and differentiation programs, alongside an accumulation of noncoding RNAs. Notably, centromere-associated noncoding RNAs were highly sensitive to DIS3 activity, and their accumulation in DIS3-deficient cells, either as transcripts or DNA-associated RNAs, correlated with the mislocalization of the centromere-specific histone variant CENP-A. We finally observed reduced physiological DNA recombination and somatic hypermutation but increased genomic instability in DIS3-deficient cells, in agreement with the higher levels of IGH translocations observed in our large cohort of DIS3-mutant MM patients. Together, these results underscore the essential role of DIS3 in regulating B cell proliferation, DNA recombination, and physiological or malignant PC differentiation in humans.

Humans↗

Molecular characterization of PL97-1, the first Korean isolate of the porcine reproductive and respiratory syndrome virus.

We determined the complete nucleotide and predicted amino acid sequence of the genomic RNA of PL97-1, the first Korean strain of porcine reproductive and respiratory syndrome virus (PRRSV), which was isolated from the serum of an infected pig in 1997. We found that the 15411-nucleotide genome of PL97-1 consisted of a 189-nucleotide 5' noncoding region (NCR), a 15071-nucleotide protein-coding region, and a 151-nucleotide 3'NCR, followed by a poly (A) tail. The 5'-end of PL97-1 began with 1ATG ACG TAT AGG12. Comparison of the PL97-1 genome with the 11 fully sequenced PRRSV genomes currently available revealed sequence divergence ranging from 0.3% (the VR-2332-derived vaccine MLV RespPRRS/Repro strain) to 38% (the Dutch Lelystad strain). To better understand the genetic relationships between these different strains, phylogenetic analyses were performed on the full-length PRRSV genomes. Significantly, the phylogenetic tree based on the ORF1b or ORF7 genes most closely resembled the tree based on the full-length genomes. Thus, these single genes will be the most useful in revealing the genetic relationships between the different strains relative to their geographical distribution. Extensive phylogenetic analyses using the ORF7 sequences of 111 PRRSV isolates available revealed that PL97-1 is most closely related to the North American genotype VR-2332, a VR-2332-derived vaccine strain, and Chinese BJ-4. It is distantly related to the European genotype Lelystad. This study provides the largest full-length genome phylogenetic analysis of PRRSV that has been published to date, and supports an earlier genetic grouping of the many temporally and geographically diverse PRRSV strains currently isolated.

Amino Acid Sequence↗

GBRAP: A Comprehensive Database and Tool for Exploring Genomic Diversity Across All Domains of Life.

Evolutionary studies require extensive examination of genomic information across all domains of life. Despite the availability of a large number of genomes through GenBank, the effective visualization or comparison of the information they contain is challenging due to many reasons, including their size. We introduce genome-based retrieval and analysis parser, a comprehensive software tool to analyze genome files, and an online database housing an extensive collection of carefully curated, high-quality genome statistics for all the organisms available in the RefSeq database of National Center for Biotechnology Information. Users can either directly search, or select from precategorized groups, the organisms of their choice and retrieve data, and the output is generated as tables containing more than 200 columns of useful genomic information (base counts, GC content, Shannon entropy, codon usage, etc.) separately calculated for different genomic elements (e.g. coding sequences, introns, transfer RNA, ribosomal RNA, noncoding RNA, etc.). The data are independently displayed (if applicable) for each chromosomal, mitochondrial, plastid, or plasmid sequence. All the data can be visualized on the database or downloaded as comma-separated value or Excel files. The genome-based retrieval and analysis parser database is free to access without any registration and is publicly available at http://tacclab.org/gbrap/.

Software↗

Gene structure of Bombyx mori larval serum protein (BmLSP).

To understand the molecular mechanisms of the larval-specific transcription of Bombyx mori larval serum protein (BmLSP), we isolated a clone of the BmLSP gene from a genomic library and sequenced a 3.5-kb fragment. An intron was found in the 5' noncoding region of the BmLSP gene. A putative transcription start point was determined by primer extension analysis. Genomic Southern hybridization showed that there is one copy of the BmLSP gene in a haploid genome. A database search revealed that the BmLSP gene has presumptive repetitive sequences found in other B. mori genes, the sequence homologous to ecdysone-responsive elements and a heptamer sequence found in storage protein genes.

Amino Acid Sequence↗

A developmentally regulated and cAMP-repressible gene of Dictyostelium discoideum: cloning and expression of the gene encoding cyclic nucleotide phosphodiesterase inhibitor.

A 1.6-kb genomic fragment containing the coding region for the inhibitor (PDI) of cyclic nucleotide phosphodiesterase (PD) was isolated and sequenced. The genomic sequence includes 510 nucleotides (nt) of 5'-noncoding sequence and the full coding sequence, which contains two small introns. From the deduced amino acid (aa) sequence we predict a 26-kDa protein that, in agreement with previous data, contains approximately 15% Cys residues. The PDI possesses a hydrophobic leader sequence, five potential glycosylation sites, and three internal repeats. Northern-blot analysis showed a single transcript of 0.95 kb. The gene encoding PDI (pdi) was expressed early in development with little transcript remaining following aggregation. The appearance of pdi transcript was inhibited by cAMP, but when cAMP was removed the transcript appeared within 30 min. When cAMP was applied to cells containing pdi mRNA, the transcript disappeared with a half-life of less than 30 min.

2',3'-Cyclic-Nucleotide Phosphodiesterases↗

Development and evaluation of an efficient 3'-noncoding region based SARS coronavirus (SARS-CoV) RT-PCR assay for detection of SARS-CoV infections.

The severe acute respiratory syndrome (SARS) epidemic originating from China in 2002 was caused by a previously uncharacterized coronavirus that could be identified by specific RT-PCR amplification. Efforts to control future SARS outbreaks depend on the accurate and early identification of SARS-CoV infected patients. A real-time fluorogenic RT-PCR assay based on the 3'-noncoding region (3'-NCR) of SARS-CoV genome was developed as a quantitative SARS diagnostic tool. The ideal amplification efficiency of a sensitive SARS-CoV RT-PCR assay should yield an E value (PCR product concentration increase per amplification cycle) equal to 2.0. It was demonstrated that the 3'-NCR SARS-CoV based RT-PCR reactions could be formulated to reach excellent E values of 1.81, or 91% amplification efficacy. The SARS-CoV cDNA preparations derived from viral RNA extract and the cloned recombinant plasmid both exhibit the identical amplification characteristics, i.e. amplification efficacy using the same PCR formulation developed in this study. The viral genomic copy (or genomic equivalences, GE) per infectious unit (GE/pfu) of SARS-CoV used in this study was also established to be approximate 1200-1600:1. The assay's detection sensitivity could reach 0.005 pfu or 6-8 GE per assay. It was preliminarily demonstrated that the assay could efficiently detect SARS-CoV from clinical specimens of SARS probable and suspected patients identified in Taiwan. The 3'-NCR based SARS-CoV assay demonstrated 100% diagnostic specificity testing samples of patients with acute respiratory disease from a non-SARS epidemic region.

3' Untranslated Regions↗

Functional characterization of lncIMF_17214 in regulating intramuscular fat deposition of yellow-feathered broilers.

Intramuscular fat (IMF) content and lipid composition are key determinants of both the nutritional value and sensory attributes of poultry meat, yet the underlying regulatory mechanisms remain insufficiently elucidated. In this study, triglyceride (TG) content was employed as a quantitative phenotypic proxy to dissect the molecular basis of IMF deposition in yellow-feathered broilers. By integrating TG phenotypic data from 315 individuals with transcriptomic profiles and whole-genome resequencing datasets, a TG-associated long noncoding RNA (lncRNA), lncIMF_17214, was identified. Functional characterization revealed that lncIMF_17214 functions as a negative regulator of lipid deposition. Specifically, its knockdown led to significant increases in TG and total cholesterol concentrations, promoted lipid droplet accumulation, and decreased shear force in breast muscle, whereas its overexpression elicited the opposite effects. Mechanistically, lncIMF_17214 interacts with the RNA-binding protein CNBP, forming a regulatory complex that inhibits lipid accumulation. Furthermore, liver-directed overexpression increased the abundance of lncIMF_17214 in plasma exosomes, while liver-directed manipulation was associated with changes in hepatic and breast-muscle lipid deposition; direct exosome-mediated transfer to intramuscular adipocytes remains to be established. Transcriptomic profiling coupled with pathway enrichment analyses demonstrated that lncIMF_17214 predominantly influences steroid biosynthesis, unsaturated fatty acid metabolism, and peroxisome proliferator-activated receptor (PPAR) signaling pathways. This suggests that it may be involved in the regulation of these pathways, although the underlying molecular mechanisms remain to be further elucidated. Collectively, these findings define a lncIMF_17214-centered regulatory axis linking intracellular and systemic lipid metabolism and provide a robust molecular framework for the targeted improvement of meat quality traits in yellow-feathered broilers.

Breast muscle↗

Determination of the leukaemogenicity of a murine retrovirus by sequences within the long terminal repeat.

Although the murine retrovirus SL3-3 is highly leukaemogenic, in both the structure of its genome and in its properties of replication in tissue culture it closely resembles the nonleukaemogenic retrovirus Akv (refs 3, 4). An earlier investigation of the properties of recombinant SL3-3-Akv viruses localized the major determinant of leukaemogenicity outside the env gene, in a region of the viral genome that includes the gag gene and the noncoding long terminal repeat (LTR). To localize the determinant of SL3-3's leukaemogenicity more precisely we have now construced a recombinant provirus containing the LTR of SL3-3 and the coding region of Akv. The leukaemogenicity of these recombinants demonstrates that the determinant of leukaemogenicity lies within the SL3-3 LTR. Nucleotide sequencing of the LTRs of SL3-3 and Akv shows that they differ by a set of changes in the region thought to contain a transcriptional enhancer element. We suggest that enhancer region sequences are the major determinants of leukaemogenicity in these viruses.

Animals↗

Natural selection on human microRNA binding sites inferred from SNP data.

A fundamental problem in biology is understanding how natural selection has shaped the evolution of gene regulation. Here we use SNP genotype data and techniques from population genetics to study an entire layer of short, cis-regulatory sites in the human genome. MicroRNAs (miRNAs) are a class of small noncoding RNAs that post-transcriptionally repress mRNA through cis-regulatory sites in 3' UTRs. We show that negative selection in humans is stronger on computationally predicted conserved miRNA binding sites than on other conserved sequence motifs in 3' UTRs, thus providing independent support for the target prediction model and explicitly demonstrating the contribution of miRNAs to darwinian fitness. Our techniques extend to nonconserved miRNA binding sites, and we estimate that 30%-50% of these are functional when the mRNA and miRNA are endogenously coexpressed. As we show that polymorphisms in predicted miRNA binding sites are likely to be deleterious, they are candidates for causal variants of human disease. We believe that our approach can be extended to studying other classes of cis-regulatory sites.

3' Untranslated Regions↗

Genetic structure and reproduction dynamics of Salix reinii during primary succession on Mount Fuji, as revealed by nuclear and chloroplast microsatellite analysis.

The early stage of volcanic desert succession is underway on the southeastern slope of Mount Fuji. We used markers of nuclear microsatellites (simple sequence repeats; SSR) and chloroplast microsatellites (cpSSR) to investigate the population genetic structure and reproduction dynamics of Salix reinii, one of the dominant pioneer shrubs in this area. The number of S. reinii genets in a patch and the area of the largest genet within the patch increased with patch area, suggesting that both clonal growth and seedling recruitment are involved in the reproduction dynamics of S. reinii. Five polymorphic cpSSR markers were developed for S. reinii by sequencing the noncoding regions between universal sequences in the chloroplast genome. Nineteen different cpSSR haplotypes were identified, indicating that S. reinii pioneer genets were created by the long-distance dispersal of seeds originating from different mother genets around the study site, where all vegetation was destroyed during the last eruption. Furthermore, the clustered distributions of different haplotypes within each patch or plot suggested that newly colonized genets tended to be generated from seeds dispersed near the initially established mother genets. These results revealed that the establishment of the S. reinii population on the southeastern slope of Mount Fuji involved two sequential modes of seed dispersal: long-distance dispersal followed by short-distance dispersal.

Cell Nucleus↗