Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Alternative splicing of the human serotonin transporter gene.

To explore the structural basis for regulation of human serotonin transporter (hSERT) gene expression, we used primer extension and 5' rapid amplification of cDNA ends (5'RACE) techniques to estimate levels of and identify 5'-noncoding elements of hSERT mRNAs and genomic cloning to place these elements within the overall map of the hSERT gene. Primer extension on JAR cell mRNA suggested the presence of significant hSERT mRNA sequence upstream of the 5' end of our cloned hSERT cDNA. Using 5'RACE and reverse transcription-PCR (RT-PCR) methodologies, we cloned these sequences from brain and placenta and found this material to be composed of alternatively spliced exons using a previously reported noncoding exon (1A) and a novel 97-bp noncoding exon (1B). RT-PCR of JAR cell mRNA blotted with exon-specific oligonucleotides revealed both exons 1A and 1B to be regulated in a cholera toxin-dependent manner. To clarify the structure of the hSERT gene including exon 1B, we isolated and characterized genomic hSERT clones from Lambda Fix II and P1 artificial chromosome libraries. In agreement with previous findings, a single hSERT gene was identified that accounts for hybridizing bands on genomic Southern blots and was found to utilize 13 exons to encode the transporter's coding sequences along with the two noncoding 5' exons. Exon 1B was identified approximately 14 kb downstream of exon 1A in the hSERT gene and 737 bp upstream of exon 2, where the initiation site for translation is located. Exon 1B is surrounded by several elements potentially suitable for regulating serotonin transporter gene expression in vivo, including consensus sites for transcription factors AP-1, AP-2, CREB/ATF, and NF-kappaB. These data reveal additional complexity in hSERT gene structure and expression that may be relevant to regulated and compromised transporter expression in vivo.

Alternative Splicing↗

Noncoding RNA transcripts.

Recent analyses of the human genome and available data about the other higher eukaryotic genomes have revealed that, in contrast to Eubacteria and Archaea, only a small fraction of the genetic material (ca 1.5%) codes for proteins. Most of genomic DNA and its RNA transcripts are involved in regulation of gene expression, which can be exerted at either the transcriptional level, controlling whether a gene is transcribed and to what extent, or at the post-translational level, regulating the fate of the transcribed RNA molecules, including their stability, efficiency of their translation and subcellular localization. Noncoding RNA genes produce functional RNA molecules (ncRNAs) rather than encoding proteins. These stable RNAs act by multiple mechanisms such as RNA-RNA base pairing, RNA-protein interactions and intrinsic RNA activity, as well as regulate diverse cellular functions, including RNA processing, mRNA stability, translation, protein stability and secretion. Non-protein-coding RNAs are known to play significant roles. Along with transfer RNAs, ribosomal RNAs and mRNAs, ncRNAs contribute to gene splicing, nucleotide modification, protein transport and regulation of gene expression.

Animals↗

Identification of putative noncoding RNAs among the RIKEN mouse full-length cDNA collection.

With the sequencing and annotation of genomes and transcriptomes of several eukaryotes, the importance of noncoding RNA (ncRNA)-RNA molecules that are not translated to protein products-has become more evident. A subclass of ncRNA transcripts are encoded by highly regulated, multi-exon, transcriptional units, are processed like typical protein-coding mRNAs and are increasingly implicated in regulation of many cellular functions in eukaryotes. This study describes the identification of candidate functional ncRNAs from among the RIKEN mouse full-length cDNA collection, which contains 60,770 sequences, by using a systematic computational filtering approach. We initially searched for previously reported ncRNAs and found nine murine ncRNAs and homologs of several previously described nonmouse ncRNAs. Through our computational approach to filter artifact-free clones that lack protein coding potential, we extracted 4280 transcripts as the largest-candidate set. Many clones in the set had EST hits, potential CpG islands surrounding the transcription start sites, and homologies with the human genome. This implies that many candidates are indeed transcribed in a regulated manner. Our results demonstrate that ncRNAs are a major functional subclass of processed transcripts in mammals.

Animals↗

Genetic variation in vivo and proposed functional domains of the 5' noncoding region of poliovirus RNA.

Poliovirus has a single-stranded RNA genome of about 7,440 nucleotides (nt) with an unusually long 750-nt noncoding region in the 5' end (5'NCR). Several regulatory functions have been assigned to the 5'NCR. We sequenced the 5'NCRs of 33 wild-type 3 poliovirus strains to study the range and distribution of naturally occurring sequence variations. In this regard, the 5'NCR can be divided into a conserved part (nt 1 to 650) and a hypervariable part (nt 651 to 750). In the conserved part, altogether 234 unevenly distributed nucleotide positions (36%) showed variation. When these positions were plotted against the predicted secondary-structure models, it was found that the existence of most of the proposed stem-loop structures was supported by extensive structure-conserving substitutions in the stems. Regions with conserved sequences, as well as mutational hot spots, were observed. The hypervariable part of the 5'NCR varied up to 56% between the strains studied. The A + U percentage was significantly higher than in the conserved part. The number of AUG codons varied between 5 and 15 in the conserved part of the 5'NCR, while none was found in the hypervariable part. These results provide information that can be used in site-directed mutagenesis and other approaches targeted to reveal the functional domains of the 5'NCR.

Base Sequence↗

Characterization of the human parathyroid hormone-like peptide gene. Functional and evolutionary aspects.

The single-copy gene coding for the human parathyroid hormone-like peptide was isolated from a human placental genomic library. The gene spans 13 kilobases and contains seven exons. Exons I and II encode 5'-noncoding regions; each has its own transcription initiation site, and the two promoters are separated by over 1000 base pairs of genomic DNA. Exon III encodes the prepro-coding region, and exon IV encodes the mature peptide sequence. At the end of exon IV the splice site interrupts codon 139 of the mature peptide. Exon V, which is contiguous with exon IV, encodes a stop codon and a 3'-noncoding region. Exon VI encodes 34 additional amino acids, a stop codon, and a second 3'-noncoding region. Exon VII encodes two extra amino acids, a stop codon, and a third 3'-noncoding region. This genomic organization reveals how the multiple human parathyroid hormone-like peptide RNA transcripts, which have been observed, arise by both alternative splicing out of exons and use of multiple promoters. The mRNAs, which can potentially be formed from the primary transcript of this gene, could have one of three different carboxyl-terminal coding regions. The use of different exons to encode the different functional domains, 5'-noncoding region, pre-pro-coding region, and mature peptide region is identical to the organization of the human parathyroid hormone gene. This strongly suggests a common evolutionary origin of the two genes.

Amino Acid Sequence↗

Structural organization of the human flavin-containing monooxygenase 3 gene (FMO3), the favored candidate for fish-odor syndrome, determined directly from genomic DNA.

The inherited metabolic disorder trimethylaminuria (fish-odor syndrome) is associated with defective hepatic N-oxidation of dietary-derived trimethylamine catalyzed by flavin-containing monooxygenase (FMO). As FMO3 encodes the major form of FMO expressed in adult human liver, it represents the best candidate gene for the disorder. The structural organization of FMO3 was determined by sequencing the products of exon-to-exon and vectorette PCR, the latter through the use of vectorette libraries constructed directly from genomic DNA. The gene contains one noncoding and eight coding exons. Knowledge of the exon/intron organization of the human FMO3 gene enabled each of the coding exons of the gene, together with their associated flanking intron sequences, to be amplified from genomic DNA and will thus facilitate the identification of mutations in FMO3 in families affected with fish-odor syndrome.

Exons↗

Association of protein tyrosine phosphatase 1B gene polymorphisms with type 2 diabetes.

The PTPN1 gene codes for protein tyrosine phosphatase 1B (PTP1B) (EC 3.1.3.48), which negatively regulates insulin signaling by dephosphorylating the phosphotyrosine residues of the insulin receptor kinase activation segment. PTPN1 is located in 20q13, a genomic region linked to type 2 diabetes in multiple genetic studies. Surveys of the gene have previously identified only a few uncommon coding single nucleotide polymorphisms (SNPs). We have carried out a detailed association analysis of 23 noncoding SNPs spanning the 161-kb genomic region, which includes the PTPN1 gene. These SNPs have been assessed for association with type 2 diabetes in two independently ascertained collections of Caucasian subjects with type 2 diabetes and two control groups. Association is observed between multiple SNPs and type 2 diabetes. The most consistent evidence for association occurred with SNPs spanning the 3' end of intron 1 of PTPN1 through intron 8 (P values ranging from 0.043 to 0.004 in one case-control set and 0.038-0.002 in a second case-control set). Analysis of the combined case-control data increased the evidence of SNP association with type 2 diabetes (P = 0.005-0.0016). All of the associated SNPs lie in a single 100-kb haplotype block that encompasses the PTPN1 gene. Analysis of haplotypes indicates a significant difference between haplotype frequencies in type 2 diabetes case and control subjects (P = 0.0035-0.0056), with one common haplotype (36%) contributing strongly to the evidence for association with type 2 diabetes. Odds ratios calculated from single SNP or haplotype data are in the proximity of 1.3. Haplotype-based calculation of population-attributable risk (PAR) results in an estimated PAR of 17-20% based on different models and assumptions. These results suggest that PTPN1 is a significant contributor to type 2 diabetes susceptibility in the Caucasian population. This risk is likely due to noncoding polymorphisms.

Diabetes Mellitus, Type 2↗

The 3'-noncoding region of the chick myosin light-chain gene hybridizes to a family of repetitive sequences in the slime mold Dictyostelium discoideum.

During studies aimed at isolating myosin-specific genomic clones in Dictyostelium, we probed a lambda genomic library with a chicken myosin light-chain sequence (pML10). Many lambda recombinant Dictyostelium clones hybridized to the pML10 cDNA insert, indicating that this sequence was reiterated in the Dictyostelium genome. It was found that the 3'-noncoding region (pML10-NC) alone was responsible for these results. Dictyostelium DNA contained approximately 65 copies of a sequence(s) similar but not identical to that of pML10-NC. Southern blot analysis showed that pML10-NC hybridized to many Dictyostelium genomic DNA fragments of varying sizes generated by digestion with EcoRI, HindIII, or AluI. In addition, each of the Dictyostelium clones was different in its size, restriction map, and flanking sequences. It seems likely, therefore, that the sequences which hybridized to pML10-NC are scattered throughout the Dictyostelium genome and similar but not identical to each other or to pML10-NC. Thus, probing with pML10-NC has allowed us to select a family of closely related but not identical sequences. These D. discoideum sequences are not found in other slime mold species. No RNA complementary to pML10-NC was found in vegetative cells, 18 h culmination stage, spores, or 1- and 2-h germinating spores. pML10-NC-related sequences were present in two other Dictyostelium species but were absent in the related genus Polysphondylium.

Animals↗

N6-methyladenosine modification of a parvovirus-encoded small noncoding RNA facilitates viral DNA replication through recruiting Y-family DNA polymerases.

Human bocavirus 1 (HBoV1) is a human parvovirus that causes lower respiratory tract infections in young children. It contains a single-stranded (ss) DNA genome of ~5.5 kb that encodes a small noncoding RNA of 140 nucleotides known as bocavirus-encoded small RNA (BocaSR), in addition to viral proteins. Here, we determined the secondary structure of BocaSR in vivo by using DMS-MaPseq. Our findings reveal that BocaSR undergoes N6-methyladenosine (m6A) modification at multiple sites, which is critical for viral DNA replication in both dividing HEK293 cells and nondividing cells of the human airway epithelium. Mechanistically, we found that m6A-modified BocaSR serves as a mediator for recruiting Y-family DNA repair DNA polymerase (Pol) η and Pol κ likely through a direct interaction between BocaSR and the viral DNA replication origin at the right terminus of the viral genome. Thus, this report represents direct involvement of a viral small noncoding RNA in viral DNA replication through m6A modification.

Humans↗

mettannotator: a comprehensive and scalable Nextflow annotation pipeline for prokaryotic assemblies.

SUMMARY: In recent years, there has been a surge in prokaryotic genome assemblies, coming from both isolated organisms and environmental samples. These assemblies often include novel species that are poorly represented in reference databases creating a need for a tool that can annotate both well-described and novel taxa, and can run at scale. Here, we present mettannotator-a comprehensive, scalable Nextflow pipeline for prokaryotic genome annotation that identifies coding and noncoding regions, predicts protein functions, including antimicrobial resistance, and delineates gene clusters. The pipeline summarizes these results in a GFF (General Feature Format) file that can be easily utilized in downstream analysis or visualized using common genome browsers. Here, we show how it works on 200 genomes from 29 prokaryotic phyla, including isolate genomes and known and novel metagenome-assembled genomes, and present metrics on its performance in comparison to other tools. AVAILABILITY AND IMPLEMENTATION: The pipeline is written in Nextflow and Python and published under an open source Apache 2.0 licence. Instructions and source code can be accessed at https://github.com/EBI-Metagenomics/mettannotator. The pipeline is also available on WorkflowHub: https://workflowhub.eu/workflows/1069.

Software↗

Extensive homologous recombination among widely divergent TT viruses.

Analyses of a collection of full-length TT virus genomes showed nearly half of them to be recombinant. The results were highly significant and revealed homologous recombination both within and among genotypes, often involving extremely divergent lineages. Recombination breakpoints were significantly more common in the noncoding region of the TT virus genome than in the coding region.

DNA Viruses↗

Interpreting mammalian evolution using Fugu genome comparisons.

Recently, it has been shown that a significant number of evolutionarily conserved human-Fugu noncoding elements function as tissue-specific transcriptional enhancers in vivo, suggesting that distant comparisons are capable of identifying a particular class of regulatory elements. We therefore hypothesized that by juxtaposing human/Fugu and human/mouse conservation patterns we can define conservation criteria for discovering transcriptional regulatory elements specific to mammals. Genome-scale comparisons of noncoding human/Fugu evolutionary conserved elements (ECRs) and their humans/mouse counterparts revealed a particular signature common to human/mouse ECRs (>or=350 bp long, >or=77% identity) that are also conserved in fishes. This newly defined threshold identifies 90% of all human/Fugu noncoding ECRs without the assistance of human-Fugu genome alignments and provides a very efficient filter for identifying functional human/mouse ECRs.

Animals↗

Polydnavirus genes and genomes: emerging gene families and new insights into polydnavirus replication.

Polydnavirus genome sequencing is providing new insights into viral genome organization and viral gene function. Sequence analyses demonstrate that the genomes of these viral mutualists are largely noncoding but maintain genes and gene families that are unrelated to other viral genes. Interestingly, these organizational patterns in polydnavirus genomes are evident in both the bracovirus and ichnovirus genera, even though these two genera are evolutionarily unrelated. The identity and function of some polydnavirus gene families are considered with some functions experimentally supported and others implied by homology relationships with known insect genes. The evidence relative to polydnavirus origins and evolution is considered but remains an area of speculation. However, sequencing of these viral genomes has been informative and provides opportunities for productive investigation of these unusual mutualistic insect viruses.

Animals↗

Gene conversion between direct noncoding repeats promotes genetic and phenotypic diversity at a regulatory locus of Zea mays (L.).

While evolution of coding sequences has been intensively studied, diversification of noncoding regulatory regions remains poorly understood. In this study, we investigated the molecular evolution of an enhancer region located 5 kb upstream of the transcription start site of the maize pericarp color1 (p1) gene. The p1 gene encodes an R2R3 Myb-like transcription factor that regulates the flavonoid biosynthetic pathway in maize floral organs. Distinct p1 alleles exhibit organ-specific expression patterns on kernel pericarp and cob glumes. A cob glume-specific regulatory region has been identified in the distal enhancer. Further characterization of 6 single-copy p1 alleles, including P1-rr (red pericarp/red cob) and P1-rw (red pericarp and white cob), reveals 3 distinct enhancer types. Sequence variations in the enhancer are correlated with the p1 gene expression patterns in cob glume. Structural comparisons and phylogenetic analyses suggest that evolution of the enhancer region is likely driven by gene conversion between long direct noncoding repeats (approximately 6 kb in length). Given that tandem and segmental duplications are common in both animal and plant genomes, our studies suggest that recombination between noncoding duplicated sequences could play an important role in creating genetic and phenotypic variations.

Alleles↗

microRNAs exhibit high frequency genomic alterations in human cancer.

MicroRNAs (miRNAs) are endogenous noncoding RNAs, which negatively regulate gene expression. To determine genomewide miRNA DNA copy number abnormalities in cancer, 283 known human miRNA genes were analyzed by high-resolution array-based comparative genomic hybridization in 227 human ovarian cancer, breast cancer, and melanoma specimens. A high proportion of genomic loci containing miRNA genes exhibited DNA copy number alterations in ovarian cancer (37.1%), breast cancer (72.8%), and melanoma (85.9%), where copy number alterations observed in >15% tumors were considered significant for each miRNA gene. We identified 41 miRNA genes with gene copy number changes that were shared among the three cancer types (26 with gains and 15 with losses) as well as miRNA genes with copy number changes that were unique to each tumor type. Importantly, we show that miRNA copy changes correlate with miRNA expression. Finally, we identified high frequency copy number abnormalities of Dicer1, Argonaute2, and other miRNA-associated genes in breast and ovarian cancer as well as melanoma. These findings support the notion that copy number alterations of miRNAs and their regulatory genes are highly prevalent in cancer and may account partly for the frequent miRNA gene deregulation reported in several tumor types.

Breast Neoplasms↗

Benchmarking tools for the alignment of functional noncoding DNA.

BACKGROUND: Numerous tools have been developed to align genomic sequences. However, their relative performance in specific applications remains poorly characterized. Alignments of protein-coding sequences typically have been benchmarked against "correct" alignments inferred from structural data. For noncoding sequences, where such independent validation is lacking, simulation provides an effective means to generate "correct" alignments with which to benchmark alignment tools. RESULTS: Using rates of noncoding sequence evolution estimated from the genus Drosophila, we simulated alignments over a range of divergence times under varying models incorporating point substitution, insertion/deletion events, and short blocks of constrained sequences such as those found in cis-regulatory regions. We then compared "correct" alignments generated by a modified version of the ROSE simulation platform to alignments of the simulated derived sequences produced by eight pairwise alignment tools (Avid, BlastZ, Chaos, ClustalW, DiAlign, Lagan, Needle, and WABA) to determine the off-the-shelf performance of each tool. As expected, the ability to align noncoding sequences accurately decreases with increasing divergence for all tools, and declines faster in the presence of insertion/deletion evolution. Global alignment tools (Avid, ClustalW, Lagan, and Needle) typically have higher sensitivity over entire noncoding sequences as well as in constrained sequences. Local tools (BlastZ, Chaos, and WABA) have lower overall sensitivity as a consequence of incomplete coverage, but have high specificity to detect constrained sequences as well as high sensitivity within the subset of sequences they align. Tools such as DiAlign, which generate both local and global outputs, produce alignments of constrained sequences with both high sensitivity and specificity for divergence distances in the range of 1.25-3.0 substitutions per site. CONCLUSION: For species with genomic properties similar to Drosophila, we conclude that a single pair of optimally diverged species analyzed with a high performance alignment tool can yield accurate and specific alignments of functionally constrained noncoding sequences. Further algorithm development, optimization of alignment parameters, and benchmarking studies will be necessary to extract the maximal biological information from alignments of functional noncoding DNA.

Animals↗

Noncoding transcription controls the developmental dynamics of long-range gene regulation.

The genomic regions regulating gene expression are often themselves transcribed into a variety of noncoding RNAs (ncRNAs). However, the regulatory roles of this noncoding transcription remain largely unknown. By using live imaging, we reveal that the sequential transcription of ncRNAs emanating from distinct regulatory elements underlies gene activation in Drosophila embryos. Single-allele co-visualization uncovers that optimal gene activation is achieved by only moderate levels of enhancer activity. Disrupting enhancer-associated ncRNAs causes precocious gene activation, providing evidence that ncRNAs control the timing of gene expression in development. We further show that enhancer transcription can regulate long-range interactions within complex regulatory landscapes. We propose that ncRNAs locally modulate regulatory element activity in cis to shape genome organization and orchestrate the temporal control of gene expression in development.

Journal Article↗