Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Marked genomic heterogeneity and frequent mixed infection of TT virus demonstrated by PCR with primers from coding and noncoding regions.

A nonenveloped, single-stranded, and circular DNA virus designated TT virus (TTV) has been reported in association with hepatitis of unknown etiology. TTV has a wide sequence divergence (approximately 52%), by which it is classified into at least 16 genotypes separated by an evolutionary distance of >0.30. Therefore, the detection of TTV DNA by polymerase chain reaction would be influenced by primers deduced from conserved or divergent regions of the genome. Of the 30 sera from healthy individuals, up to 17% tested positive with primers deduced from coding region, much less frequently than up to 93% testing positive with primers from noncoding region. These differences were not attributable to the sensitivity of detection, because a cloned TTV DNA of genotype 1a was detected sensitively (up to 1 copy per test) with primers deduced from either the coding or the noncoding region of the same genotype. Sera testing positive only with noncoding region primers, or those showing higher titers with noncoding than coding region primers, contained TTV DNA strains with sequence divergence of 47-53% from the TA278 isolate of genotype 1a within the N22 region spanning 222-231 nucleotides. Some of the sera contained two or three TTV DNA strains of distinct genotypes. These results indicate TTV strains with extremely high sequence divergence prevailing in healthy individuals and frequent mixed infection with TTV strains of distinct genotypes.

Amino Acid Sequence↗

A chloroplast DNA mutational hotspot and gene conversion in a noncoding region near rbcL in the grass family (Poaceae).

The noncoding DNA region of the chloroplast genome, flanked by the genes rbcL and psaI (ORF36), has been sequenced for seven species of the grass family (Poaceae). This region had previously been observed as a hotspot area for length mutations. Sequence comparison reveals that short duplications, probably resulting from slipped-strand mispairing, account for many small length differences between sequences but that major mutational hotspots are localized in three small areas, two of which show potential secondary structure. Mutation in one of these hotspots appears to be a result of more complex recombination events. All seven species contain a pseudogene for rpl23 and evidence is presented that this pseudogene is being maintained by gene conversion with the functional gene. Different transition/transversion biases and AT contents between the pseudogene and the surrounding noncoding sequences are noted. In the subfamily Panicoideae there is a deletion in which almost 1 kb of ancestral sequence, including the 3' end of the rpl23 pseudogene, has been replaced by a non-homologous 60-base sequence of unknown origin. Two other deletions of almost the same region have occurred in the grass family. The deleted noncoding region has mutational and compositional properties similar to the rbcL coding sequence and the rpl23 pseudogene. The three independent deletions, as well as the pattern of mutation in the localized hotspots, indicate that such noncoding DNA may be misleading for studies of phylogenetic inference.

Base Sequence↗

The sources of variation in the human genome and genome instability in human cancers.

The human genome is viewed as a stable collection of about 60,000-70,000 genes--a minority of protein--coding DNA sequences--dispersed in a large majority of noncoding DNA sequences--more than 90 per cent of the entire genome sequences. Some of these ubiquitous noncoding DNA sequences, metonymically called "parasitic DNA," "ballast DNA," "selfish DNA" or "extra DNA," especially, the repeated sequences tandemly organized, are not stable but vary with considerable frequency. Recently, the confused or inadequately known origin of native of pathological variations of these DNA sequences appears to be unravelled, with great implications in genome stability. The human chromosomes, the bearer of genome, store and carry it. Their structure is qualified to perform its fastidious functions. The chromosomal conformation, "with variable geometry," exposed to genetoxic action of different damaging factors and to torsional stress after their fast and repeated changes during mitosis. The exaggerate exceeding of the native variation of human genome in disease states, probably, generates genome instability. The chromosome fragility--the cellular phenotypic expression of these molecular instability--reflects the closely relations between the genome and its carrier. The pattern of DNA replication with asynchrony of different domains of "parcelled" genome and the results of replication, susceptible to be corrected by the action of DNA repair genes, render certain limited regions of genome more vulnerable to damaging. These "target" regions focused damaging effects and exhibit an increased susceptibility to breakage and recombination, often with chromosomal expression. The coincidence of these regions, frequently, with locations of many protooncogenes and sometimes, antioncogenes could be subsequently, starting points for a genuine chain of genomic events related to growth cell and cell division. Cancer multistage accumulation of various genomic disorders in a single cell tends to take advantage of discriminating situations of these regions, which themselves can generate other genetic disorders, involving its in carcinogenesis. The gene expression disorders or the genuine mutations of dominant protooncogenes and the recessive behaviour of antioncogenes explain the nature of human cancers--a mixture of inherited and somatically acquired gene disorders. They attest the recessive characteristic of human cell malignancy and emphasize the decisive role of cancer predisposition which operates in interaction with damaging environmental factors. Seemingly, the pivotal causes of genome instability originate from strange behaviour of certain repeated DNA sequences dispersed throughout the human genome. Perhaps they hold the key to the puzzle of cancer processes.

Chromosome Aberrations↗

An inflammatory bowel disease-linked lncRNA suppresses transcription factor T-BET expression in T cells to limit intestinal inflammation.

Among the tens of thousands of annotated long noncoding RNAs (lncRNAs) in the human genome, only a small fraction have been functionally characterized. Here, we show that a well-established inflammatory bowel disease (IBD) risk locus encoded a conserved lncRNA, lnc15 (2310015A10Rik/ENSMUSG00000097729), whose structure was destabilized by risk-associated variants, leading to its degradation. Deletion of lnc15 in mice resulted in molecular features of inflammation under steady-state conditions and conferred heightened susceptibility to experimental colitis. Lnc15 was abundantly expressed in T cells, with highest expression in regulatory T (Treg) cells. Mechanistically, lnc15 suppressed the transcription factor T-BET by recruiting the CCR4-NOT RNA degradation complex to Tbx21 mRNA. Our study identifies that lnc15 simultaneously enhances Treg cell suppressive function and impairs conventional T cell pathogenicity in the context of intestinal inflammation. Collectively, these findings identify lnc15 as a functional lncRNA that links noncoding genetic variation to immune regulation and prevention of mucosal inflammation. VIDEO ABSTRACT.

RNA, Long Noncoding↗

Compositional constraints and genome evolution.

Nucleotide sequences of all genomes are subject to compositional constraints that affect, to about the same extent, both coding and noncoding sequences; influence not only the structure and function of the genome, but also those of transcripts and proteins; are the result of environmental pressures; and largely control the fixation of mutations. These findings indicate that noncoding sequences are associated with biological functions; that the organismal phenotype comprises two components, the classical phenotype, corresponding to the "gene products," and a "genome phenotype," which is defined by the compositional constraints; and that natural selection plays a more important role in genome evolution than do random events.

Base Composition↗

Mitochondrial genome of the Komodo dragon: efficient sequencing method with reptile-oriented primers and novel gene rearrangements.

The mitochondrial genome of the Komodo dragon (Varanus komodoensis) was nearly completely sequenced, except for two highly repetitive noncoding regions. An efficient sequencing method for squamate mitochondrial genomes was established by combining the long polymerase chain reaction (PCR) technology and a set of reptile-oriented primers designed for nested PCR amplifications. It was found that the mitochondrial genome had novel gene arrangements in which genes from NADH dehydrogenase subunit 6 to proline tRNA were extensively shuffled with duplicate control regions. These control regions had 99% sequence similarity over 700 bp. Although snake mitochondrial genomes are also known to possess duplicate control regions with nearly identical sequences, the location of the second control region suggested independent occurrence of the duplication on lineages leading to snakes and the Komodo dragon. Another feature of the mitochondrial genome of the Komodo dragon was the considerable number of tandem repeats, including sequences with a strong secondary structure, as a possible site for the slipped-strand mispairing in replication. These observations are consistent with hypotheses that tandem duplications via the slipped-strand mispairing may induce mitochondrial gene rearrangements and may serve to maintain similar copies of the control region.

Animals↗

Characterization of the soluble guanylyl cyclase beta-subunit gene in the mosquito Anopheles gambiae.

Genomic DNA corresponding to the soluble guanylyl cyclase beta-subunit (GCSbeta) gene was cloned and sequenced from Anopheles gambiae. The sequence was 8103 bp long and presumably included the entire coding region. The deduced amino acid sequence was 71% and 62% similar to previously known Drosophila and vertebrate GCSbeta, while the C-terminus of A. gambiae GCSbeta was shorter. Because of the conserved characteristics in each functional domain, the high G+C% in the third codon positions compared to the introns, the lack of internal stop codons, and the fact that we identified the gene from a cDNA, we conclude that this A. gambiae gene is functional. This is the first detailed description of a guanylyl cyclase gene structure (e.g. intron-exon boundaries). Interestingly, within the fifth intron we found high similarity to the flanking regions of the Pegasus-27 transposable element and other noncoding regions of the A. gambiae genome.

Amino Acid Sequence↗

Rabies virus glycoprotein gene contains a long 3' noncoding region which lacks pseudogene properties.

Analysis of a limited number of laboratory strains of rabies virus had demonstrated the presence of a genome region bounded by two transcription termination and polyadenylation-like (TTP) signals (approximately 400 to 450 nucleotides apart) which was located between the end of the glycoprotein (G) coding sequence and the beginning of the L polymerase coding sequence. Although this region had been suggested to represent a remnant or pseudogene (psi), no detailed analysis had been carried out to examine this possibility. Here we present the nucleotide sequence analysis of this genome region for several laboratory rabies virus strains and a large number of diverse rabies viruses detected directly in brain tissue of naturally infected animals. Only one distinct lineage of the laboratory strains and none of the wild-type rabies viruses contained the upstream TTP-like signal, indicating that only the downstream TTP motif is the authentic G mRNA transcription termination and polyadenylation and signal. Phylogenetic analysis of sequence differences provided no evidence of laboratory strains containing the two TTP-like signals being ancestral to any of the viruses possessing only the downstream TTP sequence motif. These data indicate that this region of the rabies virus genome encodes a G mRNA with a long 3' noncoding region with no evidence of a pseudogene.

Animals↗

Sequence analysis of the 5' non coding region of Turkish HCV isolates: implications for PCR diagnosis.

Hepatitis C virus (HCV) is a positive-strand RNA virus related to pestiviruses and flaviviruses. The 5' noncoding region (NCR) of the virus genome consists of 324-341 nucleotides and is generally highly conserved among different HCV isolates which has made this region the choice for primer selection in amplification of HCV sequences by polymerase chain reaction (PCR). In this study, we report the partial nucleotide sequences of the 5'-NCR from type 1a (n = 4), type 1b (n = 6) and type 4 (n = 1) Turkish HCV isolates. Sequence information was obtained by direct sequencing of RT-PCR product using biotinylated primers and single strands were sequenced using T7 DNA polymerase after binding to streptavidin coated magnetic beads. In comparison to prototype type 1a consensus sequence, all type 1b sequences had A-G substitution at position - 99. Nucleotid changes from the prototype 1a sequence were found in 12 of the 174 nucleotide positions. The most variable domain spans 51 nucleotides (positions - 167 to - 117) where nine polymorphic sites were identified. Although the nucleotide sequence of the 5'-noncoding region is highly conserved there are type-specific polymorphic sites within this region that has to be taken into consideration in the design of oligonucleotide primers for reliable amplification of sequences from different HCV genotypes.

Journal Article↗

[Progress of research on epigenetic and human disease.].

In the past few fears, there has been a nascent convergence of scientific understanding of human disease with epigenetic. Identified epigenetic processes involved in human disease include chromatin remodeling, genomic imprinting, X chromosome inactivation, and noncoding RNAs regulation. These processes influence chromatin structure and thereby regulate gene expression on the chromosome level or a cluster of linked genes level. Deregulation of these processes result in lots of disease which are characterized by complex patterns of mutations and associated phenotypes affecting pre- and postnatal growth, development, and neurological function. Epigenetic diseases are illustrated by the array of multi-system disorders and neoplasias and investigations of these diseases have an impact on biomedical research and provide interesting models for functions and mechanisms of epigenetic gene control.

Chromatin Assembly and Disassembly↗

Mammalian microevolution in action: adaptive edaphic genomic divergence in blind subterranean mole-rats.

Genomic diversity of anonymous regions across the genome, most probably including coding and noncoding amplified fragment length polymorphisms (AFLPs), was examined in 20 individuals of the blind mole-rat, Spalax galili, one of four allospecies of the Spalax ehrenbergi superspecies of blind subterranean mole-rats in Israel. We compared 10 individuals from two nearby populations in Upper Galilee, separated by only a few dozen to hundreds of metres and living in two sharply contrasting ecologies: white chalk and rendzina soil with Sarcopterium spinosum and Majorana syriaca versus black volcanic basalt soil with Carlina hispanica-Psorelea bitominosa and Alhagi graecorum plant formations. The microsite tested ranged in an area of less than 10000 m2. Out of 729 AFLP loci, 433 (59.4%) were polymorphic, with 211 soil unique alleles. Genetic polymorphism was significantly higher on the ecologically more xeric and stressful chalky rendzina soil than on the neighbouring mesic basalt soil. This is a remarkable pattern for a mammal that can disperse each generation between tens to hundreds of metres. These results cannot be explained by migration (which causes homogenization) or by chance (which will exclude sharp genomic soil divergence). Natural selection is the only evolutionary adaptive force that can cause genetic divergence across the genome matching the sharp microscale ecological contrast.

Animals↗

Detection of mitochondrial tDRs in killifish embryos and other non-model organisms.

In recent years a diversity of small noncoding RNAs have been identified that originate from the mitochondrial genome. These mitosRNAs are often dominated by tRNA-derived small RNAs (mito-tDRs). Differential expression of mito-tDRs is associated with responses to stress. They also appear to be expressed differentially during development and their expression may be enriched in stress-tolerant animals. Very little is currently known about roles or modes of action of these sequences, although they are implicated in a diversity of processes such as cell cycle regulation, mRNA stability, regulation of ROS production, and import of proteins into the mitochondrion. To better understand the various roles these sequences may play, it is critical that we understand their diversity, cellular location, and the context for their expression. This protocol outlines the methodologies used to detect mitosRNAs, including mito-tDRs, in embryos and cells of the annual killifish Austrofundulus limnaeus. We highlight critical steps in the isolation of RNA, creation of sequencing libraries, bioinformatics processing of sequence data, and methods for validation of expression that support a robust discovery pipeline for mitosRNAs even from species with incomplete reference genome sequences.

Animals↗

Enhancer-like properties of an RNA element that modulates Tombusvirus RNA accumulation.

Prototypical defective interfering (DI) RNAs of the plus-strand RNA virus tomato bushy stunt virus contain four noncontiguous segments (regions I-IV) derived from the viral genome. Region I corresponds to 5'-noncoding sequence, regions II and III are derived from internal positions, and region IV represents a 3'-terminal segment. We analyzed the internally located region III in a prototypical DI RNA to understand better its role in DI RNA accumulation. Our results indicate that (1) region III is not essential for DI RNA accumulation, but molecules that lack it accumulate at significantly reduced levels ( approximately 10-fold lower), (2) region III is able to function at different positions and in opposite orientations, (3) a single copy of region III is favored over multiple copies, (4) the stimulatory effect observed on DI RNA accumulation is not due to region III-mediated RNA stabilization, (5) DI RNAs lacking region III permit the efficient accumulation of head-to-tail dimers and are less effective at suppressing helper RNA accumulation, and (6) negative-strand accumulation is also significantly depressed for DI RNAs lacking region III. Collectively, these results support a role for region III as an enhancer-like element that facilitates DI RNA replication. A scanning-type mutagenesis strategy was used to define portions of region III important for its stimulatory effect on DI RNA accumulation. Interestingly, the results revealed several differences in the requirements for activity when region III was in the forward versus the reverse orientation. In the context of the viral genome, region III was found to be essential for biological activity. This latter finding defines a critical role for this element in the reproductive cycle of the virus.

Base Sequence↗

Minimum internal ribosome entry site required for poliovirus infectivity.

Translation initiation by internal ribosome binding is a recently discovered mechanism of eukaryotic viral and cellular protein synthesis in which ribosome subunits interact with the mRNAs at internal sites in the 5' untranslated RNA sequences and not with the 5' methylguanosine cap structure present at the extreme 5' ends of mRNA molecules. Uncapped poliovirus mRNAs harbor internal ribosome entry sites (IRES) in their long and highly structured 5' noncoding regions. Such IRES sequences are required for viral protein synthesis. In this study, a novel poliovirus was isolated whose genomic RNA contains two gross deletions removing approximately 100 nucleotides from the predicted IRES sequences within the 5' noncoding region. The deletions originated from previously in vivo-selected viral revertants displaying non-temperature-sensitive phenotypes. Each revertant had a different predicted stem-loop structure within the 5' noncoding region of their genomic RNAs deleted. The mutant poliovirus (Se1-5NC-delta DG) described in this study contains both stem-loop deletions in a single RNA genome, thereby creating a minimum IRES. Se1-5NC-delta DG exhibited slow growth and a pinpoint plaque phenotype following infection of HeLa cells, delayed onset of protein synthesis in vivo, and defective initiation during in vitro translation of the mutated poliovirus mRNAs. Interestingly, the peak levels of viral RNA synthesis in cells infected with Se1-5NC-delta DG occurred at slightly later times in infection than those achieved by wild-type poliovirus, but these mutant virus RNAs accumulated in the host cells during the late phases of virus infection. UV cross-linking assays with the 5' noncoding regions of wild-type and mutated RNAs were carried out in cytoplasmic extracts from HeLa cells and neuronal cells and in reticulocyte lysates to identify the cellular factors that interact with the putative IRES elements. The cellular proteins that were cross-linked to the minimum IRES may represent factors playing an essential role in internal translation initiation of poliovirus mRNAs.

Cross-Linking Reagents↗

The contribution of slippage-like processes to genome evolution.

Simple sequences present in long (> 30 kb) sequences representative of the single-copy genome of five species (Homo sapiens, Caenorhabditis elegans, Saccharomyces cerevisiae, E. coli, and Mycobacterium leprae) have been analyzed. A close relationship was observed between genome size and the overall level of sequence repetition. This suggested that the incorporation of simple sequences had accompanied increases of genome size during evolution. Densities of simple sequence motifs were higher in noncoding regions than in coding regions in eukaryotes but not in eubacteria. All five genomes showed very biased frequency distributions of simple sequence motifs in all species, particularly in eukaryotes where AAA and TTT predominated. Interspecific comparisons showed that noncoding sequences in eukaryotes showed highly significantly similar frequency distributions of simple sequence motifs but this was not true of coding sequences. ANOVA of the frequency distributions of simple sequence motifs indicated strong contributions from motif base composition and repeat unit length, but much of the variation remained unexplained by these parameters. The sequence composition of simple sequences therefore appears to reflect both underlying sequence biases in slippage-like processes and the action of selection. Frequency distributions of simple sequence motifs in coding sequences correlated weakly or not at all with those in noncoding sequences. Selection on coding sequences to eliminate undesirable sequences may therefore have been strong, particularly in the human lineage.

Animals↗

[Genomics and evolution of cellular organelles].

The structure, functions, and evolution of cellular organelles are reviewed. The mitochondrial genomes of eukaryotes differ considerably in size and structural organization mainly due to the length variation in noncoding regions and the presence of introns. The mitochondrial genomes of angiosperms are the largest and most complicated. Gene content in eukaryotic mitochondrial genomes is similar. They usually encode all types of rRNA, a complete or partial complement of tRNA, and a limited number of proteins essential for mitochondrial functions. In all eukaryotes studied, mitochondrial genomes code for two highly hydrophobic proteins involved in respiration, cytochrome b and subunit 1 of cytochrome oxidase. Genome structure and gene content in plastids, mainly in higher plant chloroplasts, are highly conserved. Plastid genomes of algae are more variable in gene composition and contain several unique genes absent in the chloroplast DNA of higher plants. Plastid genomes encode proteins involved in transcription and translation, as well as proteins of the photosynthetic apparatus. Both types of cellular organelles are supposed to be of endosymbiotic origin. Modern plastids originate from a cyanobacterial ancestor. Alpha-proteobacteria, especially the most mitochondrion-like rickettsia, gave rise to mitochondria. The origin of plastids of higher plants and green algae as a result of primary endosymbiosis and that of other algal lineages by secondary endosymbiosis are briefly discussed.

Animals↗

Isolation of a cDNA encoding Aspergillus oryzae Taka-amylase A: evidence for multiple related genes.

Complementary and genomic DNAs encoding Aspergillus oryzae Taka-amylase A (Taa) were cloned and sequenced. The coding sequence of the cDNA comprised the signal peptide [21 amino acids (aa)] and mature Taa (478 aa). The deduced aa sequence agrees well with the published aa sequence, except for one insertion, one deletion and ten aa substitutions. These differences might be due to the difference in the strains used. Sequence comparison of the cDNA and genomic DNA indicates the presence of eight introns ranging in size from 55 to 86 bp. Southern-blot analysis showed the presence of at least two Taa genes, and the second gene (Taa-G2) was isolated. All the intron/exon junctions follow the 'GT-AG' rule, except for intron I of the first gene (Taa-G1). The 5'-noncoding region was well conserved among the genomic genes and contained sequences similar to 'CAAT' and 'TATA' boxes at nucleotides -121 and -31, counted from the transcription start point, respectively. The 3'-noncoding regions, however, differed significantly from each other. Taa-G2 contains a sequence identical to that of several independent cDNA clones, suggesting that it may be the major transcribed gene in A. oryzae.

Amino Acid Sequence↗

Analysis of the genomic sequence of a human metapneumovirus.

We recently described the isolation of a novel paramyxovirus from children with respiratory tract disease in The Netherlands. Based on biological properties and limited sequence information the virus was provisionally classified as the first nonavian member of the Metapneumovirus genus and named human metapneumovirus (hMPV). This report describes the analysis of the sequences of all hMPV open reading frames (ORFs) and intergenic sequences as well as partial sequences of the genomic termini. The overall percentage of amino acid sequence identity between APV and hMPV N, P, M, F, M2-1, M2-2, and L ORFs was 56 to 88%. Some nucleotide sequence identity was also found between the noncoding regions of the APV and hMPV genomes. Although no discernible amino acid sequence identity was found between two of the ORFs of hMPV and ORFs of other paramyxoviruses, the amino acid content, hydrophilicity profiles, and location of these ORFs in the viral genome suggest that they represent SH and G proteins. The high percentage of sequence identity between APV and hMPV, their similar genomic organization (3'-N-P-M-F-M2-SH-G-L-5'), and phylogenetic analyses provide evidence for the proposed classification of hMPV as the first mammalian metapneumovirus.

Amino Acid Sequence↗