Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Molecular taxonomy of a new potyvirus isolated from chilli pepper in Thailand.

A virus isolated from chilli pepper plants in Kamphaengsaen, Nakorn Pathom, showing vein banding mottle symptoms was classified using sequence analysis and phylogeny of the coat protein gene and 3' noncoding region (3'NCR). This virus was found to be a typical potyvirus on the basis of particle morphology, biological properties and cytopathology. The 3'terminus region of the genome of 1,309 nucleotides, representing the viral coat protein gene and 3' NCR was cloned and sequenced. Nucleotide sequence analysis indicated that the 3' region of the viral genome had a poly A tail of at least 12 nucleotides, a noncoding region of 272 nucleotides, a coat protein gene of 864 nucleotides and 161 nucleotides representing the 3' terminus of the polymerase gene. The amino acid sequence of the coat protein was compared with those of 23 distinct potyviruses, and 63.1% shown to be the highest homology. However the 3' NCR had, at most, 29.7% random homology, thus indicating that this virus is a distinct species in the genus Potyvirus in the family Potyviridae. The result is well supported by previous studies on the biology and biochemical properties of this virus.

Amino Acid Sequence↗

Multilocus patterns of nucleotide variability and the demographic and selection history of Drosophila melanogaster populations.

Uncertainty about the demographic history of populations can hamper genome-wide scans for selection based on population genetic models. To obtain a portrait of the effects of demographic history on genome variability patterns in Drosophila melanogaster populations, we surveyed noncoding DNA polymorphism at 10 X-linked loci in large samples from three African and two non-African populations. All five populations show significant departures from expectations under the standard neutral model. We detect weak but significant differentiation between East (Kenya and Zimbabwe) and West/Central sub-Saharan (Gabon) African populations. A skew toward high-frequency-derived polymorphisms, elevated levels of linkage disequilibrium (LD) and significant heterogeneity in levels of polymorphism and divergence in the Gabon sample suggest that this population is further from mutation-drift equilibrium than the two Eastern African populations. Both non-African populations harbor significantly higher levels of LD, a large excess of high-frequency-derived mutations and extreme heterogeneity among loci in levels of polymorphism and divergence. Rejections of the neutral model in D. melanogaster populations using these and similar features have been interpreted as evidence for an important role for natural selection in shaping genome variability patterns. Based on simulations, we conclude that simple bottleneck models are sufficient to account for most, if not all, polymorphism features of both African and non-African populations. In contrast, we show that a steady-state recurrent hitchhiking model fails to account for several aspects of the data. Demographic departures from equilibrium expectations in both ancestral and derived populations thus represent a serious challenge to detecting positive selection in genome-wide scans using current methodologies.

Animals↗

Genomic clones encoding chicken myosin heavy-chain genes.

A chicken genomic library was screened with a cDNA probe containing the 3' coding and noncoding portions of quail fast-twitch skeletal muscle myosin heavy chain (MHC). This probe hybridized to seven to nine bands on Southern blots of chicken genomic DNA, and 17 clones that hybridized to this probe were obtained from the genomic library. Partial restriction maps were constructed and probable orientation of transcription was determined for each of the 17 clones. These maps indicate the presence of at least 14 unique MHC genes or pseudogenes. Dot-blot hybridization analysis using DNA complementary to RNA from a variety of chicken tissues demonstrated that these genes are all related to the gene for sarcomeric MHC, and permitted tentative assignment of the tissue of expression for several of the MHC isoforms. To substantiate further the dot-blot data, a subclone of one of the genes (4b1), which showed significant homology with adult breast muscle RNA but which also showed weaker hybridization to RNA from other tissues, was sequenced. The sequence data verified that the clone contains a portion of a MHC gene, that it contains both 3' coding and noncoding regions, and that its predicted amino acid sequence is identical (with 96% nucleotide homology) to that of the 75-bp quail fast MHC cDNA clone published by Hastings and Emerson (1982). Thus, clone 4b1 contains a portion of one of the genes that is expressed in adult chicken breast skeletal muscle tissue.

Amino Acid Sequence↗

Functional phenotyping of genomic variants using joint multiomic single-cell DNA-RNA sequencing.

Genetic variants (both coding and noncoding) can impact gene function and expression, driving disease mechanisms such as cancer progression. The systematic study of endogenous genetic variants is hindered by inefficient precision editing tools, combined with technical limitations in confidently linking genotypes to gene expression at single-cell resolution. We developed single-cell DNA-RNA sequencing (SDR-seq) to simultaneously profile up to 480 genomic DNA loci and genes in thousands of single cells, enabling accurate determination of coding and noncoding variant zygosity alongside associated gene expression changes. Using SDR-seq, we associate coding and noncoding variants with distinct gene expression in human induced pluripotent stem cells. Furthermore, we demonstrate that in primary B cell lymphoma samples, cells with a higher mutational burden exhibit elevated B cell receptor signaling and tumorigenic gene expression. SDR-seq provides a powerful platform to dissect regulatory mechanisms encoded by genetic variants, advancing our understanding of gene expression regulation and its implications for disease.

Humans↗

Sequence relationships between Kirsten retrovirus genomes and the genomes of other murine retroviruses.

RNA sequence relationships between the genomes of the Kirsten murine sarcoma virus (MSV-K) complex, the Kirsten murine leukemia virus (MuLV-K) complex, the Gross murine leukemia virus (MuLV-G), and the Moloney murine leukemia virus (MuLV-M) were investigated. Sedimentation analyses revealed the expected 30 and 34 S RNA subunits in the MSV-K complex and a previously undetected 30 S RNA subunit accompanying the 34 S RNA subunit in the MuLV-K complex. Nucleic acid hybridization data indicated that each Kirsten virus 30 S RNA subunit had about 40% sequence homology with the RNA genome of MuLV-G, although these sequences were only partially homologous between the two 30 S subunits. In contrast, the MuLV-K 34 S RNA subunit had 96% sequence homology with the MuLV-G genome, whereas the MSV-K 34 S RNA subunit displayed only 71% sequence homology with the MuLV-G genome. Similar relationships were indicated by oligonucleotide fingerprinting. The oligonucleotide data, taken with published sequence data on the MuLV-G and MuLV-M genomes, enabled us to construct partial sequence maps of the MuLV-K 34 S RNA subunit and the MSV-K 34 and 30 S RNA subunits. The sequence arrangements indicated that (1) the MuLV-K 34 S RNA subunit is a variant of the MuLV-G genome; (2) the MSV-K 34 S RNA subunit is a recombinant molecule, which maintains the length of its leukemia virus parent; and (3) the MSV-K 30 S RNA subunit may have been generated from the MuLV-K 34 S genome by a two-stage process, culminating in the retention of parental sequences only within the U5 and U3 noncoding segments and within several amino-terminal coding segments. Further examination of published retrovirus genome sequences revealed several strategically situated sets of potential recognition signals for transcription and translation and suggested a model for genetic recombination based on mRNA splicing signals and areas of limited sequence homology. This model may explain how foreign gene elements can be inserted into retrovirus genomes to generate either functional or defective recombinant retroviruses.

AKR murine leukemia virus↗

Expression profiling and comparative genomics identify a conserved regulatory region controlling midline expression in the zebrafish embryo.

Differential gene transcription is a fundamental regulatory mechanism of biological systems during development, body homeostasis, and disease. Comparative genomics is believed to be a rapid means for the identification of regulatory sequences in genomes. We tested this approach to identify regulatory sequences that control expression in the midline of the zebrafish embryo. We first isolated a set of genes that are coexpressed in the midline of the zebrafish embryo during somitogenesis stages by gene array analysis and subsequent rescreens by in situ hybridization. We subjected 45 of these genes to Compare and DotPlot analysis to detect conserved sequences in noncoding regions of orthologous loci in the zebrafish and Takifugu genomes. The regions of homology that were scored in nonconserved regions were inserted into expression vectors and tested for their regulatory activity by transient transgenesis in the zebrafish embryo. We identified one conserved region from the connective tissue growth factor gene (ctgf), which was able to drive expression in the midline of the embryo. This region shares sequence similarity with other floor plate/notochord-specific regulatory regions. Our results demonstrate that an unbiased comparative approach is a relevant method for the identification of tissue-specific cis-regulatory sequences in the zebrafish embryo.

Animals↗

The effects of selection against spurious transcription factor binding sites.

Most genomes contain nucleotide sequences with no known function; such sequences are assumed to be free of constraints, evolving only according to the vagaries of mutation. Here we show that selection acts to remove spurious transcription factor binding site motifs throughout 52 fully sequenced genomes of Eubacteria and Archaea. Examining the sequences necessary for polymerase binding, we find that spurious binding sites are underrepresented in both coding and noncoding regions. The average proportion of spurious binding sites found relative to the expected is 80% in eubacterial genomes and 89% in archaeal genomes. We also estimate the strength of selection against spurious binding sites in the face of the constant creation of new binding sites via mutation. Under conservative assumptions, we estimate that selection is weak, with the average efficacy of selection against spurious binding sites, Nes, of -0.12 for eubacterial genomes and -0.06 for archaeal genomes, similar to that of codon bias. Our results suggest that both coding and noncoding sequences are constrained by selection to avoid specific regions of sequence space.

Archaea↗

Evolutionary origin of a plant mitochondrial group II intron from a reverse transcriptase/maturase-encoding ancestor.

Group II introns are widespread in plant cell organelles. In vivo, most if not all plant group II introns do not self-splice but require the assistance of proteinaceous splicing factors. In some cases, a splicing factor (also referred to as maturase) is encoded within the intronic sequence and produced by translation of the (excised) intron RNA. However, most present-day group II introns in plant organellar genomes do not contain open reading frames (ORFs) for splicing factors, and their excision may depend on proteins encoded by other organellar introns or splicing factors encoded in the nuclear genome. Whether or not the ancestors of all of these noncoding organellar introns originally contained ORFs for maturases is currently unknown. Here we show that a noncoding intron in the mitochondrial cox2 gene of seed plants is likely to be derived from an ancestral reverse transcriptase/maturase-encoding form. We detected remnants of maturase and reverse transcriptase sequences in the 2.7 kb cox2 intron of Ginkgo biloba, the only living species of an ancient gymnosperm lineage, suggesting that the intron originally harbored a splicing factor. This finding supports the earlier proposed hypothesis that the ancient group II introns that invaded organellar genomes were autonomous genetic entities in that they encoded the factor(s) required for their own excision and mobility.

Amino Acid Sequence↗

Genome DNA sequencing around the EF-1 alpha multigene locus of Arabidopsis thaliana indicates a high gene density and a shuffling of noncoding regions.

In Arabidopsis thaliana, EF-1 alpha proteins are encoded by a multigene family of four members. Three of them are clustered at the same locus, which was positioned 24 cM from the top of chromosome 1. A region of DNA spanning 63 kb around these locus was sequenced and analyzed. One main characteristic of the locus is the mosaic organization of both genes and intergenic regions. Fourteen genes were identified, among which only four were already described, and other unidentified are most likely present. Functionally diverse genes are found at close intervals. Exon and intron distribution is highly variable at this locus, one gene being split into at least 20 introns. Several duplications were found within the sequenced segment both in coding and noncoding regions, including two gene families. Moreover, a sequence corresponding to the 5' noncoding region of the EF-1 alpha genes and harboring a 5' intervening sequence is duplicated and found upstream of several genes, suggesting that noncoding regions can be shuffled during evolution.

Arabidopsis↗

Evolution of dinoflagellate unigenic minicircles and the partially concerted divergence of their putative replicon origins.

Dinoflagellate chloroplast genes are unique in that each gene is on a separate minicircular chromosome. To understand the origin and evolution of this exceptional genomic organization we completely sequenced chloroplast psbA and 23S rRNA gene minicircles from four dinoflagellates: three closely related Heterocapsa species (H. pygmaea, H. rotundata, and H. niei) and the very distantly related Amphidinium carterae. We also completely sequenced a Protoceratium reticulatum minicircle with a 23S rRNA gene of novel structure. Comparison of these minicircles with those previously sequenced from H. triquetra and A. operculatum shows that in addition to the single gene all have noncoding regions of approximately a kilobase, which are likely to include a replication origin, promoter, and perhaps segregation sequences. The noncoding regions always have a high potential for folding into hairpins and loops. In all six dinoflagellate strains for which multiple minicircles are fully sequenced, parts of the noncoding regions, designated cores, are almost identical between the psbA and 23S rRNA minicircles, but the remainder is very different. There are two, three, or four cores per circle, sometimes highly related in sequence, but no sequence identity is detectable between cores of different species, even within one genus. This contrast between very high core conservation within a species, but none among species, indicates that cores are diverging relatively rapidly in a concerted manner. This is the first well-established case of concerted evolution of noncoding regions on numerous separate chromosomes. It differs from concerted evolution among tandemly repeated spacers between rRNA genes, and that of inverted repeats in plant chloroplast genomes, in involving only the noncoding DNA cores. We present two models for the origin of chloroplast gene minicircles in dinoflagellates from a typical ancestral multigenic chloroplast genome. Both involve substantial genomic reduction and gene transfer to the nucleus. One assumes differential gene deletion within a multicopy population of the resulting oligogenic circles. The other postulates active transposition of putative replicon origins and formation of minicircles by homologous recombination between them.

Animals↗

Human and rodent DNA sequence comparisons: a mosaic model of genomic evolution.

Three patterns of DNA sequence conservation have been identified from five human and rodent genomic sequence comparisons. First, a divergent pattern was observed in the noncoding sequences of the beta-globin and gamma-crystallin gene clusters, and second, a highly conserved pattern was observed in the noncoding regions of the T cell receptor C alpha-C delta, and the alpha- and beta-myosin-heavy-chain genes. A third, mixed pattern has also been found in the immunoglobulin IgH C mu-C delta gene region. These three patterns of genomic evolution pose the fascinating possibility that large portions of the genome evolve at different rates.

Animals↗

How can we deliver the large plant genomes? Strategies and perspectives.

The first sequenced plant genome, from the small mustard plant Arabidopsis thaliana, was published at the end of 2000. The sequencing of the rice genome is well under way. The sizes of plant genomes vary by a factor of up to 1000, and many important crop plants have genomes that are several times larger than the human genome. To gain insight into the gene toolbox of plant species, numerous large-scale EST sequencing projects have been launched successfully, and analysis procedures are constantly being refined to add maximum value to the sequence data. In addition, an alternative approach to exclude repetitive noncoding DNA and to enrich sequence libraries for gene-containing genomic regions has been developed. This strategy has the potential to deliver information about both genes and regulatory regions outside the transcribed regions.

Arabidopsis↗

Investigation of the RH locus in gorillas and chimpanzees.

The human Rh blood-group system is encoded by two homologous genes, RhD and RhCE. The RH genes in gorillas and chimpanzees were investigated to delineate the phylogeny of the human RH genes. Southern blot analysis with an exon 7-specific probe suggested that gorillas have more than two RH genes, as has recently been reported for chimpanzees. Exon 7 was well conserved between humans, gorillas, and chimpanzees, although the exon 7 nucleotide sequences from gorillas were more similar to the human D gene, whereas the nucleotide sequences of this exon in chimpanzees were more similar to the human CE gene. The intron between exon 4 and exon 5 is polymorphic and can be used to distinguish the human D gene from the CE gene. Nucleotide sequencing revealed that the basis for the intron polymorphism is an Alu element in CE which is not present in the D gene. Examination of gorilla and chimpanzee genomic DNA for this intron polymorphism demonstrated that the D intron was present in all the chimpanzees and in all but one gorilla. The CE intron was found in three of six gorillas, but in none of the seven chimpanzees. Sequence data suggested that the Alu element might have previously been present in the chimpanzee RH genes but was eliminated by excision or recombination. Conservation of the RhD gene was also apparent from the complete identity between the 3'-noncoding region of the human D cDNA and a gorilla genomic clone, including an Alu element which is present in both species. The data suggest that at least two RH genes were present in a common ancestor of humans, chimpanzees, and gorillas, and that additional RH gene duplication has taken place in gorillas and chimpanzees. The RhCE gene appears to have diverged more than RhD among primates. In addition, the RhD gene deletion associated with the Rh-negative phenotype in humans seems to have occurred after speciation.

Amino Acid Sequence↗

An oligonucleotide microchip for genome-wide microRNA profiling in human and mouse tissues.

MicroRNAs (miRNAs) are a class of small noncoding RNA genes recently found to be abnormally expressed in several types of cancer. Here, we describe a recently developed methodology for miRNA gene expression profiling based on the development of a microchip containing oligonucleotides corresponding to 245 miRNAs from human and mouse genomes. We used these microarrays to obtain highly reproducible results that revealed tissue-specific miRNA expression signatures, data that were confirmed by assessment of expression by Northern blots, real-time RT-PCR, and literature search. The microchip oligolibrary can be expanded to include an increasing number of miRNAs discovered in various species and is useful for the analysis of normal and disease states.

Adult↗

Cloning and sequence analysis of the Mucor circinelloides pyrG gene encoding orotidine-5'-monophosphate decarboxylase: use of pyrG for homologous transformation.

A 3.2-kb BamHI genomic DNA fragment containing the pyrG gene of Mucor circinelloides was isolated by heterologous hybridization using a pyrG cDNA clone of Phycomyces blakesleeanus as the probe. The complete nucleotide sequence of the M. circinelloides pyrG gene encoding orotidine-5'-monophosphate decarboxylase (OMPD) was determined and the transcription start points (tsp) were mapped by primer extension analysis. The predicted amino acid sequence showed homology with the OMPD sequences reported from other filamentous fungi, with 96% similarity with the OMPD of P. blakesleeanus. Analysis of the sequence revealed the presence of two short introns whose length and location were confirmed by sequencing a cDNA clone and comparing this with its genomic counterpart. The intron splice sites and the 5'- and 3'-noncoding flanking regions show general features of fungal genes. Northern-blot hybridization revealed the pyrG transcript to be approx. 1.0 kb. The M. circinelloides pyrG cDNA clone was able to complement the pyrF::Mu-1 mutation of Escherichia coli when inserted between bacterial expression signals. Additionally, the genomic clone complemented the M. circinelloides pyrG4 mutation. When an M. circinelloides autonomous replication sequence was included in the transforming plasmid, the average transformation frequency obtained was 600 to 800 transformants per micrograms DNA and per 10(6) viable protoplasts.

Amino Acid Sequence↗

Kappa-chain constant-region gene sequences in genus Rattus: coding regions are diverging more rapidly than noncoding regions.

We have determined the nucleotide sequence of a 1,200-base pair (bp) genomic fragment that includes the kappa-chain constant-region gene (C kappa) from two species of native Australian rodents, Rattus leucopus cooktownensis and Rattus colletti. Comparison of these sequences with each other and with other rodent C kappa genes shows three surprising features. First, the coding regions are diverging at a rate severalfold higher than that of the nearby noncoding regions. Second, replacement changes within the coding region are accumulating at a rate at least as great as that of silent changes. Third, most of the amino acid replacements are localized in one region of the C kappa domain--namely, the carboxy-terminal "bends" in the alpha-carbon backbone. These three features have previously been described from comparisons of the two allelic forms of C kappa genes in R. norvegicus. These data imply the existence of considerable evolutionary constraints on the noncoding regions (based on as yet undetermined functions) or powerful positive selection to diversify a portion of the constant-region domain (whose physiological significance is not known). These surprising features of C kappa evolution appear to be characteristic only of closely related C kappa genes, since comparison of rodent with human sequences shows the expected greater conservation of coding regions, as well as a predominance of silent nucleotide substitutions within the coding regions.

Animals↗

Identification of two evolutionarily conserved and functional regulatory elements in intron 2 of the human BRCA1 gene.

Cross-species comparative genomics is a powerful strategy for identifying functional regulatory elements within noncoding DNA. In this paper, comparative analysis of human and mouse intronic sequences in the breast cancer susceptibility gene (BRCA1) revealed two evolutionarily conserved noncoding sequences (CNS) in intron 2, 5 kb downstream of the core BRCA1 promoter. The functionality of these elements was examined using homologous-recombination-based mutagenesis of reporter gene-tagged cosmids incorporating these regions and flanking sequences from the BRCA1 locus. This showed that CNS-1 and CNS-2 have differential transcriptional regulatory activity in epithelial cell lines. Mutation of CNS-1 significantly reduced reporter gene expression to 30% of control levels. Conversely mutation of CNS-2 increased expression to 200% of control levels. Regulation is at the level of transcription and shows promoter specificity. Both elements also specifically bind nuclear proteins in vitro. These studies demonstrate that the combination of comparative genomics and functional analysis is a successful strategy to identify novel regulatory elements and provide the first direct evidence that conserved noncoding sequences in BRCA1 regulate gene expression.

Alternative Splicing↗

A binary model of repetitive DNA sequence in Caenorhabditis elegans.

A great amount of genomic DNA in multicellular eukaryotic organisms is regarded as junk because it has no real function in protein coding. However, there is growing evidence that noncoding DNA can play a vital role in the regulation of gene expression during development (Lee et al., 1993). This indicates that the so-called junk DNA may have essential functions that are yet to be found (Nowak, 1994). A novel binary model of noncoding repetitive DNA sequence is proposed to illustrate its possible structure and implications in genome organization and development.

Animals↗