Search PubMedSearch

SEARCH · Search PubMed

Results for “intergenic region”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Transcription Start Regions in PTU-intergenic regions drive cell cycle-dependent transcriptional activation events in Leishmania donovani.

Leishmania displays an unconventional mode of transcription, with long clusters of genes being transcribed polycistronically from Transcription Start Regions (TSRs), being processed into monocistronic units prior to translation. It has long been believed that transcription is constitutive: failure to identify consensus sequences across TSRs (except a GT-rich motif supporting transcription in Trypanosoma brucei) and absence of canonical eukaryotic transcription factors led to the conclusion that regulation is primarily post-transcriptional, with epigenetics playing a role in triggering transcription initiation. This study stems from our previous findings identifying a few genes to be activated in a cell cycle-dependent manner. Using nuclear run-on assays to analyze nascent transcripts of two chromosomes, chromosomes 2 and 14, we find that while most genes are constitutively transcribed, a subset of genes gets activated at specific cell cycle stages. Reporter assays reveal that this transcriptional activation is driven by the regions immediately upstream of the genes. Sequence analyses of these TSRs lying in polycistronic intergenic regions (PIRs) uncovered a 10-mer GT-rich motif, in synchrony with earlier findings in T. brucei identifying a GT-rich motif at bidirectional TSRs. We also identify a second 25-mer motif at these TSRs, and deletion analyses find this motif to be critical for regulating gene expression. The findings of this study reveal that transcriptional events in these unicellular parasites are more complex than believed thus far: not all transcriptional events are constitutive, polycistronic transcription is not the only mode of transcription, and cis-acting sequence elements regulate at least some transcriptional events in these parasites.IMPORTANCEEndemic to 90 countries, Leishmania parasites cause a spectrum of diseases called Leishmaniases. No vaccines for human use are available to date, and the drugs currently used to treat the disease are expensive, have toxic side effects, and have complex administration regimens, with emerging drug resistance compounding problems. Researchers continue to investigate Leishmania cellular processes, with the hope of uncovering new therapeutic target sites. Gene regulation in these parasites is unusual, being modulated by various mechanisms, including epigenetic modifications, gene dosage, and post-transcriptional processing. Transcription is typically polycistronic and constitutive, initiating from Transcription Start Regions (TSRs) lying upstream of the first gene in the polycistronic transcription unit (PTU). The work presented here reveals that a subset of genes is transcribed monocistronically in a cell cycle-dependent manner from Transcription Start Regions lying in the PTU-intergenic regions (PIRs), underscoring the complexities of gene regulation in these parasites.

Leishmania donovani

Ribosomal RNA genes of Saccharomyces cerevisiae. II. Physical map and nucleotide sequence of the 5 S ribosomal RNA gene and adjacent intergenic regions.

A DNA fragment containing the structural gene for the 5 S ribosomal RNA and intergenic regions before and after the 35 S ribosomal RNA precursor gene of Saccharomyces cerevisiae has been amplified in a bacterial plasmid and physically mapped by restriction endonuclease cleavage and hybridization to purified yeast 5 S ribosomal RNA. The nucleotide sequence of the DNA fragments carrying the 5 S ribosomal RNA gene and adjacent regions has been determined. The sequence unambiguously identifies the 5 S ribosomal RNA gene, determines its polarity within the ribosomal DNA repeating unit, and reveals the structure of its promoter and termination regions. Partial DNA sequence of the regions near the beginning and end of the 35 S ribosomal RNA gene has also been determined as a preliminary step in establishing the structure of promoter and termination regions for the 35 S ribosomal RNA gene.

Base Sequence

Restriction-enzyme-cleavage maps of bacteriophage M13. Existence of an intergenic region on the M13 genome.

Replicative form DNA of bacteriophage M13 was cleaved into specific fragments by an endonuclease isolated from Hemophilus aegyptius (endoR.HaeII) and an endonuclease from Arthrobacter luteus (endoR.AluI). The fragments were ordered as to construct a circular map of the phage M13 genome by: (a) using each fragment as a primer for the synthesis in vitro of its respective neighbour and (b) digesting the isolated fragments with the Hemophilus aegyptius enzyme endoR.HaeII or the Hemophilus aphirophilus enzyme endoR.HapII and subsequent analysis of the overlapping sets of fragments. The resulting physical map was correlated with the M13 genetic map by marker rescue experiments with amber mutant phage DNAs and purified wild-type fragments. From the results of these analyses it has been concluded that gene II and gene V are contiguous on the genetic map. Evidence is provided that there is an internal start of RNA synthesis within the C-terminal region of gene II which then ultimately leads to the synthesis of X protein. Furthermore, we conclude that there is an intergenic space of considerable length (450-500 base pairs) which is located between gene II and gene IV on the M13 genome. The function of this intergenic region as the origin site for phage DNA replication is discussed.

Arthrobacter

Transcriptional switch of the dia1 and impA promoter during the growth/differentiation transition.

When growth stops due to the depletion of nutrients, Dictyostelium cells rapidly turn off vegetative genes and start to express developmental genes. One of the early developmental genes, dia1, is adjacent to a vegetative gene, impA, on chromosome 4. An intergenic region of 654 bp separates the coding regions of these divergently transcribed genes. Constructs carrying the intergenic region expressed a reporter gene (green fluorescent protein gene) that replaced impA in growing cells and a reporter gene that replaced dia1 (DsRed) during development. Deletion of a 112-bp region proximal to the transcriptional start site of impA resulted in complete lack of expression of both reporter genes during growth or development. At the other end of the intergenic region there are two copies of a motif that is also found in the carA regulatory region. Removing one copy of this repeat reduced impA expression twofold. Removing the second copy had no further consequences. Removing the central portion of the intergenic region resulted in high levels of expression of dia1 in growing cells, indicating that this region contains a sequence involved in repression during the vegetative stage. Gel shift experiments showed that a nuclear protein present in growing cells recognizes the sequence GAAGTTCTAATTGATTGAAG found in this region. This DNA binding activity is lost within the first 4 h of development. Different nuclear proteins were found to recognize the repeated sequence proximal to dia1. One of these became prevalent after 4 h of development. Together these regulatory components at least partially account for this aspect of the growth-to-differentiation transition.

Animals

Uncovering hidden complexity in the Apis mellifera mitotranscriptome: a polyadenylation-centered perspective.

Mitochondrial transcription is gaining increasing attention as researchers seek to better understand the full coding potential of mitochondrial DNA (mtDNA). Emerging evidence suggests that mtDNA may encode additional elements beyond classical oxidative phosphorylation genes, pointing to a more complex transcriptional architecture than previously recognized. In this study, we explored the mitochondrial transcriptome of Apis mellifera (Insecta: Hymenoptera), with a particular focus on polyadenylation-associated features. Our analysis revealed that both sense and antisense transcripts undergo polyadenylation, although transcript abundance and poly(A) tail lengths varied markedly across mitochondrial genes. Several transcripts exhibited alternative isoforms, either extended or truncated, frequently including intergenic regions. These regions may represent functional non-coding elements or structural variants rather than conventional untranslated regions (UTRs). Interestingly, some transcripts also contained non-templated nucleotide additions particularly cytosine residues immediately upstream of the poly(A) tails. Monocistronic units that included portions of downstream intergenic regions were among the most abundantly represented, suggesting a possible regulatory role for these sequences. To experimentally validate our in silico findings, we performed RT-qPCR to assess relative gene expression and applied 3' RACE-PCR to define transcript boundaries. These approaches confirmed the presence of multiple transcript isoforms and supported the involvement of polyadenylation in shaping mitochondrial RNA diversity. Together, our findings reveal a previously underappreciated level of complexity in the A. mellifera mitochondrial transcriptome and highlight the potential regulatory significance of polyadenylation dynamics and intergenic region transcription.

Animals

Nucleotide Combination Proportions Across Algae, Monocotyledons and Dicotyledons: Insights into Plant Genome Evolution.

Plant evolution started with unicellular algae, gradually evolving multicellularity and terrestrial colonization. These evolutionary events were accompanied by the interplay of chromosome polyploidization, rearrangement, gene loss, and point mutation. We counted the proportion of nucleotide combinations in the genome sequences of 64 sequenced plants, and analyzed the significant difference in these nucleotide combination proportions among algae, monocotyledons and dicotyledons. The correlation of highly significant different and no significant different nucleotide combinations was analyzed respectively. Nucleotide combinations and their reverse complementary sequence proportions were analyzed in different functional regions of the genome. These results reveal that some nucleotide combinations are subject to strict selection, and these combinations have a higher proportion in the CDS regions and lower proportion in the intergenic regions. Meanwhile, there are some nucleotide combinations that are under less selective pressure, and these combinations have a higher proportion in the intergenic regions and lower proportion in the CDS regions. Cluster analysis based on trinucleotide to octanucleotide combination proportions reveals that plant genome evolution is accompanied by clade-wide differentiation of genome-wide nucleotide composition patterns, in addition to well-documented chromosomal polyploidization, structural rearrangement and gene loss events. We analyzed the changes in the proportion of nucleotide combinations at the genome level in 64 sequenced plants, providing a new idea for studying genome evolution in the plant kingdom.

comparative genomics

Multiple Forms and Functions of Premature Termination by RNA Polymerase II.

Eukaryotic genomes are widely transcribed by RNA polymerase II (pol II) both within genes and in intergenic regions. POL II elongation complexes comprising the polymerase, the DNA template and nascent RNA transcript must be extremely processive in order to transcribe the longest genes which are over 1 megabase long and take many hours to traverse. Dedicated termination mechanisms are required to disrupt these highly stable complexes. Transcription termination occurs not only at the 3' ends of genes once a full length transcript has been made, but also within genes and in promiscuously transcribed intergenic regions. Termination at these latter positions is termed "premature" because it is not triggered in response to a specific signal that marks the 3' end of a gene, like a polyA site. One purpose of premature termination is to remove polymerases from intergenic regions where they are "not wanted" because they may interfere with transcription of overlapping genes or the progress of replication forks. Premature termination has recently been appreciated to occur at surprisingly high rates within genes where it is speculated to serve regulatory or quality control functions. In this review I summarize current understanding of the different mechanisms of premature termination and its potential functions.

RNA Polymerase II

Filamentous coliphage M13 as a cloning vehicle: insertion of a HindII fragment of the lac regulatory region in M13 replicative form in vitro.

A HindII restriction fragment comprising the Escherichia coli lac regulatory region and the genetic information for the alpha peptide of beta-galactosidase (beta-D-galactosidegalactohydrolase, EC. 3.2.1.23) has been inserted into 1 of the 10 Bsu I cleavage sites of M13 by blunt end ligation. A stable hybrid phage was isolated and identified by its ability to complement the lac alpha function. Further characterization of the hybrid phage includes retransformation studies, agarose gel electrophoresis, DNA-DNA hybridization, and heteroduplex mapping. The insertion point has been localized at 0.083 map unit on thewild-type circular map-i.e., within the intergenic region. The results prove that part of the intergenic region is nonessential and that the phage can be used as a cloning vehicle.

Coliphages

Research note: Development of a recombinant duck enteritis virus vector expressing DHAV-3 VP1 and DTMUV prM/TE genes.

Duck enteritis virus (DEV) is a promising viral vector for vaccine development. In a previous study, an HDR-CRISPR/Cas9-based strategy was used to generate a recombinant virus, rDEV-DHAV-VP1, by inserting the VP1 gene of duck hepatitis A virus type 3 (DHAV-3) into the UL27/UL26 intergenic region of DEV vaccine strain, resulting in good genetic stability and immunogenicity. In the present study, the same strategy was applied to insert the EGFP gene into the US7/US8 and LORF11/UL55 intergenic regions of a DEV vaccine strain. Among the evaluated insertion sites, the highest level of EGFP expression was observed at the US7/US8 locus, followed by the UL27/UL26 locus. Based on rDEV-DHAV-VP1, the pre-membrane (prM) and truncated envelope (TE) genes of duck Tembusu virus (DTMUV) were further inserted into the US7/US8 locus, resulting in a bivalent recombinant virus, rDEV-VP1-prM/TE. The recombinant virus exhibited growth kinetics comparable to those of the parental virus, while maintaining efficient expression and high genetic stability of the inserted genes. These findings indicate that the HDR-CRISPR/Cas9 system is an efficient strategy for generating stable DEV-based recombinant vectors and provides a promising platform for the development of multivalent vaccines against major duck viral diseases.

CRISPR/Cas9 genome editing

Experimental evolution reveals contrasting adaptive landscapes in lab and field environments.

Experimental evolution is widely used to infer microbial responses to environmental change, yet most laboratory studies impose constant, well-mixed conditions that differ fundamentally from fluctuating, spatially structured field environments. We compared genomic evolution in the leaf litter-associated bacterium Curtobacterium strain MMLR14_002 under control and warming treatments in laboratory culture and in a complementary field experiment. Laboratory-derived isolates accumulated more mutations per genome and exhibited stronger locus-level parallelism, with mutations recurring in a small number of coding loci. Field-derived isolates accumulated fewer mutations per genome, and these mutations rarely occurred in the same coding loci across replicate populations. Instead, field isolates exhibited a higher proportion of intergenic mutations, with mutations recurring in the same intergenic regions across independent field deployments. When coding mutations were detected in the field, they were distributed across functionally diffuse targets and more often involved metabolic pathways than the core cellular processes repeatedly targeted during laboratory evolution. Warming itself did not consistently influence mutation accumulation or the genomic distribution of mutations; instead, laboratory and field contexts primarily shaped the accumulation, targets, and repeatability of genomic change. These results suggest that laboratory thermal evolution identifies adaptive routes favored under sustained selection but may overestimate coding-level parallelism under heterogeneous field conditions. Bridging laboratory and field evolution will likely require experimental designs that incorporate temporal variability and spatial heterogeneity characteristic of natural systems.IMPORTANCEA central goal of experimental evolution is to infer how microbes evolve in nature from laboratory studies. Here, we evaluate this assumption by comparing genomic evolution of a leaf litter-associated Curtobacterium strain in laboratory and field warming experiments to identify broad patterns rather than isolate the contribution of any single environmental factor. We find that the strong parallelism at coding loci observed under laboratory conditions is reduced in the field, while mutations recurring in the same intergenic regions across field deployments suggest that parallel evolution in nature may more often involve regulatory noncoding regions rather than coding targets. These results show that environmental context reshapes adaptive landscapes and may limit the parallelism of coding-level genomic responses inferred from homogeneous laboratory conditions.

experimental evolution

Deep learning reveals genomic regions introgressed between two recurrently hybridizing lynx species.

Recently, diverged species with overlapping distributional ranges have high chances of hybridizing and if hybrids are viable, genomic material can be transferred between species in a process called introgression. To characterize the patterns and consequences of introgression in species with historically low population sizes and recent steep declines resulting in genetic erosion, we analyze the Iberian and Eurasian lynx (EL) as an illustrative and relevant case study. While genome-wide introgression was already detected, here we apply a method using a deep convolutional neural network to detect specific regions of the genome with signals of introgression in three populations of these two species. Over 6% of the genome of both Iberian lynx and ELw shows introgression from the other species, compared with only 2% in the ELs. This observation, along with the results from demographic modeling, suggests that the ELw population is genetically closest to the source of EL introgression, a probably now extinct group that coexisted with the Iberian lynx in Southern Europe and Northern Iberia until recently. As predicted by theory, introgression was generally higher in populations with smaller effective sizes and in genomic regions of high recombination. However, the Iberian lynx did not show higher overall introgression than the more abundant ELw, and coding regions introgressed as frequently as intergenic regions. Local genetic diversity is boosted approximately 3-fold in genomic windows where introgression occurs, potentially including the adaptively relevant and highly diverse MHC region of the Iberian lynx.

Animals

APAV: An advanced pangenome analysis and visualization toolkit.

Traditional pangenome analysis focuses on gene presence/absence variations (gene PAVs). However, the current methods for gene PAV analysis are insensitive to detect small but valuable mutations within gene regions, and they overlook variations in intergenic regions. Additionally, the visual inspection of PAVs is an important but time-consuming step for pangenome analysis and result interpretation. To address these issues, we present APAV, an advanced toolkit designed for comprehensive PAV analysis and visualization. It integrates gene element-level PAV analysis and provides PAV analysis for arbitrary given regions in a genome. The resulted PAV profile can be visualized and investigated interactively with reports in HTML format, enabling researchers to conveniently verify sequencing read depth, target region coverage, and intervals of absence for each PAV. Furthermore, APAV offers various subsequent analysis and visualization functions based on the PAV profile table, including basic statistics, sample clustering, genome size estimation, and phenotype association analysis. We demonstrated the capability of APAV with pangenome analysis of tumor genomes and rice genomes. Performing PAV analysis at the element level not only provides more accurate information about the variations but also uncovers a larger number of variations for the phenotype-genotype association studies. In the rice genome analysis, we identified over twenty thousand distributed genes and more than fifty thousand distributed genetic elements. In the tumor genome analysis, element-level analysis revealed approximately three times as many phenotype-related genes as gene-level analysis. This indicates that altering the PAV unit from genes to smaller segments or elements can lead to more biological insights.

Software

Beyond exons: Linking noncoding heritability and polygenicity across complex human traits and disorders.

The genetic architecture of complex traits spans a continuum of polygenicity, yet it remains unclear how differences in polygenicity relate to the functional localization of SNP heritability across the genome. We use a MiXeR-based framework to partition heritability across 74 functional annotations covering exonic, intronic, and intergenic regions for 34 complex traits and introduce a likelihood-based annotation contribution score that quantifies annotation-specific impact on heritability. Exons account for a minority of heritability, and their contribution decreases with increasing polygenicity, from an average of 22% in less-polygenic somatic diseases and biomarkers to 13% in highly polygenic psychiatric and cognitive phenotypes. Intergenic fractions show the opposite trend, whereas intronic fractions remain relatively stable. Analysis of the broader set of functional annotations also reveals systematic differences along the polygenicity axis: highly polygenic traits show stronger contributions from comparative genomics and variant-effect scores, whereas less-polygenic traits show stronger contributions from promoter, transcription, and chromatin annotations. Together, these results indicate that the functional partitioning of heritability systematically varies with polygenicity, shifting from gene-proximal regulatory architectures to architectures shaped by numerous dispersed regulatory effects.

MiXeR

Di-, tri-, and tetranucleotide frequencies covary with lifespan and genome size across protostome invertebrates.

Animal lifespans span orders of magnitude, yet how genome sequence covaries with lifespan remains poorly characterized outside vertebrates. Although promoter CpG density has been linked to vertebrate longevity due to its gene-regulatory function through DNA methylation, it is unclear whether such patterns are promoter- and CpG-specific, or if they reflect broader sequence evolution. We curated maximum lifespan estimates for 466 protostome species spanning eight phyla with available genome assemblies and quantified mono-, di-, tri-, and tetranucleotide composition across whole genomes, intergenic regions, and six gene-associated regions (two upstream regions, exons, introns, and two downstream regions) defined using Benchmarking Universal Single-Copy Orthologs. Dinucleotide observed/expected ratios showed significant associations with lifespan and genome size in different ways. Lifespan-associated motifs were most pronounced in gene-associated non-coding regions, especially in introns and downstream regions, whereas genome-size effects were strongest in whole-genome and intergenic sequence. Tri- and tetranucleotide observed/expected ratios broadly recapitulated this regional organization. In contrast, GC content was not associated with lifespan across regions, indicating that the observed signals are not explained by mononucleotide composition but instead by how those nucleotides are arranged into short sequence motifs. These results suggest that lifespan and genome size show distinct but overlapping associations with regional sequence composition across invertebrate species and that lifespan-associated motif evolution extends beyond vertebrate promoter methylation architectures.

CpG density

Genetic Variants Associated With the Biochemical Response to Vitamin D3 in the Multi-Ethnic Study of Atherosclerosis.

CONTEXT: The response to treatment with vitamin D varies between patients. OBJECTIVE: To identify genetic variants associated with the biochemical response to vitamin D3 supplementation. DESIGN: Randomized placebo-controlled trial conducted between 2017 and 2019. SETTING: The trial was nested in an ongoing community-based cohort study, the Multi-Ethnic Study of Atherosclerosis. INTERVENTION: 2000 International Units of vitamin D3 or placebo daily for 16 weeks. PARTICIPANTS: The analytic sample included 427 participants assigned to vitamin D3 (mean age, 73 years; 54% females) and was 36% White, 33% Black, 18% Hispanic, and 14% Chinese. MAIN OUTCOME MEASURES: The biochemical response to vitamin D3 included changes in serum concentrations of 1,25-dihydroxyvitamin D3 [1,25(OH)2D3], PTH, and 25-hydroxyvitamin D3 [25(OH)D3]. RESULTS: In genome-wide analyses, single nucleotide polymorphisms in 8 regions of the genome had significant association (P < 5E-08) with 1 of the traits (2 with change in 1,25(OH)2D3, 1 with change in PTH, and 5 with change in 25(OH)D3). rs16867276 within an intergenic region on 2q31 was associated with change in serum 1,25(OH)2D3 (+8.37&#x2005;pg/mL difference per effect allele; P = 4.93E-08) and was the only locus that achieved genome-wide significance in transethnic meta-analysis. rs114044709 adjacent to FAM20A, which encodes a protein required for biomineralization, was associated with change in PTH among Black participants (+20.32&#x2005;pg/mL difference per effect allele; P = 1.34E-08). In candidate analyses, single nucleotide polymorphisms within SULT2A1 and CYP24A1 had significant association (P < .05&#xf7;36 = .0014) with the changes in 1,25(OH)2D3 and PTH, respectively. CONCLUSION: Our results reveal potential new pathways of vitamin D regulation that require replication in other vitamin D trials.

Humans

Construction of an M13 histidine-transducing phage: a single-stranded cloning vehicle with one EcoRI site.

In order to create a ready source of single-stranded DNA for DNA sequence determination by the dideoxy chain-termination method, the promoter-proximal part of the histidine operon, the hisOGD region of Salmonella typhimurium, was cloned onto the single-stranded phage M13. Both orientations of the his DNA were cloned to supply DNA template for sequencing of each strand. Insertion was achieved at an HaeIII site in the intergenic region (IR) of M13, and a single EcoRI site was purposely regenerated at one boundary of the his DNA insert. Infected colonies, not plaques, were selected using the hisD gene as a selective marker. The single RI site and the hisD marker for auxotrophic selection represent improvements on the wild type M13 as a single-stranded vector for cloning other DNA.

Base Sequence

Nucleotide sequence of bacteriophage G4 DNA.

The 5,577 nucleotide long sequence of bacteriophage G4 DNA has been determined using the 'plus and minus' and chain termination methods of DNA sequencing. This sequence has been compared with that of the closely related bacteriophage phiX174 (refs 1, 55). In the coding regions there is an average of 33.1% nucleotide sequence differences between the two genomes, but the distribution of these changes is not random and the sequence of some genes is more conserved than others. There is less sequence similarity between the untranslated intergenic regions of G4 and phiX174, but despite this the sequences of the J/F, F/G and H/A untranslated spaces in both genomes have similar sized hairpin loops, which may be related to their function.

Bacteriophages