Search PubMedSearch

SEARCH · Search PubMed

Results for “GC-rich sequence”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Using the DNA language model, GROVER, to parse effects of sequence, chromatin and regulatory features on genome stability.

MOTIVATION: Genome stability is shaped by DNA sequence and chromatin context, but their relative contributions to double-strand break (DSB) sensitivity remain unclear. RESULTS: We show that the DNA language model, GROVER, can infer DSB location based on sequence. DSB hotspots tend to contain GC-rich sequences that belong to promoters, genes and short interspersed nuclear elements (SINEs). Additionally, we identified several specific short sequences (tokens) that are associated with modulating DSB sensitivity. Another model using chromatin and genome regulatory features outperforms the sequence-only model, highlighting complementary and cell-type specific information. Integrating sequence and genome biological features yields the best performance, demonstrating their synergy. Analyzing this model revealed that, dependent on the sample, genome stability information encoded in H3K36me3 and DNase-seq can be learned from the sequence, but not H3K27ac or H3K9me3. Embedding chromatin data directly into the GROVER architecture enabled cell-type specific modeling with performance matching the full chromatin feature model. Our results suggest that while chromatin and regulatory context provides important information, such as cell-type specificity, much of the information shaping DSB patterns is already encoded in the DNA sequence itself. Our integrative modeling approach not only reveals DSB patterns but also provides a generalizable strategy for tracing predictions in genomic data. AVAILABILITY: Data, models, and a tutorial are available on Zenodo.

Chromatin

ELYS associates with distinct DNA sequence environments during post-mitotic nuclear pore reassembly.

Nuclear pore complexes (NPCs) contribute to genome organization and cell identity, yet how post-mitotic NPC assembly is coordinated with chromatin architecture remains unclear. Here, we show that the nucleoporin ELYS preferentially associates with chromatin regions displaying distinct intrinsic DNA sequence features that are not explained by the repressive histone marks examined here. ELYS-bound regions are enriched for AT-rich sequences, whereas ELYS binding at super-enhancer-associated loci shift toward GC-rich sequence composition, revealing distinct sequence environments. These findings indicate that ELYS localization is associated with distinct intrinsic DNA sequence features and suggest a mechanism by which nuclear pore-associated architecture restores transcriptional programs after mitosis.

Journal Article

The nucleotide sequence of oocyte 5S DNA in Xenopus laevis. II. The GC-rich region.

The primary sequence of the GC-rich half of the repeating unit in X. laevis 5S DNA has been determined in both a single plasmid-cloned repeating unit and in the total population of repeatig units. The GC-rich half of the repeating unit contains a single long duplication of 174 nucleotides. The duplicated segment commences 73 nucleotides preceding the 5' end of the gene and terminates at nucleotide 101 of the gene. The duplicated portion of the gene, termed the pseudogene, differs by 10 nucleotides from the corresponding portion of the gene, and the remaining duplicated sequence of 73 nucleotides differs by 13 nucleotides. The plasmid-cloned repeating unit differs from the dominant sequence in the total population repeating units by 6 nucleotides in the GC-rich region. Evidence is provided that most of the CpG dinucleotides in 5S DNA are at least partially methylated.

Animals

In vitro reconstitution of chromatin replication recapitulates symmetric histone recycling.

Symmetric histone recycling is vital for maintaining epigenetic inheritance upon eukaryotic DNA replication. Recent genome-wide studies have uncovered key determinants of this process, but how these factors collectively support parental histone transfer remains incompletely understood. Here, we successfully reconstitute histone recycling with 24 purified proteins and analyze the products digested by Micrococcal nuclease with Repli-pore-seq, the newly developed pipeline combining nanopore sequencing and deep-learning-based classification. As a result, we identify histones symmetrically recycled as tetrasomes or hexasomes on nucleosome-favorable sequences. We also observe the discordance of the recycled position between lagging and leading strands on the GC-rich DNA sequences. Moreover, removal of Pol δ, Pol32, Dpb3/4, Ctf4, Csm3/Tof1, or Mrc1 disrupts the balance of histone recycling between the two daughter strands, whereas removal of Ctf4, Csm3/Tof1, or Mrc1 additionally alters the positions at which histones were recycled. Furthermore, addition of the lagging-strand maturation factors Fen1 and Cdc9 enhances histone recycling to the lagging strand. These findings provide critical insights into the molecular players and mechanisms underlying symmetric histone recycling.

Histones

Genome-wide etiology analysis of autoimmune hypothyroidism supports somatic mutations of at-risk DNA as the underlying cause.

Autoimmune hypothyroidism (AIHT) is the most common autoimmune disease. Through an unidentified mechanism, the immune system attacks the thyroid gland, destroys thyroid follicular cells, and causes hypothyroidism. A new theory poses that all DNA is continuously damaged and, as a result, is exposed to somatic mutations at a constant rate. Based on this theory, several assumptions related to epidemiology and DNA sequence can be made. These have been summarized as a method called genome-wide etiology analysis (GWEA) to facilitate the interpretation of GWAS results of autoimmune diseases. Here, GWEA is applied to AIHT. The results show that existing epidemiological and genomic data of AIHT adhere to the principles of GWEA. Therefore, AIHT appears to be the result of somatic mutations in people at risk for the disease. AIHT develops once sufficient mutations create a new "autoimmune pathway" driven by non-self-signal and supported by neopeptide formation and signal amplification. Given the random nature of somatic mutations throughout life, the new theory explains why some people with AIHT develop additional autoimmune diseases, why family members may develop a range of non-AIHT autoimmune diseases, why the age of onset cannot be predicted, and why AIHT is transferred to the following generations through dominant inheritance with delayed, incomplete penetrance.

Humans

The nucleotide sequence of oocyte 5S DNA in Xenopus laevis. I. The AT-rich spacer.

The primary sequence of the principal spacer region in X. laevis oocyte 5S DNA has been determined. The spacer is AT-rich and comprises half or more of each repeating unit. The sequence is internally repetitious; most of it can be represented by the following set of oligonucleotides: CAACAGTTTTCAAAAGGTTTCGAAGTTTTT(T). The spacer, which varies in length from about 360 to 570 or more nucleotides, can be subdivided into a region (A2) which is variable in length in different repeating units, flanked by regions (A1, A3, B1) which are relatively constant in length. The A2 region consists, on the average, of 5-6 tandem copies of the oligonucleotide CAAAGTTTGAGTTTT; variation in the redundancy of this oligonucleotide accounts for much of the repeat length variation in the genomic 5S DNA. Most copies of this oligonucleotide are identical, although several differing by 1 or 2 nucleotides have been detected in plasmid-cloned 5S DNA fragments. Regions A1 and A3 comprise a linear array of similar, but not identical, oligonucleotides; most repeating units contain very similar A1 and A3 sequences. Region B1 is a sequence of 49 nucleotides immediately adjacent to the 5' terminus of the 5S rRNA sequence. It is GC-rich, much less repetitive than the remainder of the spacer and contains several palindromes, but no regions of dyad symmetry. This sequence is identical in all six of the single cloned repeating units of 5S DNA analyzed.

Adenine

Integration of bovine leukemia virus DNA in the bovine genome.

DNA preparations from circulating leukocytes, lymph node tumors, and spleens of three bovine leukemia virus-infected cattle were fractionated by Cs2SO4/3,6-bis(acetatomercurimethyl)dioxane density gradient centrifugation. Bovine leukemia virus proviral sequences were found in large GC-rich fragments having a buoyant density in CsCl close to 1.708 g/cm3. Provirus integration, therefore, does not take place at random locations in the host genome, but in a specific class of DNA segments. Hybridization of cDNA synthesized on viral RNA to EcoRI and Xba I restriction fragments of the DNA from infected cells showed that: (i) only one copy of proviral DNA is integrated per haploid genome; (ii) different restriction patterns were found in the proviral DNAs present in the genomes of different animals, providing evidence for the existence of several strains or mutants; and (iii) different integration sites for the proviral DNA were found in the genome of different animals and of different infected cells in the same animal. The latter finding strongly suggests a polyclonal origin of bovine leukemia virus-infected cells.

Animals

Sequence and secondary structure of Drosophila melanogaster 5.8S and 2S rRNAs and of the processing site between them.

Drosophila melanogaster 5.8S and 2S rRNAs were end-labeled with 32p at either the 5' or 3' end and were sequenced. 5.8S rRNA is 123 nucleotides long and homologous to the 5' part of sequenced 5.8S molecules from other species. 2S rRNA is 30 nucleotides long and homologous to the 3' part of other 5.8S molecules. The 3' end of the 5.8S molecule is able to base-pair with the 5' end of the 2S rRNA to generate a helical region equivalent in position to the "GC-rich hairpin" found in all previously sequenced 5.8S molecules. Probing the structure of the labeled Drosophila 5.8S molecule with S1 nuclease in solution verifies its similarity to other 5.8S rRNAs. The 2S rRNA is shown to form a stable complex with both 5.8S and 26S rRNAs separately and together. 5.8S rRNA can also form either binary or ternary complexes with 2S and 26S rRNA. It is concluded that the 5.8S rRNA in Drosophila melanogaster is very similar both in sequence and structure to other 5.8 rRNAs but is split into two pieces, the 2S rRNA being the 3' part. 2S anchors the 5.8S and 26S rRNA. The order of the rRNA coding regions in the ribosomal DNA repeating unit is shown to be 18S - 5.8S - 2S - 26S. Direct sequencing of ribosomal DNA shows that the 5.8S and 2S regions are separated by a 28 nucleotide spacer which is A-T rich and is presumably removed by a specific processing event. A secondary structure model is proposed for the 26S-5.8S ternary complex and for the presumptive precursor molecule.

Animals

Denaturation map of the ribosomal DNA of Lytechinus variegatus sperm.

Electron microscopy of the partially heat denatured ribosomal DNA (rDNA) from sea urchin (Lytechinus variegatus) sperm has demonstrated that it consists of repeating units of 3.6 +/- 0.2 micron, corresponding to a mol.wt of 7.2 +/- 0.4 x 10(6). Based on differential denaturability, each repeat unit is divided into 2 regions. The larger region of 2.47 +/- 0.11 micron (mol.wt 4.9 +/- 0.22 x 10(6)) corresponds in length to the ribosomal precursor RNA of sea urchins and the smaller, GC-rich, subunit of 1.16 +/- 0.09 micrometer (mol.wt 2.3 +/- 0.18 x 10(6)) is presumed to contain non-transcribed spacer sequences.

Animals

Mithramycin and DIPI: a pair of fluorochromes specific for GC-and AT-rich DNA respectively.

The AT specificity of the fluorochromes DIPI and DAPI and the GC specificity of mithramycin are evidenced by observations in human, mouse, and bovine chromosomes. DIPI and DAPI produce a pattern similar to Hoechst 33258 in all three species, whereas mithramycin results in a reverse pattern. The AT-rich centromeric heterochromatin in mouse is brilliantly stained by DIPI or DAPI and remains nearly invisible after mithramycin staining. In the GC-rich centromeric heterochromatin of cattle the opposite behavior is observed.

Animals

Preferential in vitro assembly of nucleosome cores on some AT-rich regions of SV40 DNA.

We have found that nucleosomes reconstituted from histone octamers and SV40 DNA Form I by progressively decreasing the salt concentration from 2 M NaCl are formed preferentially around 0.27, 0.37, 0.50 and 0.85 on SV40 DNA (relative to the EcoRI site). When SV40 DNA Form III is used, the nucleosomes form mainly at 0.28, 0.38, 0.61 and 0.83. These sites are very close to both the sites of RNA chain initiation by calf thymus RNA polymerase B on SV40 DNA Form I (0.25, 0.35, 0.42 and 0.88) and the regions of the supercoiled DNA which are readily denaturable by T4 gene 32 protein (0.25, 0.47 and 0.88), and correspond to AT-rich regions as deduced from the nucleotide sequence of SV40 DNA. The physiologically important region around 0.67 is an unfavourable site for all three types of proteins, and corresponds to a GC-rich region surrounding a 17 base pair AT cluster.

Base Composition

Nucleotide sequence of Xenopus borealis oocyte 5S DNA: comparison of sequences that flank several related eucaryotic genes.

Genomic Xenopus borealis oocyte-specific 5S DNA (Xbo) contains clusters of 5S rRNA genes. The number of genes varies among clusters, and the distance between genes within a cluster is about 80 nucleotides. The spacer DNA between gene clusters is AT-rich and heterogeneous in length due in part to variable numbers of a tandemly repeated 21 nucleotide sequence. A cloned fragment of Xbo 5S DNA (Xbo1) containing three 5S rRNA genes has been sequenced. The sequences of Xbo1 genes 1 and 2 are very similar to the dominant 5S RNA sequence, whereas 15 of the 120 residues in the third gene are different. The sequence of gene 3 is as different from the dominant gene sequence as the X. laevis pseudogene is from the 5S RNA gene. Sequence analysis of genomic DNA shows that gene 3 is an abundant component of the multigene family. All three genes are transcribed when added to an extract of X. laevis oocyte nuclei, and a fragment of Xbo1 lacking the AT-rich spacer DNA and the 5' end of the first gene supports transcription of genes 2 and 3 in this in vitro system. Thus the 80 nucleotides preceding each 5S gene are sufficient for promoter function. Nucleic acid sequences preceding several eucaryotic genes that are transcribed by RNA polymerase III were analyzed and the following common features were found: a purine-rich region; at least one direct repeat; the absence of dyad symmetry; transcription beginning with a purine; a pyrimidine residue immediately preceding the first nucleotide of the gene; and the oligonucleotides AAAAG, AGAAG and GAC, located approximately 15, 25 and 35 nucleotides, respectively, before the start of transcription. The 10 base pair (bp) spacing between the homologous oligonucleotides is that expected for a recognition signal on one face of a DNA double helix. The extensive sequence differences between most of the spacers that precedes these genes make the three conserved oligonucleotides more striking. Parts of the 5' flanking regions of the three Xbo1 gene (-12 to -40), which include the conserved oligonucleotides, are identical. In contrast, 7 of the first 11 nucleotides that precede the third 5S RNA gene in Xbo1 differ from those that precede the first gene. The sequences following the X. borealis oocyte and somatic 5S genes are identical in 12 of the first 14 residues and contain two or more T clusters, as does the corresponding region of X. laevis oocyte 5S DNA. The 3' sequences of the Xenopus 5S rRNA genes and several other eucaryotic genes contain features in common with procaryotic transcription termination sites. The 3' end of the gene is GC-rich and contains a dyad symmetry. Termination occurs in an AT-rich region containing one or more T clusters on the noncoding strand.

Animals

Characterization of in vitro transcription initiation and termination sites in Col E1 DNA.

Overlapping restriction fragments from the region between the single Eco R1 site and the origin of replication of the plasmid, Col E1, have been utilised as templates in an in vitro transcription assay using E. coli RNA polymerase. Transcription towards the single Eco R1 site is initiated at a point 415 bp to the origin side of that site. In vivo, transcription starting at this point probably produces the mRNA for the colicin immunity protein. Transcription away from the Eco R1 site is initiated at a point 140 bp to the origin side of that site and terminated 30 bp further on. This terminator is probably the point at which transcription of the colicin gene is terminated in vivo. DNA sequence analysis in both these regions demonstrated several similarities to other prokaryotic regulatory regions. 50% homology between the putative immunity promoter and other prokaryotic promoters is apparent, so are similarities in AT-content. Upstream of the ATG start codon the sequence PuPuTTTPuPu and a termination codon (TAA) appear; both are typical of prokaryotic ribosome binding sites. The colicin terminator demonstrated similarities to other rho-independent prokaryotic terminators: a GC-rich region with termination in an adjacent AT-rich region containing T clusters on the non-coding strand. The possible role of initiation upstream from the colicin terminator is discussed.

Base Sequence

Analysis of targeted and whole genome sequencing of PacBio HiFi reads for a comprehensive genotyping of gene-proximal and phenotype-associated Variable Number Tandem Repeats.

Variable Number Tandem repeats (VNTRs) refer to repeating motifs of size greater than five bp. VNTRs are an important source of genetic variation, and have been associated with multiple Mendelian and complex phenotypes. However, the highly repetitive structures require reads to span the region for accurate genotyping. Pacific Biosciences HiFi sequencing spans large regions and is highly accurate but relatively expensive. Therefore, targeted sequencing approaches coupled with long-read sequencing have been proposed to improve efficiency and throughput. In this paper, we systematically explored the trade-off between targeted and whole genome HiFi sequencing for genotyping VNTRs. We curated a set of 10&#xa0;,&#xa0;787 gene-proximal (G-)VNTRs, and 48 phenotype-associated (P-)VNTRs of interest. Illumina reads only spanned 46% of the G-VNTRs and 71% of P-VNTRs, motivating the use of HiFi sequencing. We performed targeted sequencing with hybridization by designing custom probes for 9,999 VNTRs and sequenced 8 samples using HiFi and Illumina sequencing, followed by adVNTR genotyping. We compared these results against HiFi whole genome sequencing (WGS) data from 28 samples in the Human Pangenome Reference Consortium (HPRC). With the targeted approach only 4,091 (41%) G-VNTRs and only 4 (8%) of P-VNTRs were spanned with at least 15 reads. A smaller subset of 3,579 (36%) G-VNTRs had higher median coverage of at least 63 spanning reads. The spanning behavior was consistent across all 8 samples. Among 5,638 VNTRs with low-coverage (&#xa0;<&#xa0;15), 67% were located within GC-rich regions (&#xa0;>&#xa0;60%). In contrast, the 40X WGS HiFi dataset spanned 98% of all VNTRs and 49 (98%) of P-VNTRs with at least 15 spanning reads, albeit with lower coverage. Spanning reads were sufficient for accurate genotyping in both cases. Our findings demonstrate that targeted sequencing provides consistently high coverage for a small subset of low-GC VNTRs, but WGS is more effective for broad and sufficient sampling of a large number of VNTRs.

Minisatellite Repeats

Codon Composition in Human Oocytes Reveals Age-Associated Defects in mRNA Decay.

Oocytes from women of advanced reproductive age exhibit diminished developmental potential, but the underlying mechanisms remain incompletely defined. Oocyte maturation depends on translational control of maternal mRNA synthesized during growth. We performed a computational analysis on human oocytes from women <30 versus &#x2265;40 years and observed that mRNA GC content correlates negatively with half-life in oocytes from young (<30 yr) but positively with oocytes from aged (>40 yr) women. In young oocytes, longer mRNA half-life is associated with lower protein abundance, whereas in aged oocytes GC content correlates positively with protein abundance. During the GV-to-MII transition, codon composition stratifies stability: codons that support rapid translation (optimal) stabilize mRNA, while slow-translating codons (non-optimal) promote decay. With reproductive aging, GC-containing codons become more optimal and align with increased protein abundance. These findings indicate that reproductive aging remodels codon-optimality-linked, translation-coupled mRNA decay, stabilizing a subset of GC-rich maternal mRNA that may be prone to excess translation during maturation. Our analysis is explicitly within human reproductive aging; it does not revisit cross-species stability rules. Instead, it shows that sequence-stability relations are reprogrammed with age within human oocytes, including an inversion of the GC-stability association during GV-to-MII transition. Disruption of the normal mRNA clearance program in aged oocytes may compromise oocyte competence and alter maternal mRNA dosage, with downstream consequences for early embryonic development.

Humans

SV40 DNA sequences as an example of the structure of genes functioning in animal cell nuclei.

Recent studies of the structure of messenger RNA have demonstrated the existence of untranslated sequences of the 3' and 5' end of the messages. In addition analysis of transcription in vitro has indicated that the nucleotide sequence U6 purine may be part of a transcription termination signal in prokaryotes. Recently it has been possible to determine the sequence of extensive portions of the DNA of SV40 virus. This article reviews the analogies between certain of these sequences and sequences available from prokaryotic messengers and DNAs. Unusual structures, including blocks of AT-rich and GC-rich segment sections and symmetric regions in the DNA near the origin of DNA replication, have been demonstrated and the distribution of stretches of 6 or more deoxyadenylic acids in the DNA of SV40 is consistent with some rho for these sequences in animal cells, either as terminators of transcription or as sites where degradation of transcripts is initiated or sites related to the selective rejection or degradation of transcipts.

Animals

Improving spliced alignment by modeling splice sites with deep learning.

MOTIVATION: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands sophisticated modeling of splice sites, but current aligners use simple models, which may affect their accuracy given dissimilar sequences. RESULTS: We implemented minisplice to learn splice signals with a one-dimensional convolutional neural network (1D-CNN) and trained a model with 7,026 parameters for vertebrate and insect genomes. It captures conserved splice signals across phyla and reveals GC-rich introns specific to mammals and birds. We used this model to estimate the empirical splicing probability for every GT and AG in genomes, and modified minimap2 and miniprot to leverage pre-computed splicing probability during alignment. Evaluation on human long-read RNA-seq data and cross-species protein datasets showed our method greatly improves the junction accuracy especially for noisy long RNA-seq reads and proteins of distant homology. AVAILABILITY AND IMPLEMENTATION: https://github.com/lh3/minisplice.

Journal Article

TAp73beta and DNp73beta activate the expression of the pro-survival caspase-2S.

p73, the p53 homologue, exists as a transactivation-domain-proficient TAp73 or deficient deltaN(DN)p73 form. Expectedly, the oncogenic DNp73 that is capable of inactivating both TAp73 and p53 function, is over-expressed in cancers. However, the role of TAp73, which exhibits tumour-suppressive properties in gain or loss of function models, in human cancers where it is hyper-expressed is unclear. We demonstrate here that both TAp73 and DNp73 are able to specifically transactivate the expression of the anti-apoptotic member of the caspase family, caspase-2(S). Neither p53 nor TAp63 has this property, and only the p73beta form, but not the p73alpha form, has this competency. Caspase-2 promoter analysis revealed that a non-canonical, 18 bp GC-rich Sp-1-binding site-containing region is essential for p73beta-mediated activation. However, mutating the Sp-1-binding site or silencing Sp-1 expression did not affect p73beta's transactivation ability. In vitro DNA binding and in vivo chromatin immunoprecipitation assays indicated that p73beta is capable of directly binding to this region, and consistently, DNA binding p73 mutant was unable to transactivate caspase-2(S). Finally, DNp73beta over-expression in neuroblastoma cells led to resistance to cell death, and concomitantly to elevated levels of caspase-2(S.) Silencing p73 expression in these cells led to reduction of caspase-2(S) expression and increased cell death. Together, the data identifies caspase-2(S) as a novel transcriptional target common to both TAp73 and DNp73, and raises the possibility that TAp73 may be over-expressed in cancers to promote survival.

Binding Sites