Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “GC-rich sequence”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Characterization of a 12-mer duplex d(GGCGGAGTTAGG).d(CCTAACTCCGCC) containing a highly reactive (+)-CC-1065 sequence by 1H and 31P NMR, hydroxyl-radical footprinting, and NOESY restrained molecular dynamics calculations.

The solution structure of the GC-rich non-self-complementary DNA 12-mer duplex (I), which contains a (+)-CC-1065 highly reactive bonding sequence 5'AGTTA* (where * denotes the [formula: see text] covalent modification site), has been examined thoroughly by one- and two-dimensional proton and phosphorus NMR spectroscopy, hydroxyl-radical footprinting, and NOESY restrained molecular mechanics and dynamics calculations. The assignments of the nonexchangeable proton resonances (except some of the H5' and H5" protons due to severe resonance overlap), phosphorus resonances, and the exchangeable resonances (except amino protons of adenosine and guanosine) of this 12-mer duplex have been made. The results show that this 12-mer duplex maintains an overall B-form DNA with all anti base orientation throughout in aqueous solution at room temperature. Hydroxyl-radical footprinting experiments on a 21-mer sequence that contains this 12-mer duplex used for NMR studies showed that the minor groove is somewhat narrowed at the 7G-8T and 17A-18C steps, as indicated by the inhibition of cleavage at these locations. Although both high-field NMR and hydroxyl-radical footprinting experiments supported a bent-like structure for this 12-mer duplex, nondenaturing gel electrophoresis on the ligated 21-mer sequence that contains this 12-mer duplex did not show the abnormally slow migration characteristic of a bent DNA duplex. Analysis of the NMR data sets reveals several local structural perturbations similar to those found on an (A)n tract DNA duplex. For example, the existence of a propeller twist was detected within the A.T-rich region for both the 12-mer and the (A)n tract DNA duplexes. The 18CH5 aromatic resonance that is directly adjacent to the 3' side of the 5'TAA segment was significantly shifted upfield with a chemical shift of 5.10 ppm, which is almost within the region normally associated with sugar H3' protons. The sugar geometries for 18C and 7G, which are located to the 3' side of the 5'TAA segment, are proposed to be in the neighborhood of C3'-endo and O1'-endo in equilibrium C3'-endo, respectively. We propose that this unusually upfield-shifted resonance signal for 18CH5 and the average C3'-endo sugar geometry for 18C nucleotide on the 12-mer duplex is connected with the peculiar conformation, possibly a transient kink, within the 5'AC/GT step. The results of the NOESY restrained molecular mechanics and dynamics calculations on the 12-mer sequence reveal two kinks, which are located on either side of the 18C nucleotide that has an average C3'-endo sugar geometry.(ABSTRACT TRUNCATED AT 400 WORDS)

Base Sequence↗

Molecular cloning of mouse connexins26 and -32: similar genomic organization but distinct promoter sequences of two gap junction genes.

Connexins26 and -32 are subunit proteins of gap junctions that are coexpressed in hepatocytes and several tissues but individually expressed in other cells. Molecular cloning of both corresponding mouse genes revealed similar genomic organization, i.e., each gene consists of two exons with the complete coding region located in the second exon. The first exon of each gene is preceded by a TATA-less promoter region. The promoter of the mouse Cx26 gene has at least two transcription start sites and is located in a very GC-rich region which is reminiscent of promoters of house-keeping genes. Putative consensus sequences for a metal response element, the transcription factor NFkappaB, and several GC-boxes were found within 600 bp upstream of the Cx26 transcription start sites. The promoter region of the mouse Cx32 gene contains two putative binding sites for the transcription factor HNF-1 and consensus motifs for NF-1 as well as NFkappaB within 680 bp upstream of the main transcription start site. Thus the sequence comparison of mouse Cx26 and Cx32 promoter regions provides hints for possible consensus elements that could control individual expression as well as common regulation of these gap junction genes in various tissues. Cx26 mRNA is much more abundant in adult mouse skin than in adult kidney and liver where Cx32 transcripts are relatively strongly expressed.

Amino Acid Sequence↗

A novel GC-rich human macrosatellite VNTR in Xq24 is differentially methylated on active and inactive X chromosomes.

A new X chromosome-specific repetitive sequence, a 3 kilobase HindIII clone with a base composition of 63% C+G, has been isolated. The sequence is organized as a hypervariable tandem repeat cluster ranging in size from 150-350 kilobases, with outlying single copies. This locus, designated DXZ4 and mapped to chromosome band Xq24, may consist of as many as 50 variable-length alleles. It represents a class of variable number of tandem repeat polymorphism which may be termed 'macrosatellite'. The cluster is highly methylated on the active X chromosome and hypomethylated on the inactive X.

Base Composition↗

Denaturation map of the ribosomal DNA of Lytechinus variegatus sperm.

Electron microscopy of the partially heat denatured ribosomal DNA (rDNA) from sea urchin (Lytechinus variegatus) sperm has demonstrated that it consists of repeating units of 3.6 +/- 0.2 micron, corresponding to a mol.wt of 7.2 +/- 0.4 x 10(6). Based on differential denaturability, each repeat unit is divided into 2 regions. The larger region of 2.47 +/- 0.11 micron (mol.wt 4.9 +/- 0.22 x 10(6)) corresponds in length to the ribosomal precursor RNA of sea urchins and the smaller, GC-rich, subunit of 1.16 +/- 0.09 micrometer (mol.wt 2.3 +/- 0.18 x 10(6)) is presumed to contain non-transcribed spacer sequences.

Animals↗

Mithramycin and DIPI: a pair of fluorochromes specific for GC-and AT-rich DNA respectively.

The AT specificity of the fluorochromes DIPI and DAPI and the GC specificity of mithramycin are evidenced by observations in human, mouse, and bovine chromosomes. DIPI and DAPI produce a pattern similar to Hoechst 33258 in all three species, whereas mithramycin results in a reverse pattern. The AT-rich centromeric heterochromatin in mouse is brilliantly stained by DIPI or DAPI and remains nearly invisible after mithramycin staining. In the GC-rich centromeric heterochromatin of cattle the opposite behavior is observed.

Animals↗

Preferential in vitro assembly of nucleosome cores on some AT-rich regions of SV40 DNA.

We have found that nucleosomes reconstituted from histone octamers and SV40 DNA Form I by progressively decreasing the salt concentration from 2 M NaCl are formed preferentially around 0.27, 0.37, 0.50 and 0.85 on SV40 DNA (relative to the EcoRI site). When SV40 DNA Form III is used, the nucleosomes form mainly at 0.28, 0.38, 0.61 and 0.83. These sites are very close to both the sites of RNA chain initiation by calf thymus RNA polymerase B on SV40 DNA Form I (0.25, 0.35, 0.42 and 0.88) and the regions of the supercoiled DNA which are readily denaturable by T4 gene 32 protein (0.25, 0.47 and 0.88), and correspond to AT-rich regions as deduced from the nucleotide sequence of SV40 DNA. The physiologically important region around 0.67 is an unfavourable site for all three types of proteins, and corresponds to a GC-rich region surrounding a 17 base pair AT cluster.

Base Composition↗

Nucleotide sequence of Xenopus borealis oocyte 5S DNA: comparison of sequences that flank several related eucaryotic genes.

Genomic Xenopus borealis oocyte-specific 5S DNA (Xbo) contains clusters of 5S rRNA genes. The number of genes varies among clusters, and the distance between genes within a cluster is about 80 nucleotides. The spacer DNA between gene clusters is AT-rich and heterogeneous in length due in part to variable numbers of a tandemly repeated 21 nucleotide sequence. A cloned fragment of Xbo 5S DNA (Xbo1) containing three 5S rRNA genes has been sequenced. The sequences of Xbo1 genes 1 and 2 are very similar to the dominant 5S RNA sequence, whereas 15 of the 120 residues in the third gene are different. The sequence of gene 3 is as different from the dominant gene sequence as the X. laevis pseudogene is from the 5S RNA gene. Sequence analysis of genomic DNA shows that gene 3 is an abundant component of the multigene family. All three genes are transcribed when added to an extract of X. laevis oocyte nuclei, and a fragment of Xbo1 lacking the AT-rich spacer DNA and the 5' end of the first gene supports transcription of genes 2 and 3 in this in vitro system. Thus the 80 nucleotides preceding each 5S gene are sufficient for promoter function. Nucleic acid sequences preceding several eucaryotic genes that are transcribed by RNA polymerase III were analyzed and the following common features were found: a purine-rich region; at least one direct repeat; the absence of dyad symmetry; transcription beginning with a purine; a pyrimidine residue immediately preceding the first nucleotide of the gene; and the oligonucleotides AAAAG, AGAAG and GAC, located approximately 15, 25 and 35 nucleotides, respectively, before the start of transcription. The 10 base pair (bp) spacing between the homologous oligonucleotides is that expected for a recognition signal on one face of a DNA double helix. The extensive sequence differences between most of the spacers that precedes these genes make the three conserved oligonucleotides more striking. Parts of the 5' flanking regions of the three Xbo1 gene (-12 to -40), which include the conserved oligonucleotides, are identical. In contrast, 7 of the first 11 nucleotides that precede the third 5S RNA gene in Xbo1 differ from those that precede the first gene. The sequences following the X. borealis oocyte and somatic 5S genes are identical in 12 of the first 14 residues and contain two or more T clusters, as does the corresponding region of X. laevis oocyte 5S DNA. The 3' sequences of the Xenopus 5S rRNA genes and several other eucaryotic genes contain features in common with procaryotic transcription termination sites. The 3' end of the gene is GC-rich and contains a dyad symmetry. Termination occurs in an AT-rich region containing one or more T clusters on the noncoding strand.

Animals↗

Characterization of in vitro transcription initiation and termination sites in Col E1 DNA.

Overlapping restriction fragments from the region between the single Eco R1 site and the origin of replication of the plasmid, Col E1, have been utilised as templates in an in vitro transcription assay using E. coli RNA polymerase. Transcription towards the single Eco R1 site is initiated at a point 415 bp to the origin side of that site. In vivo, transcription starting at this point probably produces the mRNA for the colicin immunity protein. Transcription away from the Eco R1 site is initiated at a point 140 bp to the origin side of that site and terminated 30 bp further on. This terminator is probably the point at which transcription of the colicin gene is terminated in vivo. DNA sequence analysis in both these regions demonstrated several similarities to other prokaryotic regulatory regions. 50% homology between the putative immunity promoter and other prokaryotic promoters is apparent, so are similarities in AT-content. Upstream of the ATG start codon the sequence PuPuTTTPuPu and a termination codon (TAA) appear; both are typical of prokaryotic ribosome binding sites. The colicin terminator demonstrated similarities to other rho-independent prokaryotic terminators: a GC-rich region with termination in an adjacent AT-rich region containing T clusters on the non-coding strand. The possible role of initiation upstream from the colicin terminator is discussed.

Base Sequence↗

Analysis of targeted and whole genome sequencing of PacBio HiFi reads for a comprehensive genotyping of gene-proximal and phenotype-associated Variable Number Tandem Repeats.

Variable Number Tandem repeats (VNTRs) refer to repeating motifs of size greater than five bp. VNTRs are an important source of genetic variation, and have been associated with multiple Mendelian and complex phenotypes. However, the highly repetitive structures require reads to span the region for accurate genotyping. Pacific Biosciences HiFi sequencing spans large regions and is highly accurate but relatively expensive. Therefore, targeted sequencing approaches coupled with long-read sequencing have been proposed to improve efficiency and throughput. In this paper, we systematically explored the trade-off between targeted and whole genome HiFi sequencing for genotyping VNTRs. We curated a set of 10&#xa0;,&#xa0;787 gene-proximal (G-)VNTRs, and 48 phenotype-associated (P-)VNTRs of interest. Illumina reads only spanned 46% of the G-VNTRs and 71% of P-VNTRs, motivating the use of HiFi sequencing. We performed targeted sequencing with hybridization by designing custom probes for 9,999 VNTRs and sequenced 8 samples using HiFi and Illumina sequencing, followed by adVNTR genotyping. We compared these results against HiFi whole genome sequencing (WGS) data from 28 samples in the Human Pangenome Reference Consortium (HPRC). With the targeted approach only 4,091 (41%) G-VNTRs and only 4 (8%) of P-VNTRs were spanned with at least 15 reads. A smaller subset of 3,579 (36%) G-VNTRs had higher median coverage of at least 63 spanning reads. The spanning behavior was consistent across all 8 samples. Among 5,638 VNTRs with low-coverage (&#xa0;<&#xa0;15), 67% were located within GC-rich regions (&#xa0;>&#xa0;60%). In contrast, the 40X WGS HiFi dataset spanned 98% of all VNTRs and 49 (98%) of P-VNTRs with at least 15 spanning reads, albeit with lower coverage. Spanning reads were sufficient for accurate genotyping in both cases. Our findings demonstrate that targeted sequencing provides consistently high coverage for a small subset of low-GC VNTRs, but WGS is more effective for broad and sufficient sampling of a large number of VNTRs.

Minisatellite Repeats↗

Capture of a cellular transcriptional unit by a retrovirus: mode of provirus activation in embryonal carcinoma cells.

The expression of murine leukemia provirus in embryonal carcinoma (EC) cells is blocked by a mechanism still incompletely understood. The blockage is not overcome by deleting a large portion of the enhancer region (in U3) in recombinant retroviruses (M-MuLVneo delta Enh). This confirms the presence of negative elements outside the viral 82-bp repeats. However, a few sites in the genomes of EC cells permit M-MuLVneo delta Enh proviral expression. One such site, identified in PCC4, PCC3, and LT, was studied. The complete analysis of the mechanism of activation by Northern (RNA) blotting, cloning, and sequencing of partial cDNA copies of the viral transcript and of the site of integration establishes that viral transcripts are initiated from an upstream host-cell promoter and are spliced from a host donor to a cryptic viral acceptor at position 542 in the Moloney murine leukemia virus (M-MuLV) genome. In consequence, the mature transcripts are host cell-virus fusion transcripts from which M-MuLV sequences, including the cis-active negative elements of the 5' long terminal repeat-containing region, are absent. The provirus integrates apparently randomly into any of the three most proximal introns of the transcriptional unit. The host cell promoter contains a TATA box and 14 potential SpI binding sites included in a 1.0-kb GC-rich island. These elements promote gene expression of recombinant vectors in EC and differentiated cells. The mechanism described points to a mechanism by which retroviruses can be transcribed from upstream nonviral elements and can acquire host genes by 5' annexation of exons.

3T3 Cells↗

Identification and characterization of the GC-rich and cyclic adenosine 3',5'-monophosphate (cAMP)-inducible promoter of the type II beta cAMP-dependent protein kinase regulatory subunit gene.

A rat genomic clone containing 4.5 kilobases of 5'-flanking DNA and the first exon of the type II beta regulatory subunit (RII beta) of cAMP-dependent protein kinase was isolated, restriction mapped, and sequenced. The proximal 400-basepair promoter region was GC rich, lacked TATA/CAAT box motifs, and initiated transcription at multiple sites. Bandshifting and DNase-I footprinting experiments using this region of the RII beta promoter detected several related specific DNA-protein complexes formed using crude and fractionated nuclear extracts from rat ovary, brain, adrenal gland, and liver. All binding in these experiments mapped to a domain within the same region found to confer cAMP inducibility to a chloramphenicol acetyltransferase (CAT) reporter gene when transfected into primary cultures of rat granulosa cells. Although GC boxes (putative SP1-binding sites) and activator protein-2 (AP-2) elements were present in this functional region, and although expression vectors containing AP-2 sites conferred high levels of cAMP regulation of the CAT gene in cultured ovarian cells, neither the GC boxes nor the AP-2 sites were protected by footprint analyses or required for band shift activity of nuclear extract protein. These known regulatory elements, therefore, may be involved in functional activity of the RII beta promoter, but additional cis-acting DNA and trans-acting factors (yet to be characterized) also appear to interact with the functional promoter of the RII beta gene and regulate the hormone-specific expression of the A-kinase subunit in ovarian and neuronal cells.

Amino Acid Sequence↗

Codon Composition in Human Oocytes Reveals Age-Associated Defects in mRNA Decay.

Oocytes from women of advanced reproductive age exhibit diminished developmental potential, but the underlying mechanisms remain incompletely defined. Oocyte maturation depends on translational control of maternal mRNA synthesized during growth. We performed a computational analysis on human oocytes from women <30 versus &#x2265;40 years and observed that mRNA GC content correlates negatively with half-life in oocytes from young (<30 yr) but positively with oocytes from aged (>40 yr) women. In young oocytes, longer mRNA half-life is associated with lower protein abundance, whereas in aged oocytes GC content correlates positively with protein abundance. During the GV-to-MII transition, codon composition stratifies stability: codons that support rapid translation (optimal) stabilize mRNA, while slow-translating codons (non-optimal) promote decay. With reproductive aging, GC-containing codons become more optimal and align with increased protein abundance. These findings indicate that reproductive aging remodels codon-optimality-linked, translation-coupled mRNA decay, stabilizing a subset of GC-rich maternal mRNA that may be prone to excess translation during maturation. Our analysis is explicitly within human reproductive aging; it does not revisit cross-species stability rules. Instead, it shows that sequence-stability relations are reprogrammed with age within human oocytes, including an inversion of the GC-stability association during GV-to-MII transition. Disruption of the normal mRNA clearance program in aged oocytes may compromise oocyte competence and alter maternal mRNA dosage, with downstream consequences for early embryonic development.

Humans↗

Electron microscopic characterization of Rhizobium bacteriophage 16-6-12 and its isolated deoxyribonucleic acid.

Bacteriophage 16-6-12 of Rhizobium lupini has a long, non-contractile tail and a head which is hexagonal in outline. The tail is 140 nm in length, 11 nm in diameter, and carries a short term fiber. Analysis of the tail structure by optical diffraction indicates that it is of the helical "stacked disc" type. After phenol-extraction from purified particles, the DNA of phage 16-6-12 can circularize in vitro. No significant difference in contour length was observed between the linear (14.34 plus or minus 0.28 mum) and circular (14.44 plus or minus 0.24 mum) forms of molecules. After partial denaturation with alkali an AT-GC-map was constructed, which shows an asymmetric distribution of AT- and GC-rich regions. It is concluded that this phage DNA can circularize due to the presence of cohesive ends and that it is not circularly permuted.

Bacteriophages↗

SV40 DNA sequences as an example of the structure of genes functioning in animal cell nuclei.

Recent studies of the structure of messenger RNA have demonstrated the existence of untranslated sequences of the 3' and 5' end of the messages. In addition analysis of transcription in vitro has indicated that the nucleotide sequence U6 purine may be part of a transcription termination signal in prokaryotes. Recently it has been possible to determine the sequence of extensive portions of the DNA of SV40 virus. This article reviews the analogies between certain of these sequences and sequences available from prokaryotic messengers and DNAs. Unusual structures, including blocks of AT-rich and GC-rich segment sections and symmetric regions in the DNA near the origin of DNA replication, have been demonstrated and the distribution of stretches of 6 or more deoxyadenylic acids in the DNA of SV40 is consistent with some rho for these sequences in animal cells, either as terminators of transcription or as sites where degradation of transcripts is initiated or sites related to the selective rejection or degradation of transcipts.

Animals↗

Improving spliced alignment by modeling splice sites with deep learning.

MOTIVATION: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands sophisticated modeling of splice sites, but current aligners use simple models, which may affect their accuracy given dissimilar sequences. RESULTS: We implemented minisplice to learn splice signals with a one-dimensional convolutional neural network (1D-CNN) and trained a model with 7,026 parameters for vertebrate and insect genomes. It captures conserved splice signals across phyla and reveals GC-rich introns specific to mammals and birds. We used this model to estimate the empirical splicing probability for every GT and AG in genomes, and modified minimap2 and miniprot to leverage pre-computed splicing probability during alignment. Evaluation on human long-read RNA-seq data and cross-species protein datasets showed our method greatly improves the junction accuracy especially for noisy long RNA-seq reads and proteins of distant homology. AVAILABILITY AND IMPLEMENTATION: https://github.com/lh3/minisplice.

Journal Article↗

Construction of a chimeric ArsA-ArsB protein for overexpression of the oxyanion-translocating ATPase.

Resistance to toxic oxyanions of arsenic and antimony in Escherichia coli is conferred by the conjugative R-factor R773, which encodes an ATP-driven anion extrusion pump. The ars operon is composed of three structural genes, arsA, arsB, and arsC. Although transcribed as a single unit, the three genes are differentially expressed as a result of translational differences, such that the ArsA and ArsC proteins are produced in high amounts relative to the amount of ArsB protein made. Consequently, biochemical characterization of the ArsB protein, which is an integral membrane protein containing the anion-conducting pathway, has been limited, precluding studies of the mechanism of this oxyanion pump. To overexpress the arsB gene, a series of changes were made. First, the second codon, an infrequently used leucine codon, was changed to a more frequently utilized codon. Second, a GC-rich stem-loop (delta G = -17 kcal/mol) between the third and twelfth codons was destabilized by changing several of the bases of the base-paired region. Third, the re-engineered arsB gene was fused 3' in frame to the first 1458 base pairs of the arsA gene to encode a 914-residue chimeric protein (486 residues of the ArsA protein plus 428 residues of the mutated ArsB protein) containing the entire re-engineered ArsB sequence except for the initiating methionine. The ArsA-ArsB chimera has been overexpressed at approximately 15-20% of the total membrane proteins. Cells producing the chimeric ArsA-ArsB protein with an arsA gene in trans excluded 73AsO2- from cells, demonstrating that the chimera can function as a component of the oxyanion-translocating ATPase.

Adenosine Triphosphatases↗

TAp73beta and DNp73beta activate the expression of the pro-survival caspase-2S.

p73, the p53 homologue, exists as a transactivation-domain-proficient TAp73 or deficient deltaN(DN)p73 form. Expectedly, the oncogenic DNp73 that is capable of inactivating both TAp73 and p53 function, is over-expressed in cancers. However, the role of TAp73, which exhibits tumour-suppressive properties in gain or loss of function models, in human cancers where it is hyper-expressed is unclear. We demonstrate here that both TAp73 and DNp73 are able to specifically transactivate the expression of the anti-apoptotic member of the caspase family, caspase-2(S). Neither p53 nor TAp63 has this property, and only the p73beta form, but not the p73alpha form, has this competency. Caspase-2 promoter analysis revealed that a non-canonical, 18 bp GC-rich Sp-1-binding site-containing region is essential for p73beta-mediated activation. However, mutating the Sp-1-binding site or silencing Sp-1 expression did not affect p73beta's transactivation ability. In vitro DNA binding and in vivo chromatin immunoprecipitation assays indicated that p73beta is capable of directly binding to this region, and consistently, DNA binding p73 mutant was unable to transactivate caspase-2(S). Finally, DNp73beta over-expression in neuroblastoma cells led to resistance to cell death, and concomitantly to elevated levels of caspase-2(S.) Silencing p73 expression in these cells led to reduction of caspase-2(S) expression and increased cell death. Together, the data identifies caspase-2(S) as a novel transcriptional target common to both TAp73 and DNp73, and raises the possibility that TAp73 may be over-expressed in cancers to promote survival.

Binding Sites↗

Variable substructure in the secondary constriction of the human chromosome 1.

The secondary constriction in human chromosome 1 consists of a proximal segment stained by the GC-specific fluorochrome mithramycin and a distal segment stained by such fluorochromes as DAPI or DIPI, which show enhanced fluorescence intensities in AT-rich regions of the chromosomes. A study involving 21 individuals revealed that both parts are independently involved in length variability. In two cases, two GC-rich regions separated by an AT-rich segment and an additional distal AT-rich part were found.

Base Sequence↗