Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “GC-rich sequence”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Codon Composition in Human Oocytes Reveals Age-Associated Defects in mRNA Decay.

Oocytes from women of advanced reproductive age exhibit diminished developmental potential, but the underlying mechanisms remain incompletely defined. Oocyte maturation depends on translational control of maternal mRNA synthesized during growth. We performed a computational analysis on human oocytes from women <30 versus &#x2265;40 years and observed that mRNA GC content correlates negatively with half-life in oocytes from young (<30 yr) but positively with oocytes from aged (>40 yr) women. In young oocytes, longer mRNA half-life is associated with lower protein abundance, whereas in aged oocytes GC content correlates positively with protein abundance. During the GV-to-MII transition, codon composition stratifies stability: codons that support rapid translation (optimal) stabilize mRNA, while slow-translating codons (non-optimal) promote decay. With reproductive aging, GC-containing codons become more optimal and align with increased protein abundance. These findings indicate that reproductive aging remodels codon-optimality-linked, translation-coupled mRNA decay, stabilizing a subset of GC-rich maternal mRNA that may be prone to excess translation during maturation. Our analysis is explicitly within human reproductive aging; it does not revisit cross-species stability rules. Instead, it shows that sequence-stability relations are reprogrammed with age within human oocytes, including an inversion of the GC-stability association during GV-to-MII transition. Disruption of the normal mRNA clearance program in aged oocytes may compromise oocyte competence and alter maternal mRNA dosage, with downstream consequences for early embryonic development.

Humans↗

Electron microscopic characterization of Rhizobium bacteriophage 16-6-12 and its isolated deoxyribonucleic acid.

Bacteriophage 16-6-12 of Rhizobium lupini has a long, non-contractile tail and a head which is hexagonal in outline. The tail is 140 nm in length, 11 nm in diameter, and carries a short term fiber. Analysis of the tail structure by optical diffraction indicates that it is of the helical "stacked disc" type. After phenol-extraction from purified particles, the DNA of phage 16-6-12 can circularize in vitro. No significant difference in contour length was observed between the linear (14.34 plus or minus 0.28 mum) and circular (14.44 plus or minus 0.24 mum) forms of molecules. After partial denaturation with alkali an AT-GC-map was constructed, which shows an asymmetric distribution of AT- and GC-rich regions. It is concluded that this phage DNA can circularize due to the presence of cohesive ends and that it is not circularly permuted.

Bacteriophages↗

SV40 DNA sequences as an example of the structure of genes functioning in animal cell nuclei.

Recent studies of the structure of messenger RNA have demonstrated the existence of untranslated sequences of the 3' and 5' end of the messages. In addition analysis of transcription in vitro has indicated that the nucleotide sequence U6 purine may be part of a transcription termination signal in prokaryotes. Recently it has been possible to determine the sequence of extensive portions of the DNA of SV40 virus. This article reviews the analogies between certain of these sequences and sequences available from prokaryotic messengers and DNAs. Unusual structures, including blocks of AT-rich and GC-rich segment sections and symmetric regions in the DNA near the origin of DNA replication, have been demonstrated and the distribution of stretches of 6 or more deoxyadenylic acids in the DNA of SV40 is consistent with some rho for these sequences in animal cells, either as terminators of transcription or as sites where degradation of transcripts is initiated or sites related to the selective rejection or degradation of transcipts.

Animals↗

Improving spliced alignment by modeling splice sites with deep learning.

MOTIVATION: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands sophisticated modeling of splice sites, but current aligners use simple models, which may affect their accuracy given dissimilar sequences. RESULTS: We implemented minisplice to learn splice signals with a one-dimensional convolutional neural network (1D-CNN) and trained a model with 7,026 parameters for vertebrate and insect genomes. It captures conserved splice signals across phyla and reveals GC-rich introns specific to mammals and birds. We used this model to estimate the empirical splicing probability for every GT and AG in genomes, and modified minimap2 and miniprot to leverage pre-computed splicing probability during alignment. Evaluation on human long-read RNA-seq data and cross-species protein datasets showed our method greatly improves the junction accuracy especially for noisy long RNA-seq reads and proteins of distant homology. AVAILABILITY AND IMPLEMENTATION: https://github.com/lh3/minisplice.

Journal Article↗

A REB1-binding site is required for GCN4-independent ILV1 basal level transcription and can be functionally replaced by an ABF1-binding site.

The ILV1 gene of Saccharomyces cerevisiae encodes the first committed step in isoleucine biosynthesis and is regulated by general control of amino acid biosynthesis. Deletion analysis of the ILV1 promoter revealed a GC-rich element important for the basal level expression. This cis-acting element, called ILV1BAS, is functional independently of whether GCN4 protein is present. Furthermore, unlike the situation at HIS4, the magnitude of GCN4-mediated derepression is independent of ILV1BAS. The element has homology to the consensus REB1-binding sequence CGGGTARNNR. Gel retardation assays showed that REB1 binds specifically to this element. We show that REB1-binding sites normally situated in the SIN3 promoter and in the 35S rRNA promoter can substitute for the ILV1 REB1 site. Furthermore, a SIN3 REB1 site containing a point mutation that abolishes REB1 binding does not support ILV1 basal level expression, suggesting that binding of REB1 is important for the control of ILV1 basal level expression. Interestingly, an ABF1-binding site can also functionally replace the ILV1 REB1-binding site. A mutated ABF1 site that displays a very low affinity for ABF1 does not functionally replace the ILV1 REB1 site. This suggests that ABF1 and REB1 may have related functions within the cell. Although the REB1-binding site is required for the ILV1 basal level expression, the site on its own stimulates transcription only slightly when combined with the CYC1 downstream promoter elements, indicating that another ILV1 promoter element functions in combination with the REB1 site to control high basal level expression.

Base Sequence↗

FGF-mediated aspects of skeletal muscle growth and differentiation are controlled by a high affinity receptor, FGFR1.

Fibroblast growth factors (FGFs) and FGF receptors (FGFRs) play major roles in vertebrate embryogenesis, including control of skeletal muscle growth and differentiation. Understanding their roles requires delineating the specific FGF and FGFR isoforms involved. This study analyzes the FGFR transcripts found in a model mouse skeletal myoblast cell line (MM14) during growth and terminal differentiation. MM14 cells express transcripts for FGFR1 (flg) but not FGFR2 (bek). The predominate FGFR1 transcript contains three immunoglobulin (Ig)-like domains in the extracellular ligand binding region. Approximately one-fourth of the three Ig-like domain transcripts possess a 6-nt deletion between the first and second Ig-like domains which after translation would result in deletion of an Arg-Arg pair. Cloning of mouse genomic DNA surrounding the region of the FGFR1 6-nt deletion indicates that the deletion is derived by alternative splicing of FGFR1 transcripts. Transcripts containing two Ig-like domains account for less than 5% of total FGFR1 mRNA in MM14 cells. A survey of RNA from mouse tissues indicated that two Ig-like domain FGFR1 transcripts are rare in all tissues except in lung, in which the two Ig-like domain form accounts for roughly 70% of the lung FGFR1 mRNA. PCR RACE cloning studies disclosed 162 nt of additional FGFR1 5'-flanking RNA which was highly GC-rich. FGFR1 transcripts decline 8- to 10-fold during low serum, (-)FGF-mediated differentiation of MM14 cultures. The kinetics of the FGFR1 mRNA decline is similar to the previously described differentiation-dependent decrease in cell surface FGF receptors.

Amino Acid Sequence↗

Screening for mutations in human HPRT cDNA using the polymerase chain reaction (PCR) in combination with constant denaturant gel electrophoresis (CDGE).

Previously, we reported the modification of denaturing gradient gel electrophoresis called constant denaturant gel electrophoresis (CDGE). CDGE separates mutant fragments in specific melting domains. CDGE seems to be a useful tool in mutation detection. Since the hypoxanthine phosphoribosyltransferase (HPRT) gene is widely used as target locus for mutation studies in vitro and in vivo, we have examined the approach of analyzing human HPRT cDNA by polymerase chain reaction (PCR) and CDGE. All nine HPRT exons are included in a 716-bp cDNA fragment obtained by PCR using HPRT cDNA as template. When the full-length cDNA fragment was examined by CDGE, it was possible to detect mutations only in the last part of exon 8 and exon 9. However, digestion of the cDNA fragment with the restriction enzyme AvaI prior to CDGE enabled us to detect point mutations in most of exon 2, the beginning of exon 3, the last part of exon 8 and exon 9. With the use of two internal primer sets, including a GC-rich clamp on one of the primers in each pair, a region containing most of exon 3 through exon 6 was amplified and we were able to resolve fragments with point mutations in this region from wild-type DNA. The approach described here allows for rapid screening of point mutations in about two thirds of the human HPRT cDNA sequence. In a test of this approach, we were able to resolve 12 of 13 known mutants. The mutant panel included one single-base deletion, one two-base deletion and 11 single-base substitutions.

Base Sequence↗

Analysis of the regulation of the pelBC genes in Erwinia chrysanthemi 3937.

Erwinia chrysanthemi secretes five major isoenzymes of pectate lyases encoded by the pelABCDE genes. The nucleotide sequence of the region surrounding the pelB gene of E. chrysanthemi 3937 was determined, including the regulatory regions involved in pelB and pelC expression. Analysis of the transcripts showed that transcription of pelB or pelC gave, in both cases, only one transcript. The transcription initiation sites of both pelB and pelC were precisely determined as well as the position of the transcription termination of pelB. The pelB and pelC promoters are very similar, showing a good homology with the -35 consensus region but low homology with the -10 consensus. In both cases a KdgR-box overlaps the -35 region. The pelC gene may have two KdgR operators. Moreover, the pelB and pelC genes are preceded by other sequences presenting the typical symmetry of operator sites that could be involved in more specific regulations. Comparison of E. chyrsanthemi pel regulatory regions revealed three classes of homology: pelA, pelB-pelC and pelD-pelE. The sole regulatory sequence conserved among the three classes corresponds to the KdgR-binding site. Moreover, all the pel regulatory regions are AT-rich in contrast to the coding regions which are GC-rich. Gel retardation experiments with fragments overlapping the pelB or pelC regulatory regions demonstrated that the KdgR protein specifically binds to these regions. Other proteins probably also interact with these DNA fragments. Transcription of pelB terminates in a region corresponding to a GC-rich inverted repeat followed by a run of T residues, typical of rho-independent transcription termination sites. Moreover, preliminary results imply that a region adjacent to pelC provoke, directly or indirectly, the repression of pelB and pelC expression.

Amino Acid Sequence↗

Drug-DNA sequence-dependent interactions analysed by electric linear dichroism.

The interactions between 20 drugs and a variety of synthetic DNA polymers and natural DNAs were studied by electric linear dichroism (ELD). All compounds tested, including several clinically used antitumour agents, are thought to exert their biological activities mainly by virtue of their abilities to bind to DNA. The selected drugs include intercalating agents with fused and unfused aromatic structures and several groove binders. To examine the role of base composition and base sequence in the binding of these drugs to DNA, ELD experiments were carried out with natural DNAs of widely differing base composition as well as with polynucleotides containing defined alternating and non-alternating repeating sequences, poly(dA).poly(dT), poly(dA-dT).poly(dA-dT),poly(dG).poly(dC) and poly(dG-dC).poly(dG-dC). Among intercalating agents, actinomycin D was found to be by far the most GC-selective. GC selectivity was also observed with an amsacrine-4-carboxamide derivative and to a lesser extent with methylene blue. In contrast, the binding of amsacrine and 9-aminoacridine was practically unaffected by varying the GC content of the DNAs. Ethidium bromide, proflavine, mitoxantrone, daunomycin and an ellipticine derivative were found to bind best to alternating purine-pyrimidine sequences regardless of their nature. ELD measurements provided evidence for non-specific intercalation of amiloride. A significant AT selectivity was observed with hycanthone and lucanthone. The triphenyl methane dye methyl green was found to exhibit positive and negative dichroism signals at AT and GC sites, respectively, showing that the mode of binding of a drug can change markedly with the DNA base composition. Among minor groove binders, the N-methylpyrrole carboxamide-containing antibiotics netropsin and distamycin bound to DNA with very pronounced AT specificity, as expected. More interestingly the dye Hoechst 33258, berenil and a thiazole-containing lexitropsin elicited negative reduced dichroism in the presence of GC-rich DNA which is totally inconsistent with a groove binding process. We postulate that these three drugs share with the trypanocide 4',6-diamidino-2-phenylindole (DAPI) the property of intercalating at GC-rich sites and binding to the minor groove of DNA at other sites. Replacement of guanines by inosines (i.e., removal of the protruding exocyclic C-2 amino group of guanine) restored minor groove binding of DAPI, Hoechst 33258 and berenil. Thus there are several cases where the mode of binding to DNA is directly dependent on the base composition of the polymer. Consequently the ELD technique appears uniquely valuable as a means of investigating the possibility of sequence-dependent recognition of DNA by drugs.

Base Sequence↗

Construction of a chimeric ArsA-ArsB protein for overexpression of the oxyanion-translocating ATPase.

Resistance to toxic oxyanions of arsenic and antimony in Escherichia coli is conferred by the conjugative R-factor R773, which encodes an ATP-driven anion extrusion pump. The ars operon is composed of three structural genes, arsA, arsB, and arsC. Although transcribed as a single unit, the three genes are differentially expressed as a result of translational differences, such that the ArsA and ArsC proteins are produced in high amounts relative to the amount of ArsB protein made. Consequently, biochemical characterization of the ArsB protein, which is an integral membrane protein containing the anion-conducting pathway, has been limited, precluding studies of the mechanism of this oxyanion pump. To overexpress the arsB gene, a series of changes were made. First, the second codon, an infrequently used leucine codon, was changed to a more frequently utilized codon. Second, a GC-rich stem-loop (delta G = -17 kcal/mol) between the third and twelfth codons was destabilized by changing several of the bases of the base-paired region. Third, the re-engineered arsB gene was fused 3' in frame to the first 1458 base pairs of the arsA gene to encode a 914-residue chimeric protein (486 residues of the ArsA protein plus 428 residues of the mutated ArsB protein) containing the entire re-engineered ArsB sequence except for the initiating methionine. The ArsA-ArsB chimera has been overexpressed at approximately 15-20% of the total membrane proteins. Cells producing the chimeric ArsA-ArsB protein with an arsA gene in trans excluded 73AsO2- from cells, demonstrating that the chimera can function as a component of the oxyanion-translocating ATPase.

Adenosine Triphosphatases↗

Structure and expression of the Kas12 gene encoding a beta-ketoacyl-acyl carrier protein synthase I isozyme from barley.

The beta-ketoacyl-acyl carrier protein (ACP) synthase I in the plant fatty acid synthetase catalyzes the condensations of acetate units to a growing acyl-ACP leading to the synthesis of palmitoyl-ACP. Barley chloroplasts contain three cerulenin sensitive beta-ketoacyl-ACP synthase I isoforms, alpha 2, alpha beta, and beta 2. The Kas12 gene encoding the beta 2 isozyme has been isolated and sequenced. The gene spans 3.8 kilobases and contains seven exons separated by six intervening sequences varying from 75 to 1008 base pairs in length. The mosaic gene structure is different compared with that of the beta-ketoacyl synthase in the multifunctional rat and goose fatty acid synthetases. Southern blot analyses of genomic DNA from barley, wheat, and the barley-wheat chromosome addition lines indicate that Kas12 is a single copy gene located on chromosome 2. Primer extension analyses identified four transcription start sites located 168-171 nucleotides upstream from the translation initiation codon. The Kas12 promoter lacks an appropriately positioned TATA box and contains a GC-rich region including two GC elements similar to the Sp1 transcription factor-binding site. In this regard Kas12 closely resembles a set of ubiquitously expressed eucaryotic genes. In accord with this deduction, polymerase chain reaction analysis showed that the Kas12 transcript is present in barley roots, germinating embryos, developing kernels, and leaves.

3-Oxoacyl-(Acyl-Carrier-Protein) Synthase↗

[Chiasma distribution in the lampbrush chromosomes of the chicken Gallus gallus domesticus: hot spots of recombination and their possible role in proper dysjunction of homologous chromosomes at the first meiotic division].

Chiasma distribution in the lambrush chromosomes of the chicken Gallus gallus domesticus was studied. The data of the authors show that the general pattern of chiasmata in the interstitional region of chromosomes corresponds to the Poisson distribution. However, in the telomeric and subtelomeric regions of all chicken macrochromosomes one can see chiasma as a rule. In the half of 140 microchromosomes from 24 different oocytes, there are also the telomeric chiasmata. On the basis of this observation, it may be predicted that there are hot spots of recombination near or into the telomeric GC-rich heterochromatic bands of chicken chromosomes. We suggest that these hot spots of recombination near the telomeres are a necessary facility for not only macrochromosomes but all microchromosomes as well to have at least one chiasma. The constant presence of at least one chiasma in a bivalent in needed for correct disjunction of homologous chromosomes at the first meiotic division.

Animals↗

Cloning and characterization of the 5'-flanking region of the human topoisomerase II alpha gene.

Topoisomerases are essential enzymes for DNA metabolism in prokaryotes and eukaryotes. In human cells, DNA topoisomerase II enzyme activity can be modulated by both viral transformation and changes in proliferation status. To identify elements important for regulation of topoisomerase II alpha gene expression, genomic DNA clones covering the 5'-end of the gene were isolated. The intron/exon structure of a 2.5-kilobase region encompassing the translation start site was determined. Transcription was found to initiate at multiple sites clustered around 90 base pairs 5' to the ATG initiation codon. Transient expression of chimeric topoisomerase II-reporter gene constructs in HeLa cells revealed that the 5'-flanking region exhibited promoter activity. The region -90 to -1 upstream of the major transcription start site was shown by deletion analysis to include a promoter. This minimal promoter lacks a TATA box, is moderately GC-rich, and contains a high frequency of CpG dinucleotides; characteristic of a "housekeeping" gene promoter. Maximal promoter activity was observed using a fragment extending to position -562. Putative regulatory elements are contained within and immediately upstream of the minimal promoter region. The regulatory region of the topoisomerase II alpha gene identified here is similar in basic structure to those of the human thymidine kinase and DNA polymerase alpha genes, which are also controlled by proliferation-specific factors.

Base Sequence↗

Polymerase chain reaction analysis of fragile X mutations.

The mutation that underlies the fragile X syndrome is presumed to be a large expansion in the number of CGG repeats within the gene FMR-1. The unusually GC-rich composition of the expanded region has impeded attempts to amplify it by the polymerase chain reaction (PCR). We have developed a PCR protocol that successfully amplifies the (CGG)n region in normal, carrier and affected individuals. The PCR analysis of several large fragile X families is presented. The PCR results agree with those obtained by direct genomic Southern blot analyses. These favorable comparisons suggest that the PCR assay may be suitable for rapid testing for fragile X mutations and premutations and genetic screening of at-risk individuals.

Base Sequence↗

A novel myoblast enhancer element mediates MyoD transcription.

The MyoD gene can orchestrate the expression of the skeletal muscle differentiation program. We have identified the regions of the gene necessary to reproduce transcription specific to skeletal myoblasts and myotubes. A proximal regulatory region (PRR) contains a conserved TATA box, a CCAAT box, and a GC-rich region that includes a consensus SP1 binding site. The PRR is sufficient for high levels of skeletal muscle-specific activity in avian muscle cells. In murine cells the PRR alone has only low levels of activity and requires an additional distal regulatory region to achieve high levels of muscle-specific activity. The distal regulatory region differs from a conventional enhancer in that chromosomal integration appears necessary for productive interactions with the PRR. While the Moloney leukemia virus long terminal repeat can enhance transcription from the MyoD PRR in both transient and stable assays, the simian virus 40 enhancer cannot, suggesting that specific enhancer-promoter interactions are necessary for PRR function.

Animals↗

TAp73beta and DNp73beta activate the expression of the pro-survival caspase-2S.

p73, the p53 homologue, exists as a transactivation-domain-proficient TAp73 or deficient deltaN(DN)p73 form. Expectedly, the oncogenic DNp73 that is capable of inactivating both TAp73 and p53 function, is over-expressed in cancers. However, the role of TAp73, which exhibits tumour-suppressive properties in gain or loss of function models, in human cancers where it is hyper-expressed is unclear. We demonstrate here that both TAp73 and DNp73 are able to specifically transactivate the expression of the anti-apoptotic member of the caspase family, caspase-2(S). Neither p53 nor TAp63 has this property, and only the p73beta form, but not the p73alpha form, has this competency. Caspase-2 promoter analysis revealed that a non-canonical, 18 bp GC-rich Sp-1-binding site-containing region is essential for p73beta-mediated activation. However, mutating the Sp-1-binding site or silencing Sp-1 expression did not affect p73beta's transactivation ability. In vitro DNA binding and in vivo chromatin immunoprecipitation assays indicated that p73beta is capable of directly binding to this region, and consistently, DNA binding p73 mutant was unable to transactivate caspase-2(S). Finally, DNp73beta over-expression in neuroblastoma cells led to resistance to cell death, and concomitantly to elevated levels of caspase-2(S.) Silencing p73 expression in these cells led to reduction of caspase-2(S) expression and increased cell death. Together, the data identifies caspase-2(S) as a novel transcriptional target common to both TAp73 and DNp73, and raises the possibility that TAp73 may be over-expressed in cancers to promote survival.

Binding Sites↗

Variable substructure in the secondary constriction of the human chromosome 1.

The secondary constriction in human chromosome 1 consists of a proximal segment stained by the GC-specific fluorochrome mithramycin and a distal segment stained by such fluorochromes as DAPI or DIPI, which show enhanced fluorescence intensities in AT-rich regions of the chromosomes. A study involving 21 individuals revealed that both parts are independently involved in length variability. In two cases, two GC-rich regions separated by an AT-rich segment and an additional distal AT-rich part were found.

Base Sequence↗

Compositional bimodality and evolution of retroviral genomes.

The compositional distributions of genomes, genes (and their third codon positions) and long terminal repeats from retroviruses of warm-blooded vertebrates are characterized by a striking bimodality which is accompanied by a remarkable compositional homogeneity within each retroviral genome. A first, major class of retroviral genomes is GC-rich, whereas a second, minor class is GC-poor. Representative expressed viral genomes from the two classes integrate in GC-rich and GC-poor isochores, respectively, of host genomes. The first class comprises all oncoviruses (except B-types and some D-types), the second, lentiviruses, spumaviruses, as well as B-type and some D-type oncoviruses (e.g., mouse mammary tumor virus and simian retroviruses type D, respectively). The compositional bimodal distribution of retroviral genomes and the accompanying compositional homogeneity within each retroviral genome appear to be the result of the compositional evolution of retroviral genomes in their integrated form.

Animals↗