Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Inverted Repeat Sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Generation of authentic 3' termini of an H2A mRNA in vivo is dependent on a short inverted DNA repeat and on spacer sequences.

We have determined what sequences are required to generate the authentic 3' termini of a sea urchin H2A histone mRNA. We have constructed a series of deletion and insertion mutants in the cloned histone repeat unit h22 of Psammechinus miliaris and have analyzed the transcripts of both wild-type and mutant DNAs produced in the frog oocyte. The protein-coding sequences of the H2A gene can be removed without any deleterious effects on transcription initiation or termination. A 12 bp deletion, which removes a highly conserved inverted DNA repeat immediately preceding the H2A mRNA 3' terminus, elicits read-through of the polymerase into the spacer DNA further downstream. However, the inverted repeat and the sequence coding for the 3' terminus of the mRNA are by themselves not sufficient to generate faithful 3' ends. Our data suggest that spacer sequences downstream of the 3' mRNA terminus are required as well.

Animals↗

Characteristic enrichment of DNA repeats in different genomes.

Using computer programs developed for this purpose, we searched for various repeated sequences including inverted, direct tandem, and homopurine-homopyrimidine mirror repeats in various prokaryotes, eukaryotes, and an archaebacterium. Comparison of observed frequencies with expectations revealed that in bacterial genomes and organelles the frequency of different repeats is either random or enriched for inverted and/or direct tandem repeats. By contrast, in all eukaryotic genomes studied, we observed an overrepresentation of all repeats, especially homopurine-homopyrimidine mirror repeats. Analysis of the genomic distribution of all abundant repeats showed that they are virtually excluded from coding sequences. Unexpectedly, the frequencies of abundant repeats normalized for their expectations were almost perfect exponential functions of their size, and for a given repeat this function was indistinguishable between different genomes.

Algorithms↗

An inverted repeat triggers cytosine methylation of identical sequences in Arabidopsis.

The Wassilewskija (WS) strain of Arabidopsis has four PAI genes at three sites: an inverted repeat at one locus plus singlet genes at two unlinked loci. These four genes are methylated over their regions of DNA identity. In contrast, the Columbia (Col) strain has three singlet PAI genes with no methylation. To test the hypothesis that the WS inverted repeat locus triggers methylation of unlinked identical sequences, we introduced this locus into the Col background by genetic crosses. The inverted repeat induced de novo methylation of all three unmethylated Col PAI genes, with methylation efficiency varying with the position of the target locus. These results, plus results with inverted repeat transgenes, show that methylation is communicated by a DNA/DNA pairing mechanism.

Arabidopsis↗

Nucleotide sequence of 42 kbp of vaccinia virus strain WR from near the right inverted terminal repeat.

The nucleotide sequence of 42090 bp of vaccinia virus strain WR is presented. The sequence includes the SalI L, F, G and I fragments and starts near the centre of the HindIII A fragment and extends rightwards towards the genomic terminus, finishing approximately 0.5 kb internal of the inverted terminal repeat (ITR). Translation of this region has identified 65 open reading frames (ORFs) of greater than 65 amino acids in length. Fifty-one of these which do not extensively overlap other larger ORFs have been subjected to further analysis; the other 14 are termed minor ORFs. In the rightmost 28.7 kb, the genes are, with one exception, transcribed towards the genomic terminus, similar to the arrangement of genes at the left end of the virus genome. Internal of this region the genes are expressed off either DNA strand but still predominately rightwards. ORFs are tightly packed with few intergenic non-coding regions of greater than 250 bp. Protein sequence comparisons have established a remarkably high number of homologies with entries in existing protein databases. Of these, DNA ligase, thymidylate kinase, two serine-threonine protein kinases, two serine proteinase inhibitors (serpins), two interleukin-1 receptor homologous and a discontinuous ORF related to tumour necrosis factor receptor have been reported. Other homologies include lectins, profilin, 3 beta-hydroxy steroid dehydrogenase, superoxide dismutase, guanylate kinase, ankyrin and complement factor H. In addition, there are a number of polypeptides with predicted properties of membrane-associated, secretory or glyco-proteins. Twelve gene families are described here and elsewhere. There is considerable similarity between genes from the right and left end of the virus genome that may have arisen by terminal transposition events. Several differences from the corresponding region of vaccinia virus strain Copenhagen sequence are noted. Near the right terminus the sequences diverge completely, and internal of this there are multiple examples of deletion of short sequences (eight to 10 nucleotides) that lie within penta- or hexanucleotide direct repeats.

Amino Acid Sequence↗

Cloning and nucleotide sequence analysis of the Streptococcus mutans membrane-bound, proton-translocating ATPase operon.

The function of the membrane-bound ATPase in S. mutans is to regulate cytoplasmic pH values for the purpose of maintaining delta pH. Previous studies have shown that as part of its acid-adaptive ability, S. mutans is able to increase H(+)-ATPase levels in response to acidification. As part of the study of ATPase regulation in S. mutans, we have cloned the ATPase operon and determined its genetic organization. The structural genes from S. mutans were found to be in the order: c, a, b, delta, alpha, gamma, beta, and epsilon; where c and a were reversed from the more typical bacterial organization. The operon contained no I gene homologue but was preceded by a 239-bp intergenic space. Deduced aa sequences from open reading frames indicated that genes encoding homologues of glycogen phosphorylase and nonphosphorylating, NADP-dependent glyceraldehyde-3-phosphate dehydrogenase flank the H(+)-ATPase operon, 5' and 3' respectively. Sequence analysis indicated the presence of three inverted-repeat nt sequences in the glgP-uncE intergenic space. Primer extension analysis of mRNAs prepared from batch-grown or steady-state cultures demonstrated that the transcriptional start site did not change as a function of culture pH value. The data suggest that potential stem-and-loop structures in the promoter region of the operon do not function to alter the starting position of ATPase-specific mRNA transcription.

Amino Acid Sequence↗

An unusual nucleoporin-related messenger ribonucleic acid is present in the germ cells of rat testis.

An mRNA with a substantial similarity to the rat p62 mRNA that encodes a nucleoporin was cloned from the rat testis. A probe derived from a unique sequence in the nucleoporin-related (NPR) cDNA revealed a novel mRNA of 1.3 kb, different from the 2.7-kb transcript attributed to the p62 gene. This 1.3-kb transcript was not detected in Sertoli cells; it was found primarily in the haploid germ cells of the adult testis. The DNA sequencing revealed that the central region of the NPR cDNA sequence was identical to the 3' portion of the p62 cDNA containing heptad repeat sequences. However, the 5' region and the extreme 3' region of the NPR cDNA sequence were different from the p62 cDNA. Interestingly, the extreme 3' untranslated region (UTR) contained a 212-bp inverted repeat of a sequence located in the middle of the NPR cDNA that is identical to the p62 sequence. The inverted repeats of the NPR sequence could potentially hybridize, leading to the formation of circular transcripts. Using antibodies specific for the C-terminal regions of p62, a 26-kDa protein was detected from NPR cDNA hybrid-arrested translational products, and a 28-kDa protein was detected from the testis germ cell extracts but not from Sertoli cell extracts.

Amino Acid Sequence↗

Chlorobenzoate catabolic transposon Tn5271 is a composite class I element with flanking class II insertion sequences.

The structure of a transposon specifying the biodegradation of chlorobenzoate contaminants is described. Tn5271 is a 17-kilobase (kb) transposon that resides in the plasmid or chromosome of Alcaligenes sp. strain BR60 and allows this organism to grow on 3- and 4-chlorobenzoate. The transposon is flanked by a directly repeated sequence of 3201 base pairs (bp), which in turn is flanked by 110-bp inverted repeats. The 3.2-kb repeated sequence, designated IS1071, exists in multiple copies in the genome of Alcaligenes sp. strain BR60 and is involved in recombination of the catabolic genes into the chromosome of this strain. Sequence analysis revealed that the inverted repeat of IS1071 and the derived amino acid sequence of the single open reading frame within IS1071 are related to the inverted repeats and transposase (TnpA) proteins of the class II (Tn3 family) transposable elements. The absence of a resolvase gene within IS1071 suggests that this element is capable of determining the first step in class II transposition only. This was confirmed by observations on the IS1071-dependent formation of stable cointegrates in a recombination-deficient Escherichia coli. These results support an evolutionary scheme in which the class II transposable elements descended from simple insertion sequences.

Alcaligenes↗

A novel IS element, IS621, of the IS110/IS492 family transposes to a specific site in repetitive extragenic palindromic sequences in Escherichia coli.

An Escherichia coli strain, ECOR28, was found to have insertions of an identical sequence (1,279 bp in length) at 10 loci in its genome. This insertion sequence (named IS621) has one large open reading frame encoding a putative protein that is 326 amino acids in length. A computer-aided homology search using the DNA sequence as the query revealed that IS621 was homologous to the piv genes, encoding pilin gene invertase (PIV). A homology search using the amino acid sequence of the putative protein encoded by IS621 as the query revealed that the protein also has partial homology to transposases encoded by the IS110/IS492 family elements, which were known to have partial homology to PIV. This indicates that IS621 belongs to the IS110/IS492 family but is most closely related to the piv genes. In fact, a phylogenetic tree constructed on the basis of amino acid sequences of PIV proteins and transposases revealed that IS621 belongs to the piv gene group, which is distinct from the IS110/IS492 family elements, which form several groups. PIV proteins and transposases encoded by the IS110/IS492 family elements, including IS621, have four acidic amino acid residues, which are conserved at positions in their N-terminal regions. These residues may constitute a tetrad D-E(or D)-D-D motif as the catalytic center. Interestingly, IS621 was inserted at specific sites within repetitive extragenic palindromic (REP) sequences at 10 loci in the ECOR28 genome. IS621 may not recognize the entire REP sequence in transposition, but it recognizes a 15-bp sequence conserved in the REP sequences around the target site. There are several elements belonging to the IS110/IS492 family that also transpose to specific sites in the repeated sequences, as does IS621. IS621 does not have terminal inverted repeats like most of the IS110/IS492 family elements. The terminal sequences of IS621 have homology with the 26-bp inverted repeat sequences of pilin gene inversion sites that are recognized and used for inversion of pilin genes by PIV. This suggests that IS621 initiates transposition through recognition of their terminal regions and cleavage at the ends by a mechanism similar to that used for PIV to promote inversion at the pilin gene inversion sites.

Amino Acid Sequence↗

Inhibition of transpositional recombination by OrfA and OrfB proteins encoded by insertion sequence IS3.

BACKGROUND: An insertion element IS3 is flanked by terminal inverted repeat (IR) sequences. IS3 encodes two, out-of-phase, overlapping open reading frames, orfA and orfB, from which three proteins are produced. OrfAB is a transframe protein produced by -1 translational frameshifting between orfA and orfB, and it is known to be IS3 transposase. OrfA and OrfB are the proteins produced without frameshifting, but their functions have not been elucidated. RESULTS: A plasmid carrying an IS3 mutant that produces only transposase generates miniplasmids--which are the IS3-mediated intramolecular transposition products--as well as characteristic IS3 circles and linear IS3 molecules. OrfA inhibited the generation of these small molecules to a lesser degree, but OrfB did not. OrfB, together with OrfA, however, inhibited the generation more strongly than OrfA alone. OrfA also inhibited the intermolecular transposition of mini-IS3 with the chloramphenicol-resistance gene flanked by IRs to a reduced frequency, and OrfB together with OrfA inhibited it almost completely. OrfA and/or OrfB did not, however, repress transcription from the promoter in the left-terminal region preceding orfA. CONCLUSIONS: The results obtained above show that OrfA and OrfB are not repressors but are inhibitors of transpositional recombination promoted by transposase. OrfA with an alpha helix-turn-alpha helix DNA-binding motif may compete with transposase to bind to terminal IRs. OrfA, together with OrfB that has a DDE motif conserved in retroviral integrases, may inhibit the formation of an active transpososome consisting oftransposase, two terminal IRs and target DNA for the strand transfer reaction. IS3 with a limited size, 1258 bp in length, uses strategies of translational frameshifting and coupling to produce transposase as well as negative regulators to make its copies at a low level, which minimizes a deleterious effect of transposition on bacterial hosts.

Bacterial Proteins↗

Characterization of genomic poly(dT-dG).poly(dC-dA) sequences: structure, organization, and conformation.

Hybridization studies suggest the abundant presence of poly(dT-dG).poly(dC-dA) (TG-element), a potential Z-DNA sequence, in eucaryotic genomes. We have isolated and characterized TG-elements from different locations in the human genome: from randomly isolated clones, associated with the actin gene family, and linked to another repeated element. The results indicate that the following features are typical of these TG-elements: the elements consist of 20 to 60 base pairs of (dT-dG)n.(dC-dA)n, the sequences characterized in our study were not flanked by direct or inverted repeats, the sequences are interspersed rather than in satellite blocks, the elements are not usually associated with other repeated elements, and some of the elements are found near coding sequences or in introns. Studies on the conformation of a genomic TG-element in a supercoiled plasmid indicate several distinct properties of the TG-element: it is in the Z-form only at low ionic strength, S1 nuclease recognizes its Z-form with a marked preference for one of the B-Z junctions, and the sensitive region extends for 20 base pairs near the B-Z junction. In contrast to the result with the supercoiled plasmid, S1 nuclease failed to recognize the TG-element in minichromosomes.

Actins↗

Identification and distribution of new insertion sequences in the genome of the extremely halotolerant and alkaliphilic Oceanobacillus iheyensis HTE831.

Six kinds of new insertion sequences (ISs), IS667 to IS672, a group II intron (Oi.Int), and an incomplete transposon (Tn852loi) were identified in the 3,630,528-bp genome of the extremely halotolerant and alkaliphilic Oceanobacillus iheyensis HTE831. Of 19 ISs identified in the HTE831 genome, 7 were truncated, indicating the occurrence of internal rearrangement of the genome. All ISs except IS669 generated a 4- to 8-bp duplication of the target site sequence, and these ISs carried 23- to 28-bp inverted repeats (IRs). Sequence analysis revealed that four ISs (IS669, IS670, IS671, and IS672) were newly identified as belonging to separate IS families (IS200/IS605, IS30, IS5, and IS3, respectively). IS667 and IS668 were also characterized as new members of the ISL3 family. Tn8521oi, which belongs to the Tn3 family as a new member, generated a 5-bp duplication of the target site sequence and carried complete 38-bp IRs. Of the eight protein-coding sequences (CDSs) identified in Tn8521oi, three CDSs (OB481, OB482, and OB483) formed a ger gene cluster, and two other paralogous gene clusters were found in the HTE831 genome. Most of the ISs and the group II intron widely distributed throughout the genome were inserted in noncoding regions, while two ISs (IS667-08 and IS668-02) and Oi.Int-04 were inserted in the coding regions.

Amino Acid Sequence↗

The human chorionic somatomammotropin gene enhancer is composed of multiple DNA elements that are homologous to several SV40 enhansons.

Previous studies indicate that a human chorionic somatomammotropin (hCS) gene enhancer (CSEn) associated with the growth hormone (hGH) gene locus is involved in directing cell-specific expression of the hCS genes in placenta. In the current studies, we report a detailed structural analysis of this enhancer. CSEn stimulated transcription of a variety of promoters, including the hCS, human growth hormone, thymidine kinase, and Rous sarcoma virus promoters, in human choriocarcinoma cell lines (BeWo and JEG-3) but not HeLa cells or rat somatolactotrophes (GC). Maximal enhancer activity was confined to a 242-base pair DNA segment. Of several CSEn subfragments, only the En 57/242 subfragment retained activity (33.5% wild-type). The CSEn DNA sequence contained direct and inverted repeat motifs and sequences related to the SV40 enhansons, GT-IIC, GT-I, and SphI/SphII. DNase I footprint analysis revealed that most of these sites were protected by nuclear proteins derived from BeWo, JEG-3, HeLa, and GC cells. Site-specific block mutation of the GT-IIC-related and inverted repeat motifs virtually abolished enhancer activity, and mutation of all but the GT-I-related motif resulted in significant loss (30-60%) of activity. These data demonstrate that the CS enhancer is comprised of multiple elements related to SV40 enhansons that interact cooperatively to generate enhancer function.

Animals↗

Gene-sized DNA molecules of the Oxytricha macronucleus have the same terminal sequence.

The DNA in the macronucleus of the ciliated protozoan Oxytricha exists as small linear molecules with a number average size of about 3000 base pairs. Most, and possibly all, of these DNA molecules contain the same inverted terminal repeat sequence of approximately 26 base pairs. In addition to its terminal location, two inverted copies of this same sequence surround single-strand interruptions within these DNA molecules. This sequence arrangement may function in the processing of these molecules from large chromosomal precursors or in the subsequent replication of these small linear DNAs during cell reproduction.

Base Sequence↗

Regulation of the Dha operon of Lactococcus lactis: a deviation from the rule followed by the Tetr family of transcription regulators.

Dihydroxyacetone (Dha) kinases are a novel family of kinases with signaling and metabolic functions. Here we report the x-ray structures of the transcriptional activator DhaS and the coactivator DhaQ and characterize their function. DhaQ is a paralog of the Dha binding Dha kinase subunit; DhaS belongs to the family of TetR repressors although, unlike all known members of this family, it is a transcriptional activator. DhaQ and DhaS form a stable complex that in the presence of Dha activates transcription of the Lactococcus lactis dha operon. Dha covalently binds to DhaQ through a hemiaminal bond with a histidine and thereby induces a conformational change, which is propagated to the surface via a cantilever-like structure. DhaS binding protects an inverted repeat whose sequence is GGACACATN6ATTTGTCC and renders two GC base pairs of the operator DNA hypersensitive to DNase I cleavage. The proximal half-site of the inverted repeat partially overlaps with the predicted -35 consensus sequence of the dha promoter.

Amino Acid Sequence↗

Escherichia coli heat-labile enterotoxin genes are flanked by repeated deoxyribonucleic acid sequences.

The enterotoxin regions of the heat-labile and heat-stable enterotoxin (LT+ ST+) plasmid, pJY11, originating in a clinically isolated Escherichia coli strain, have been isolated as various-sized deoxyribonucleic acid (DNA) fragments by using cloning vehicles. The structure of the LT+ region and its neighboring DNA regions was studied by utilizing these recombinant plasmids. The LT+ region consisted of at least two genes, toxA and toxB, which could complement each other in trans. The toxA- and toxB-encoded polypeptides (LT subunits A and B, respectively) were identified by their immunological cross-reactivity with Vibrio cholerae enterotoxin subunit A or B. These tox genes and the promoter(s) were localized with respect to the restriction endonuclease cleavage map. The LT+ region was flanked by repeated DNA sequences (designated as beta). Another tox gen(s), encoding ST (designated as toxS), which was also flanked by inverted, repeated DNA sequences (designated as alpha), was located between one of the beta sequences and the LT+ region. These novel DNA structures (beta-alpha-toxS-alpha-toxA-toxB-beta) suggest the possibility that the LT+ region is on a transposon containing an ST transposon within the structure.

Bacterial Toxins↗

Evolutionary comparisons of the S segments in the genomes of herpes simplex virus type 1 and varicella-zoster virus.

The genomes of herpes simplex virus type 1 (HSV-1) and varicella-zoster virus (VZV) consist of two covalently joined segments, L and S. Each segment comprises an unique sequence flanked by inverted repeats. We have reported previously the DNA sequences of the S segments in these two genomes, and have identified protein-coding regions therein. In HSV-1, the unique sequence of S contains ten entire genes plus the major parts of two more, and each inverted repeat contains one entire gene; in VZV, the unique sequence of S contains two entire genes plus the major parts of two more, and each inverted repeat contains three entire genes. In this report, an examination of polypeptide sequence homology has shown that each VZV gene has an HSV-1 counterpart, but that six of the HSV-1 genes have no VZV homologues. Thus, although these regions of the two genomes differ in gene layout, they are related to a significant degree. The analysis indicates that the inverted repeats are evidently capable of large-scale expansion or contraction during evolution. The differences in gene layout can be understood as resulting from a small number of recombinational events during the descent of HSV-1 and VZV from a common ancestor.

Amino Acid Sequence↗

Macronuclear and micronuclear configurations of a gene encoding the protein synthesis elongation factor EF 1 alpha in Stylonychia lemnae.

The micronuclear and macronuclear configurations of a gene encoding the protein synthesis elongation factor EF 1 alpha in the hypotrich ciliate Stylonychia lemnae were compared. The two sequences are generally colinear. The coding sequence of the micronuclear gene is, however, interrupted by a 64 bp insert flanked by a 2 bp direct repeat in a gene region which is moderately conserved among EF 1 alpha genes of different organisms. The insertion site is distinct from known intron positions in eukaryotic EF 1 alpha genes. The insert sequence shows inverted repeats at its ends and thus exhibits typical features of an internal eliminated sequence (IES). Comparison with other such sequences in the related organism Oyxtricha nova shows that the IES falls into a new group of such elements. The macronuclear gene exhibits a strikingly limited codon usage, which cannot be simply explained by the overall base composition of the DNA but probably also relates to the very high copy number of the macronuclear gene and the putative high amount of the gene product.

Amino Acid Sequence↗