Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “tandem repeat”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Effect of nuclear protein HMG1 on in vitro slippage synthesis of the tandem repeat dTG x dCA.

Tandem repeats of simple doublet and triplet sequences occur with high frequency in the DNA of eucaryotes. Among the most frequent is the repeat of dTG, which has unusual structural properties. We show here that HMG1 (modeled by the second HMG box motif from HMG1 of the rat, HMGb) binds to complexes formed from annealing unequal lengths of dTG x dCA and inhibits the in vitro elongation of these complexes by the Klenow fragment of DNA polymerase I at 37 degrees C. At 46 degrees C, HMGb enhances the elongation. Polylysine inhibits elongation at both temperatures. These results show that the stability of this repeat in vivo can be influenced by the presence of basic proteins in general, and more selectively by the abundant nuclear protein HMG1.

Animals↗

A V3 loop haptenic peptide sequence, when tandemly repeated, enhances immunogenicity by facilitating helper T-cell responses to a covalently linked carrier protein.

Subunit immunogens containing tandemly repeated copies of T- and B-cell epitopes have been shown to be more immunogenic than the respective immunogen containing only a single copy of the sequence. It has been unclear, however, whether the increased immunogenicity of a tandemly repeated B-cell epitope necessarily results from increased helper T-cell responses to intrinsic T-cell epitopes in the tandemly repeated sequences, or to neodeterminant T-cell epitopes created at the junction of adjacent repeat sequences. We examined this question by comparing the immunogenicity in mice of two immunogens containing one or eight tandemly repeated copies of an HIV-1 V3 loop haptenic sequence. Our results show that the tandemly repeated haptenic sequence potentiates the immunogenicity of the protein construct, likely through the facilitation of enhanced B-cell interaction with the tandem repeat construct and the consequent elicitation of increased carrier protein-specific helper T-cell responses.

Amino Acid Sequence↗

Centromere 3 specific tandem repeat from Chironomus pallidivittatus.

A 155-bp tandem repeat was previously reported to be present in all centromeric regions of the dipteran Chironomus pallidivittatus. We have now isolated a second centromere specific tandem repeat, 375 bp long. Two blocks were found of the new unit, differing in size, probably representing allelic forms. The repeat is present only in chromosome 3, bordering 155-bp repeat arrays. There are about 100 repeats per genome, compared to 1300 units for the 155-bp repeat. The two units contain an identical 9-bp sequence which can form target-site duplications flanking a short mobile element, Cp1. An inversion within the tandem array was isolated, the breakpoint of which is within the 9-bp target sequence. Another short shared motif, 10-bp long, is also present at the insertion site for a mobile element. The two repeat units are similar in having long regions with more than 80% AT and an overall high AT content.

Animals↗

Biophysical characterization of one-, two-, and three-tandem repeats of human mucin (muc-1) protein core.

Until recently mucin tandem repeat protein cores were believed to exist in random-coil conformations and to attain structure solely by the addition of carbohydrates to serine and threonine residues. Matsushima et al. (Proteins Struct. Funct. Genet., 7: 125-155, 1990) recently proposed a model of the secondary structure of proline rich tandem repeat proteins that has challenged this idea, especially for the case of the human polymorphic epithelial mucin encoded by the muc-1 gene. We report here results of structural analyses of the muc-1 protein core by using synthetic peptide analogues. Synthetic peptides were prepared to correspond to one-, two-, and three-tandem repeats of muc-1. Results of one- and two-dimensional 1H NMR correlation spectroscopy on these peptides confirm that the muc-1 protein core is not a random-coil secondary structure. Long-lived amide protons are protected in D2O, and increasing spectral complexity in the region of the beta-protons of Asp2 and His 15 reveals that structural changes are occurring as the number of repeats increases. The greatest changes occur when the number of repeats increases from one to two. These results are supported by the reactivity of a panel of monoclonal antibodies raised against tumor associated muc-1 with these synthetic peptides in enzyme-linked immunosorbent assay. The primary immunodominant mucin epitope, PDTRP, does not appear to attain a native conformation in the single repeat peptide (20 amino acids, starting with P), but is expressed on peptides with multiple repeats. Intrinsic viscosity measurements of the peptide containing three repeats indicate that an ordered structure present in solution is rod shaped. The circular dichroism spectrum of the same peptide is dominated by proline in the trans conformation. These results are all consistent with the prediction that the muc-1 tandem repeat polypeptide core forms a polyproline beta-turn helix.

Amino Acid Sequence↗

A tandem repeats database for bacterial genomes: application to the genotyping of Yersinia pestis and Bacillus anthracis.

BACKGROUND: Some pathogenic bacteria are genetically very homogeneous, making strain discrimination difficult. In the last few years, tandem repeats have been increasingly recognized as markers of choice for genotyping a number of pathogens. The rapid evolution of these structures appears to contribute to the phenotypic flexibility of pathogens. The availability of whole-genome sequences has opened the way to the systematic evaluation of tandem repeats diversity and application to epidemiological studies. RESULTS: This report presents a database (http://minisatellites.u-psud.fr) of tandem repeats from publicly available bacterial genomes which facilitates the identification and selection of tandem repeats. We illustrate the use of this database by the characterization of minisatellites from two important human pathogens, Yersinia pestis and Bacillus anthracis. In order to avoid simple sequence contingency loci which may be of limited value as epidemiological markers, and to provide genotyping tools amenable to ordinary agarose gel electrophoresis, only tandem repeats with repeat units at least 9 bp long were evaluated. Yersinia pestis contains 64 such minisatellites in which the unit is repeated at least 7 times. An additional collection of 12 loci with at least 6 units, and a high internal conservation were also evaluated. Forty-nine are polymorphic among five Yersinia strains (twenty-five among three Y. pestis strains). Bacillus anthracis contains 30 comparable structures in which the unit is repeated at least 10 times. Half of these tandem repeats show polymorphism among the strains tested. CONCLUSIONS: Analysis of the currently available bacterial genome sequences classifies Bacillus anthracis and Yersinia pestis as having an average (approximately 30 per Mb) density of tandem repeat arrays longer than 100 bp when compared to the other bacterial genomes analysed to date. In both cases, testing a fraction of these sequences for polymorphism was sufficient to quickly develop a set of more than fifteen informative markers, some of which show a very high degree of polymorphism. In one instance, the polymorphism information content index reaches 0.82 with allele length covering a wide size range (600-1950 bp), and nine alleles resolved in the small number of independent Bacillus anthracis strains typed here.

Bacillus anthracis↗

STRING: finding tandem repeats in DNA sequences.

MOTIVATION AND RESULTS: The importance of Tandem Repeats in some genomes is now well established. We have reported elsewhere some interesting new results obtained by means of a preliminary program for finding Tandem Repeats in DNA sequences, together with a brief description of the basic ideas of the algorithm. We describe here a completely new program based only in part on those ideas, we briefly discuss the interpretation of the results, and, by way of example, we provide a few novel results relative to the parasites responsible of two re-emerging diseases, Plasmodium falciparum and Mycobacterium tuberculosis. Our program is portable, effective, powerful and fast: it can run on current desktop computers, and it finds all significant Tandem Repeats also in the longest segments of sequences in databases (up to millions of bases), in short times (minutes). AVAILABILITY: An academic version of the algorithm (full source listing in standard C language) can be freely downloaded (http://www.caspur.it/~castri/STRING/). SUPPLEMENTARY INFORMATION: Some illustrative figures and some sample results are provided as supplementary material at: http://www.caspur.it/~castri/STRING/

Algorithms↗

Long stretches of short tandem repeats are present in the largest replicons of the Archaea Haloferax mediterranei and Haloferax volcanii and could be involved in replicon partitioning.

We report the presence of long stretches of tandem repeats in the genome of the halophilic Archaea Haloferax mediterranei and Haloferax volcanii. A 30 bp sequence with dyad symmetry (including 5 bp inverted repeats) was repeated in tandem, interspersed with 33-39 bp unique sequences. This structure extends for long stretches--1.4 kb at one location in H. mediterranei chromosome and about 3 kb in the H. volcanii chromosome. The tandem repeats (designated TREPs) show a similar distribution in both organisms, appearing once or twice in the H. volcanii and H. mediterranei chromosomes, and once in the largest, probably essential megaplasmid of each organism but not in the smaller replicons. Sequencing of the structures in both H. volcanii replicons revealed an extremely high sequence conservation in both replicons within the species, as well as in the different organisms. Homologous sequences have also been found in other more distantly related halophilic members of the Archaea. Transformation of H. volcanii with a recombinant plasmid containing a 1.1 kb fragment of the TREPs produced significant alterations in the host cells, particularly in terms of cell viability. The introduction of extra copies of TREPs within the vector significantly alters the distribution of the genome among the daughter cells, as observed by DAPI staining. Although the precise biological role cannot be completely ascertained, all the data conform with the tandem repeats being involved in replicon partitioning in halobacteria.

Base Sequence↗

Stabilization of perfect and imperfect tandem repeats by single-strand DNA exonucleases.

Rearrangements between tandemly repeated DNA sequences are a common source of genetic instability. Such rearrangements underlie several human genetic diseases. In many organisms, the mismatch-repair (MMR) system functions to stabilize repeats when the repeat unit is short or when sequence imperfections are present between the repeats. We show here that the action of single-stranded DNA (ssDNA) exonucleases plays an additional, important role in stabilizing tandem repeats, independent of their role in MMR. For perfect repeats of approximately 100 bp in Escherichia coli that are not susceptible to MMR, exonuclease (Exo)-I, ExoX, and RecJ exonuclease redundantly inhibit deletion. Our data suggest that >90% of potential deletion events are avoided by the combined action of these three exonucleases. Imperfect tandem repeats, less prone to rearrangements, are stabilized by both the MMR-pathway and ssDNA-specific exonucleases. For 100-bp repeats containing four mispairs, ExoI alone aborts most deletion events, even in the presence of a functional MMR system. By genetic analysis, we show that the inhibitory effect of ssDNA exonucleases on deletion formation is independent of the MutS and UvrD proteins. Exonuclease degradation of DNA displaced during the deletion process may abort slipped misalignment. Exonuclease action is therefore a significant force in genetic stabilization of many forms of repetitive DNA.

Base Pair Mismatch↗

Tandemly repeated DNA: why should anyone care?

Recent excitement over SNPs has tended to obscure the real advantages of studying tandemly repeated loci. In this commentary, I make the case for studying tandem repeats, concentrating on two major arguments. Firstly, tandemly repeated loci are unrivalled as a source of detailed mechanistic information in studies of variation and mutation, and are highly informative reporters of genomic instability in studies of induced mutation. Secondly, changes at many tandem repeats have important functional consequences, and in addition to examples of "strong" single-gene effects such as those at the triplet repeat disease loci, there may well be a much larger number of loci at which subtler functional effects remain to be discovered.

Alu Elements↗

Cloning and characterization of a core histone gene tandem repeat in Urechis caupo.

A Urechis caupo histone gene tandem repeat has been isolated from a 5.0-kilobase EcoRI genomic library in lambda gtWES.lambda B. Genomic reconstruction experiments indicate that the cloned sequence is repeated approximately 100 times per haploid genome. Unique restriction fragments from the cloned sequence hybridize with individual core histone genes from a histone gene tandem repeat of the sea urchin, Strongylocentrotus purpuratus. No hybridization is detected when restriction digests are probed with a sea urchin H1 histone gene. Hybrid selection and in vitro translation of embryo mRNAs demonstrate that the clone contains sequences complementary to all four core histones; however, no H1 histone is detected among the translation products. Based on a restriction site map of the clone and the subcloned sequences which hybridize to the histone mRNAs, the order of the core histone genes in the clone is shown to be H3 H2A H2B H4. S1 nuclease hybrid protection mapping is used to locate the coding regions and to determine the transcript lengths of the core histone mRNAs. The transcript lengths of H2A, H2B, H3, and H4 mRNAs are approximately 464, 438, 494, and 397 bases, respectively. The S1 nuclease mapping also demonstrates that H2A and H4 are transcribed from one DNA strand while H2B and H3 are transcribed from the other strand. In the tandem repeat, the genes are organized so that transcription of the H2A-H2B and H3-H4 gene pairs is divergent.

Animals↗

A new look at the challenging world of tandem repeats.

Recent research has shown a correlation between some genetic diseases and genomic sequences tandemly repeated a variable and excessive number of times. The excessive number of tandem repeats is usually caused by a progressive expansion, generally considered as purely harmful. We put forward a number of hypotheses: the main one is that the number of repeats has normally a specific significance, and that there exist purposive mechanisms having as a primary function the management of tandem repeats length; such a function is generally useful and only rarely may it become harmful, because of some malfunctioning. These hypotheses are suggested by plausibility arguments, and are supported by a number of recent experimental results. They could provide a simple and unifying explanation of many pathological and non-pathological phenomena replacing many ad hoc assumptions. We finally propose to call the study of the above tandem repeat managing mechanisms 'dynamical genetics'.

Aging↗

Detection of sequence variability of the collagen type IIalpha 1 3' variable number of tandem repeat.

The variable number of tandem repeat (VNTR) 3' of the collagen type II (COL2A1) gene has been shown to be highly variable with a complex molecular structure. In a previous pilot experiment we observed discordance between methods to genotype this informative marker. To further investigate the extent and molecular nature of this discordance, we genotyped a random sample of 207 Caucasian individuals with two genotyping methods and sequenced new alleles. We compared single-strand (SS) analysis, which is based on detection of size differences between the different alleles, and heteroduplex analysis (HA), which is sensitive to both size and sequence differences. Overall, 26% of discordance between the two methods was detected. Approximately two thirds of this discordance was caused by subdivision of SS-alleles 13R1 and 14R2 into HA-alleles 4A + 4B and 3B + 3C, respectively. Sequence analysis of the COL2A1 VNTR alleles 4B and 3C showed that these alleles differed in sequence, but not in size, from already described SS-alleles, which explains why they escape detection by SS. The 4B allele is a frequent allele in the population (14%) and is, therefore, important to distinguish in association studies. We conclude that HA is a reliable method when the described optimized electrophoretic conditions are used. HA is a sensitive genotyping method to document allelic diversity at this locus, which can distinguish more alleles compared to the SS method.

Aged↗

Least-square deconvolution: a framework for interpreting short tandem repeat mixtures.

Interpreting mixture short tandem repeat DNA data is often a laborious process, involving trying different genotype combinations mixed at assumed DNA mass proportions, and assessing whether the resultant is supported well by the relative peak-height information of the mixture sample. If a clear pattern of major-minor alleles is apparent, it is feasible to identify the major alleles of each locus and form a composite genotype profile for the major contributor. When alleles are shared between the two contributors, and/or heterozygous peak imbalance is present, it becomes complex and difficult to deduce the profile of the minor contributor. The manual trial and error procedures performed by an analyst in the attempt to resolve mixture samples have been formalized in the least-square deconvolution (LSD) framework reported here for two-person mixtures, with the allele peak height (or area) information as its only input. LSD operates on the peak-data information of each locus separately, independent of all other loci, and finds the best-fit DNA mass proportions and calculates error residual for each possible genotype combination. The LSD mathematical result for all loci is then to be reviewed by a DNA analyst, who will apply a set of heuristic interpretation guidelines in an attempt to form a composite DNA profile for each of the two contributors. Both simulated and forensic peak-height data were used to support this approach. A set of heuristic guidelines is to be used in forming a composite profile for each of the mixture contributors in analyzing the mathematical results of LSD. The heuristic rules involve the checking of consistency of the best-fit mass proportion ratios for the top-ranked genotype combination case among all four- and three-allele loci, and involve assessing the degree of fit of the top-ranked case relative to the fit of the second-ranked case. A different set of guidelines is used in reviewing and analyzing the LSD mathematical results for two-allele loci. Resolution of two-allele loci is performed with less confidence than for four- and three-allele loci. This paper gives a detailed description of the theory of the LSD methodology, discusses its limitations, and the heuristic guidelines in analyzing the LSD mathematical results. A 13-loci sample case study is included. The use of the interpretation guidelines in forming composite profiles for each of the two contributors is illustrated. Application of LSD in this case produced correct resolutions at all loci. Information on obtaining access to the LSD software is also given in the paper.

Algorithms↗

Tandem repeat deletion in the alpha C protein of group B streptococcus is recA independent.

Group B streptococci (GBS) contain a family of protective surface proteins characterized by variable numbers of repeating units within the proteins. The prototype alpha C protein of GBS from the type Ia/C strain A909 contains a series of nine identical 246-bp tandem repeat units. We have previously shown that deletions in the tandem repeat region of the alpha C protein affect both the immunogenicity and protective efficacy of the protein in animal models, and these deletions may serve as a virulence mechanism in GBS. The molecular mechanism of tandem repeat deletion is unknown. To determine whether RecA-mediated homologous recombination is involved in this process, we identified, cloned, and sequenced the recA gene homologue from GBS. A strain of GBS with recA deleted, A909DeltarecA, was constructed by insertional inactivation in the recA locus. A909DeltarecA demonstrated significant sensitivity to UV light, and the 50% lethal dose of the mutant strain in a mouse intraperitoneal model of sepsis was 20-fold higher than that of the parent strain. The spontaneous rate of tandem repeat deletion in the alpha C protein in vitro, as well as in our mouse model of immune infection, was studied using A909DeltarecA. We report that tandem repeat deletion in the alpha C protein does occur in the absence of a functional recA gene both in vitro and in vivo, indicating that tandem repeat deletion in GBS occurs by a recA-independent recombinatorial pathway.

Animals↗

High density O-glycosylation of the MUC2 tandem repeat unit by N-acetylgalactosaminyltransferase-3 in colonic adenocarcinoma extracts.

A synthetic peptide corresponding to the human MUC2 tandem repeat unit was glycosylated in vitro using UDP-GalNAc and extracts of colonic adenocarcinoma and paired normal mucosa, followed by fractionation of the products by reverse phase high-performance liquid chromatography. Several peaks of glycopeptides with different numbers of GalNAc residues attached were detected. It is notable that the adenocarcinoma extract was capable of glycosylating peptides to a much greater extent than was normal mucosa. The levels of mRNA for N-acetylgalactosaminyltransferases-1, -2, and -3 were determined by reverse transcription-PCR. Only N-acetylgalactosaminyltransferase-3 mRNA was expressed at a higher level in the adenocarcinoma than in the normal tissue. When the MUC2 tandem repeat peptide was glycosylated with a mixture of the normal mucosa extract and recombinant N-acetylgalactosaminyltransferase-3, larger amounts of glycopeptides with higher contents of GalNAc residues were produced. The MUC2 tandem repeat peptides glycosylated extensively by recombinant N-acetylgalactosaminyltransferase-1, -2, or -3 were prepared and characterized. Substitution at each Thr residue, as revealed by Edman degradation sequencing, in conjunction with evidence obtained on mass spectrometry indicated a heterogeneous pattern of site-specific glycosylation within the MUC2 tandem repeat. It was found that maximum numbers of 6, 8, and 11 GalNAc residues were incorporated by N-acetylgalactosaminyltransferases-1, -2, and -3, respectively, and that only N-acetylgalactosaminyltransferase-3 could completely glycosylate both consecutive sequences composed of three and five Thr residues in the MUC2 tandem repeat unit. These results suggest that O-glycosylation of the clustered Thr residues is a selective process controlled by N-acetylgalactosaminyltransferase-3 in the synthesis of clustered carbohydrate antigens.

Acetylgalactosamine↗

mreps: Efficient and flexible detection of tandem repeats in DNA.

The presence of repeated sequences is a fundamental feature of genomes. Tandemly repeated DNA appears in both eukaryotic and prokaryotic genomes, it is associated with various regulatory mechanisms and plays an important role in genomic fingerprinting. In this paper, we describe mreps, a powerful software tool for a fast identification of tandemly repeated structures in DNA sequences. mreps is able to identify all types of tandem repeats within a single run on a whole genomic sequence. It has a resolution parameter that allows the program to identify 'fuzzy' repeats. We introduce main algorithmic solutions behind mreps, describe its usage, give some execution time benchmarks and present several case studies to illustrate its capabilities. The mreps web interface is accessible through http://www.loria.fr/mreps/.

Algorithms↗

Variable numbers of simple tandem repeats make birds of the order ciconiiformes heteroplasmic in their mitochondrial genomes.

We have analyzed a variable domain of the mitochondrial DNA control region of 18 avian species. Intra-individual length variation was identified and characterized in 15 species. The occurrence of heteroplasmy among species is phylogenetically consistent with a current classification of birds. Polymerase chain reaction amplifications, direct sequencing, and Southern analysis of mitochondrial DNA showed that the heteroplasmy is due to variable numbers of direct repeats in a tandem organization, located in the control region close to the tRNAPhe gene. The tandem repeats consist of short sequence motifs that vary in size from 4 to 32 base pairs between species. Sequence complexity of the repeat motifs was low, with almost exclusively Ts and Gs in the heavy-strand. Extensive variation in the copy number of the repeats was seen both intra-specifically and within individuals. This is the first report of mitochondrial heteroplasmy characterized at the sequence level in birds.

Animals↗

Molecular characterization and distribution of a 145-bp tandem repeat family in the genus Populus.

This report aims to describe the identification and molecular characterization of a 145-bp tandem repeat family that accounts for nearly 1.5% of the Populus genome. Three members of this repeat family were cloned and sequenced from Populus deltoides and P. ciliata. The dimers of the repeat were sequenced in order to confirm the head-to-tail organization of the repeat. Hybridization-based analysis using the 145-bp tandem repeat as a probe on genomic DNA gave rise to ladder patterns which were identified to be a result of methylation and (or) sequence heterogeneity. Analysis of the methylation pattern of the repeat family using methylation-sensitive isoschizomers revealed variable methylation of the C residues and lack of methylation of the A residues. Sequence comparisons between the monomers revealed a high degree of sequence divergence that ranged between 6% and 11% in P. deltoides and between 4.2% and 8.3% in P. ciliata. This indicated the presence of sub-families within the 145-bp tandem family of repeats. Divergence was mainly due to the accumulation of point mutations and was concentrated in the central region of the repeat. The 145-bp tandem repeat family did not show significant homology to known tandem repeats from plants. A short stretch of 36 bp was found to show homology of 66.7% to a centromeric repeat from Chironomus plumosus. Dot-blot analysis and Southern hybridization data revealed the presence of the repeat family in 13 of the 14 Populus species examined. The absence of the 145-bp repeat from P. euphratica suggested that this species is relatively distant from other members of the genus, which correlates with taxonomic classifications. The widespread occurrence of the tandem family in the genus indicated that this family may be of ancient origin.

Base Sequence↗