Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “tandem repeat”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Haplotype studies support slippage as the mechanism of germline mutations in short tandem repeats.

Germline mutations of human short tandem repeat (STR) loci are expansions or contractions of repeat arrays which are not well understood in terms of the mechanism(s) underlying such mutations. Although polymerase slippage is generally accepted as a mechanism capable to explain most features of such mutations, it is still possible that unequal crossing over plays some role in those events, as most studies in humans could not exclude unequal crossing over (UCO). Crossing over can be studied by analyzing haplotypes using flanking markers. To check for UCO in mutations, we have analyzed 150 paternity cases for which more than the usual trio (mother, child, and father) were available for testing by analyzing 16 STR loci. In a total of 4900 parent-child allele transfers four mutations were observed at different loci (D8S1179, D18S51, D21S11, and SE33/ACTBP2). To identify the mutated allele and to check for UCO, we typed at least four informative loci flanking the mutated locus and used the pedigree data to establish haplotypes. By doing so we were able to exclude UCO in each case. Moreover, we were able to identify the mutations as one-repeat contractions/expansions. Our data thus support slippage as the mechanism of germline mutations in STRs.

Crossing Over, Genetic↗

A polymorphic X-linked tetranucleotide repeat locus displaying a high rate of new mutation: implications for mechanisms of mutation at short tandem repeat loci.

We report a high rate of new mutation at a short tandem repeat sequence polymorphism (STR, microsatellite) at locus DXS981 on the proximal long arm of the human X chromosome. Among individuals of the CEPH pedigrees, new allele lengths are detected at this tetranucleotide repeat with a frequency of approximately 1.5%. In cases where the origin of the new allele was traceable, new mutant alleles at DXS981 varied by exactly one repeat length (4 bp) relative to that on the originating parental chromosome. Complete linkage disequilibrium between two additional insertion/deletion polymorphisms which closely flank the variation at the tetranucleotide repeat suggests that, to the extent that these new mutants are germline in origin, they are not generated by unequal exchange between homologues. Considered in light of the types of new mutations detected and the substantial linkage disequilibrium at this locus, these data have implications for the mechanism of variation at other loci containing short tandemly repeated sequences.

Alleles↗

D18S535, D1S1656 and D10S2325: three efficient short tandem repeats for forensic genetics.

Three short tandem repeat (STR) polymorphisms characterized by PCR product length < 175 bp were investigated. D18S535 and D1S1656 contained a 4 bp unit as basic repeat motif, D10S2325 a 5 bp unit. The heterozygosity rates were 0.76 (D18S535), 0.88 (D10S2325) and 0. 90 (D1S1656), leading to a combined discrimination power of 0.9999. In contrast to D10S2325 and D18S535, which showed a homogeneous repeat array without any variation in the repeat motifs, repeat length and sequence variation was found for D1S1656. Robust typing results could be observed for all three STRs using highly degraded DNA.

Alleles↗

Analysis of the renin gene intron A tandem repeat region of Milan and Lyon hypertensive rat strains.

The region of intron A of the rat renin gene containing a unique tandemly repeated sequence was analysed in the Milan and Lyon hypertensive rat strains and their controls, and in several Sprague-Dawley rats, using an oligonucleotide probe complementary to the tandemly repeated sequence and a renin complementary DNA probe. In the Milan rats, the size of the Bgl II DNA fragment encompassing the tandem repeat region was the same in the hypertensive (MHS) and normotensive (MNS) strains. In the Lyon model, a difference of 1.1 kilobase (equivalent to about 28 copies of the 38 basepair tandem repeat sequence) was observed in the size of the Bgl II fragment of the hypertensive (LH) and normotensive (LN) strains. However, the finding that the size of the fragment in the Lyon low-blood-pressure (LL) strain was the same as that in the LH strain rather than the LN strain suggests that the difference between the two latter strains is not by itself a major cause of the blood pressure difference between them in the intron A tandem region. An analysis of Sprague-Dawley rats, from which the Lyon strains are derived, showed that at least three different renin gene alleles, two with Bgl II fragments of the same size as those seen in the Lyon strains, are randomly segregating in this population.

Alleles↗

Tandemly repeated sequences in the mitochondrial DNA control region and phylogeography of the Pike-Perches Stizostedion.

DNA sequences from the mitochondrial DNA control region are used to test the phylogeographic relationships among the pike-perches, Stizostedion (Teleostei: Percidae) and to examine patterns of variation. Sequences reveal two types of variability: single nucleotide polymorphisms and 6 to 14 copies of 10- to 11-base-pair tandemly repeated sequences. Numbers of copies of the tandem repeats are found to evolve too rapidly to detect phylogenetic signal at any taxonomic level, even among populations. Sequence similarities of the tandem repeats among Stizostedion and other percids suggest concerted evolutionary processes. Predicted folding of the tandem repeats and their proximity to termination-associated sequences indicate that secondary structure mediates slipped-strand mispairing among the d-loop, heavy, and light strands. Neighbor-joining and maximum parsimony analyses of sequences indicate that the genus is divided into clades on the continents of North America and Eurasia. Calibrating genetic distances with divergence times supports the hypothesis that Stizostedion dispersed from Eurasia to North America across a North Pacific Beringial land bridge approximately 4 million years before present, near the beginning of the Pliocene Epoch. The North American S. vitreum and S. canadense appear separated by about 2.75 million years, and the Eurasian S. lucioperca and S. volgensis are diverged by about 1.8 million years, suggesting that speciation occurred during the late Pliocene Epoch.

Animals↗

Short tandem repeat analysis in Japanese population.

Short tandem repeats (STRs), known as microsatellites, are one of the most informative genetic markers for characterizing biological materials. Because of the relatively small size of STR alleles (generally 100-350 nucleotides), amplification by polymerase chain reaction (PCR) is relatively easy, affording a high sensitivity of detection. In addition, STR loci can be amplified simultaneously in a multiplex PCR. Thus, substantial information can be obtained in a single analysis with the benefits of using less template DNA, reducing labor, and reducing the contamination. We investigated 14 STR loci in a Japanese population living in Sendai by three multiplex PCR kits, GenePrint PowerPlex 1.1 and 2.2. Fluorescent STR System (Promega, Madison, WI, USA) and AmpF/STR Profiler (Perkin-Elmer, Norwalk, CT, USA). Genomic DNA was extracted using sodium dodecyl sulfate (SDS) proteinase K or Chelex 100 treatment followed by the phenol/chloroform extraction. PCR was performed according to the manufacturer's protocols. Electrophoresis was carried out on an ABI 377 sequencer and the alleles were determined by GeneScan 2.0.2 software (Perkin-Elmer). In 14 STRs loci, statistical parameters indicated a relatively high rate, and no significant deviation from Hardy-Weinberg equilibrium was detected. We apply this STR system to paternity testing and forensic casework, e.g., personal identification in rape cases. This system is an effective tool in the forensic sciences to obtain information on individual identification.

Alleles↗

Mutation patterns of amino acid tandem repeats in the human proteome.

BACKGROUND: Amino acid tandem repeats are found in nearly one-fifth of human proteins. Abnormal expansion of these regions is associated with several human disorders. To gain further insight into the mutational mechanisms that operate in this type of sequence, we have analyzed a large number of mutation variants derived from human expressed sequence tags (ESTs). RESULTS: We identified 137 polymorphic variants in 115 different amino acid tandem repeats. Of these, 77 contained amino acid substitutions and 60 contained gaps (expansions or contractions of the repeat unit). The analysis showed that at least about 21% of the repeats might be polymorphic in humans. We compared the mutations found in different types of amino acid repeats and in adjacent regions. Overall, repeats showed a five-fold increase in the number of gap mutations compared to adjacent regions, reflecting the action of slippage within the repetitive structures. Gap and substitution mutations were very differently distributed between different amino acid repeat types. Among repeats containing gap variants we identified several disease and candidate disease genes. CONCLUSION: This is the first report at a genome-wide scale of the types of mutations occurring in the amino acid repeat component of the human proteome. We show that the mutational dynamics of different amino acid repeat types are very diverse. We provide a list of loci with highly variable repeat structures, some of which may be potentially involved in disease.

Amino Acid Substitution↗

A variable number of tandem repeats locus within the human complement C2 gene is associated with a retroposon derived from a human endogenous retrovirus.

We have previously described multiallelic restriction fragment length polymorphisms of the C2 gene, suggesting the presence of a variable number of tandem repeats (VNTR) locus. We report here the cloning and sequencing of the polymorphic fragments from the two most common alleles of the gene, a and b. The results confirm the presence of a VNTR locus consisting of a nucleotide sequence, 41 bp in average length, repeated tandemly 23 and 17 times in alleles a and b, respectively. The difference in the number of repeats between the two alleles is due to the deletion/insertion of two noncontiguous segments, 143 and 118 bp long, of allele a, and of a 40-bp segment of allele b. The VNTR region is associated with a SINE (short interspersed sequence)-type retroposon, SINE-R.C2, located within the third intron of the C2 gene. SINE-R.C2 is a member of a previously described large retroposon family of the human genome, apparently derived from the human endogenous retrovirus, (HERV) K10, which is homologous to the mouse mammary tumor virus.

Alleles↗

Highly discriminating heptaplex short tandem repeat PCR system for forensic identification.

We describe a highly discriminating multiplex short tandem repeat PCR human identification system that gives a matching probability for Caucasians of European ancestry of 2.94 x 10(-8) or 5.66 x 10(-10) when used in combination with a previously described system. The system produces discrimination equal to or greater than four single locus probes (restriction fragment length polymorphism [RFLP] typing of variable nucleotide tandem repeat [VNTR] loci). The test is robust and reproducible and works with 1-10 ng of template DNA, using fluorescent detection of PCR products from either 4 or 6 short tandem repeat loci and the X-Y homologous gene amelogenin, giving simultaneous sex diagnosis.

Alleles↗

AniAnn's: alignment-free annotation of tandem repeat arrays using fast average nucleotide identity estimates.

MOTIVATION: Satellite DNA has long posed challenges for genome assembly and analysis due to its low sequence complexity and poor mappability. These large heterochromatic arrays of tandem repeats are ubiquitous across eukaryotic genomes, yet remain understudied. Current methods for annotating satellite regions, and other classes of tandem repeat arrays, are limited in their ability to annotate divergent or novel sequences. RESULTS: In this work, we introduce AniAnn's, an algorithm for annotating large blocks of tandemly repeating DNAs. AniAnn's exploits the high Average Nucleotide Identity (ANI) shared between repeat units of the same array to quickly and accurately infer the boundaries of such arrays. We show that AniAnn's improves the annotation of satellites and other tandem repeats within a variety of plant and animal genomes, while requiring only a fraction of the runtime compared to previous approaches. We conclude by exploring several use cases of AniAnn's as a lightweight method for masking repeats prior to whole-genome alignment as well as the de novo annotation and classification of satellite repeats. AVAILABILITY: AniAnn's is open source software and available at github.com/marbl/anianns.

Algorithms↗

Influence of acceptor substrate primary amino acid sequence on the activity of human UDP-N-acetylgalactosamine:polypeptide N-acetylgalactosaminyltransferase. Studies with the MUC1 tandem repeat.

Synthetic peptides (30 and 20 residues long) corresponding to the native MUC1 tandem repeat sequence (20 residues long) were glycosylated in vitro using UDP-[3H]GalNAc and lysates from the human breast tumor cell line MCF7. Purified glycopeptides were sequenced on a gas-phase sequenator, and glycosylated positions were determined by measuring the incorporated radioactivity in fractions collected following each round of Edman degradation. The results showed that 2 of 3 threonines on the MUC1 tandem repeat peptides were glycosylated at the following positions: GVTSAPDTRPAPGSTAPPAH (underlined Thr residues indicate positions of GalNAc attachment); no glycosylation of serine residues was detected. Determination of the mass of the glycopeptides by mass spectrometry showed that a maximum of two molecules of GalNAc were covalently linked to each 20-residue repeat unit in the peptides. The influence of substrate primary amino acid sequence in determining the substrate specificity of UDP-N-acetylgalactosamine:polypeptide N-acetylgalactosaminyl-transferase activity was evaluated using as acceptor substrates a series of overlapping 9-residue peptides that represent a moving set through the tandem repeat of the MUC1 mucin. In addition, the influence of primary amino acid sequence on acceptor substrate activity was evaluated using several peptides that contained single or double amino acid substitutions (relative to the native human MUC1 sequence). These included substitutions in the residues that were glycosylated and substitutions in the surrounding primary amino acid sequence. This study demonstrates that primary amino acid sequence, length, and relative position of the residue to be glycosylated dramatically affect the ability of peptides to serve as acceptor substrates for UDP-N-acetylgalactosamine:polypeptide N-acetylgalactosaminyltransferase.

Amino Acid Sequence↗

Identification and functional analysis of single nucleotide polymorphism in the tandem repeat sequence of thymidylate synthase gene.

The variable number of tandem repeat (VNTR) of thymidylate synthase (TS) gene, mainly 2 repeat (2R) and 3 repeat (3R), is one of the genetic variations that can potentially predict the effectiveness of 5-fluorouracil-based chemotherapy. In this study we identified an additional single nucleotide polymorphism (SNP) in the VNTR of TS, followed by functional and clinical analysis of the SNP. Two-hundred fifty eight tumor samples were obtained from patients with primary colorectal adenocarcinoma. We observed three different patterns of electrophoresis by analysis of the VNTR with 2R/3R heterozygote. The sequencing results revealed a SNP, G/C polymorphism, within the 28-bp repeat component of TS VNTR. Each polymorphic allele was assigned as 2G, 2C, 3G, or 3C according to the combination of SNP and VNTR. Functional analysis showed that the plasmid construct with 3G sequence had three to four times greater efficiency of translation than other polymorphic sequences. 3R allele in colorectal cancer was subdivided into around half by the SNP, indicating its commonness among Japanese. TS genotypes of the patients with colorectal cancer were classified into high expression type (2R/3G, 3C/3G, and 3G/3G) and low expression type (2R/2R, 2R/3C, and 3C/3C). The patients who received oral fluoropyrimedines survived longer than the patients with no treatment in the group of low expression type. No benefit of oral fluoropyrimedines was observed in the group of high expression type. These results suggest that the double polymorphism in the TS tandem repeat sequence, the SNP and the VNTR, may provide a potential for more effective prediction of the clinical outcome of 5-fluorouracil-based chemotherapy.

Adenocarcinoma↗

The largest variant of platelet glycoprotein Ib alpha has four tandem repeats of 13 amino acids in the macroglycopeptide region and a genetic linkage with methionine145.

Platelet membrane glycoprotein Ib alpha (GPIb alpha) bears the human platelet alloantigen (HPA)-2 and molecular weight (MW) polymorphisms on sodium dodecyl sulfate-polyacrylamide gels. HPA-2 arises from a threonine/methionine dimorphism at residue 145 of the GPIb alpha sequence, whereas different numbers of tandem repeats of a 39-bp sequence encoding 13-amino acids corresponding to a region between serine399 and threonine411 of the GPIb alpha account for the latter. To identify the genetic basis of the MW polymorphism among Japanese, we counted the tandem repeats in 103 individuals. In addition to the reported three variants with one, two, or three tandem repeats, we identified a new variant with four perfect tandem repeats of the 39-bp sequence that corresponded to the largest phenotype. Phenotypic analysis of the MW polymorphism on 12 individuals including all four phenotypes completely accorded in the genotype. We also determined the genotype of HPA-2 and found that methionine145 was in complete linkage disequilibrium, with the larger variants containing three or four tandem repeats. These results imply a model of evolutionary steps in the gene encoding GPIb alpha.

Alleles↗

Highly constrained proteins contain an unexpectedly large number of amino acid tandem repeats.

Single-amino-acid tandem repeats are very common in mammalian proteins but their function and evolution are still poorly understood. Here we investigate how the variability and prevalence of amino acid repeats are related to the evolutionary constraints operating on the proteins. We find a significant positive correlation between repeat size difference and protein nonsynonymous substitution rate in human and mouse orthologous genes. This association is observed for all the common amino acid repeat types and indicates that rapid diversification of repeat structures, involving both trinucleotide slippage and nucleotide substitutions, preferentially occurs in proteins subject to low selective constraints. However, strikingly, we also observe a significant negative correlation between the number of repeats in a protein and the gene nonsynonymous substitution rate, particularly for glutamine, glycine, and alanine repeats. This implies that proteins subject to strong selective constraints tend to contain an unexpectedly high number of repeats, which tend to be well conserved between the two species. This is consistent with a role for selection in the maintenance of a significant number of repeats. Analysis of the codon structure of the sequences encoding the repeats shows that codon purity is associated with high repeat size interspecific variability. Interestingly, polyalanine and polyglutamine repeats associated with disease show very distinctive features regarding the degree of repeat conservation and the protein sequence selective constraints.

Amino Acid Sequence↗

The central domain of bovine submaxillary mucin consists of over 50 tandem repeats of 329 amino acids. Chromosomal localization of the BSM1 gene and relations to ovine and porcine counterparts.

We previously elucidated five distinct protein domains (I-V) for bovine submaxillary mucin, which is encoded by two genes, BSM1 and BSM2. Using Southern blot analysis, genomic cloning and sequencing of the BSM1 gene, we now show that the central domain (V) consists of approximately 55 tandem repeats of 329 amino acids and that domains III-V are encoded by a 58.4-kb exon, the largest exon known for all genes to date. The BSM1 gene was mapped by fluorescence in situ hybridization to the proximal half of chromosome 5 at bands q2. 2-q2.3. The amino-acid sequence of six tandem repeats (two full and four partial) were found to have only 92-94% identities. We propose that the variability in the amino-acid sequences of the mucin tandem repeat is important for generating the combinatorial library of saccharides that are necessary for the protective function of mucins. The deduced peptide sequences of the central domain match those determined from the purified bovine submaxillary mucin and also show 68-94% identity to published peptide sequences of ovine submaxillary mucin. This indicates that the core protein of ovine submaxillary mucin is closely related to that of bovine submaxillary mucin and contains similar tandem repeats in the central domain. In contrast, the central domain of porcine submaxillary mucin is reported to consist of 81-amino-acid tandem repeats. However, both bovine submaxillary mucin and porcine submaxillary mucin contain similar N-terminal and C-terminal domains and the corresponding genes are in the conserved linkage regions of the respective genomes.

Amino Acid Sequence↗

Generating tandem repeats by cloning with double initiator fragments.

The ability to generate tandem repeats of a DNA sequence has proven important for a large variety of studies of DNA structure and function. The most commonly used method to produce tandem repeats involves cloning of an oligomerized monomer sequence that contains asymmetric overlapping ends, but, in practice, this approach is inefficient because of the circularization of oligomers before they ligate into vector. Described here is a method that circumvents this problem by the use of two separate oligomerization reactions, each containing an initiator fragment onto which monomer polymerizes without circularization. Subsequent mixing of the two reactions permits circularization, generating a viable plasmid containing the sum of the added repeats from each reaction. A variation of this method is also demonstrated that permits the synthesis of constructs with a defined number of repeats.

Cloning, Molecular↗

Functional analysis and DNA polymorphism of the tandemly repeated sequences in the 5'-terminal regulatory region of the human gene for thymidylate synthase.

Triple tandemly repeated sequences and the corresponding complementary sequence are known to exist in the 5'-terminal regulatory region of the human gene for thymidylate synthase (TS). To examine the function of these sequences, a set of deletion mutants was prepared and used in a transient expression assay. The results showed that at least one repeated sequence and its complementary sequence were necessary for the efficient expression of the gene. As another approach to understanding the function of this unique structure, DNA polymorphism in the same region was analyzed. In addition to the TS gene with the triple tandem repeat, the TS gene with a double tandem repeat was found in genomes of normal human subjects at an estimated frequency of 19% when genomes of 21 unrelated Japanese were analyzed. The expression activity of a reporter gene linked to the promoter region of the human TS genes with the two types of repeated sequence was examined and the result showed that the expression activity of the gene with the double repeat was lower than that of the gene with the triple repeat in the transient expression assay. Thus, it appears that the unique repeated sequences in the 5'-terminal region of the human TS gene are polymorphic and contribute to the efficiency of expression of the gene.

Base Sequence↗

Analysis of intrachromosomal homologous recombination in mammalian cell, using tandem repeat sequences.

In all the organisms, homologous recombination (HR) is involved in fundamental processes such as genome diversification and DNA repair. Several strategies can be devised to measure homologous recombination in mammalian cells. We present here the interest of using intrachromosomal tandem repeat sequences to measure HR in mammalian cells and we discuss the differences with the ectopic plasmids recombination. The present review focuses on the molecular mechanisms of HR between tandem repeats in mammalian cells. The possibility to use two different orientations of tandem repeats (direct or inverted repeats) in parallel constitutes also an advantage. While inverted repeats measure only events arising by strand exchange (gene conversion and crossing over), direct repeats monitor strand exchange events and also non-conservative processes such as single strand annealing or replication slippage. In yeast, these processes depend on different pathways, most of them also existing in mammalian cells. These data permit to devise substrates adapted to specific questions about HR in mammalian cells. The effect of substrate structures (heterologies, insertions/deletions, GT repeats, transcription) and consequences of DNA double strand breaks induced by ionizing radiation or endonuclease (especially the rare-cutting endonuclease ISce-I) on HR are discussed. Finally, transgenic mouse models using tandem repeats are briefly presented.

Animals↗