Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “tandem repeat”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Identification of polymorphic tandem repeats by direct comparison of genome sequence from different bacterial strains: a web-based resource.

BACKGROUND: Polymorphic tandem repeat typing is a new generic technology which has been proved to be very efficient for bacterial pathogens such as B. anthracis, M. tuberculosis, P. aeruginosa, L. pneumophila, Y. pestis. The previously developed tandem repeats database takes advantage of the release of genome sequence data for a growing number of bacteria to facilitate the identification of tandem repeats. The development of an assay then requires the evaluation of tandem repeat polymorphism on well-selected sets of isolates. In the case of major human pathogens, such as S. aureus, more than one strain is being sequenced, so that tandem repeats most likely to be polymorphic can now be selected in silico based on genome sequence comparison. RESULTS: In addition to the previously described general Tandem Repeats Database, we have developed a tool to automatically identify tandem repeats of a different length in the genome sequence of two (or more) closely related bacterial strains. Genome comparisons are pre-computed. The results of the comparisons are parsed in a database, which can be conveniently queried over the internet according to criteria of practical value, including repeat unit length, predicted size difference, etc. Comparisons are available for 16 bacterial species, and the orthopox viruses, including the variola virus and three of its close neighbors. CONCLUSIONS: We are presenting an internet-based resource to help develop and perform tandem repeats based bacterial strain typing. The tools accessible at http://minisatellites.u-psud.fr now comprise four parts. The Tandem Repeats Database enables the identification of tandem repeats across entire genomes. The Strain Comparison Page identifies tandem repeats differing between different genome sequences from the same species. The "Blast in the Tandem Repeats Database" facilitates the search for a known tandem repeat and the prediction of amplification product sizes. The "Bacterial Genotyping Page" is a service for strain identification at the subspecies level.

Bacteria↗

Identification and characterization of tandem repeats in exon III of dopamine receptor D4 (DRD4) genes from different mammalian species.

In this study we have identified and characterized dopamine receptor D4 (DRD4) exon III tandem repeats in 33 public available nucleotide sequences from different mammalian species. We found that the tandem repeat in canids could be described in a novel and simple way, namely, as a structure composed of 15- and 12- bp modules. Tandem repeats composed of 18-bp modules were found in sequences from the horse, zebra, onager, and donkey, Asiatic bear, polar bear, common raccoon, dolphin, harbor porpoise, and domestic cat. Several of these sequences have been analyzed previously without a tandem repeat being found. In the domestic cow and gray seal we identified tandem repeats composed of 36-bp modules, each consisting of two closely related 18-bp basic units. A tandem repeat consisting of 9-bp modules was identified in sequences from mink and ferret. In the European otter we detected an 18-bp tandem repeat, while a tandem repeat consisting of 27-bp modules was identified in a sequence from European badger. Both these tandem repeats were composed of 9-bp basic units, which were closely related with the 9-bp repeat modules identified in the mink and ferret. Tandem repeats could not be identified in sequences from rodents. All tandem repeats possessed a high GC content with a strong bias for C. On phylogenetic analysis of the tandem repeats evolutionary related species were clustered into the same groups. The degree of conservation of the tandem repeats varied significantly between species. The deduced amino acid sequences of most of the tandem repeats exhibited a high propensity for disorder. This was also the case with an amino acid sequence of the human DRD4 exon III tandem repeat, which was included in the study for comparative purposes. We identified proline-containing motifs for SH3 and WW domain binding proteins, potential phosphorylation sites, PDZ domain binding motifs, and FHA domain binding motifs in the amino acid sequences of the tandem repeats. The numbers of potential functional sites varied pronouncedly between species. Our observations provide a platform for future studies of the architecture and evolution of the DRD4 exon III tandem repeat, and they suggest that differences in the structure of this tandem repeat contribute to specialization and generation of diversity in receptor function.

Amino Acid Motifs↗

Identification and characterization of two tandem repeat sequences (TrsB and TrsC) and a retrotransposon (RIRE1) as genome-general sequences in rice.

Three kinds of DNA sequences (here called TrsB, TrsC and RIRE1) have been previously reported to be those repeated in tandem specifically in the wild rice species with FF, CC or EE genome, respectively. To characterize these genome type-specific sequences, we carried out PCR using a pair of primers, which hybridize to a restricted region in the repeating unit sequence and prime DNA synthesis in both directions. Gel electrophoresis and DNA sequencing revealed that PCR using primers for TrsB (or TrsC) amplified the fragments with an integral series of a unit length not only from total DNA of the rice strain with FF (or CC) genome, but also from those of the rice strains with non-FF (or non-CC) genome. TrsB or TrsC was, however, found to be repeated in an extraordinary number of copies in the species with FF or CC genome, respectively, in which the TrsB (or TrsC) sequence has been originally identified. PCR using primers for RIRE1 produced various sizes of fragments from total DNA of the rice strains with EE genome. The fragments, however, showed no progression at interval of the unit length characteristic for tandem repeats. Nucleotide sequencing of the amplified fragments revealed that they were not the sequences repeated in tandem, but were those interspersed as an element having partial homology with the LTR sequences of retrotransposons, Wis-2-1A in wheat and BARE-1 in barley. RIRE1 was present in the rice species with any types of genomes, but in the species with EE genome in an extraordinary number of copies.

Base Sequence↗

Detection of single and multiple polymorphic loci by synthetic tandem repeats of short oligonucleotides.

Loci containing tandem repeats of short sequences are sometimes associated with a high level of polymorphism due to variations in the number of repeats. The different variants can be easily characterized by Southern blotting when the repeats span a range from a few hundred bases to a few kilobases, and probes derived from such tandem repeats constitute convenient genetic markers. These structures, usually called minisatellites, are best documented in the human genome, where their number has been estimated to be at least 1500. However, their role and mode of evolution are poorly understood. We are developing tools to evaluate the number of such redundant sequences in a genome and to gain access to new polymorphic loci. Our strategy is based on the use of polymers of oligonucleotides as DNA probes for hybridization on Southern blots. In a previous report, we made polymers with random units of 14 bp and showed that they detect multiple polymorphic loci on human genomic DNA. At present, we are testing the effect of an increase in the complexity of the polymer, as obtained by the use of a longer random unit, and the effect of slight sequence modifications to a particular tandem repeat sequence. In addition, some of these synthetic probes can detect a single polymorphic locus and directly provide new genetic markers.

Base Sequence↗

TROLL--tandem repeat occurrence locator.

SUMMARY: Tandem Repeat Occurrence Locator (TROLL), is a light-weight Simple Sequence Repeat (SSR) finder based on a slight modification of the Aho-Corasick algorithm. It is fast and only requires a standard Personal Computer (PC) to operate. We report running times of 127 s to find all SSRs of length 20 bp or more on the complete Arabdopsis genome--approx. 130 Mbases divided in five chromosomes--using a PC Athlon 650 MHz with 256 MB of RAM. AVAILABILITY: TROLL is an open source project and is available at http://finder.sourceforge.net.

Algorithms↗

Human mucin gene MUC4: organization of its 5'-region and polymorphism of its central tandem repeat array.

In a previous study we isolated a partial cDNA with a tandem repeat of 48 bp, which allowed us to map a novel human mucin gene named MUC4 to chromosome 3q29. Here we report the organization and sequence of the 5'-region and its junction with the tandem repeat array of MUC4. Analysis of three overlapping genomic clones allowed us to obtain a partial restriction map of MUC4 and to locate the complete 48 bp tandem repeat domain on a PstI/EcoRI genomic fragment that exhibits a very large variation in number of tandem repeats (7-19 kb). cDNA clonal extension allowed us to obtain the entire 5' coding region of MUC4. Exon 1 consists of a 5' untranslated region and an 82 bp fragment encoding the signal peptide. This latter shows a high degree of similarity to the signal peptide of another apomucin, ASGP-1. Exon 2 is extremely large and contains a unique sequence that is followed by the whole tandem repeat domain. It encodes only one cysteine residue, making MUC4 different from mucin genes belonging to the 11p15.5 family. Moreover, an intron downstream from the tandem repeat array consists mainly of a 15 bp tandem repeat that exhibits a polymorphism in having a variable number of tandem repeats.

Amino Acid Sequence↗

Characterization of a Plasmodium chabaudi gene encoding a protein with glutamate-rich tandem repeats.

Several highly antigenic proteins containing tandem repeats rich in glutamic acid residues have been described in Plasmodium falciparum. However, relatively little information is available about analogous genes in rodent parasites. This report describes a 4.2-kb genomic DNA fragment from P. chabaudi with a deduced amino acid sequence that is predominantly glutamate-rich tandem repeats. Several different monoclonal antibodies raised against a 93-kDa P. chabaudi protein, which does not correspond to the cloned DNA fragment, recognize a recombinant protein expressed from the 4.2-kb DNA fragment. The only sequence similarities between these two genes are tandem repeats with a predominance of glutamate pairs followed by a hydrophobic residue. This repetitious-sequence motif may be the basis for the observed cross-reactivity. A similar motif has been demonstrated to be the basis for antibody cross-reactivity between glutamate-rich proteins of P. falciparum. The expression of multiple glutamate-rich proteins with cross-reacting epitopes may be a general phenomenon in Plasmodium species.

Amino Acid Sequence↗

Length of uninterrupted repeats determines instability at the unstable mouse expanded simple tandem repeat family MMS10 derived from independent SINE B1 elements.

Mouse expanded simple tandem repeats (ESTRs) provide highly informative loci for analyzing spontaneous and induced germline mutation. We have conducted an extensive sequence database search and identified 17 new members of the highly unstable rodent-specific ESTR family called MMS10. This family has arisen by independent expansions of a common GGCAGA repeat unit from within a subset of both ancestral and modern SINE B1 elements during the course of mouse evolution. Analysis of the interspersion patterns of variant repeats along alleles of 20 of these MMS10 loci revealed two distinct classes of tandem arrays: one composed of uninterrupted GGCAGA repeats and the second with generally larger arrays interrupted by variant units. Surveys of allelic diversity at 11 representative members of these two classes of loci in various laboratory strains and BXD recombinant inbred lines revealed that the level of repeat instability was positively correlated with the length of uninterrupted repeats. Turnover processes at MMS10 loci, therefore, appear similar to the type of mechanism observed at human microsatellites. The MMS10 family thus provides a potentially useful murine model for studying dynamic mutation at simple tandem repeats.

Alleles↗

GeomeTRe: accurate calculation of geometrical descriptors of tandem repeat proteins.

MOTIVATION: Structured tandem repeat proteins (STRPs) are characterized by preserved structural motifs arranged in a modular way. The structural and functional diversity of STRPs makes them particularly important for studying evolution and novel structure-function relationships, and ultimately for designing new synthetic proteins with specific functions. One crucial aspect of their classification is the estimation of geometrical parameters, which can provide better insight into their properties and the relationship between the spatial arrangement of repeated units and protein function. Calculating geometric descriptors for STRPs is challenging because naturally occurring repeats are not "perfect" and often contain insertions and deletions. Existing tools for predicting structural symmetry work well on simple cases but often fail for most natural proteins. RESULTS: Here, we present GeomeTRe, an algorithm that calculates geometrical descriptors such as curvature (yaw), twist (roll), and pitch for a protein structure with known repeat unit positions. The algorithm simulates the movement of consecutive units, identifies rotational axes, and calculates the corresponding Tait-Bryan angles. GeomeTRe's parameters can enhance STRP annotation and classification by identifying variations in geometric arrangements among different functional groups. The package is fast and suitable for processing large protein structure datasets when repeat region information (e.g. from RepeatsDB) is available. AVAILABILITY AND IMPLEMENTATION: GeomeTRe is available as a Python package; source code and documentation can be found at https://github.com/BioComputingUP/GeomeTRe.

Algorithms↗

Intragenic tandem repeats generate functional variability.

Tandemly repeated DNA sequences are highly dynamic components of genomes. Most repeats are in intergenic regions, but some are in coding sequences or pseudogenes. In humans, expansion of intragenic triplet repeats is associated with various diseases, including Huntington chorea and fragile X syndrome. The persistence of intragenic repeats in genomes suggests that there is a compensating benefit. Here we show that in the genome of Saccharomyces cerevisiae, most genes containing intragenic repeats encode cell-wall proteins. The repeats trigger frequent recombination events in the gene or between the gene and a pseudogene, causing expansion and contraction in the gene size. This size variation creates quantitative alterations in phenotypes (e.g., adhesion, flocculation or biofilm formation). We propose that variation in intragenic repeat number provides the functional diversity of cell surface antigens that, in fungi and other pathogens, allows rapid adaptation to the environment and elusion of the host immune system.

Antigens, Surface↗

Evolution of tandemly repeated sequences: What happens at the end of an array?

Tandemly repeated sequences are a major component of the eukaryotic genome. Although the general characteristics of tandem repeats have been well documented, the processes involved in their origin and maintenance remain unknown. In this study, a region on the paternal sex ratio (PSR) chromosome was analyzed to investigate the mechanisms of tandem repeat evolution. The region contains a junction between a tandem array of PSR2 repeats and a copy of the retrotransposon NATE, with other dispersed repeats (putative mobile elements) on the other side of the element. Little similarity was detected between the sequence of PSR2 and the region of NATE flanking the array, indicating that the PSR2 repeat did not originate from the underlying NATE sequence. However, a short region of sequence similarity (11/15 bp) and an inverted region of sequence identity (8 bp) are present on either side of the junction. These short sequences may have facilitated nonhomologous recombination between NATE and PSR2, resulting in the formation of the junction. Adjacent to the junction, the three most terminal repeats in the PSR2 array exhibited a higher sequence divergence relative to internal repeats, which is consistent with a theoretical prediction of the unequal exchange model for tandem repeat evolution. Other NATE insertion sites were characterized which show proximity to both tandem repeats and complex DNAs containing additional dispersed repeats. An "accretion model" is proposed to account for this association by the accumulation of mobile elements at the ends of tandem arrays and into "islands" within arrays. Mobile elements inserting into arrays will tend to migrate into islands and to array ends, due to the turnover in the number of intervening repeats.

Base Sequence↗

A novel single molecule analysis of spontaneous and radiation-induced mutation at a mouse tandem repeat locus.

Expanded simple tandem repeat (ESTR) loci include some of the most unstable DNA in the mouse genome and have been extensively used in pedigree studies of germline mutation. We now show that repeat DNA instability at the mouse ESTR locus Ms6-hm can also be monitored by single molecule PCR analysis of genomic DNA. Unlike unstable human minisatellites which mutate almost exclusively in the germline by a meiotic recombination-based process, mouse Ms6-hm shows repeat instability both in germinal (sperm) DNA and in somatic (spleen, brain) DNA. There is no significant variation in mutation frequency between mice of the same inbred strain. However, significant variation occurs between tissues, with mice showing the highest mutation frequency in sperm. The size spectra of somatic and sperm mutants are indistinguishable and heavily biased towards gains and losses of only a few repeat units, suggesting repeat turnover by a mitotic replication slippage process operating both in the soma and in the germline. Analysis of male mice following acute pre-meiotic exposure to X-rays showed a significant increase in sperm but not somatic mutation frequency, though no change in the size spectrum of mutants. The level of radiation-induced mutation at Ms6-hm was indistinguishable from that established by conventional pedigree analysis following paternal irradiation. This confirms that mouse ESTR loci are very sensitive to ionizing radiation and establishes that induced germline mutation results from radiation-induced mutant alleles being present in sperm, rather than from unrepaired sperm DNA lesions that subsequently lead to the appearance of mutants in the early embryo. This single molecule monitoring system has the potential to substantially reduce the number of mice needed for germline mutation monitoring, and can be used to study not only germline mutation but also somatic mutation in vivo and in cell culture.

Animals↗

NIST mixed stain study 3: signal intensity balance in commercial short tandem repeat multiplexes.

Short-tandem repeat (STR) allelic intensities were collected from more than 60 forensic laboratories for a suite of seven samples as part of the National Institute of Standards and Technology-coordinated 2001 Mixed Stain Study 3 (MSS3). These interlaboratory challenge data illuminate the relative importance of intrinsic and user-determined factors affecting the locus-to-locus balance of signal intensities for currently used STR multiplexes. To varying degrees, seven of the eight commercially produced multiplexes used by MSS3 participants displayed very similar patterns of intensity differences among the different loci probed by the multiplexes for all samples, in the hands of multiple analysts, with a variety of supplies and instruments. These systematic differences reflect intrinsic properties of the individual multiplexes, not user-controllable measurement practices. To the extent that quality systems specify minimum and maximum absolute intensities for data acceptability and data interpretation schema require among-locus balance, these intrinsic intensity differences may decrease the utility of multiplex results and surely increase the cost of analysis.

Alleles↗

Approximating edit distances between complex tandem repeats efficiently.

MOTIVATION: Extended tandem repeats (TRs) have been associated with 60 or more diseases over the past 30 years. Although most TRs have single repeat units (or motifs), complex TRs with different units have recently been correlated with some brain disorders. Of note, a population-scale analysis shows that complex TRs at one locus can be divergent, and different units are often expanded between individuals. To understand the evolution of high TR diversity, it is informative to visualize a phylogenetic tree. To do this, we need to measure the edit distance between pairs of complex TRs by considering duplication and contraction of units created by replication slippage. However, traditional rigorous algorithms for this purpose are computationally expensive. RESULTS: We here propose an efficient heuristic algorithm to estimate the edit distance with duplication and contraction of units (EDDC, for short). We select a set of frequent units that occur in given complex TRs, encode each unit as a single symbol, compress a TR into an optimal series of unit symbols that partially matches the original TR with the minimum Levenshtein distance, and estimate the EDDC between a pair of complex TRs from their compressed forms. Using substantial synthetic benchmark datasets, we demonstrate that the estimated EDDC is highly correlated with the accurate EDDC, with a Pearson correlation coefficient of >0.983, while the heuristic algorithm achieves orders of magnitude performance speedup. AVAILABILITY AND IMPLEMENTATION: The software program hEDDC that implements the proposed algorithm is available at https://github.com/Ricky-pon/hEDDC (DOI: 10.5281/zenodo.14732958).

Algorithms↗

Maximum likelihood estimates of admixture in Northeastern Mexico using 13 short tandem repeat loci.

Tetrameric short tandem repeat (STR) polymorphisms are widely used in population genetics, molecular evolution, gene mapping and linkage analysis, paternity tests, forensic analysis, and medical applications. This article provides allelic distributions of the STR loci D3S1358, vWA, FGA, D8S1179, D21S11, D18S51, D5S818, D13S317, D7S820, CSF1PO, TPOX, TH01, and D16S539 in 143 Mestizos from Northeastern Mexico, estimates of contributions of genes of European (Spanish), American Indian and African origin in the gene pool of this admixed Mestizo population (using 10 of these loci); and a comparison of the genetic admixture of this population with the previously reported two polymorphic molecular markers, D1S80 and HLA-DQA1 (n = 103). Genotype distributions were in agreement with Hardy-Weinberg expectations (HWE) for almost all 13 STR markers. Maximum likelihood estimates of admixture components yield a trihybrid model with Spanish, Amerindian, and African ancestry with the admixture proportions: 54.99% +/- 3.44, 39.99% +/- 2.57, and 5.02% +/- 2.82, respectively. These estimates were not significantly different from those obtained using D1S80 and HLA-DQA1 loci (59.99% +/- 5.94, 36.99% +/- 5.04, and 3.02% +/- 2.76). In conclusion, Mestizos of Northeastern Mexico showed a similar ancestral contribution independent of the markers used for evolutionary purposes. Further validation of this database supports the use of the 13 STR loci along with D1S80 and HLA-DQA1 as a battery of efficient DNA forensic markers in Northeastern Mestizo populations of Mexico.

Genetic Markers↗

Exact tandem repeats analyzer (E-TRA): a new program for DNA sequence mining.

Exact Tandem Repeats Analyzer 1.0 (E-TRA) combines sequence motif searches with keywords such as 'organs', 'tissues', 'cell lines' and 'development stages' for finding simple exact tandem repeats as well as non-simple repeats. E-TRA has several advanced repeat search parameters/options compared to other repeat finder programs as it not only accepts GenBank, FASTA and expressed sequence tags (EST) sequence files, but also does analysis of multiple files with multiple sequences. The minimum and maximum tandem repeat motif lengths that E-TRA finds vary from one to one thousand. Advanced user defined parameters/options let the researchers use different minimum motif repeats search criteria for varying motif lengths simultaneously. One of the most interesting features of genomes is the presence of relatively short tandem repeats (TRs). These repeated DNA sequences are found in both prokaryotes and eukaryotes, distributed almost at random throughout the genome. Some of the tandem repeats play important roles in the regulation of gene expression whereas others do not have any known biological function as yet. Nevertheless, they have proven to be very beneficial in DNA profiling and genetic linkage analysis studies. To demonstrate the use of E-TRA, we used 5,465,605 human EST sequences derived from 18,814,550 GenBank EST sequences. Our results indicated that 12.44% (679,800) of the human EST sequences contained simple and non-simple repeat string patterns varying from one to 126 nucleotides in length. The results also revealed that human organs, tissues, cell lines and different developmental stages differed in number of repeats as well as repeat composition, indicating that the distribution of expressed tandem repeats among tissues or organs are not random, thus differing from the un-transcribed repeats found in genomes.

Cells, Cultured↗

Beyond tandem repeats: complex pattern structures and distant regions of similarity.

MOTIVATION: Tandem repeats (TRs) are associated with human disease, play a role in evolution and are important in regulatory processes. Despite their importance, locating and characterizing these patterns within anonymous DNA sequences remains a challenge. In part, the difficulty is due to imperfect conservation of patterns and complex pattern structures. We study recognition algorithms for two complex pattern structures: variable length tandem repeats (VLTRs) and multi-period tandem repeats (MPTRs). RESULTS: We extend previous algorithmic research to a class of regular tandem repeats (RegTRs). We formally define RegTRs, as well as two important subclasses: VLTRs and MPTRs. We present algorithms for identification of TRs in these classes. Furthermore, our algorithms identify degenerate VLTRs and MPTRs: repeats containing substitutions, insertions and deletions. To illustrate our work, we present results of our analysis for two difficult regions in cattle and human data which reflect practical occurrences of these subclasses in GenBank sequence data. In addition, we show the applicability of our algorithmic techniques for identifying Alu sequences, gene clusters and other distant regions of similarity. We illustrate this with an example from yeast chromosome I.

Algorithms↗

Identification and characterization of a tandem repeat in exon III of the dopamine receptor D4 (DRD4) gene in cetaceans.

A large number of mammalian species harbor a tandem repeat in exon III of the gene encoding dopamine receptor D4 (DRD4), a receptor associated with cognitive functions. In this study, a DRD4 gene exon III tandem repeat from the order Cetacea was identified and characterized. Included in our study were samples from 10 white-beaked dolphins (Lagenorhynchus albirostris), 10 harbor porpoises (Phocoena phocoena), eight sperm whales (Physeter macrocephalus), and five minke whales (Balaenoptera acutorostrata). Using enzymatic amplification followed by sequencing of amplified fragments, a tandem repeat composed of 18-bp basic units was detected in all of these species. The tandem repeats in white-beaked dolphin and harbor porpoise were both monomorphic and consisted of 11 and 12 basic units, respectively. In contrast, the sperm whale harbored a polymorphic tandem repeat with size variants composed of three, four, and five basic units. Also the tandem repeat in minke whale was polymorphic; size variants composed of 6 or 11 basic units were found in this species. The consensus sequences of the basic units were identical in the closely related white-beaked dolphin and harbor porpoise, and these sequences differed by a maximum of two changes when compared to the remaining species. There was a high degree of similarity between the cetacean basic unit consensus sequences and those from members of the horse family and domestic cow, which also harbor a tandem repeat composed of 18-bp basic units in exon III of their DRD4 gene. Consequently, the 18-bp tandem repeat appears to have originated prior to the differentiation of hoofed mammals into odd-toed and even-toed ungulates. The composition of the tandem repeat in cetaceans differed markedly from that in primates, which is composed of 48-bp repeat basic units.

Animals↗