Search PubMedSearch

SEARCH · Search PubMed

Results for “tandem repeat”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Polyoma virus giant RNAs contain tandem repeats of the nucleotide sequence of the entire viral genome.

The bulk of late virus-specific RNA synthesized in polyoma virus-infected mouse cells is larger than a single strand of poloma DNA. The arrangement of viral nucleotide sequences in these giant polyoma RNAs was studied by electron microscopy of hybrids between purified high molecular weight viral RNA and the HindII-1 fragment of polyoma DNA, which contains 91% of the viral genome. Hybrid molecules containing a short single-stranded gap (corresponding to the 9% of viral sequences not present in HindII-1), flanked by double-stranded regions, were photographed and measured. The majority of hybrid molecules contained no single-stranded loops or branches, showing that all viral sequences are transcribed contiguously and that no nonviral sequences are present in the RNA. Hybrid molecules, containing RNA up to 3.5 times the genome length, had a repeating structure of single-stranded gaps 8% of genome length interspersed with double-stranded regions 89% of genome length, showing that giant polyoma RNAs contain tandem repeats of the nucleotide sequence of the entire viral DNA. A small proportion of hybrid molecules contained single-stranded branches or deletion loops in characteristic positions, indicating that RNA "splicing" may occur on high molecular weight nuclear polyoma RNA.

Cell Nucleus

Heterochromatin-based silencing of a foreign tandem repeat in Drosophila melanogaster shows unusual biochemistry and temperature sensitivity.

Eukaryotic genomes are packaged into chromatin, a regulatory nucleoprotein assembly. Establishment, maintenance, and interconversion of chromatin states is required for correct patterns of gene expression, genome integrity, and survival. Transcriptionally repressive heterochromatin minimizes mobilization of transposable elements and limits expansion of other repetitive DNA, but mechanisms for recognition of the latter sequences are not well established. We previously demonstrated in Drosophila melanogaster that transcripts derived from 1360 and Invader4 transposon insertions can trigger local conversion of transcriptionally permissive euchromatin to heterochromatin through the piRNA system, but only in a subset of genomic locations near existing blocks of heterochromatin. Here we show that a ~9 kb tandem array of the 36-nucleotide lac operator (lacO) sequence of Escherichia coli can form ectopic heterochromatin at a similar subset of sites, resulting in variegating expression of an adjacent reporter gene. Heterochromatin Protein 1a (HP1a) and histone deacetylation are required for lacO repeat-induced silencing, but, contrasting with previously described Position Effect Variegation (PEV), we do not observe increased histone H3 lysine 9 methylation. Silencing is effective at 25°C and suppressed at 18°C (in contrast to canonical PEV, which is enhanced at 18°C), indicating involvement of a temperature-sensitive component. Temperature switching experiments show that lacO repeat-induced heterochromatin formation is reversible throughout larval development following an HP1a-dependent initiation step in the early embryo. We conclude that the Drosophila nucleus can recognize a completely foreign tandem repeat as a target for heterochromatin formation, and that the heterochromatin structure established is distinct from that of endogenous tandem arrays.

HP1a

The genes for 18S, 5.8S and 28S ribosomal RNA of Bombyx mori are organized into tandem repeats of uniform length.

The organization of the multiple genes for 18S, 5.8S and 28S rRNA in the genome of the silkworm, Bombyx mori was determined by restriction endonuclease digestion and Southern blot hybridization. The ribosomal genes (rDNA) are tandemly reiterated, with a uniform repeat length of 6.9 . 10(6) daltons. Each rDNA repeat has a single site for EcoRI, HindIII, HpaI and SmaI and each of these sites has been mapped with respect to the others and to the rRNA genes; each repeat consists of a transcribed region (6 . 10(6)daltons) containing the 18S, 5.8S and 28S rRNA genes (5' leads to 3') and also a small non-transcribed spacer (approximately 10(6) daltons). Complete rDNA repeats were cloned using the vector RSF2124 and grown in Escherichia coli. Characterization of the rDNA plasmids confirmed the conclusions from studies of the total rDNA. The organization of B. mori rDNA is similar to that of other eukaryotes, except for the absence of heterogeneity in the rDNA repeat length; thus, there is neither variation in the length of the non-transcribed spacer nor the presence of inserts in a detectable portion of the rDNA. The utility of this map, and particularly of the rDNA plasmids, for detailed studies of rRNA transcription and processing is discussed.

Animals

Synthesis of hybrid bacterial plasmids containing highly repeated satellite DNA.

Hybrid plasmid molecules containing tandemly repeated Drosophila satellite DNA were constructed using a modification of the (dA)-(dT) homopolymer procedure of Lobban and Kaiser (1973). Recombinant plasmids recovered after transformation of recA bacteria contained 10% of the amount of satellite DNA present in the transforming molecules. The cloned plasmids were not homogenous in size. Recombinant plasmids isolated from a single colony contained populations of circular molecules which varied both in the length of the satellite region and in the poly(dA)-(dt) regions linking satellite and vector. While subcloning reduced the heterogeneity of these plasmid populations, continued cell growth caused further variations in the size of the repeated regions. Two different simple sequence satellites of Drosophila melanogaster (1.672 and 1.705 g/cm3) were unstable in both recA and recBC hosts and in both pSC101 and pCR1 vectors. We propose that this recA-independent instability of tandemly repeated sequences is due to unequal intramolecular recombination events in replicating DNA molecules, a mechanism analogous to sister chromatid exchange in eucaryotes.

DNA

Characterisation of ribosomal satellite in total nuclear DNA from Physarum polycephalum.

The distinctive properties of satellite DNA molecules containing the genes for ribosomal RNA in Physarum polycephalum permits their identification in total, unfractionated nuclear DNA in the foldback form, after denaturation and fast annealing. Using the electron microscope the location and properties of three characteristic regions containing tandemly-repeated, inverted sequences have been investigated. At least two additional regions, also containing tandem repeats, are shown to be present and located towards each end of the rDNA molecule, at a site adjacent to the segment coding for the 26 S rRNA. All the regions which contain tandem repeats are composed of sequences which, within experimental error, appear to share a common unit repeat length of about 90 nucleotides.

Base Sequence

Nucleotide sequence of the region of an origin of replication of the antibiotic resistance plasmid R6K.

A 2.1-kilobase segment of the antibiotic resistance plasmid R6K carries sufficient information to replicate as a plasmid in Escherichia coli. This segment contains a functional origin of replication and a structural gene for a protein, designated pi, that is required for the initiation of R6K replication. The nucleotide sequence of a 520-base-pair portion of this 2.1-kilobase segment that includes the functional origin of replication and the region adjacent to the start of the pi structural gene was determined. A striking feature of the sequence is the presence of seven 22-base-pair direct repeats joined in tandem in the region adjacent to the start of the pi gene. A possible role of the tandem repeats in the regulation of expression of the pi protein and the control of initiation of replication of the plasmid R6K is discussed.

Anti-Bacterial Agents

Assembly and comparative analysis of the mitochondrial genome of Pleione yunnanensis: genome structure and evolutionary insights.

BACKGROUND: Pleione yunnanensis a terrestrial or semi-epiphytic herbaceous plant belonging to the Orchidaceae family, is valued for both its medicinal uses and ornamental appeal. Although its chloroplast genomes have been sequenced, its complete mt genome had not previously been resolved, limiting genetic and evolutionary studies of the species. RESULTS: In this work, we assembled and characterized the first complete mt genome of P. yunnanensis, revealing a structurally complex, multibranched system composed of 14 circular-mapping molecules totaling 468,176 bp with a GC content of 44.32%. The genome encodes 44 annotated genes, including 28 protein-coding genes (PCGs), 15 tRNAs, and one rRNA. The multibranched architecture provides new evidence supporting the dynamic and recombinational nature of plant mt genomes. Repeat analysis uncovered 29 simple sequence repeats (SSRs), 19 tandem repeats, and 118 dispersed repeats, indicating a comparatively lower repeat abundance than that found in closely related orchids with similar mt genome sizes. Codon-usage profiling of PCGs showed a marked bias toward A/T-ending codons. Prediction of RNA editing sites identified 4,708 putative edits across mitochondrial PCGs. Most mitochondrial genes displayed Ka/Ks ratios close to 1.0, suggesting relaxed selective constraints or lineage-specific evolutionary patterns rather than strong positive selection. Moreover, we detected 69 chloroplast-derived homologous fragments, including 15 intact genes, suggesting ongoing plastid-mitochondrial DNA transfer. Phylogenetic reconstruction and collinearity comparisons demonstrated that P. yunnanensis clustered closely with Dendrobium species, including D. amplum and D. hancockii, within the Orchidaceae clade. CONCLUSIONS: This study provides the first complete mt genome of P. yunnanensis, providing a foundational genomic resource for the genus Pleione. The results not only improve our understanding of mt genome structure and evolution in Orchidaceae, but also offer valuable molecular evidence for phylogenetic inference, germplasm identification, and conservation of this endangered medicinal species.

Orchidaceae

Assembly and characterization of the first complete mitochondrial genome of Epimedium sagittatum (Sieb. et Zucc.) Maxim (Berberidaceae):an invaluable traditional Chinese medicine.

BACKGROUND: Epimedium sagittatum (Sieb. et Zucc.) Maxim is an invaluable traditional Chinese medicine plant known for its properties of tonifying kidney yang, strengthening bones and muscles, and dispelling rheumatism. The chloroplast (cp) genome of E. sagittatum have been sequenced, offering critical insights for breeding and phylogenetic research. However, the mitochondrial (mt) genome of E. sagittatum remains uncharacterized, limiting comprehensive insights into its genomic evolution. RESULTS: In this study, we assembled the first complete mt genome of E. sagittatum employing Illumina and Nanopore sequencing technology and subsequently investigated comparative analysis with its closely related species. The mt genome of E. sagittatum was assembled as a multi-branched structure with a length of 339,191 bp, within a GC content of 46.91%. Our annotation results have shown 39 protein-coding genes (PCGs), 22 tRNA genes, three rRNA genes and four pseudogenes in the E. sagittatum mt genome. The analysis of sequence repeats has detected 79 simple sequence repeats (SSRs), 10 tandem repeats and 255 dispersed repeats in the E. sagittatum mt genome. A total of 720 C to U RNA editing sites of the 34 PCGs was predicted in E. sagittatum. The codons exhibited a strong preference for A or U bases in the E. sagittatum mt genome. The analysis of nucleotide diversity (Pi) highlighted differences in genetic variability across the tested genes, with atp9 gene exhibiting the highest genetic variation. Selection pressure analysis showed that most genes were affected by negative selection during evolution, whereas ccmB, rps10, and rps12 underwent positive selection in different plants. Additionally, a Bayesian phylogenetic tree showed that E. sagittatum was closely related to E. wushanense and E. pubescens. In total of 14 homologous fragments totaling 8,954 bp were identified between the cp and mt genomes of E. sagittatum. CONCLUSIONS: This study presents the first assembled and annotated mt genome of E. sagittatum, which provides a valuable genetic resource for the Epimedium genus and lays the foundation for investigating the phylogenetic relationship and genetic variation of this invaluable medicinal plant.

Epimedium

Molecular epidemiological surveillance for non-tuberculous mycobacterial pulmonary disease: a single-center prospective cohort study.

UNLABELLED: Bacterial species cultured from sputum change during treatment or observation for non-tuberculous mycobacterial pulmonary disease; however, strain-level changes remain unrecognized. Variable number tandem repeat typing is a standard technique for strain identification; nonetheless, its labor-intensive and time-consuming nature limits routine clinical use. Therefore, we aimed to elucidate species-subspecies and strain dynamics in non-tuberculous mycobacteria and develop a simple sequence-based strain-level determination method. We performed a single-center prospective cohort study of 112 patients with non-tuberculous mycobacterial pulmonary disease. Whole-genome sequencing was performed on two sputum samples collected at enrollment and at the end of follow-up, followed by variable number tandem repeat (VNTR) typing. We also developed a simple long-read sequencing-based digital VNTR (dVNTR) typing method and evaluated its efficacy. Our results demonstrate that core genome multi-locus sequencing typing revealed species/subspecies changes in 13 patients (11.6%); VNTR typing detected strain changes in 16 patients (14.3%) without species/subspecies changes. Overall, pathogen shifts occurred in 29 patients (shift [+] group, 25.9%), whereas 83 had no detectable pathogen shift (shift [-] group, 74.1%). Interestingly, macrolide and amikacin susceptibility changed in both groups, but resistance remained higher in shift (-) patients. dVNTR results aligned with those of conventional VNTR typing. In conclusion, since susceptibility factors remain unclear, routine species/subspecies identification and molecular typing, such as VNTR, are optimal for patient care. Core genome multi-locus sequencing typing with a dVNTR identified pathogen shifts, innovating non-tuberculous mycobacterial pulmonary disease management.Clinical TrialsThis study is registered with UMIN as UMIN 000056067. IMPORTANCE: Pulmonary non-tuberculous mycobacterial disease is a chronic infection in which the causative pathogens may change at the species, subspecies, or strain level over time. Accurate tracking of these changes is essential for optimizing treatment; however, conventional clinical practice lacks efficient methods for monitoring such dynamics. Our study revealed pathogen changes in approximately one-quarter of patients over 1.5 years, prompting the development of a novel surveillance system that integrates next-generation sequencing for both species-subspecies identification and strain-level molecular epidemiology. This innovation enables real-time monitoring of pathogen dynamics, allowing clinicians to promptly adjust treatment strategies and improve patient care through more informed decision-making.

Humans

Occurrence of reiterated sequences in an untranslated region of Simian virus 40 DNA determined by nucleotide sequence analysis.

An earlier report (Subramanian, Dhar, and Weissman, 1977c) presented the nucleotide sequence of Eco RII-G fragment of SV40 DNA, which contains the origin of DNA replication. The nucleotide sequence of Eco RII-N fragment located next to Eco RII-G on the physical map of SV40 DNA is presented in this report. Eco RII-N is found to be a tandem duplication of the last 55 nucleotides of Eco RII-G. This tandem repeat is immediately preceded by two other reiterated sequences occurring within Eco RII-G, one of them being a tandem repeat of 21 nucleotides and the other a nontandem repeat of 10 nucleotides. These repetitive sequences occur in close proximity to the origin of DNA replication which is known to contain other specialized sequences such as a few palindromes (one of which is 27 long and possesses a perfect 2-fold axis of symmetry), one "true" palindrome, and a long A/T-rich cluster. The repeats (and the replication origin) occur within an untranslated region of SV40 DNA flanked by (the few) structural genes coding for the "late" proteins on the one side and that (those) coding for the "early" protein(s) on the other side. The reiterated sequences are comparable in some respects to repetitive sequences occurring in eucaryotic DNAs. Possible biological functions of the repeats are discussed.

Base Sequence

Organization of the 5S RNA genes in macro- and micronuclei of Tetrahymena pyriformis.

The organization of the 5S genes in macro- and micronuclei of Tetrahymena pyriformis was studied using restriction endonucleases. After complete digestion of macronuclear DNA with BamH-I or Hpa I, 5S RNA hybridized to a DNA fragment of approximately 280 base pairs (bp). When macronuclear DNA was only partially digested with these enzymes, hybridization with 32P-5S RNA demonstrated an oligomeric series with a spacing of 280 bp. These results indicate that the 5S genes are tandemly repeated in macronuclei and that the repeating unit is 280 bp (or 180,000 daltons). Since 5S RNA is 120 nucleotides, we conclude that the 5S repeat units contain a 120 bp transcribed region and a 160 bp spacer region. When macronuclear DNA was digested with Eco RI, Bgl I, or Eco RI + Bgl I, 5S RNA hybridized to DNA of molecular weight 3--4 X 10(6), suggesting that these enzymes do not cleave within a 5S repeat. These 3--4 X 10(6) dalton fragments define the maximum size of an average cluster of 5S repeated units. Assuming the size of the 5S repeat to be 0.18 X 10(6) daltons, there are about 15--20 5S repeats per average tanden cluster, and since there are 350 5S-genes per haploid genome, there must be approximately 15--20 tandem arrays. Results obtained using micronuclear DNA suggest that organization of the 5S-genes is very similar in macro- and micronuclei. Macronuclear rRNA genes are extrachromosomal palindromic dimers. In contrast, 5S genes in Tetrahymena were found to be integrated within the genomes of both macro- and micronuclei and not linked to the rRNA genes. Moreover, it is unlikely that they are palindromes; rather they appear to be tandemly repeated in "head-to-tail" linkages. Thus the organization of the 5S genes in Tetrahymena is similar to that of higher eukaryotes.

Animals

Deletions, rearrangements and tandem duplications of recombinant plasmids containing yeast ribosomal DNA.

Recombinant DNA plasmids, formed by the insertion of yeast ribosomal DNA into Escherichia coli plasmids pSC101 or pMB9, underwent deletions, rearrangements and tandem duplications. Independently derived deletion products of plasmids, constructed using pMB9 as a vector, are indistinguishable from each other. These deletion products from multimers composed of tandem repeats. In at least one case a plasmid, constructed by inserting an SmaI fragment of yeast ribosomal DNA into pSC101, underwent deletion and rearrangement to form a product in which a segment, consisting of part of the pSC101 sequence and part of the yeast ribosomal DNA sequence, was duplicated to form a tandem repeat. Deletion and rearrangement take place in Rec+, recA- and recB- recC- cells. The rate of deletion in Rec+ cells is higher than in recA- cells. The rate of deletion in minicell-producing, X-ray resistant strains is much higher than in other Rec+ strains.

Chromosome Deletion

Nucleotide sequence of the Hind-C fragment of simian virus 40 DNA. Comparison of the 5'-untranslated region of wild-type virus and of some deletion Mutants.

We report here the nucleotide sequence of the wild-type simian virus 40 (strain 776) restriction fragment Hind-C-P1 DNA and of the homologous region of various mutant DNAs which lack part of this fragment. During this work, we detected between EcoRII fragments N and G an additional, 17-base-pair EcoRII fragment, fragment P, which had previously been overlooked. Also, an additional dTpdG dinucleotide at residues L 339--340 was observed by sequence analysis of the DNA minus (E) strand; the presence of this dinucleotide was masked on sequencing patterns of the plus strand due to the persistence (during gel electrophoresis) of some secondary structures in the strand's 5'-terminal region. These nucleotide additions raise the total length of SV40 DNA to 5243 base pairs. The longest tandemly repeated segment in SV40 DNA now extends over 72 base pairs. SV40 deletion mutants dl 893 and dl 894 and SV40 strains Rh 911 and 1801 all lack an identical 72-base-pair-long DNA segment in the Hind-C region. This deletion corresponds precisely to one of the two aforementioned large tandemly repeated sequences. Mutant dl 895 lacks 66 base pairs, 63 of which are part of the former repetition. All these mutants, except dl 895, very probably were generated by an intramolecular, homologous recombination event. The 40-base-pair deletion in mutant dl 1811 includes the major capping site of SV40 late RNA. dl 1812 lacks only three base pairs, which are part of the overlapping HhaI and HpaII restriction sites at position 0.725--0.726.

Base Sequence

A genome-wide approach for the discovery of novel repeat expansion disorders in the Undiagnosed Diseases Network cohort.

PURPOSE: The Undiagnosed Diseases Network is a National Institutes of Health funded research study that aims to solve a broad clinical spectrum of challenging rare disease cases. Participants receive care from multiple clinical specialists, who collaborate to perform deep phenotyping and state-of-the-art multiomics analyses. As bioinformatics of short-read sequencing has matured, the discovery of repeat expansion disorders (REDs) is accelerating. REDs comprise approximately 60 characterized disorders, which exhibit a broad spectrum of phenotypes. Thus, a largely unbiased genome-wide approach in a phenotypically diverse sample will add to the diagnostic depth, explore the limits of short-read genome analysis, and establish novel candidate RED loci. METHODS: Here, we present a genome-wide analysis of repeat expansions conducted on 1018 genomes from the Undiagnosed Diseases Network. By leveraging 2 distinct bioinformatics tools, ExpansionHunter Denovo and STRling, we showed that repeat expansions can be accurately detected in short-read genomes. RESULTS: We demonstrated that a genotype-first approach can diagnose atypical cases of known REDs and provide valuable clinical insights. We present clinical details on participants with expansions in ATXN7, DMPK, FMR1, GLS, HTT, RFC1, AFF3, and MARCH6. Importantly, we highlight 2 cases of juvenile Huntington disease that were discovered through our analysis. Finally, we present a list of novel candidate short tandem repeats (TR) that could potentially be pathogenic if expanded. CONCLUSION: Importantly, our approach showcases the bioinformatic advancements in genome analysis for RED detection and highlights its practical applications.

Humans

Mitogenome assembly and phylogenetic relationships of Phalaris arundinacea.

INTRODUCTION: As a perennial herb of Poaceae, Phalaris arundinacea plays key roles in grazing, production, and soil and water conservation because of its well-developed rhizomes and seed dispersal. We assembled and annotated the first mitogenome of P. arundinacea to support evolutionary and taxonomic research. METHODS: We assembled and annotated the first complete mitochondrial genome of P. arundinacea by integrating Illumina short reads with Nanopore long reads via a hybrid assembly strategy. The genome architecture was comprehensively characterized, encompassing codon usage bias, repetitive sequence organization, and inter-organellar genetic exchange with the chloroplast genome. RESULTS AND DISCUSSION: Assembly of the P. arundinacea mitogenome revealed two circular structures with a combined length of 526,717 bp. The genome comprised a set of 37 protein-coding genes (PCGs), 27 tRNAs, and 8 rRNAs, with the rRNA genes exhibiting full assembly (100% coverage). The mitochondrial genome contained 154 forward and 164 palindromic repeats, along with 25 tandem repeats and 124 simple sequence repeats (SSRs). Notably, 102 SSRs were distributed on contig1, predominantly in tetrameric form. Furthermore, 376 RNA editing sites were predicted. A total of 104 fragments were integrated into the mitochondrial genome from the chloroplast, amounting to 55,866 bp of transferred sequence. Finally, phylogenetic analysis of 28 plant mitogenomes placed P. arundinacea closest to species within the genus Poa (P. chaixii and P. pratensis). Comparative analysis of non-synonymous-to-synonymous substitution rate (Ka/Ks) ratios across divergent species revealed that the mitochondrial genome of P. arundinacea underwent stabilizing evolutionary dynamics, characterized by predominant purifying selection with several lineage-specific variations in selective pressure. Our findings support the close phylogenetic relationship between P. arundinacea and species of the genus Poa and provide a reference mitochondrial genome resource for future comparative studies within Phalaris that incorporate broader taxon sampling. These results support deeper phylogenetic investigations of P. arundinacea and facilitate future work on its germplasm characterization and applied use.

Phalaris arundinacea

The 5S genes of Drosophila melanogaster.

We have cloned embryonic Drosophila DNA using the poly (dA-DT) connector method (Lobban and Kaiser, 1973) and the ampicillin-resistant plasmid pSF2124 (So, Gill and Falkow, 1975) as a cloning vehicle. Two clones, containing hybrid plasmids with sequences complementary to a 5S RNA probe isolated from Drosophila tissue culture cells, were identified by the Grunstein and Hogness (1975) colony hybridization procedure. One hybrid plasmid has a Drosophila insert which is comprised solely of tandem repeats of the 5S gene plus spacer sequences. The other plasmid contains an insert which has about 20 tandem 5S repeat units plus an additional 4 kilobases of adjacent sequences. The size of the 5S repeat unit was determined by gel electrophoresis and was found to be approximately 375 base pairs. We present a restriction map of both plasmids, and a detailed map of of the5S repeat unit. The 5S repat unit shows slight length and sequence heterogeneity. We present evidence suggesting that the 5S genes in Drosophila melanogaster may be arranged in a single continuous cluster.

Base Sequence