Search PubMedSearch

SEARCH · Search PubMed

Results for “simple sequence repeat”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

The simple repeat poly(dT-dG).poly(dC-dA) common to eukaryotes is absent from eubacteria and archaebacteria and rare in protozoans.

Genomic DNA from a wide variety of prokaryotic and eukaryotic organisms has been assayed for the simple repeat sequence poly(dT-dG).poly(dC-dA) by Southern blotting and DNA slot blot hybridizations. Consistent with findings of others, we have found the simple alternating sequence to be present in multiple copies in all organisms in the animal kingdom (e.g., mammals, reptiles, amphibians, fish, crustaceans, insects, jellyfish, nematodes). The TG element was also found in lower eukaryotes (Saccharomyces cerevisiae, Neurospora crassa, and Dictyostelium discoideum) and at a much lower frequency in protozoans (Oxytricha fallux and Tetrahymena thermophila). The sequence was also repeated in high copy number in a higher plant (Zea mays) as well as at very high levels in a unicellular green alga (Chlamydomonas reinhardi). Although the copy number of the repeat per haploid genome was generally proportional to genome size, there was a greater-than-1,000-fold variation in the number of (TG)25/100-kb genomic DNA. By contrast, no eu-or archaebacterium--including Myxococcus xanthus, whose life cycle is very similar to that of the slime mold Dictyostelium discoideum, and Halobacter volcanii, whose genome contains other repeated sequences--was found whose genomic DNA contained this sequence in detectable amounts. A computer search also failed to find the TG element in human mitochondrial DNA.

Animals

Evolutionary dynamics of the chloroplast genome in Abutilon (Malvoideae, Malvaceae).

The genus Abutilon Mill. (Malvaceae) comprises approximately 178 species distributed across tropical and subtropical regions, many of which hold significant ornamental, economic, and medicinal value; yet its taxonomic classification remains challenging. In this study, six species were sequenced from herbarium specimens, and the chloroplast (cp.) genomes of ten additional species were assembled de novo from publicly available raw data. Three previously reported cp. genomes were also incorporated to characterise cp. genome structure, identify polymorphic loci, and perform phylogenetic analyses. The cp. genomes ranged from 159,458 to 160,454 bp and exhibited the typical quadripartite structure, with each genome containing 112 unique genes (78 protein-coding, 30 tRNA, and 4 rRNA) that showed conserved content and organisation. These genomes exhibited high similarity in GC content, inverted repeat boundaries, relative synonymous codon usage, amino acid frequencies, and substitution patterns. However, notable variation was observed in the total number of simple sequence repeats, ranging from 70 to 97 per genome. Selection analyses indicated predominant purifying selection, with evidence of episodic positive selection detected in rpoC2, rbcL, and ycf1. Two codons in rbcL were clade-specific and provided phylogenetic signal distinguishing Australian and Old World pantropical species. Nucleotide diversity analysis identified six highly polymorphic intergenic spacers (trnH-psbA, rps19-rpl2, psbT-pbf1, psaC-ndhD, trnR-atpA, and ndhJ-ndhK) that may be suitable for taxonomic studies. The phylogeny from maximum likelihood (ML) and Bayesian inference (BI) resolved two major clades: one comprising an exclusively Australian lineage occurring predominantly in arid and semi-arid environments, and the other a pantropical lineage spanning multiple continents. Abutilon grandifolium was recovered as sister to the remaining sampled Abutilon taxa in both ML and BI analyses, although no biogeographic origin inference can be drawn from this placement pending broader taxon sampling and integration of nuclear genomic data. These findings provide insights into the evolutionary dynamics of the cp. genome in Abutilon and offer a foundational genomic framework for refining Abutilon taxonomy.

Genome, Chloroplast

The Complete Chloroplast Genome and the Phylogenetic Analysis of Panicum bisulcatum (Thumb.) (Poaceae).

The chloroplast (cp) genome of Panicum bisulcatum (Thumb.), a significant agricultural weed, was sequenced and characterized to elucidate its genomic architecture, evolutionary dynamics, and phylogenetic relationships. The complete cp genome was assembled as a circular DNA molecule of 138,489 bp, exhibiting a typical quadripartite structure comprising a large single-copy (LSC, 82,260 bp), a small single-copy (SSC, 12,569 bp), and a pair of inverted repeats (IR, 21,830 bp each) regions. It encodes 135 genes, including 89 protein-coding genes, 49 tRNAs, and 8 rRNAs. Functional annotation revealed that most genes are involved in photosynthesis and genetic system. A total of 51 simple sequence repeats (SSRs) and 62 long repeats (LRs) were identified, providing potential molecular markers. Comparative analysis of IR boundaries highlighted both conserved features and species-specific expansion/contraction events among Panicum species. Phylogenomic analysis robustly placed P. bisulcatum within the genus Panicum, showing a closest relationship with P. incomtum and confirming the monophyly of the genus. Furthermore, single nucleotide polymorphism (SNP) analysis with its closest relative, P. incomtum, revealed 4659 SNPs, with a dominance of synonymous substitutions, indicating the action of purifying selection. This study provides the first comprehensive cp genomic resource for P. bisulcatum, which will facilitate future studies in species identification, phylogenetic reconstruction, population genetics, and the development of sustainable management strategies for this weed.

Phylogeny

Organization, structure, and function of 95 kb of DNA spanning the murine T-cell receptor C alpha/C delta region.

We have analyzed the organization, structure, and function of the murine T-cell receptor C alpha/C delta region. This region spans 94.6 kb of DNA and contains the C alpha and C delta genes, as well as the V delta 5, J delta 2, and 50 different J alpha gene segments. Within this sequence we have identified 15 new J alpha gene segments, 40 new 5' RNA splice signals, and 40 new DNA rearrangement signals for the J alpha gene segments. The murine C alpha/C delta sequence contains an exceptionally high level of coding sequence with over 5.7% of the total sequence found in the exons. This is much more than that found in the beta-globin locus and the HPRT locus. Using the sequence data obtained from the C alpha/C delta region, we have designed simple assays to test for J alpha gene segment transcription and to determine the level of polymorphism for simple repeat sequences among different inbred strains of mice using the polymerase chain reaction. Furthermore, comparisons of this 95 kb of sequence with the available sequence from homologous regions of other species have led to the identification of a highly conserved sequence that is present throughout vertebrates and in the mouse binds lymphocyte-specific nuclear proteins. Comparisons of a 10-kb region, which includes the C alpha gene in human and mouse, average 66% sequence similarity. These studies support the contention that large-scale DNA sequencing projects of homologous regions of mouse and human will provide powerful new tools for studying the biology and evolution of loci such as the T-cell receptor and for identifying and posing new questions about the functions of conserved sequences.

Amino Acid Sequence

Transcription of a satellite DNA on two Y chromosome loops of Drosophila melanogaster.

Primary spermatocyte nuclei of Drosophila melanogaster exhibit three giant lampbrush-like loops formed by the kl-5, kl-3 and ks-1 Y chromosome fertility factors. Detailed mapping of satellite DNA sequences along the Y chromosome has recently shown that AA-GAC satellite repeats are a significant component of the kl-5 and ks-1 loop-forming regions. To determine whether these simple repeated sequences are transcribed on the loop structures we performed a series of DNA-RNA in situ hybridization experiments to fixed loop preparations using as a probe cloned AAGAC repeats. These experiments showed that the probe hybridizes with homologous transcripts specifically associated with the kl-5 and ks-1 loops. These transcripts are detected at all stages of development of these two loops, do not appear to migrate to the cytoplasm and are degraded when loops disintegrate during the first meiotic prophase. Moreover, an examination of the testes revealed that the transcription of the AAGAC sequences is restricted to the loops of primary spermatocytes; the other cell types of D. melanogaster spermatogenesis do not exhibit nuclear or cytoplasmic labeling. These experiments were confirmed by RNA blotting analysis which showed that transcription of the AAGAC sequences occurs in wild-type testes but not in X/O testes. The patterns of hybridization to the RNA blots indicated that the transcripts are highly heterogeneous in size, from large (migration at limiting mobility) to less than 1 kb. We discuss the possible function of the AAGAC satellite transcripts, in the light of the available information on the Y chromosome loops of D. melanogaster.

Animals

Multiple forms of male-specific simple repetitive sequences in the genus Mus.

Previous reports indicate that in laboratory strains of mice, males are distinct from females in possession of repetitive DNA, notably devoid of Eco RI and Hae III sites and rich in the simple tetranucleotides GATA/GACA. We report here that such sequences originated in an ancestor common to laboratory mice, Mus hortulanus, M. spretus, and possibly also M. cookii. Interestingly, other male-specific satellite sequences were detected in M. caroli, M. cookii, M. saxicola, and M. minutoides. This novel satellite is also likely to be composed of simple repetitious sequences, but does not contain GATA and GACA. Thus, the Y chromosome appears to contain a disproportionately large amount of simple repetitious DNA. An attractive explanation for these results is that long tandem arrays of simple repeated sequences are generated at high frequency throughout the genome and that they are retained for a longer time on the Y chromosome due to the absence of homologous pairing at meiosis.

Animals

Towards construction of a high resolution map of the mouse genome using PCR-analysed microsatellites.

Fifty sequences from the mouse genome database containing simple sequence repeats or microsatellites have been analysed for size variation using the polymerase chain reaction and gel electrophoresis. 88% of the sequences, most of which contain the dinucleotide repeat, CA/GT, showed size variations between different inbred strains of mice and the wild mouse, Mus spretus. 62% of sequences had 3 or more alleles. GA/CT and AT/TA-containing sequences were also variable. About half of these size variants were detectable by agarose gel electrophoresis. This simple approach is extremely useful in linkage and genome mapping studies and will facilitate construction of high resolution maps of both the mouse and human genomes.

Animals

PCR-analyzed microsatellites of the mouse genome--additional polymorphisms among ten inbred mouse strains.

Eighty sequences from the mouse genome database containing microsatellites (simple sequence repeats) have been analyzed for size variation among ten different inbred strains of mice; 62/80 (77.5%) showed polymorphism of at least three alleles. We have been able to detect all the polymorphisms by agarose gel electrophoresis, often running the gels for up to 3 h. Between individual pairs of mouse strains to be used in chromosomal mapping studies in our laboratory, 35-60% polymorphism occurred. There are potentially enough microsatellites within the mouse and human genome to have a marker at every 1-cM distance. This simple approach will, therefore, continue to be useful in genome mapping studies, leading eventually to high-resolution maps of both the mouse and human genomes; this should allow for physical mapping and cloning of specific genes.

Animals

Intramolecular DNA triplexes, bent DNA and DNA unwinding elements in the initiation region of an amplified dihydrofolate reductase replicon.

The nucleotide sequence of 6.2 kb (1 kb = 10(3) base-pairs) of DNA that encompasses the earliest replicating portion of the amplified dihydrofolate reductase domains of CHOC 400 cells has been determined. Origin region DNA contains two AluI family repeats, a novel repetitive element (termed ORR-1), a TGGGT-rich region, and several homopurine/homopyrimidine and alternating purine/pyrimidine tracts, including an unusual cluster of simple repeating sequences composed of (G-C)5, (A-C)18, (A-G)21, (G)9, (CAGA)4, GAGGGAGAGAGGCAGAGAGGG, (A-G)27. Recombinant plasmids containing origin region sequences were examined for DNA structural conformations previously implicated in origin activation. Mung bean nuclease sensitivity assays for DNA unwinding elements show the preferred order of nuclease cleavage at neutral pH in supercoiled origin plasmids to be: (A-T)23 much greater than the (A-G) cluster much greater than (A)38 much greater than vector = (AATT)n. At acid pH, the hierarchy of cleavage preferences changes to: the (A-G) cluster much greater than (A-T)23 much greater than (AATT)n greater than vector = (A)38. A region of stably bent DNA was identified and shown not to be reactive in the mung bean nuclease unwinding assay at either acid or neutral pH. Intermolecular hybridization studies show that, in the presence of torsional stress at pH 5.2, the (A-G) cluster forms triple-stranded DNA. These results show that the origin region of an amplified chromosomal replicon contains a novel repetitive element and multiple sequence elements that facilitate DNA bending, DNA unwinding and the formation of intramolecular triple-stranded DNA.

Animals

Differential methylation of a retrotransposon upstream of a MYB gene causes variegation of lettuce leaves, which is abolished by the presence of an (AT)5 repeat in the promoter.

Variegation, a common phenomenon in plants, can be the result of several genetic, developmental, and physiological factors. Leaves of some lettuce cultivars exhibit dramatic red variegation; however, the genetic mechanisms underlying this variegation remain unknown. In this study, we cloned the causal gene for variegation on lettuce leaves and elucidated the underlying molecular mechanisms. Genetic analysis revealed that the polymorphism of variegated versus uniformly red leaves is caused by an "AT" repeat in the promoter of the RLL2A gene encoding a MYB transcription factor. Complementation tests demonstrated that the RLL2A allele (RLL2AV) with (AT)n repeat numbers other than five led to variegated leaves. RLL2AV was expressed in the red spots but not in neighboring green regions. This expression pattern was in concert with a relatively low level of methylation in a retrotransposon inserted in -761 bp of the gene in the red spots compared to high methylation of the retrotransposon in the green region. The presence of (AT)5 in the promoter region, however, stabilized the expression of RLL2A, resulting in uniformly red leaves. In summary, we identified a novel promoter mechanism controlling variegation through inconsistent levels of methylation and showed that the presence of a simple sequence repeat of specific size could stabilize gene expression.

Promoter Regions, Genetic

Structure of the major block of alphoid satellite DNA on the human Y chromosome.

Alphoid DNA is a family of tandemly repeated simple sequences found mainly at the centromeres of the chromosomes of many primates. This paper describes the structure of the alphoid DNA at the centromere of the human Y chromosome. We have used pulsedfield gradient gel electrophoresis, cosmid cloning and DNA sequencing to determine the organization of the alphoid DNA on each of the Y chromosomes present in two somatic cell hybrids. In each case there is a single major block of alphoid DNA. This is approximately 470,000 bases (475 kb) long on one chromosome and approximately 575 kb long on the other. Apart from the size difference, the structures of the two blocks and the surrounding sequences are very similar. However, one restriction enzyme, AvaII, detects two clusters of sites within one block but does not cleave the other. The alphoid DNA within each block is organized into tandemly repeating units, most of which are about 5.7 kb long. A few variant units present on one chromosome are about 6.0 kb long. These variants, like the AvaII site variants, are clustered. The 5.7 kb and 6.0 kb units themselves consist of tandemly repeating 170 base-pair subunits. The 6.0 kb unit has two more of these subunits than the 5.7 kb unit. Our results provide a basis for further structural analysis of the human Y chromosome centromeric region, and suggest that long-range structural polymorphisms of tandemly repeated sequence families may be frequent.

Base Sequence

Generation and characterization of a human chromosome 9 cosmid library.

A cosmid library has been constructed from the hamster-human hybrid cell line PK-87-9, which contains chromosome 9 as its sole known human component. Ten thousand colonies were produced, of which approximately 200, or 2%, contain human material. Fifty of these 200 were regionally mapped by an Alu-primed PCR product hybridization procedure. These cosmids were localized to all regions of chromosome 9, but were especially concentrated in the distal portion of 9q. The map location derived by the Alu-primed PCR product hybridization procedure was compared to the map location derived by fluorescent in situ hybridization. Assignment of chromosomal location by the two methods was correspondent in all but a few cases. The presumptive presence of HTF islands was investigated for 130 cosmids by digestion with the restriction enzyme NotI. Twenty percent of cosmids contained at least one NotI site. A number of simple sequence repeat polymorphisms identified from the cosmid set were characterized and will provide a link between the genetic and physical maps for this chromosome.

Animals

Heavy metal stress in native plant species: investigating phytoremediation potential through physiological and ISSR/SCoT molecular assessments.

In emerging countries, increased industrial activity has a significant impact on economic growth and urban development. However, the acceleration of industrial processes is accompanied by the release of contaminants such as heavy metals. According to the World Health Organization, one-fourth of all human diseases are caused by environmental contaminants, including heavy metals, which can impair numerous organs such as the neurological system, liver, and reproductive systems. This increased efforts to find effective and sustainable methods to remove heavy metals. Phytoremediation is an environmentally benign method of removing heavy metals using specific plants. Thus, from industrially contaminated locations, common native plant species of Lactuca serriola, Sisymbrium irio, Chenopodium murale, and Cynanchum acutum were selected for this study to assess the mechanisms of their molecular and physiological tolerance. Soil and plants were tested for heavy metals (Cd, Pb, and Cu), and contaminated locations were classified as low and highly polluted. Measurements were made of soluble sugar, protein, secondary metabolites, malondialdehyde, and H2O2. Additionally, inter simple sequence repeat (ISSR), start codon targeted (SCoT), and genomic template stability GTS were used. In heavily polluted areas, all plant species exhibit elevated amounts of sugar, proteins, H2O2, MDA, and secondary metabolites, while total phenolics showed a unique significant interaction (plant-location), where Cynanchum exhibited a hyper-stress phenolic accumulation to cope with toxicity, whereas Chenopodium maintained genomic stability with balanced phenolic level. Based on these findings, both Cynanchum acutum and Chenopodium murale demonstrate superior potential for phytoremediation and warrant further investigation for ecological restoration.

Heavy metal

The Drosophila fsh locus, a maternal effect homeotic gene, encodes apparent membrane proteins.

The maternal effect gene fsh is involved in the establishment of segments and the specification of their identities; the progeny of mutant females are missing portions of thoracic and abdominal segments, and may have homeotic transformations of third thoracic segments to second thoracic segments. The fsh locus interacts synergistically with loci such as Ubx and trx in the production of homeotic transformations. We have characterized cDNA clones corresponding to the major fsh transcripts expressed in ovaries and early embryos, and to a pupal transcript. The expression of fsh transcripts in ovaries is restricted to the germline; in developing embryos, transcripts are found throughout the cytoplasm. The different ovarian/embryonic transcripts (7.6 and 5.9 kb) are generated by use of alternative polyadenylation and splice sites. These transcripts encode two large predicted proteins of 110 and 205 kDa that have unusual amino acid compositions: 40% of the residues are glycine, alanine, or serine, and there are several regions of homopolymers and simple sequence repeats. Hydropathy analysis indicates that these proteins span the membrane. We suggest that the expression of fsh proteins in the membrane of the embryo is required for proper functioning of genes such as Ubx in the specification of segmental identity.

Amino Acid Sequence

Rapid genetic analysis of families with polycystic kidney disease 1 by means of a microsatellite marker.

Presymptomatic diagnosis of polycystic kidney disease 1 (PKD1) is possible by genetic linkage analysis with markers from both sides of the disease locus. The existing proximal markers are not informative in many families, so such analysis is difficult and time-consuming. We sought more useful length polymorphisms on the proximal side of the locus among simple sequence repeats (microsatellites). We identified two microsatellite polymorphisms that lie closer to the PKD1 locus than any previously described highly variable marker. One, SM7, is especially informative; we have found fourteen alleles and the observed heterozygosity in caucasians is 62.7%. Genetic linkage analysis in PKD1 families suggests that both of the markers lie proximal to the disease gene, closer than existing flanking markers. These polymorphisms can be simply assayed by polymerase chain reaction amplification of the variable regions, which generates DNA fragments that can be separated on non-denaturing acrylamide gels and directly examined after gel staining. This rapid, inexpensive, and non-radioactive method of linkage analysis allows the complete study of DNA samples within 8 h.

Alleles

Structure and in vitro transcription of a mouse B1 cluster containing a unique B1 dimer.

A highly repetitive DNA element located 950 bp upstream from a mouse U2 small nuclear RNA gene has been cloned and characterized. The repetitive element is composed of a simple sequence repeat and a cluster of three B1 sequences. Two of these B1 elements are arranged head-to-tail and are joined by an oligo(dA)-rich linker. This unique B1 dimer, comprised of 339 bp, resembles the dimeric structure of primate Alu-family sequences, particularly that of a prototypic human Alu element. The other B1 element within the mouse cluster is a typical monomeric unit. Transcription studies performed in HeLa cell extracts with deletion mutants of the B1 cluster reveal that the single B1 unit is expressed at least 50 times more efficiently than the B1 dimer region. Furthermore, the B1 dimer which contains mutations in the first polymerase III promoter region is not transcribed end-to-end. We conclude that this B1 dimer is unlikely to give rise to a new dimeric retroposon family in the mouse genome.

Animals

Human ribosomal DNA: conserved sequence elements in a 4.3-kb region downstream from the transcription unit.

The sequence of 4366 bp of nontranscribed spacer (NTS) human ribosomal DNA (rDNA) located downstream from the 3' end of the transcription unit has been determined. The NTS rDNA is rich in pyrimidine nucleotides (31% T and 30% C) that tend to occur on the coding strand in runs of simple sequence repeats. Other highly repetitive sequence elements are also represented, including tracts of (dA-dC)26 and (dG-dT)29 on the coding strand downstream from the putative termination of transcription. Still farther downstream, two Alu repeat sequences are found. Such sequences are also found in rat DNA at comparable locations, consistent with the possibility of a comparable functional role.

Base Composition

The generation of a library of PCR-analyzed microsatellite variants for genetic mapping of the mouse genome.

Forty-three sequences containing simple sequence repeats or microsatellites were generated from an M13 library of total genomic mouse DNA. These sequences were analyzed for size variation using the polymerase chain reaction and gel electrophoresis without the need for radiolabeling. Seventy-two percent of the sequences showed allelic size variations between different inbred strains of mouse and the wild mouse, Mus spretus; and 53% showed variation between inbred strains. Thirty-seven percent were variant between B6/J and DBA/2J, and 81% of these were resolved using minigel agarose electrophoresis alone. This approach is a useful way of generating the large number of variants that are needed to create high resolution maps of the mouse genome.

Animals