Search PubMedSearch

SEARCH · Search PubMed

Results for “Short tandem repeat”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Genetic inference in social insects: The continued utility of microsatellites in the sociogenomic era.

Social insects differ from many other biological systems because colonies function as integrated reproductive, ecological, and evolutionary units, often conceptualized as superorganisms. This organization makes genetic inference inherently hierarchical, often requiring genotyping across multiple biological levels: the colony, the population, the individual, and, in some cases, the cellular level. Although whole-genome sequencing and single-nucleotide polymorphism (SNP)-based approaches are now widely used in population genomics, microsatellites or short tandem repeats (STRs) remain a useful approach for cost-effective, low-input, and highly replicated genotyping, particularly in the hierarchical sampling designs common in social insect studies. Here, we review the utility and limitations of microsatellites in social insect research using a three-tiered framework spanning colony-, population-, and individual- or cellular-level analyses. Across these scales, microsatellites are especially valuable for colony delimitation, kinship inference, diagnostic screening of known reproductive systems, and low-input genotyping. By comparing the suitability of microsatellites with that of SNP-based and broader genomic approaches across these applications, this review links marker choice to biological scale, sampling design, and inferential goal in studies of social insects.

Journal Article

HLA-Typing of Donor-Origin Cells Enriched From Urine Cell Culture of Kidney Transplanted Recipients.

The incomplete or lack of histocompatibility information constitutes a barrier for the early detection and management of de novo donor-specific antibodies (DSA). To improve the quantity and quality of DNA materials for HLA typing, we developed a non-invasive culture-based method, using DNA extracted from enriched donor-derived kidney stem cells (DKSC) selectively cultured from the urine of kidney transplant receipients (KTR) to allow high-resolution typing by next-generation sequencing. This prospective proof-of-concept study evaluated the feasibility and performance of this approach. DKSC were enriched from the urine of 60 KTRs. DNA extracted from culture-enriched DKSC showed significantly higher concentration and better quality than unbound cells, and with identical short tandem repeat (STR) and 100% concordance compared to that obtained from peripheral blood. Our results suggest that cultured-enriched DKSC are non-invasive and useful for determining HLA and other genes for KTRs where donor information is limited or lacking.

Humans

Development of Microsatellite Marker System to Determine the Genetic Diversity of Experimental Chicken, Duck, Goose, and Pigeon Populations.

Poultries including chickens, ducks, geese, and pigeons are widely used in the biological and medical research in many aspects. The genetic quality of experimental poultries directly affects the results of the research. In this study, following electrophoresis analysis and short tandem repeat (STR) scanning, we screened out the microsatellite loci for determining the genetic characteristics of Chinese experimental chickens, ducks, geese, and pigeons. The panels of loci selected in our research provide a good choice for genetic monitoring of the population genetic diversity of Chinese native experimental chickens, ducks, geese, and ducks.

Animals

Identification of a novel non-coding deletion in Allan-Herndon-Dudley syndrome by long-read HiFi genome sequencing.

BACKGROUND: Allan-Herndon-Dudley syndrome (AHDS) is an X-linked disorder caused by pathogenic variants in the SLC16A2 gene. Although most reported variants are found in protein-coding regions or adjacent junctions, structural variations (SVs) within non-coding regions have not been previously reported. METHODS: We investigated two male siblings with severe neurodevelopmental disorders and spasticity, who had remained undiagnosed for over a decade and were negative from exome sequencing, utilizing long-read HiFi genome sequencing. We conducted a comprehensive analysis including short-tandem repeats (STRs) and SVs to identify the genetic cause in this familial case. RESULTS: While coding variant and STR analyses yielded negative results, SV analysis revealed a novel hemizygous deletion in intron 1 of the SLC16A2 gene (chrX:74,460,691 - 74,463,566; 2,876 bp), inherited from their carrier mother and shared by the siblings. Determination of the breakpoints indicates that the deletion probably resulted from Alu/Alu-mediated rearrangements between homologous AluY pairs. The deleted region is predicted to include multiple transcription factor binding sites, such as Stat2, Zic1, Zic2, and FOXD3, which are crucial for the neurodevelopmental process, as well as a regulatory element including an eQTL (rs1263181) that is implicated in the tissue-specific regulation of SLC16A2 expression, notably in skeletal muscle and thyroid tissues. CONCLUSIONS: This report, to our knowledge, is the first to describe a non-coding deletion associated with AHDS, demonstrating the potential utility of long-read sequencing for undiagnosed patients. Although interpreting variants in non-coding regions remains challenging, our study highlights this region as a high priority for future investigation and functional studies.

Humans

[Empirical Classification of Tri-Allelic Genotype Cases and Parentage Index Calculation].

OBJECTIVES: To standardize the calculation method of the parentage index (PI) for short tandem repeat (STR) tri-allelic genotypes, thereby ensuring the accuracy and reliability of parentage tes‑ ting conclusions. METHODS: A systematic analysis of 160 real cases was conducted. A classification system was constructed based on the occurrence mechanisms and inheritance patterns of STR tri-alleles, and the PI calculation method was optimized by integrating previous research findings with empirical data. RESULTS: A mechanism-based classification system for STR tri-allelic genotypes was established, comprising Type I (2 subtypes), Type II (6 subtypes), and the trisomic type (2 subtypes). On this basis, a standardized PI calculation method covering all categories of STR tri-allelic genotypes was developed. CONCLUSIONS: This study provides methodological guidance for the scientific and standardized calculation of PI for STR tri-allelic genotypes and offers an important reference for the formulation and refinement of relevant industry standards.

Humans

Invasive Wickerhamomyces anomalus Infections among Injecting Drug Users, France, 2012-20241.

Wickerhamomyces anomalus is a yeast rarely involved in human invasive fungal diseases (IFD). We retrospectively analyzed 44 episodes of W. anomalus IFD in France during 2012-2024. Injecting drug use (IDU) was the main risk factor among 26/35 (74.3%) incident cases. Most infections were community acquired; overall 3-month mortality rate was 1/30 (3.3%). Short tandem repeat (STR) genotyping and whole-genome sequencing analyses revealed substantial genetic diversity among isolates. However, 1 STR genotype was shared by 2 IDU patients, suggesting common exposure. In addition, 1 isolate obtained from a cotton filter used for drug preparation was identical by STR genotyping to the bloodstream isolate from the same patient, indicating direct inoculation via contaminated material or poor injection practices. Our findings highlight the increased risk for W. anomalus IFD among IDU patients and emphasize the importance of targeted preventive measures within that population.

Humans

Bridging the gap between legacy polymerase chain reaction-based microsatellite data with high-throughput sequencing data for conservation genomics.

Microsatellites are powerful markers for tracking genetic variation in wildlife populations due to their high polymorphism and genome-wide abundance. While polymerase chain reaction (PCR)-based fragment size analysis has been the standard for genotyping microsatellites, high-throughput sequencing offers greater resolution and the opportunity to sync historical datasets with modern analyses. We evaluated how genotypes from whole-genome sequencing align with PCR data for 15 microsatellite loci in 11 North American brown bears (Ursus arctos). Brown bear populations in the 48 contiguous United States have declined from approximately 50,000 to fewer than 2,000 over the past decades. Their endangered status has prompted extensive research and genetic monitoring, yielding large, multiyear microsatellite datasets upon which future conservation efforts can build. We achieved an overall microsatellite genotype concordance rate of 94.5% comparing high-throughput sequencing results to PCR based-fragment size results. All discrepancies occurred at complex loci containing multiple insertions and/or deletions (indels). Physically linked indels or single nucleotide polymorphisms (SNPs) occurring within the loci were misinterpreted as independent insertions, underscoring the need for genotyping tools that incorporate phasing when genotyping. To evaluate coverage effects, we downsampled high-throughput sequence data from 30x to 2x. Concordance remained high at 20 to 30x but dropped sharply at 10x, with 5x and 2x having discordant genotypes or insufficient coverage for genotyping. Accurate genotyping required both sufficient depth and number of reads spanning the entire repeat regions. Our results show that short-read whole-genome sequencing can recover microsatellite genotypes with high accuracy when paired with careful variant interpretation. By aligning historical PCR datasets with modern sequencing data, we can preserve decades of genetic insight and strengthen long-term monitoring of at-risk populations.

Animals

Complex structural variation, phylogeny, and disease associations of the mucin pangenome.

Mucins are large glycoproteins that provide hydration and barrier function to epithelial tissues. Although genetically heterogeneous, all mucins harbor a large exon composed of variable number tandem repeats (VNTRs). Short-read sequencing has limited our understanding of mucin VNTR diversity and makes disease association studies challenging. We leverage 296 long-read phased genome assemblies to characterize 14 mucin family members, achieving &#x2265;97% accuracy across 572 haplotypes. Phylogenetic haplogroup analysis reveals extraordinary structural heterozygosity, with MUC4 harboring the greatest allelic diversity (n=240 distinct lengths) and MUC12 the greatest size range (&#x394; = 55,233 bp; 23,080 amino acids). Ten mucins show significant population stratification (pFDR < 0.05). At the MUC4/MUC20 locus, we characterize higher-order structural variation, including a recurrent inversion, copy number variation, and interlocus gene conversion. Optimized genotyping achieves &#x2265;95% haplogroup concordance across 10 loci. We apply this to 4,637 deeply phenotyped cystic fibrosis patients and identify a significant association between short MUC1 VNTRs and severe disease (p=0.0056), demonstrating the pangenome's utility for complex locus genotyping and disease discovery.

Journal Article

Population-scale disease-associated tandem repeat analysis reveals locus and ancestry-specific insights.

Tandem repeat (TR) expansions, including short TRs (motifs &#x2264;6&#x2009;bp) and variable number TRs (motifs >6&#x2009;bp), underlie many monogenic disorders, with variable length and sequence influencing pathogenicity, penetrance, severity, and onset. Accurate genotype-phenotype correlation and disease prevalence estimation require characterization beyond repeat length. Here we present a population-scale analysis of 66 disease-associated TR loci using long-read assemblies from 2530 diverse haplotypes from 1265 unaffected donors. Integrating repeat length, motif composition, local ancestry, linkage disequilibrium, and phylogenetic analyses, we reveal extensive locus-, population-, and allele-specific variation shaping disease risk. Up to 8.5% of individuals carry expansions above established pathogenic thresholds, many containing interrupting motifs or sequence structures that attenuate pathogenicity. After excluding alleles from loci with uncertain disease association, non-pathogenic interrupted expansions, and carrier states inconsistent with inheritance patterns, ~4% carried expansions predicted to confer disease risk, largely at adult-onset loci with reduced penetrance. Ancestry-resolved analyses uncover population-specific TR architectures contributing to epidemiological disparities in repeat expansion disorders. Phylogenetic analyses identify conserved ancestral alleles and loci with recent instability. We describe variable linkage disequilibrium patterns and recombination signatures around specific disease-associated TR loci. Our findings emphasize integrating sequence, ancestry, and evolutionary context to understand the complex landscape of disease-associated TRs.

Humans

Tandem and inverted repeats of arginine genes in Escherichia coli: structural and evolutionary considerations.

Duplications of arg genes produced in the Rec+ and in the recA genetic backgrounds are shown by heteroduplex analysis to be strictly tandem at the level of resolution of this technique. The formation of these particular rearrangements therefore does not require the inclusion of transposons or other sequences of an appreciable size in their final structure. Duplications of short segments (about 2,000 nucleotides) appear unexpectedly stable when compared with duplications of longer segments (about 10,000 nucleotides). One of the structures analyzed displays two inversely repeated argE genes rearranged into an artificial divergent operon. The bearing of this observation on the origin of bipolar operons, of "mirror-image" map symmetries and on the production of inverted repeats in general, is discussed.

Arginine

Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing.

Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.

Journal Article

Polyoma virus giant RNAs contain tandem repeats of the nucleotide sequence of the entire viral genome.

The bulk of late virus-specific RNA synthesized in polyoma virus-infected mouse cells is larger than a single strand of poloma DNA. The arrangement of viral nucleotide sequences in these giant polyoma RNAs was studied by electron microscopy of hybrids between purified high molecular weight viral RNA and the HindII-1 fragment of polyoma DNA, which contains 91% of the viral genome. Hybrid molecules containing a short single-stranded gap (corresponding to the 9% of viral sequences not present in HindII-1), flanked by double-stranded regions, were photographed and measured. The majority of hybrid molecules contained no single-stranded loops or branches, showing that all viral sequences are transcribed contiguously and that no nonviral sequences are present in the RNA. Hybrid molecules, containing RNA up to 3.5 times the genome length, had a repeating structure of single-stranded gaps 8% of genome length interspersed with double-stranded regions 89% of genome length, showing that giant polyoma RNAs contain tandem repeats of the nucleotide sequence of the entire viral DNA. A small proportion of hybrid molecules contained single-stranded branches or deletion loops in characteristic positions, indicating that RNA "splicing" may occur on high molecular weight nuclear polyoma RNA.

Cell Nucleus

The DNA sequence of sea urchin (S. purpuratus) H2A, H2B and H3 histone coding and spacer regions.

The DNA sequence of two cloned segments of the histone gene repeat unit of the sea urchin S. purpuratus has been determined. One sequence contains the contiguous H2B and H3 genes and their interdigitated spacer regions; the other comprises the H2A gene and flanking spacer sequences. Analysis of the coding regions reveals a methionine residue within the H2A protein. H2A, which generally lacks this amino acid, contains methionine only in a protein variant which is synthesized in early sea urchin embryogenesis. We thus conclude that the cloned DNA represents a set of genes which is active early in development. Codon selection is markedly skewed and similar for each of the three genes. The DNA sequences are co-linear with known histone protein sequences and-unlike several other eucaryotic genes-do not show any insertions in the coding regions. The spacer regions are relatively AT-rich although GC cluster are scattered throughout. Several short stretches of homology are found in regions both upstream and downstream from the protein coding segments. The conservation of these sequences and their location at analogous sites suggest that they are involved in gene transcription or in mRNA translation. No tandem or dispersed repeats were found, with the exception of the remarkable sequence having the structure located in the spacer between the H2A and H1 genes.

Animals

Natural self-attenuation of pathogenic viruses by deleting the silencing suppressor coding sequence for long-term plant-virus coexistence.

Potyviridae is the largest family of plant-infecting RNA viruses. All members of the family (potyvirids) have single-stranded positive-sense RNA genomes, with polyprotein processing as the expression strategy. The 5'-proximal regions of all potyvirids, except bymoviruses, encode two types of leader proteases: the serine protease P1 and the cysteine protease HCPro. However, their arrangement and sequence composition vary greatly among genera or even species. The leader proteases play multiple important roles in different potyvirid-host combinations, including RNA silencing suppression and virus transmission. Here, we report that viruses in the genus Arepavirus, which encode two HCPro leader proteases in tandem (HCPro1-HCPro2), can naturally lose the coding sequences for these two proteins during infection. Notably, this loss is associated with a shift in foliage symptoms from severe necrosis to mild chlorosis or even asymptomatic infections. Further analysis revealed that the deleted region is flanked by two short repeated sequences in the parental isolates, suggesting that recombination during virus replication likely drives this genomic deletion. Reverse genetic approaches confirmed that the loss of leader proteases weakens RNA silencing suppression and other critical functions. A field survey of areca palm trees displaying varied symptom severity identified a transitional stage in which full-length viruses and deletion mutants coexist in the same tree. Based on these findings, we propose a scenario in which full-length isolates drive robust infections and facilitate plant-to-plant transmission, eventually giving rise to leader protease-less variants that mitigate excessive damage to host trees, allowing long-term coexistence with the perennial host. To our knowledge, this is the first report of potyvirid self-attenuation via coding sequence loss.

Plant Diseases

Effects of pentobarbital and d-amphetamine on the repeated acquisition of response sequences by pigeons.

Pigeons were trained to acquire a new 4-response sequence in each session by pecking three keys in a predetermined order. The key color varied for each step under the chained schedule, but there was only one key color under the tandem schedule. Under the reset contingency, incorrect responses produced a reset of the 4-response sequence to its beginning and a short timeout. In the non-resent contingency, only the timeout was produced by incorrect responses. Under both contingencies of both schedules, low doses (3-10 mg/kg) of pentobarbital increased the response rate and the total number of errors, although the rate increases usually occurred at lower doses than did the increases in errors. A dose of 17.5 mg/kg pentobarbital eliminated almost all responding. Injection of low doseas (0.1 -0.3 mg/kg) of d-amphetamine decreased the total number of errors under both contingencies of both the chained and the tandem schedules. Higher doses of d-amphetamine sometimes increased the total number of errors and decreased the response rate.

Animals

Approximating edit distances between complex tandem repeats efficiently.

MOTIVATION: Extended tandem repeats (TRs) have been associated with 60 or more diseases over the past 30&#x2009;years. Although most TRs have single repeat units (or motifs), complex TRs with different units have recently been correlated with some brain disorders. Of note, a population-scale analysis shows that complex TRs at one locus can be divergent, and different units are often expanded between individuals. To understand the evolution of high TR diversity, it is informative to visualize a phylogenetic tree. To do this, we need to measure the edit distance between pairs of complex TRs by considering duplication and contraction of units created by replication slippage. However, traditional rigorous algorithms for this purpose are computationally expensive. RESULTS: We here propose an efficient heuristic algorithm to estimate the edit distance with duplication and contraction of units (EDDC, for short). We select a set of frequent units that occur in given complex TRs, encode each unit as a single symbol, compress a TR into an optimal series of unit symbols that partially matches the original TR with the minimum Levenshtein distance, and estimate the EDDC between a pair of complex TRs from their compressed forms. Using substantial synthetic benchmark datasets, we demonstrate that the estimated EDDC is highly correlated with the accurate EDDC, with a Pearson correlation coefficient of >0.983, while the heuristic algorithm achieves orders of magnitude performance speedup. AVAILABILITY AND IMPLEMENTATION: The software program hEDDC that implements the proposed algorithm is available at https://github.com/Ricky-pon/hEDDC (DOI: 10.5281/zenodo.14732958).

Algorithms

Multiplexed genome editing by CRISPR-Un1Cas12f1 restores dystrophin expression in a mouse model of Duchenne muscular dystrophy.

The compact type V clustered regularly interspaced short palindromic repeats (CRISPR) nuclease Un1Cas12f1 is compatible with adeno-associated virus (AAV)-mediated genome editing, although the protospacer adjacent motif (PAM) requirements and capacity for multiplexed genome editing remain undefined. Here, we show that Un1Cas12f1 exhibits a broad tolerance for non-canonical PAMs, including Y-rich motifs with a preference for TTCR and TCTA PAMs, thereby expanding the genomic targeting range. We further demonstrate that a tandem sgRNA array expressed from a single transcript supports Un1Cas12f1-mediated multiplexed genome editing at up to five distinct genomic loci. Leveraging this multiplexing capability, we achieved targeted excision of the Dmd exon 23 through intramuscular delivery of an all-in-one AAV vector encoding Un1Cas12f1 and a CRISPR array. This treatment restored the disrupted open reading frame and dystrophin expression in a mouse model of Duchenne muscular dystrophy (DMD). Together, these findings establish Un1Cas12f1 as a compact CRISPR system capable of multiplexed genome editing and demonstrate its therapeutic potential for DMD.

Journal Article

Genomic and evolutionary analysis reveals dynamic variations of MKK3 gene, a key regulator for seed dormancy in barley.

Barley (Hordeum vulgare L.) is an important crop in the world, and its seed dormancy is primarily controlled by a mitogen-activated protein kinase kinase 3 (MKK3) gene. Although kinase activity of MKK3 and its roles in barley post-domestication have been widely studied, the pre-domestication evolution of MKK3 and the spread of nondormant alleles among global barley varieties remain largely unexplored. In this study, we analyzed MKK3 sequences in barley and its wild progenitor (Hordeum spontaneum K. Koch) and identified two polymorphic miniature inverted-repeat transposable elements&#xa0;(MITEs). Comparative analyses indicated that the insertions/excision of the MITEs predated the current estimates of barley domestication. Examination of the barley pangenomes coupled with droplet digital polymerase chain reaction revealed extensive copy number variation of MKK3 and suggested that transposons likely contributed to tandem amplification of the MKK3 gene on chromosome 5H. Additionally, approximately 1-Kb MKK3 sequences were found on chromosomes 1H and 6H. Further analysis indicated that these short MKK3 sequences were captured by a CACTA transposon that also contained fragments from four other expressed genes. The acquisition of MKK3 was estimated to be between 1.9 and 2.5 million years ago. Together, these findings illuminate the dynamic pre-domestication evolution of the MKK3 gene and identify three divergent MKK3 haplotype groups including a unique lineage predominant in Ethiopian germplasm. This study highlights the contribution of transposons to structural diversification and evolutionary differentiation of the MKK3 locus and provides helpful information for understanding the complex history of MKK3 gene in barley and also for improving preharvest sprouting tolerant varieties under distinct natural conditions.

Hordeum