Search PubMedSearch

SEARCH · Search PubMed

Results for “tandem repeat”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Accurate detection of tandem repeats exposes ubiquitous reuse of biological sequences.

Tandem repetition is one of the major processes underlying genome evolution and phenotypic diversification. While newly formed tandem repeats are often easy to identify, it is more challenging to detect repeat copies as they diverge over evolutionary timescales. Existing programs for finding tandem repeats return markedly different results, and it is unclear which predictions are more correct and how much room remains for improvement. Here, we introduce DetectRepeats, a new method that uses empirical information about structural repeats to improve the accuracy of repeat detection. We show that DetectRepeats advances the state-of-the-art by finding highly divergent repeats with relatively few false positive detections. We apply DetectRepeats to genomes across the tree of life to discover an enrichment of detectable tandem repeats within different genes, genome regions, and taxa. Furthermore, we use phylogenetic reconciliation to determine that some tandem repeats continue to evolve through intra-repeat unit replacement. In this manner, tandem repeats serve as a renewable genetic resource offering a bountiful source of alternative genetic material. Our work unlocks the confident detection of ancient tandem repeats, opening a doorway to future discoveries. DetectRepeats is part of the DECIPHER package for the R programming language and available via Bioconductor.

Tandem Repeat Sequences

Analysis of targeted and whole genome sequencing of PacBio HiFi reads for a comprehensive genotyping of gene-proximal and phenotype-associated Variable Number Tandem Repeats.

Variable Number Tandem repeats (VNTRs) refer to repeating motifs of size greater than five bp. VNTRs are an important source of genetic variation, and have been associated with multiple Mendelian and complex phenotypes. However, the highly repetitive structures require reads to span the region for accurate genotyping. Pacific Biosciences HiFi sequencing spans large regions and is highly accurate but relatively expensive. Therefore, targeted sequencing approaches coupled with long-read sequencing have been proposed to improve efficiency and throughput. In this paper, we systematically explored the trade-off between targeted and whole genome HiFi sequencing for genotyping VNTRs. We curated a set of 10&#xa0;,&#xa0;787 gene-proximal (G-)VNTRs, and 48 phenotype-associated (P-)VNTRs of interest. Illumina reads only spanned 46% of the G-VNTRs and 71% of P-VNTRs, motivating the use of HiFi sequencing. We performed targeted sequencing with hybridization by designing custom probes for 9,999 VNTRs and sequenced 8 samples using HiFi and Illumina sequencing, followed by adVNTR genotyping. We compared these results against HiFi whole genome sequencing (WGS) data from 28 samples in the Human Pangenome Reference Consortium (HPRC). With the targeted approach only 4,091 (41%) G-VNTRs and only 4 (8%) of P-VNTRs were spanned with at least 15 reads. A smaller subset of 3,579 (36%) G-VNTRs had higher median coverage of at least 63 spanning reads. The spanning behavior was consistent across all 8 samples. Among 5,638 VNTRs with low-coverage (&#xa0;<&#xa0;15), 67% were located within GC-rich regions (&#xa0;>&#xa0;60%). In contrast, the 40X WGS HiFi dataset spanned 98% of all VNTRs and 49 (98%) of P-VNTRs with at least 15 spanning reads, albeit with lower coverage. Spanning reads were sufficient for accurate genotyping in both cases. Our findings demonstrate that targeted sequencing provides consistently high coverage for a small subset of low-GC VNTRs, but WGS is more effective for broad and sufficient sampling of a large number of VNTRs.

Minisatellite Repeats

GeomeTRe: accurate calculation of geometrical descriptors of tandem repeat proteins.

MOTIVATION: Structured tandem repeat proteins (STRPs) are characterized by preserved structural motifs arranged in a modular way. The structural and functional diversity of STRPs makes them particularly important for studying evolution and novel structure-function relationships, and ultimately for designing new synthetic proteins with specific functions. One crucial aspect of their classification is the estimation of geometrical parameters, which can provide better insight into their properties and the relationship between the spatial arrangement of repeated units and protein function. Calculating geometric descriptors for STRPs is challenging because naturally occurring repeats are not "perfect" and often contain insertions and deletions. Existing tools for predicting structural symmetry work well on simple cases but often fail for most natural proteins. RESULTS: Here, we present GeomeTRe, an algorithm that calculates geometrical descriptors such as curvature (yaw), twist (roll), and pitch for a protein structure with known repeat unit positions. The algorithm simulates the movement of consecutive units, identifies rotational axes, and calculates the corresponding Tait-Bryan angles. GeomeTRe's parameters can enhance STRP annotation and classification by identifying variations in geometric arrangements among different functional groups. The package is fast and suitable for processing large protein structure datasets when repeat region information (e.g. from RepeatsDB) is available. AVAILABILITY AND IMPLEMENTATION: GeomeTRe is available as a Python package; source code and documentation can be found at https://github.com/BioComputingUP/GeomeTRe.

Algorithms

Approximating edit distances between complex tandem repeats efficiently.

MOTIVATION: Extended tandem repeats (TRs) have been associated with 60 or more diseases over the past 30&#x2009;years. Although most TRs have single repeat units (or motifs), complex TRs with different units have recently been correlated with some brain disorders. Of note, a population-scale analysis shows that complex TRs at one locus can be divergent, and different units are often expanded between individuals. To understand the evolution of high TR diversity, it is informative to visualize a phylogenetic tree. To do this, we need to measure the edit distance between pairs of complex TRs by considering duplication and contraction of units created by replication slippage. However, traditional rigorous algorithms for this purpose are computationally expensive. RESULTS: We here propose an efficient heuristic algorithm to estimate the edit distance with duplication and contraction of units (EDDC, for short). We select a set of frequent units that occur in given complex TRs, encode each unit as a single symbol, compress a TR into an optimal series of unit symbols that partially matches the original TR with the minimum Levenshtein distance, and estimate the EDDC between a pair of complex TRs from their compressed forms. Using substantial synthetic benchmark datasets, we demonstrate that the estimated EDDC is highly correlated with the accurate EDDC, with a Pearson correlation coefficient of >0.983, while the heuristic algorithm achieves orders of magnitude performance speedup. AVAILABILITY AND IMPLEMENTATION: The software program hEDDC that implements the proposed algorithm is available at https://github.com/Ricky-pon/hEDDC (DOI: 10.5281/zenodo.14732958).

Algorithms

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats

AniAnn's: alignment-free annotation of tandem repeat arrays using fast average nucleotide identity estimates.

MOTIVATION: Satellite DNA has long posed challenges for genome assembly and analysis due to its low sequence complexity and poor mappability. These large heterochromatic arrays of tandem repeats are ubiquitous across eukaryotic genomes, yet remain understudied. Current methods for annotating satellite regions, and other classes of tandem repeat arrays, are limited in their ability to annotate divergent or novel sequences. RESULTS: In this work, we introduce AniAnn's, an algorithm for annotating large blocks of tandemly repeating DNAs. AniAnn's exploits the high Average Nucleotide Identity (ANI) shared between repeat units of the same array to quickly and accurately infer the boundaries of such arrays. We show that AniAnn's improves the annotation of satellites and other tandem repeats within a variety of plant and animal genomes, while requiring only a fraction of the runtime compared to previous approaches. We conclude by exploring several use cases of AniAnn's as a lightweight method for masking repeats prior to whole-genome alignment as well as the de novo annotation and classification of satellite repeats. AVAILABILITY: AniAnn's is open source software and available at github.com/marbl/anianns.

Algorithms

Population-scale disease-associated tandem repeat analysis reveals locus and ancestry-specific insights.

Tandem repeat (TR) expansions, including short TRs (motifs &#x2264;6&#x2009;bp) and variable number TRs (motifs >6&#x2009;bp), underlie many monogenic disorders, with variable length and sequence influencing pathogenicity, penetrance, severity, and onset. Accurate genotype-phenotype correlation and disease prevalence estimation require characterization beyond repeat length. Here we present a population-scale analysis of 66 disease-associated TR loci using long-read assemblies from 2530 diverse haplotypes from 1265 unaffected donors. Integrating repeat length, motif composition, local ancestry, linkage disequilibrium, and phylogenetic analyses, we reveal extensive locus-, population-, and allele-specific variation shaping disease risk. Up to 8.5% of individuals carry expansions above established pathogenic thresholds, many containing interrupting motifs or sequence structures that attenuate pathogenicity. After excluding alleles from loci with uncertain disease association, non-pathogenic interrupted expansions, and carrier states inconsistent with inheritance patterns, ~4% carried expansions predicted to confer disease risk, largely at adult-onset loci with reduced penetrance. Ancestry-resolved analyses uncover population-specific TR architectures contributing to epidemiological disparities in repeat expansion disorders. Phylogenetic analyses identify conserved ancestral alleles and loci with recent instability. We describe variable linkage disequilibrium patterns and recombination signatures around specific disease-associated TR loci. Our findings emphasize integrating sequence, ancestry, and evolutionary context to understand the complex landscape of disease-associated TRs.

Humans

Functions of tandem-repeat galectins and domain coordination governs galectin-4 activity in grass carp (Ctenopharyngodon idella).

Galectins are &#x3b2;-galactoside-binding lectins that play essential roles in innate immunity. Among them, tandem-repeat galectins (TrGals), typically composed of two distinct carbohydrate-recognition domains (CRDs) connected by a linker peptide, are well established as key regulators of pathogen recognition and host defense in mammals. However, their structural diversity and immunological functions in teleost fish remain poorly understood. In this study, five TrGals (Gal-4, Gal-8a, Gal-8b, Gal-9, and Gal-9like) were identified in grass carp. Sequence and structural analysis revealed that Gal-8a/b, Gal-9, and Gal-9like possess the canonical two-CRD architecture, whereas Gal-4 uniquely contains four highly similar tandem-repeat domains. All five TrGals were broadly expressed across examined tissues, with predominant expression in the liver. Upon Aeromonas hydrophila infection, Gal-4, Gal-8a, Gal-8b, and Gal-9 were rapidly up-regulated at early time points (3-6&#x202f;h). To elucidate the functional significance of CRD number, recombinant full-length CiGal-4 (CiGal4-full) and three truncated variants containing one, two, or three CRDs (CiGal4-1CRD, CiGal4-2CRD, and CiGal4-3CRD) were generated and systematically characterized. All recombinant proteins contained the conserved &#x3b2;-sheet structure typical of galectin CRDs. Functional assays revealed that CiGal4-full displayed the strongest growth-inhibitory activity against all tested bacteria, whereas CiGal4-1CRD showed the weakest effect. Notably, CiGal4-2CRD exhibited the most potent bactericidal activity, surpassing the full-length protein, while CiGal4-3CRD showed no further enhancement. CiGal4-full and CiGal4-2CRD showed superior carbohydrate-binding activities compared with the other variants. Collectively, these results reveal that CRD copy number alone does not linearly determine galectin function. Instead, domain organization and conformational coordination are critical for optimizing antimicrobial activity. This study provides new insights into the structure-function relationships and evolutionary diversification of galectins in teleosts and highlights their potential as novel antimicrobial and immunomodulatory agents in aquaculture.

Animals

Detection of short tandem repeats in the cattle genome: a comparison of bioinformatic tools.

BACKGROUND: Short tandem repeats (STRs) are repetitive DNA sequences with 1&#x2013;6 nucleotide repeat units, exhibiting high polymorphism due to varying repeat counts. STRs are more variable than SNPs and can cause genetic disorders. With population-scale cattle whole-genome sequencing data available, whole-genome STR identification has attracted new interest, but challenges remain due to the lack of standardized methods, sequencing data limitations, and the diversity of STR-calling tools. This study compared six STR-calling tools: HipSTR, GangSTR, and ExpansionHunter for short-read data, and Straglr, RepeatHMM, and LongTR for Oxford Nanopore (ONT) long-read data&#x2014;using sequences from five Holstein cattle (two parent&#x2013;offspring trios with a shared sire). This is the first cattle study to evaluate short- and long-read STR callers using both data types from the same animals. RESULTS: In short-read data, ExpansionHunter identified the highest number of polymorphic STRs (pSTRs) (327,690), followed by HipSTR (205,900) and GangSTR (110,680), with 93,023 loci detected by all three tools. In long-read data, LongTR detected 470,250 pSTRs, RepeatHMM 224,185, and Straglr 90,275, with only 33,253 loci shared among them. Mendelian consistency of STR genotypes in the trio offspring was high (>&#x2009;0.8) for all short-read tools, with HipSTR and GangSTR highest at 0.98. LongTR was the only long-read tool with high consistency (0.88). Short-read tools also showed higher concordance in STR genotypes among themselves than was observed among long-read tools. However, long-read tools had a clear advantage in detecting large STRs. Relative to computational efficiency, HipSTR and GangSTR (short-reads), and LongTR (long-reads) required less memory and shorter runtimes than the other tools. CONCLUSIONS: Tool selection is critical for accurate whole-genome STR identification in cattle. For short-read data, HipSTR showed relatively high Mendelian consistency and concordance compared to the other tools, while ExpansionHunter was able to detect longer STRs but with lower Mendelian consistency. For long-read data, LongTR demonstrated higher consistency and computational efficiency relative to the other tools. Based on these results, HipSTR and LongTR are suggested as preferred options for short-read and ONT long-read datasets, respectively, in cattle STR analysis. These recommendations are based on the metrics observed in this study, and confirmatory analyses across additional breeds, larger sample sizes, and validated truth sets are encouraged.

Animals

The first complete mitochondrial genome of Strigea falconis (Digenea: Strigeidae) reveals six tandemly repeated trnE-containing units and provides mt evidence for the non-monophyly of the family Strigeidae.

BACKGROUND: Phylogenetic relationships among members in the order Diplostomida remain contentious, with mitochondrial (mt) and nuclear genomic data often yielding conflicting topologies. A major limitation is the availability of only a few mt genomes from the type genus Strigea, hindering a robust test of the monophyly of the family Strigeidae and the order Diplostomida. RESULTS: The mt genome of S. falconis was completely sequenced for the first time, which was a circular molecule of 16,872&#xa0;bp in length, encoding the typical set of 36 mt genes and six duplicate tRNA-Glu genes. Notably, there were seven identical and consecutive tandem repeat units each consist of a 169&#xa0;bp non-coding region followed by a trnE gene in the newly assembled genome. Phylogenomic analyses based on concatenated predicted amino acid sequences of 12 proteins robustly placed S. falconis in the same clade as Apharyngostrigea pipientis. Crucially, the family Strigeidae was not recovered as monophyletic. Instead, two species within Strigeidae, Cardiocephaloides medioconiger and Cotylurus marcogliesei, clustered with representatives of Diplostomidae, providing mt evidence for the paraphyly of Strigeidae under the current sampling. CONCLUSIONS: The newly sequenced mt genome of S. falconis reveals a previously unreported six-copy tandem repeat of trnE-containing units among currently available diplostomoid mt genomes. Phylogenetic analyses based on mt protein-coding genes provide additional mt evidence that the family Strigeidae was not recovered as monophyletic under the present taxon sampling. However, because mt genomes represent a single maternally inherited linkage group, broader taxon sampling, independent nuclear phylogenomic data, and explicit sensitivity analyses will be required to confirm these relationships and guide any formal systematic revision.

Animals

Atypical short tandem repeat allelic patterns in a sexual assault case involving hematopoietic stem cell transplantation.

Short tandem repeat (STR) profiling is a cornerstone of forensic DNA analysis, particularly during criminal investigations. However, certain clinical conditions, such as bone marrow transplantation, can complicate interpretation. To illustrate the impact of allogeneic bone marrow transplantation on forensic STR analysis, this case report details a sexual assault investigation involving a female victim who had previously received a bone marrow transplant from a female donor. A female sexual assault victim underwent forensic examination, during which multiple biological swabs were collected. STR profiling was conducted on the victim's blood, fingernail, buccal, breast, hip, vulvar, vaginal, cervical samples, panty, and brassiere. As conflicting profiles were found, a detailed medical history was collected, and hair follicle analysis was performed to confirm the origin of the STR profiles. Blood STR profiling revealed a single female genotype, while multiple swabs, including vaginal and cervical samples, showed a second female STR profile alongside the first. Notably, the consistency and distribution of the mixed profile across samples reduced the likelihood of laboratory contamination. The medical history of the victim revealed a prior history of allogeneic bone marrow transplant. Hair follicle analysis identified the recipient's original genotype, confirming that the secondary STR profile originated from the donor. Bone marrow transplantation may result in a mixed STR profile, potentially leading to misidentification or misinterpretation of forensic evidence. Awareness of transplant history is crucial during forensic evaluations. However, such clinical history is currently not included in standard sexual-assault evidence forms, underscoring the need for procedural updates.

Female

pSTRminer: integrated bioinformatic software for genome-wide identification and population-scale evaluation of polymorphic short tandem repeats.

Animal forensic genetics plays a critical role in criminal investigations by providing crucial evidence through domestic animal individualization and wildlife species identification. While human forensic genetics benefits from standardized short tandem repeats (STR) genotyping systems, animal forensic applications encounter significant challenges, including the limited availability of validated STR markers, the prevalence of error-prone dinucleotide STRs (di-STRs), and insufficient integration of population data. To address these challenges, we developed pSTRminer, an integrated bioinformatic tool that automates genome-wide STR mining and polymorphism evaluation. By applying pSTRminer to domestic cattle (Bos taurus), we identified 775,444 STRs de novo from the reference genome and genotyped them using whole-genome sequencing data from 60 Chinese and 111 African cattle to evaluate polymorphism across diverse genetic backgrounds. This led to the development of the cattle STR database (CSDB), comprising loci with a genotyping success rate&#x2009;&#x2265;&#x2009;40% and polymorphism information content (PIC)&#x2009;&#x2265;&#x2009;0.5. Experimental validation of 30 randomly selected tetranucleotide STRs (tetra-STRs) and 33 di-STRs via next-generation sequencing in a local Chinese cattle population (n&#x2009;=&#x2009;145) confirmed marker reliability. Although tetra-STRs had lower average polymorphism levels, they exhibited significantly lower stutter ratios (p&#x2009;<&#x2009;0.05), providing a viable path for identifying discriminative markers with fewer artifacts. Systematic screening revealed that certain tetra-STRs could surpass di-STRs in polymorphism. In conclusion, pSTRminer provides a scalable framework for developing standardized STR panels, facilitating the identification of robust and informative markers in forensic applications.

Bioinformatic software

A novel relationship between time offsets in capillary electrophoresis and DNA sequence variations in short tandem repeats.

Next-generation sequencing (NGS) provides increased discriminatory power in forensic DNA analysis due to the detection of isoalleles. Differences in sequences between alleles allow for a second layer of differentiation between DNA contributors beyond the number of short tandem repeat (STR) repeat units. However, because NGS is a more time and resource-intensive analysis than conventional capillary electrophoresis (CE), laboratories may benefit from indicators that suggest NGS is likely to provide added value. This study examined whether CE migration offsets, measured as residuals in the OSIRIS analysis software, can differ significantly among STR isoalleles. Residuals represent the time offset between a sample allele peak and its corresponding allelic ladder peak. Paired CE and NGS data from 95 single source samples were analyzed for CE-based residual differences, as the NGS data provided the sequence information of the corresponding isoalleles. Residual values differed significantly among isoalleles at several STR loci. Statistically significant differences were identified at D16S539 and D3S1358, as well as at specific allele lengths within D12S391, D13S317, and D8S1179. These findings demonstrate that CE residual variation can reflect underlying STR sequence differences between contributors. In practice, residual-based metrics could help laboratories to identify casework reference samples where NGS is likely to provide additional discrimination, without the need for processing outside of a routine CE workflow. Due to the potentially large number of isoalleles, community wide efforts to aggregate CE residual differences versus isoallele sequences may be useful in the validation and implementation of this approach to add value to forensic DNA analyses.

Electrophoresis, Capillary

A comparison of software for analysis of rare and common short tandem repeat (STR) variation using human genome sequences from clinical and population-based samples.

Short tandem repeat (STR) variation is an often overlooked source of variation between genomes. STRs comprise about 3% of the human genome and are highly polymorphic. Some cause Mendelian disease, and others affect gene expression. Their contribution to common disease is not well-understood, but recent software tools designed to genotype STRs using short read sequencing data will help address this. Here, we compare software that genotypes common STRs and rarer STR expansions genome-wide, with the aim of applying them to population-scale genomes. By using the Genome-In-A-Bottle (GIAB) consortium and 1000 Genomes Project short-read sequencing data, we compare performance in terms of sequence length, depth, computing resources needed, genotyping accuracy and number of STRs genotyped. To ensure broad applicability of our findings, we also measure genotyping performance against a set of genomes from clinical samples with known STR expansions, and a set of STRs commonly used for forensic identification. We find that HipSTR, ExpansionHunter and GangSTR perform well in genotyping common STRs, including the CODIS 13 core STRs used for forensic analysis. GangSTR and ExpansionHunter outperform HipSTR for genotyping call rate and memory usage. ExpansionHunter denovo (EHdn), STRling and GangSTR outperformed STRetch for detecting expanded STRs, and EHdn and STRling used considerably less processor time compared to GangSTR. Analysis on shared genomic sequence data provided by the GIAB consortium allows future performance comparisons of new software approaches on a common set of data, facilitating comparisons and allowing researchers to choose the best software that fulfils their needs.

Humans

Sequencing the orthologs of human autosomal forensic short tandem repeats provides individual- and species-level identification in African great apes.

BACKGROUND: Great apes are a global conservation concern, with anthropogenic pressures threatening their survival. Genetic analysis can be used to assess the effects of reduced population sizes and the effectiveness of conservation measures. In humans, autosomal short tandem repeats (aSTRs) are widely used in population genetics and for forensic individual identification and kinship testing. Traditionally, genotyping is length-based via capillary electrophoresis (CE), but there is an increasing move to direct analysis by massively parallel sequencing (MPS). An example is the ForenSeq DNA Signature Prep Kit, which amplifies multiple loci including 27 aSTRs, prior to sequencing via Illumina technology. Here we assess the applicability of this human-based kit in African great apes. We ask whether cross-species genotyping of the orthologs of these loci can provide both individual and (sub)species identification. RESULTS: The ForenSeq kit was used to amplify and sequence aSTRs in 52 individuals (14 chimpanzees; 4 bonobos; 16 western lowland, 6 eastern lowland, and 12 mountain gorillas). The orthologs of 24/27 human aSTRs amplified across species, and a core set of thirteen loci could be genotyped in all individuals. Genotypes were individually and (sub)species identifying. Both allelic diversity and the power to discriminate (sub)species were greater when considering STR sequences rather than allele lengths. Comparing human and African great-ape STR sequences with an orangutan outgroup showed general conservation of repeat types and allele size ranges. Variation in repeat array structures and a weak relationship with the known phylogeny suggests stochastic origins of mutations giving rise to diverse imperfect repeat arrays. Interruptions within long repeat arrays in African great apes do not appear to reduce allelic diversity. CONCLUSIONS: Orthologs of most human aSTRs in the ForenSeq DNA Signature Prep Kit can be analysed in African great apes. Primer redesign would reduce observed variability in amplification across some loci. MPS of the orthologs of human loci provides better resolution for both individual and (sub)species identification in great apes than standard CE-based approaches, and has the further advantage that there is no need to limit the number and size ranges of analysed loci.

Animals

Polygenic variants in DNA repair genes are associated with neurodevelopmental disorders, regression and increased burdens of somatic variants and short tandem repeat expansions.

PURPOSE: Developmental regression, characterized by the loss of acquired milestones, occurs in some individuals with neurodevelopmental disorders (NDDs); yet, its molecular basis remains unclear. Studies suggest that DNA damage repair (DDR) genes, such as FAN1, may protect against neurological dysfunction by modulating the somatic stability of short tandem repeats (STRs). This study explores the contribution of DDR gene variants in NDD cases presenting with regression. METHODS: We analyzed 1087 NDD patients, focusing on those carrying variants in DDR genes and presenting regression. We assessed the sensitivity to DNA damage using mitomycin C on lymphoblastoid cells. Somatic variants and STR expansions were evaluated through high-depth short-read genome sequencing. To further investigate the pathogenetic role of STR expansions, we performed long-read genome sequencing on the most severely affected proband. RESULTS: Probands with regression carried multiple DDR gene variants, several within the Fanconi anemia pathway. Their lymphoblastoid cells showed increased sensitivity to mitomycin C-induced cytotoxicity compared with parental and control samples. Probands with severe phenotypes and regression exhibited an accumulation of somatic variants and STR instability, enriched in neurodevelopmental genes. CONCLUSION: Our findings suggest that polygenic DDR gene variants may contribute to developmental regression in NDDs by promoting the accumulation of somatic variants and STR expansions.

Humans

Heterochromatin-based silencing of a foreign tandem repeat in Drosophila melanogaster shows unusual biochemistry and temperature sensitivity.

Eukaryotic genomes are packaged into chromatin, a regulatory nucleoprotein assembly. Establishment, maintenance, and interconversion of chromatin states is required for correct patterns of gene expression, genome integrity, and survival. Transcriptionally repressive heterochromatin minimizes mobilization of transposable elements and limits expansion of other repetitive DNA, but mechanisms for recognition of the latter sequences are not well established. We previously demonstrated in Drosophila melanogaster that transcripts derived from 1360 and Invader4 transposon insertions can trigger local conversion of transcriptionally permissive euchromatin to heterochromatin through the piRNA system, but only in a subset of genomic locations near existing blocks of heterochromatin. Here we show that a ~9 kb tandem array of the 36-nucleotide lac operator (lacO) sequence of Escherichia coli can form ectopic heterochromatin at a similar subset of sites, resulting in variegating expression of an adjacent reporter gene. Heterochromatin Protein 1a (HP1a) and histone deacetylation are required for lacO repeat-induced silencing, but, contrasting with previously described Position Effect Variegation (PEV), we do not observe increased histone H3 lysine 9 methylation. Silencing is effective at 25&#xb0;C and suppressed at 18&#xb0;C (in contrast to canonical PEV, which is enhanced at 18&#xb0;C), indicating involvement of a temperature-sensitive component. Temperature switching experiments show that lacO repeat-induced heterochromatin formation is reversible throughout larval development following an HP1a-dependent initiation step in the early embryo. We conclude that the Drosophila nucleus can recognize a completely foreign tandem repeat as a target for heterochromatin formation, and that the heterochromatin structure established is distinct from that of endogenous tandem arrays.

HP1a

Assembly and comparative analysis of the mitochondrial genome of Pleione yunnanensis: genome structure and evolutionary insights.

BACKGROUND: Pleione yunnanensis a terrestrial or semi-epiphytic herbaceous plant belonging to the Orchidaceae family, is valued for both its medicinal uses and ornamental appeal. Although its chloroplast genomes have been sequenced, its complete mt genome had not previously been resolved, limiting genetic and evolutionary studies of the species. RESULTS: In this work, we assembled and characterized the first complete mt genome of P. yunnanensis, revealing a structurally complex, multibranched system composed of 14 circular-mapping molecules totaling 468,176&#xa0;bp with a GC content of 44.32%. The genome encodes 44 annotated genes, including 28 protein-coding genes (PCGs), 15 tRNAs, and one rRNA. The multibranched architecture provides new evidence supporting the dynamic and recombinational nature of plant mt genomes. Repeat analysis uncovered 29 simple sequence repeats (SSRs), 19 tandem repeats, and 118 dispersed repeats, indicating a comparatively lower repeat abundance than that found in closely related orchids with similar mt genome sizes. Codon-usage profiling of PCGs showed a marked bias toward A/T-ending codons. Prediction of RNA editing sites identified 4,708 putative edits across mitochondrial PCGs. Most mitochondrial genes displayed Ka/Ks ratios close to 1.0, suggesting relaxed selective constraints or lineage-specific evolutionary patterns rather than strong positive selection. Moreover, we detected 69 chloroplast-derived homologous fragments, including 15 intact genes, suggesting ongoing plastid-mitochondrial DNA transfer. Phylogenetic reconstruction and collinearity comparisons demonstrated that P. yunnanensis clustered closely with Dendrobium species, including D. amplum and D. hancockii, within the Orchidaceae clade. CONCLUSIONS: This study provides the first complete mt genome of P. yunnanensis, providing a foundational genomic resource for the genus Pleione. The results not only improve our understanding of mt genome structure and evolution in Orchidaceae, but also offer valuable molecular evidence for phylogenetic inference, germplasm identification, and conservation of this endangered medicinal species.

Orchidaceae