Search PubMedSearch

SEARCH · Search PubMed

Results for “Repetitive Sequences, Nucleic Acid”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Neotelomeres and telomere-spanning chromosomal arm fusions in cancer genomes revealed by long-read sequencing.

Alterations in the structure and location of telomeres are pivotal in cancer genome evolution. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeats, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. These results provide a framework for the systematic study of telomeric repeats in cancer genomes, which could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Humans

Satellite DNA evolution in Tytonidae (Aves: Strigiformes): dynamic repeat landscapes despite conserved karyotypes.

The elevated chromosome numbers observed in Tytonidae relative to the putative ancestral avian karyotype suggest that lineage-specific chromosomal fissions may have played an important role in the evolutionary history of this family. Here, we provide the first cytogenetic characterization of the American barn owl (Tyto furcata) and performs a comparative repeatome analysis across members of the Tytonidae, including other two species, the Western barn owl (Tyto alba), and the Oriental bay owl (Phodilus badius). The karyotype of T. furcata showed a 2n = 92, closely resembling that previously described for T. alba, indicating a high degree of chromosomal conservation within Tytonidae. Although T. furcata and T. alba exhibit similar karyotypic organization, comparative repeatomic analyses revealed differences in their composition, including variation in satellite DNA (satDNA) repertoires and abundance. Eight satDNA families were identified in T. furcata, nine in T. alba, and 28 in P. badius, highlighting the dynamic evolution of repetitive sequences. Several satDNA families were shared between T. furcata and T. alba, whereas some appeared species-specific, supporting the library hypothesis of satDNA evolution. In P. badius, multiple satDNAs exhibited similarity to transposable elements, suggesting that mobile elements contributed to their diversification. Cytogenetic analyses demonstrated centromeric heterochromatin distribution in T. furcata, as well as a large heterochromatic W chromosome enriched in DNA repeats. The localization of satDNAs in centromeric regions and the apparent accumulation of repeats on the W chromosome reinforce the role of repetitive sequences in chromosome organization and sex chromosome differentiation. Together, these findings reveal repeatome diversification despite conserved macrochromosomal structure and provide new insights into genome evolution and chromosomal dynamics in birds.

Animals

The genetic control of rapid genome content divergence in Arabidopsis thaliana.

Genome evolution in eukaryotes is predominantly driven by the dynamics of repetitive sequences, which vary widely in both copy number and sequence composition. Rates of repeat evolution differ between and within species and are likely modulated by both genetics and environment. To uncover factors shaping the rate of genome content evolution, we analyzed 1043 resequenced Arabidopsis thaliana genomes using a novel K-mer-based approach to characterize genome content variation and identify hypervariable regions underlying differences in repeat abundance. We next treated repeat abundance as a quantitative trait and performed genome-wide association analyses across more than 400 repeat families to identify the genetic basis of copy number variation. Integrating these results through a meta-GWAS approach revealed both cis-acting variants and more than 50 candidate trans-acting loci associated with repeat abundance genome-wide. Cis-acting variation was predominantly localized to pericentromeric and centromeric regions, whereas trans-acting loci were enriched for candidate genes involved in DNA replication, DNA repair, and DNA methylation regulation. The results are consistent with purifying selection acting against mutations that accelerate genome content divergence, favoring alleles that constrain repeat expansion. Together, these findings provide new insights into the genetic architecture and evolutionary forces shaping genome evolution in A. thaliana and establish a framework for investigating these processes in other plant species.

Arabidopsis

Repeats mimic pathogen-associated patterns across a vast evolutionary landscape.

An emerging hallmark of many human diseases is transcription of typically silenced repetitive DNA containing pathogen-associated molecular patterns (PAMPs). These PAMPs engage the innate immune system via pattern recognition receptors (PRRs)-a phenomenon known as viral mimicry. We propose a statistical physics framework to quantify viral mimicry by measuring "selective forces" that enrich PAMPs compared to a genome-wide reference distribution. We validate our predictions by identifying repeats that bind different PRRs and show potential viral mimics in different repeat families across eukaryotic genomes, suggesting shared mechanisms drive emergence and retention. We propose two non-exclusive evolutionary hypotheses. The first "repeat-centric" hypothesis posits PAMPs are integral to the repeat life cycle and are therefore enriched as they mediate repeat expansion. The second "organism-centric" hypothesis proposes viral mimicry functions as a cell-intrinsic feedback mechanism for sensing and reacting to transcriptional dysregulation, which provides a selective pressure to maintain PAMPs in genomes.

Humans

Repeat region engineering of Cas13a crRNA enables conformational gating-based autocatalytic CRISPR biosensing.

CrRNA engineering has emerged as a pivotal strategy for extending CRISPR-Cas13a biosensing. However, structural modulation of the direct repeat (DR) region remains exceptionally challenging due to its intricate architecture and the high energetic barrier of the Cas13a-crRNA interface, which is conventionally viewed as a rigid and immutable scaffold. Here, we demonstrate that the DR region is instead a programmable topological element with unexpected structural plasticity. By systematically engineering the DR through sequence insertion and structural splitting, we identified multiple DR variants that retain robust catalytic activity. Crucially, this topological reconfiguration enables Cas13a activity to be precisely gated by unmodified nucleic acid blockers, a level of regulation unattainable with the wild-type crRNA. Building on this flexible modulation, we developed Dre-CRISPR, a DR-engineered platform that couples target-triggered DR restoration to a self-reinforcing autocatalytic loop. This self-amplifying system provides a 2 × 106-fold sensitivity enhancement over nonamplified systems. Furthermore, the Dre-CRISPR platform extends the diagnostic scope of Cas13a to a broader spectrum of analytes, ranging from microRNAs to enzymatic activities and heavy metal ions. Our findings redefine the crRNA scaffold as a versatile signaling node and provide a generalizable framework for developing high-sensitivity, self-amplifying CRISPR biosensors through topology-driven guide RNA engineering.

CRISPR-Associated Proteins

Chromosome-level genome assembly with telomeric repeats at scaffold ends for Rhabdosargus sarba.

Rhabdosargus sarba, the goldlined seabream, is a euryhaline marine fish of great aquaculture potential. Genome sequencing and assembly of R. sarba was carried utilizing a multi-platform sequencing strategy that included long-read sequencing (PacBio HiFi), short-read sequencing (Illumina), and chromatin interaction mapping (Hi-C). The final genome assembly size after scaffolding was 764.59 Mb in 31 scaffolds with an N50 length of 33.98 Mb. Repeat profiling of primary assembly showed that 28.71% of the genome comprises of repeat elements. Gene prediction utilising the evidence from ab initio prediction and transcriptome data revealed 26,913 protein encoding genes and functional annotation and pathway analysis showed their participation in 332 pathways. This genome is an excellent resource for future research on genetic improvement and molecular breeding programmes for R. sarba.

Animals

Chromosome-level genome assembly of the ornamental plant Alcea rosea.

Alcea rosea, a member of the Malvaceae family, is celebrated for its rich floral palette and global horticultural significance. Here, we present a high-quality reference genome for A. rosea, achieving a genome assembly size of 1.01 Gbp, with a Contig N50 length of 36.61 Mbp. The genome sequence was successfully mapped to 21 chromosomes, and the scaffold N50 length reached 52.57 Mbp, with a scaffold genome completeness of 99.6%. A total of 565.84 Mbp (comprising 56% of the genome) of repetitive sequences were identified, with transposable elements being predominant, particularly long terminal repeat (LTR) elements, which accounted for 48.44% of the genome. 51,436 genes were annotated. Among these predicted genes, the average gene length and coding sequence (CDS) length were 2739.92 bp and 1242.54 bp, respectively.

Genome, Plant

A telomere-to-telomere gap-free genome assembly of the endangered humphead wrasse (Cheilinus undulatus).

Humphead wrasse, Cheilinus undulatus, is an endangered fish species with high economic and ecological value as well as natural sex change from female to male, while sexual selection occurs in breeding aggregations. In our present study, we constructed the first gap-free telomere-to-telomere (T2T) genome assembly for humphead wrasse, by integration of PacBio HiFi, ONT Ultra-long and Hi-C sequencing techniques. With 99% of the entire sequences anchored into 24 chromosomes, this haplotypic genome assembly spans approximately 1.25 Gb and presents a complete set of 48 telomeres and 24 centromeres. In terms of correctness (quality value QV: 53.447) and completeness (BUSCO score: 99.3%), this chromosome-scale assembly is indeed of high quality. We predicted 658.03 Mb of repetitive sequences and annotated 26,609 protein-coding genes in the assembled genome. This high-quality T2T genome assembly not only facilitates the genetic conservation of humphead wrasse, but also offers fundamental genomic data for supporting in-depth investigations on functional genomics, genetic diversity, and selective breeding for this economically important teleost.

Animals

Chromosome-level genome assembly of Nothapodytes nimmoniana.

Nothapodytes nimmoniana is a plant species belonging to the genus Nothapodytes in the family Icacinaceae. This species holds significant medicinal value due to its camptothecin content. In this study, we present the first chromosome-level genome assembly of N. nimmoniana constructed using NGS, Hi-C, and HiFi sequencing technologies. The assembled genome spans 3.53 Gb across 14 chromosomes, with an N50 length of 248.74 Mb. Genome annotation revealed that repetitive sequences constitute 80.82% of the genome size, and 83,269 protein-coding genes were predicted. Additionally, 4,360,538 bp of non-coding RNA were annotated. This genomic resource provides a foundation for further investigation into camptothecin biosynthesis pathways and plant phylogeny in N. nimmoniana.

Genome, Plant

Chromosome-level genome assembly of starry flounder (Platichthys stellatus).

Starry flounder (Platichthys stellatus) is widely distributed along the coastlines of the North Pacific. As an euryhaline flatfish, it can adapt to a wide range of environmental salinity ranging from freshwater to seawater, and is a promising aquaculture flatfish species in Korea and North China. However, no high-quality starry flounder reference genome has been reported to date, which greatly limits the studies of genetics and functional genomics. Here, we obtained a high-quality chromosome-level starry flounder genome assembly with a length of 643.56 Mb (scaffold N50: 26.19 Mb, contig N50: 10.00 Mb) combining short-reads sequencing, PacBio HiFi sequencing, and Hi-C sequencing. Approximately 94.02% of assembled sequences were anchored into 24 pseudochromosomes, and a total of 18 telomeres were detected. Totally 22,835 protein-coding genes and 227.87 Mb repetitive sequences were identified. In summary, the high-quality chromosome-level genome assembly not only provides valuable resources for genetic research in starry flounder, but also advances the development of molecular breeding technology of starry flounder.

Animals

Chromosome-level Genome Assembly of the Halophytic Turfgrass Zoysia macrostachya.

Zoysia macrostachya Franch. & Sav. is a halophytic perennial turfgrass in the Poaceae family, commonly found in the coastal regions of Korea, Japan, and East Asia. Z. macrostachya thrives in high-salinity environments, making it an excellent model for studying abiotic stress resilience. In this study, we present a chromosome-level genome assembly of Z. macrostachya, constructed using Oxford Nanopore long reads, Illumina short reads, and Omni-C sequencing data. The assembly spans 329.78 Mb across 20 chromosomes, with a scaffold N50 of 19.24 Mb, and includes complete telomeric sequences at both ends. The assembly showed 97.8% complete BUSCOs, indicating high genome completeness. Repeat element and gene annotation identified 44.03% of the genome as repetitive elements and 33,474 protein-coding genes. The gene annotation showed 97.1% complete BUSCOs and 86.92% functionally characterized genes. Macrosynteny analysis highlighted highly collinear relationships with related species, providing a foundational understanding of the Z. macrostachya genomic structure. This high-quality genome serves as a valuable resource for advancing salinity tolerance research and improving the genetic diversity of Zoysia species.

Genome, Plant

tidk: a toolkit to rapidly identify telomeric repeats from genomic datasets.

SUMMARY: "tidk" (short for telomere identification toolkit) uses a simple, fast algorithm to scan long DNA reads for the presence of short tandemly repeated DNA in runs, and to aggregate them based on canonical DNA string representation. These are telomeric repeat candidates. Our algorithm is shown to be accurate in genomes for which the telomeric repeat unit is known and is tested across a wide variety of newly assembled genomes to uncover new telomeric repeat units. Tools are provided to identify telomeric repeats de novo, scan genomes for known telomeric repeats, and to visualize telomeric repeats on the assembly. "tidk" is implemented in Rust and is available as a command line tool which can be compiled using the Rust toolchain or downloaded as a binary from bioconda. AVAILABILITY AND IMPLEMENTATION: The "tidk" Rust crate is freely available under the MIT license (https://crates.io/crates/tidk), and the source code is available at https://github.com/tolkit/telomeric-identifier.

Telomere

ECHO: a nanopore sequencing-based workflow for (epi)genetic profiling of the human repeatome.

SUMMARY: The human genome is dominated by repetitive DNA, whose genetic and epigenetic variation plays a key role in gene regulation, genome stability, and disease. Recent advances in long-read sequencing now enable large-scale, haplotype-resolved, and DNA methylation-informative analysis of the human genome, including on previously inaccessible complex and repetitive regions. However, the comprehensive, simultaneous characterisation of the "human repeatome" remains challenging, largely due to the lack of comprehensive tools integrated in a single pipeline that can capture the full spectrum of variation across diverse types of DNA repeats. Here, we present ECHO, a user-friendly, Snakemake-based pipeline for the "(Epi)genomic Characterisation of Human Repetitive Elements using Oxford Nanopore Sequencing." ECHO provides a reproducible and scalable framework for end-to-end analysis of whole-genome nanopore sequencing data, enabling integrative but also tailored (epi)genetic analyses of the human repeatome. AVAILABILITY AND IMPLEMENTATION: ECHO is freely available at Github: https://github.com/leenput/ECHO-pipeline, with the archived version at Zenodo: https://zenodo.org/records/19068468.

Humans

The dark genome in cardiovascular medicine.

Only ∼1%-2% of the human genome directly codes for proteins. The remainder consists of non-coding DNA, often referred to as the 'dark genome'. This includes regulatory elements, transposable and repetitive sequences, structural genomic features, pseudogenes, intronic and intergenic regions, and non-coding RNA (ncRNA) genes. These components are increasingly recognized as major regulators of gene expression, cell identity, and disease susceptibility. Currently, dark genome elements, particularly ncRNAs are increasingly recognized as important regulators of cardiovascular health and disease. Advances in genome analysis technologies have greatly improved our understanding of these non-coding regions and revealed clearer connections between the dark genome and cardiovascular traits. This review highlights major parts of the dark genome involved in cardiovascular disease, with emphasis on those for which mechanistic understanding and translational relevance are beginning to emerge. As mechanistic insight into individual and collective components of the dark genome advances, it increasingly enables the development of new opportunities for targeted therapeutics for cardiovascular prevention and disease management.

Humans

A high resolution A-to-I editing map in the mouse identifies editing events controlled by pre-mRNA splicing.

Pre-mRNA-splicing and adenosine to inosine (A-to-I) RNA-editing occur mostly cotranscriptionally. During A-to-I editing, a genomically encoded adenosine is deaminated to inosine by adenosine deaminases acting on RNA (ADARs). Editing-competent stems are frequently formed between exons and introns. Consistently, studies using reporter assays have shown that splicing efficiency can affect editing levels. Here, we use Nascent-seq and identify ∼90,000 novel A-to-I editing events in the mouse brain transcriptome. Most novel sites are located in intronic regions. Unlike previously assumed, we show that both ADAR (ADAR1) and ADARB1 (ADAR2) can edit repeat elements and regular transcripts to the same extent. We find that inhibition of splicing primarily increases editing levels at hundreds of sites, suggesting that reduced splicing efficiency extends the exposure of intronic and exonic sequences to ADAR enzymes. Lack of splicing factors NOVA1 or NOVA2 changes global editing levels, demonstrating that alternative splicing factors can modulate RNA editing. Finally, we show that intron retention rates correlate with editing levels across different brain tissues. We therefore demonstrate that splicing efficiency is a major factor controlling tissue-specific differences in editing levels.

Adenosine Deaminase

The 22q11 low copy repeats are characterized by unprecedented size and structural variability.

Low copy repeats (LCRs) are recognized as a significant source of genomic instability, driving genome variability and evolution. The Chromosome 22 LCRs (LCR22s) mediate nonallelic homologous recombination (NAHR) leading to the 22q11 deletion syndrome (22q11DS). However, LCR22s are among the most complex regions in the genome, and their structure remains unresolved. The difficulty in generating accurate maps of LCR22s has also hindered localization of the deletion end points in 22q11DS patients. Using fiber FISH and Bionano optical mapping, we assembled LCR22 alleles in 187 cell lines. Our analysis uncovered an unprecedented level of variation in LCR22s, including LCR22A alleles ranging in size from 250 to 2000 kb. Further, the incidence of various LCR22 alleles varied within different populations. Additionally, the analysis of LCR22s in 22q11DS patients and their parents enabled further refinement of the rearrangement site within LCR22A and -D, which flank the 22q11 deletion. The NAHR site was localized to a 160-kb paralog shared between the LCR22A and -D in seven 22q11DS patients. Thus, we present the most comprehensive map of LCR22 variation to date. This will greatly facilitate the investigation of the role of LCR variation as a driver of 22q11 rearrangements and the phenotypic variability among 22q11DS patients.

22q11 Deletion Syndrome

Discovering human transcription factor physical interactions with genetic variants, novel DNA motifs, and repetitive elements using enhanced yeast one-hybrid assays.

Identifying transcription factor (TF) binding to noncoding variants, uncharacterized DNA motifs, and repetitive genomic elements has been technically and computationally challenging. Current experimental methods, such as chromatin immunoprecipitation, generally test one TF at a time, and computational motif algorithms often lead to false-positive and -negative predictions. To address these limitations, we developed an experimental approach based on enhanced yeast one-hybrid assays. The first variation of this approach interrogates the binding of >1000 human TFs to repetitive DNA elements, while the second evaluates TF binding to single nucleotide variants, short insertions and deletions (indels), and novel DNA motifs. Using this approach, we detected the binding of 75 TFs, including several nuclear hormone receptors and ETS factors, to the highly repetitive Alu elements. Further, we identified cancer-associated changes in TF binding, including gain of interactions involving ETS TFs and loss of interactions involving KLF TFs to different mutations in the TERT promoter, and gain of a MYB interaction with an 18-bp indel in the TAL1 superenhancer. Additionally, we identified TFs that bind to three uncharacterized DNA motifs identified in DNase footprinting assays. We anticipate that these enhanced yeast one-hybrid approaches will expand our capabilities to study genetic variation and undercharacterized genomic regions.

Algorithms

Strain-specific differences in Neisseria gonorrhoeae associated with the phase variable gene repertoire.

BACKGROUND: There are several differences associated with the behaviour of the four main experimental Neisseria gonorrhoeae strains, FA1090, FA19, MS11, and F62. Although there is data concerning the gene complements of these strains, the reasons for the behavioural differences are currently unknown. Phase variation is a mechanism that occurs commonly within the Neisseria spp. and leads to switching of genes ON and OFF. This mechanism may provide a means for strains to express different combinations of genes, and differences in the strain-specific repertoire of phase variable genes may underlie the strain differences. RESULTS: By genome comparison of the four publicly available neisserial genomes a revised list of 64 genes was created that have the potential to be phase variable in N. gonorrhoeae, excluding the opa and pilC genes. Amplification and sequencing of the repeat-containing regions of these genes allowed determination of the presence of the potentially unstable repeats and the ON/OFF expression state of these genes. 35 of the 64 genes show differences in the composition or length of the repeats, of which 28 are likely to be associated with phase variation. Two genes were expressed differentially between strains causing disseminated infection and uncomplicated gonorrhoea. Further study of one of these in a range of clinical isolates showed this association to be due to sample size and is not maintained in a larger sample. CONCLUSION: The results provide us with more evidence as to which genes identified through comparative genomics are indeed phase variable. The study indicates that there are large differences between these four N. gonorrhoeae strains in terms of gene expression during in vitro growth. It does not, however, identify any clear patterns by which previously reported behavioural differences can be correlated with the phase variable gene repertoire.

Bacterial Proteins