Search PubMedSearch

SEARCH · Search PubMed

Results for “repetitive element”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Discovering human transcription factor physical interactions with genetic variants, novel DNA motifs, and repetitive elements using enhanced yeast one-hybrid assays.

Identifying transcription factor (TF) binding to noncoding variants, uncharacterized DNA motifs, and repetitive genomic elements has been technically and computationally challenging. Current experimental methods, such as chromatin immunoprecipitation, generally test one TF at a time, and computational motif algorithms often lead to false-positive and -negative predictions. To address these limitations, we developed an experimental approach based on enhanced yeast one-hybrid assays. The first variation of this approach interrogates the binding of >1000 human TFs to repetitive DNA elements, while the second evaluates TF binding to single nucleotide variants, short insertions and deletions (indels), and novel DNA motifs. Using this approach, we detected the binding of 75 TFs, including several nuclear hormone receptors and ETS factors, to the highly repetitive Alu elements. Further, we identified cancer-associated changes in TF binding, including gain of interactions involving ETS TFs and loss of interactions involving KLF TFs to different mutations in the TERT promoter, and gain of a MYB interaction with an 18-bp indel in the TAL1 superenhancer. Additionally, we identified TFs that bind to three uncharacterized DNA motifs identified in DNase footprinting assays. We anticipate that these enhanced yeast one-hybrid approaches will expand our capabilities to study genetic variation and undercharacterized genomic regions.

Algorithms

Nuclear body assembly by a viral repeat RNA promotes Kaposi's sarcoma-associated herpesvirus gene expression.

Kaposin is the most abundantly expressed viral RNA in tumors caused by the oncogenic virus Kaposi's sarcoma-associated herpesvirus (KSHV); however, its role in viral replication is not understood. Here, we show that kaposin, previously viewed as a protein-coding transcript, exists primarily as a nuclear viral long non-coding RNA (lncRNA) that rebuilds cellular nuclear speckles (NSs) adjacent to the viral genome to enhance viral gene expression. Kaposin is both necessary and sufficient to drive substantial NS remodeling, and this effect depends on repetitive elements within the RNA. Absence of kaposin-mediated NS remodeling, depletion of the essential NS protein, serine/arginine repetitive matrix 2 (SRRM2), or steric blocking of the kaposin repetitive elements impair viral gene expression. This work defines kaposin as a viral architectural RNA that drives nuclear speckle seeding beside the viral genome and reframes our understanding of lncRNA function and the spatial organization of transcription in the infected cell nucleus.

Kaposi's sarcoma-associated herpesvirus

Enhancer activation from transposable elements in extrachromosomal DNA.

Extrachromosomal DNA (ecDNA) drives oncogene amplification and intratumoral heterogeneity in aggressive cancers. While transposable element (TE) reactivation is common in cancer, its role on ecDNA remains unexplored. Here, we map the 3D architecture of MYC-amplified ecDNA in colorectal cancer cells and identify 68 ecDNA-interacting elements (EIEs)-genomic loci enriched for TEs that are frequently integrated onto ecDNA. We focus on an L1M4a1#LINE/L1 fragment co-amplified with MYC, which functions only in the ecDNA amplified context. Using CRISPR-CATCH, CRISPR interference, and reporter assays, we confirm its presence on ecDNA, enhancer activity, and essentiality for cancer cell fitness. These findings reveal that repetitive elements can be reactivated and co-opted as functional rather than inactive sequences on ecDNA, potentially driving oncogene expression and tumor evolution. Our study uncovers a mechanism by which ecDNA harnesses repetitive elements to shape cancer phenotypes, with implications for diagnosis and therapy.

Journal Article

Nanobioreactor detection of space-associated hematopoietic stem and progenitor cell aging.

Human hematopoietic stem and progenitor cell (HSPC) fitness declines following exposure to stressors that reduce survival, dormancy, telomere maintenance, and self-renewal, thereby accelerating aging. While previous National Aeronautics and Space Administration (NASA) research revealed immune dysfunction in low-earth orbit (LEO), the impact of spaceflight on human HSPC aging had not been studied. To study HSPC aging, our NASA-supported Integrated Space Stem Cell Orbital Research (ISSCOR) team developed bone marrow niche nanobioreactors with lentiviral bicistronic fluorescent, ubiquitination-based cell-cycle indicator (FUCCI2BL) reporter for real-time HSPC tracking in artificial intelligence (AI)-driven CubeLabs. In month-long International Space Station (ISS) missions (SpX-24, SpX-25, SpX-26, and SpX-27) compared with ground controls, FUCCI2BL reporter, whole-genome and transcriptome sequencing, and cytokine arrays demonstrated cell-cycle, inflammatory cytokine, mitochondrial gene, human repetitive element, and apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like 3 (APOBEC3) deregulation together with clonal hematopoietic mutations. Furthermore, HSPC functionally organized multi-omics aging (HSPC-FOMA) analyses revealed reduced telomere maintenance, adenosine deaminase acting on RNA1 (ADAR1) p150 self-renewal gene expression, and replating capacity indicative of space-associated HSPC aging that may limit long-duration spaceflight.

Humans

Chromosome-level Genome Assembly of the Halophytic Turfgrass Zoysia macrostachya.

Zoysia macrostachya Franch. & Sav. is a halophytic perennial turfgrass in the Poaceae family, commonly found in the coastal regions of Korea, Japan, and East Asia. Z. macrostachya thrives in high-salinity environments, making it an excellent model for studying abiotic stress resilience. In this study, we present a chromosome-level genome assembly of Z. macrostachya, constructed using Oxford Nanopore long reads, Illumina short reads, and Omni-C sequencing data. The assembly spans 329.78 Mb across 20 chromosomes, with a scaffold N50 of 19.24 Mb, and includes complete telomeric sequences at both ends. The assembly showed 97.8% complete BUSCOs, indicating high genome completeness. Repeat element and gene annotation identified 44.03% of the genome as repetitive elements and 33,474 protein-coding genes. The gene annotation showed 97.1% complete BUSCOs and 86.92% functionally characterized genes. Macrosynteny analysis highlighted highly collinear relationships with related species, providing a foundational understanding of the Z. macrostachya genomic structure. This high-quality genome serves as a valuable resource for advancing salinity tolerance research and improving the genetic diversity of Zoysia species.

Genome, Plant

Transposable Element-Mediated Cis-Regulation Drives the Evolution of dmrt1 as a Candidate Master Sex-Determining Gene in Black Carp.

Sex determination in vertebrates exhibits remarkable evolutionary plasticity, with diverse mechanisms and master sex-determining (MSD) genes arising independently across lineages. Among these, dmrt1, a dosage-sensitive gene, has repeatedly been recruited as an MSD gene through gene duplication or allelic diversification. However, the biochemical basis of such evolutionary transitions, particularly those driven by allelic diversification, remains largely unexplored. Here, we generated haplotype-resolved genome assemblies for both XX and XY black carp (Mylopharyngodon piceus) and identified a ∼40-kb region on chromosome 4, containing only dmrt1, as the candidate sex-determining locus. We discovered two Y-specific insertions in the dmrt1 promoter: a 13.4-kb highly repetitive element and an 11-bp motif. Functional assays revealed that these insertions act as enhancer and a promoter element, respectively, driving early, allele-specific upregulation of dmrt1 prior to gonadal differentiation. Notably, the 13.4-kb insertion contains transposable elements (TEs) functioning as cis-regulatory modules with transcription factor binding sites that mediate Y-specific activation. Our findings reveal a TE-mediated regulatory innovation that promoted dmrt1's evolution as a male-determining gene via allelic diversification, providing new insights into how mobile genetic elements drive the origin and diversification of sex-determining systems in vertebrates.

Animals

ECHO: a nanopore sequencing-based workflow for (epi)genetic profiling of the human repeatome.

SUMMARY: The human genome is dominated by repetitive DNA, whose genetic and epigenetic variation plays a key role in gene regulation, genome stability, and disease. Recent advances in long-read sequencing now enable large-scale, haplotype-resolved, and DNA methylation-informative analysis of the human genome, including on previously inaccessible complex and repetitive regions. However, the comprehensive, simultaneous characterisation of the "human repeatome" remains challenging, largely due to the lack of comprehensive tools integrated in a single pipeline that can capture the full spectrum of variation across diverse types of DNA repeats. Here, we present ECHO, a user-friendly, Snakemake-based pipeline for the "(Epi)genomic Characterisation of Human Repetitive Elements using Oxford Nanopore Sequencing." ECHO provides a reproducible and scalable framework for end-to-end analysis of whole-genome nanopore sequencing data, enabling integrative but also tailored (epi)genetic analyses of the human repeatome. AVAILABILITY AND IMPLEMENTATION: ECHO is freely available at Github: https://github.com/leenput/ECHO-pipeline, with the archived version at Zenodo: https://zenodo.org/records/19068468.

Humans

Aging and Reproductive Cancers: An Integrative View on Cell-Free DNA and Transposable Elements.

Aging is one of the strongest risk factors for cancer, and its impact is particularly evident in malignancies of the reproductive system. Ovarian, endometrial, cervical, vulvar, prostate, and penile cancers are mainly diagnosed in older adults and often show different clinical and biological features compared with the same tumors in younger patients. Aging is associated with hormonal changes, immune decline, epigenetic alterations, and accumulation of DNA damage, all of which contribute to cancer development and progression. At the same time, many older patients have frailty and multiple comorbidities, which can limit the use of screening programs and invasive diagnostic procedures. This often leads to delayed diagnosis and worse outcomes. Cell-free DNA (cfDNA) is a minimally invasive biomarker that can be obtained from blood samples and provides molecular information on both tumor and host tissues. Circulating DNA reflects tumor-specific alterations but is also influenced by aging-related changes in DNA release, fragmentation, and methylation. For this reason, aging must be considered when cfDNA-based biomarkers are applied in clinical practice. In this review, we describe how aging influences the biology of reproductive system cancers and how these processes are mirrored in cfDNA profiles. We focus on the clinical use of cfDNA for cancer detection and monitoring in older and fragile patients. Special attention is given to repetitive elements in cfDNA, which are strongly affected by aging and tumor-related epigenetic changes and can be detected with high sensitivity even when the tumor fraction is low. We propose an integrative mechanistic framework in which age-related epigenetic and genomic changes influence both tumor biology and cfDNA composition, with transposable elements acting as a central link between aging and cancer.

Humans

De Novo Assembly and Comparative Analysis of the Complete Mitochondrial Genome of Mesenchytraeus (Annelida, Enchytraeidae).

The Changbai Mountain range is one of the key glacial refugia in Northeast Asia. Mesenchytraeus exhibits high species diversity, strong endemism, and widespread cryptic species in this region, for which mitogenomes provide useful molecular markers for exploring cryptic species complexes. This makes Mesenchytraeus an ideal model for studying mitogenome evolution among closely related lineages; however, no mitogenome data have been reported for this genus to date. In this study, we performed de novo assembly, annotation, and comparative analysis of the mitogenomes of 13 Mesenchytraeus species (14 individuals) from Changbai Mountain. All mitogenomes are typical circular molecules containing 37 genes, but putative control regions are rearranged and consistently located between ATP6 and trnR. All species exhibit annelid-specific strand nucleotide biases, characterized by negative GC skew and near-zero AT skew. Codon usage analysis reveals that codon families with wobble U are significantly biased toward mtDNA codons, whereas those with wobble C or G are biased toward non-mtDNA codons, suggesting a conserved mitochondrial codon usage pattern in annelids. All tRNAs form typical cloverleaf secondary structures except trnS2, which lacks the D-stem and the dihydrouridine (DHU) arm in some species. The putative control regions commonly contain complex palindromic repeats, hairpins, and repetitive elements, and may harbor dual replication origins. Phylogenetic analyses support the monophyly of Mesenchytraeus and reveal significant molecular divergence among morphologically cryptic species. This study provides the first mitogenome dataset for Mesenchytraeus and offers new insights into the evolution and replication mechanisms of mitogenomes in Clitellata and broader Annelida.

Mesenchytraeus

Reference-Guided Chromosome-Scale Genome Assembly With Insights on Population Genomics of the Atlantic Goliath Grouper (Epinephelus itajara), Islas del Rosario, Colombia.

Epinephelus itajara, commonly known as the Atlantic Goliath grouper, is the largest species among the western North Atlantic groupers and is critically endangered. This species plays a crucial ecological, cultural, and economic role and has been the focus of captive breeding efforts at the Oceanario of the Rosario Islands, Colombia. However, despite its ecological and conservation importance, genomic resources and population genomic data for E. itajara remain scarce, particularly in the Colombian Caribbean. This study presents a reference-guided chromosome-scale genome assembly and an analysis of the population genomic structure of E. itajara using PacBio HiFi sequencing and Illumina technologies. The assembled genome has a total size of 1.12 Gb, with a contig N50 of 42.69 Mb and a scaffold N50 of 46.30 Mb. A total of 22,692 protein-coding genes were identified after masking 46% of the genome, which consists of repetitive elements. Comparative genomic analyses revealed a high degree of collinearity with closely related Epinephelus species and identified E. lanceolatus as the closest relative, supporting recent divergence and conserved genome architecture within the genus. Additionally, a population genomics analysis was conducted using 7706 high-quality SNPs to assess the genomic structure of captive populations. The results revealed four distinct genomic lineages, with moderate genetic differentiation among the sampled individuals. In the Colombian Caribbean, two unique lineages were identified, associated with the localities of Bahía Cispatá and Bahía Barbacoas, suggesting possible geographic isolation. These genomic resources provide valuable tools and new opportunities to better understand the genomic diversity, evolutionary history, and reproductive mechanisms of E. itajara. Moreover, they serve as a foundation for conservation strategies, including selective breeding programs aimed at increasing genomic diversity in captive populations and guiding restoration efforts in its natural habitat.

Epinephelus itajara

Deoxyribonucleic acid methylation abnormalities at imprinted loci in oligospermic and azoospermic men.

OBJECTIVE: To assess deoxyribonucleic acid (DNA) methylation at imprinted and repetitive genomic regions in ejaculated and testicular sperm from men with oligospermia, azoospermia, and those undergoing vasectomy reversal (VR), compared with fertile controls. DESIGN: Observational case-control study. SUBJECTS: Samples were obtained from 68 men, including 29 infertile men (18 oligospermic [5-15 million/mL], 11 severely oligospermic [<5 million/mL]) and 20 fertile controls with confirmed natural conceptions. Testicular tissue was collected from 11 azoospermic men (7 obstructive azoospermia [OA], 4 nonobstructive azoospermia [NOA]) and 8 men with prior paternity undergoing VR. EXPOSURE: Sperm DNA methylation at four imprinted genes (H19, GTL2, MEST, and LIT1) and one repetitive element (LINE1). MAIN OUTCOME MEASURES: Methylation levels at CpG sites determined by bisulfite pyrosequencing. RESULTS: The H19 was significantly hypomethylated only in severe oligospermia compared with fertile controls, whereas other groups showed nonsignificant trends with overlapping distributions. In contrast, MEST was significantly hypermethylated in oligospermic, azoospermic, and VR groups compared with fertile controls. No significant differences were observed for GTL2, LIT1, or LINE1. CONCLUSION: The DNA methylation abnormalities in sperm are locus-specific and vary across infertility phenotypes. MEST alterations were consistent across groups, whereas H19 changes were limited to severe oligospermia. Similar methylation patterns in testicular sperm from azoospermic and VR groups suggest that epigenetic alterations may reflect the testicular environment or obstruction, or differences between testicular and ejaculated sperm.

Humans

Mudskipper detects combinatorial RNA binding protein interactions in multiplexed CLIP data.

The uncovering of protein-RNA interactions enables a deeper understanding of RNA processing. Recent multiplexed crosslinking and immunoprecipitation (CLIP) technologies such as antibody-barcoded eCLIP (ABC) dramatically increase the throughput of mapping RNA binding protein (RBP) binding sites. However, multiplex CLIP datasets are multivariate, and each RBP suffers non-uniform signal-to-noise ratio. To address this, we developed Mudskipper, a versatile computational suite comprising two components: a Dirichlet multinomial mixture model to account for the multivariate nature of ABC datasets and a softmasking approach that identifies and removes non-specific protein-RNA interactions in RBPs with low signal-to-noise ratio. Mudskipper demonstrates superior precision and recall over existing tools on multiplex datasets and supports analysis of repetitive elements and small non-coding RNAs. Our findings unravel splicing outcomes and variant-associated disruptions, enabling higher-throughput investigations into diseases and regulation mediated by RBPs.

RNA-Binding Proteins

Chromosome-level genome assembly of an Arctic fish species pale eelpout (Lycodes pallidus).

Eelpouts (Zoarcidae) are known for their bipolar distributions and distinctive biogeographic histories. However, limited genomic data have hindered our understanding of their adaptive evolution. In this study, we present a thoroughly annotated chromosome-level genome assembly of pale eelpout (Lycodes pallidus) generated through the integration of Illumina, PacBio circular consensus, and Hi-C sequencing techniques. The final assembly spans 753.4&#x2009;Mb, with its high quality confirmed by a scaffold N50 of 28.6&#x2009;Mb and a Benchmarking Universal Single-Copy Ortholog (BUSCO) completeness of 99.3%. In comparison to other eelpouts and related fishes, the L. pallidus genome is larger and exhibits greater repetitive element content, accounting for approximately 45% of its total length. We annotated 21,419 protein-coding genes, a significant proportion of which are involved in signal transduction mechanisms and transcription. These findings provide valuable genetic resources for elucidating the evolutionary mechanisms underlying polar fish adaptation.

Animals

Chromosome-level genome assembly of Ceroplastes pseudoceriferus Green, 1935 (Hemiptera: Coccidae).

Soft scales (Hemiptera: Coccidae) are significant polyphagous pests and majority of which are invasive species. The 364.14&#x2009;Mb chromosome-level genome of Ceroplastes pseudoceriferus was assembled in this work, with a contig N50 length of 6.16&#x2009;Mb and scafold N50 length of 21.24&#x2009;Mb. Approximately 99.89% of assembled sequences were anchored into 18 chromosomes with the assistance of Hi-C reads. Furthermore, approximately 53.98% of the genome was composed of repetitive elements. In total, 10,475 protein-coding genes were predicted, of which 9503 (90.72%) genes were functionally annotated. The BUSCO analysis demonstrated the completeness of the genome annotation is 92.54%. This genome represents first high-quality chromosome level assembly of Coccidae, thereby advancing our knowledge of Coccidae insects and developing effective management strategies that protect crops, forests, and natural ecosystems.

Animals

Chromosome-level assembly and annotation of the yellow-shelled fish (Barbodes Wynaadensis).

Barbodes wynaadensis, a unique cyprinid species native to Yunnan Province in China, stands out as an allotetraploid (AABB) fish with a complex evolutionary history. Leveraging a multi-platform sequencing strategy combining MGI short-read, PacBio long-read, and Hi-C scaffolding technologies, we assembled the first chromosome-level genome for B. wynaadensis. The final assembled genome spans 1.76&#x2009;Gb in length with a contig N50 of 33.53&#x2009;Mb, demonstrating high assembly continuity. Hi-C scaffolding enabled the reconstruction of 50 pseudochromosomes, representing 99.94% of the total genome assembly. Genome annotation identified 46,121 protein-coding genes, with a functional annotation rate of 99.76%. Repetitive elements constituted 48.26% of the genomic sequences, including lineage-specific expansions of DNA transposons (29.26%) and LTRs (6.36%). This high-quality assembly resolves challenges in polyploid genome reconstruction and provides a critical resource for investigating Cyprinidae evolution, particularly subgenome divergence and adaptation. The dataset also enables practical applications, such as molecular marker development for population monitoring, supporting conservation efforts for this threatened endemic species amid habitat degradation in the Nujiang River basin.

Animals

Chromosome-level genome assembly of Elaeocarpus petiolatus (Elaeocarpaceae).

Elaeocarpus petiolatus is an ecologically and economically important species in tropical and subtropical forests. Despite its significance, the lack of genomic resources has hindered research on the genetic diversity and adaptive traits of E. petiolatus. To address this gap, we present a comprehensive chromosome-level genome assembly of E. petiolatus generated using advanced PacBio high-fidelity (HiFi) long-read sequencing and Hi-C technology. The assembly spans 322.45&#x2009;Mb, with a scaffold N50 of 20.58&#x2009;Mb, indicating that 37.11% of the genome is composed of repetitive elements. We identified 25,295 protein-coding genes, of which 96.74% were functionally annotated. This high-quality genome provides a critical resource for understanding the genetic mechanisms underlying environmental adaptability and biosynthesis of bioactive compounds in E. petiolatus, thereby supporting conservation efforts and sustainable forest management. The assembled genome and associated sequencing data are publicly available, facilitating further evolutionary and functional studies on the Elaeocarpaceae family.

Chromosomes, Plant

A chromosomal-level genome assembly of Odontolabis cuvera Hope, 1842 (Coleoptera: Lucanidae).

The stag beetle (Coleoptera: Lucanidae) represents a captivating and evolutionarily significant group, regarded as one of the most basal lineages within the superfamily Scarabaeoidea. Despite their importance for studying beetle evolution and ecology, genomic resources for this family remain scarce. Here, we report a chromosome-level genome assembly of Odontolabis cuvera, generated by integrating PacBio HiFi, Illumina, and Hi-C data. The genome assembly spans 908.07&#x2009;Mb, comprising 66 scaffolds (scaffold N50: 65.36&#x2009;Mb) and 147 contigs (contig N50: 16.39&#x2009;Mb). A total of 99.58% (904.22&#x2009;Mb) of the assembly was anchored to 14 chromosomes. BUSCO analysis (insecta_odb10 dataset, n&#x2009;=&#x2009;1,367) demonstrated high completeness, with 99.1% of conserved insect orthologs identified (98.3% single-copy, 0.8% duplicated). Repetitive elements accounted for 53.00% (281.28&#x2009;Mb) of the genome, and a total of 18,332 protein-coding genes were annotated. This high-contiguity genome provides a critical foundation for uncovering the evolutionary mechanisms and ecological adaptations unique to Lucanidae.

Animals

Chromosome-level genome assembly of the large carpenter bee Xylocopa dejeanii Lepeletier, 1841 (Hymenoptera: Apidae).

Xylocopinae, a diverse bee subfamily comprising over 1,000 bee species, and also a major model system for studying the pollination and evolution of sociality. The lack of chromosome-level genome assembly resources for the Xylocopinae limits our research of their biology and evolution. Here, we provided the first pseudo-chromosomes genome assembly of the Xylocopa dejeanii combined PacBio CLR long reads, Illumina sequences, and Hi-C data. The final genome is 194.44&#x2009;Mb located in 16 chromosomes. Our assembly includes 141 scaffolds, with a scaffold N50 length of 13.15&#x2009;Mb. BUSCO analysis revealed 99.00% completeness. Genome annotation identified 28.27&#x2009;Mb of repetitive elements, 10,970 protein-coding genes, and 432 ncRNAs. This high-quality X. dejeanii assembly advances our understanding of Xylocopinae genomics and provides new insights into bee evolution.

Animals