Search PubMedSearch

SEARCH · Search PubMed

Results for “Tandem Repeat Sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

234 records · Page 3Linked to original sources

Targeted sequencing reveals a distinct genetic alteration landscape in oral multiple primary squamous cell carcinomas.

OBJECTIVE: Oral multiple primary cancers (MPCs) are associated with poor clinical outcomes, yet their genomic characteristics remain insufficiently understood. DESIGN: Fifty-four formalin-fixed paraffin-embedded (FFPE) tumor samples from 30 patients with oral MPCs were analyzed using high-depth targeted sequencing of a customized 14-gene panel derived from prior whole-exome sequencing data. Detected alterations were analyzed after removal of synonymous mutations. RESULTS: Non-silent genomic alterations were identified in 59.3% (32/54) of samples, involving 19 patients. A total of 70 variant loci across 13 genes were detected. AKAP13 was the most frequently mutated gene at both the sample (22.2%, 12/54), with recurrent mutations observed across multiple patients. In contrast, TP53 mutations occurred at a substantially lower frequency (11.1%, 6/54). Marked inter- and intra-patient mutational heterogeneity was observed. CONCLUSIONS: FFPE-based targeted sequencing enabled an initial characterization of genomic alterations in oral MPCs. Recurrent alterations in AKAP13, GLI2, JMJD1C, and DNAH8, together with the relatively low frequency of TP53 alterations, identify candidate genomic features for further investigation and provide a basis for future studies of the molecular basis of oral MPCs.

Humans

Diagnostic and clinical utility of exome sequencing and chromosomal microarray in children with GDD/iD: a meta-analysis.

BACKGROUND: Global developmental delay/intellectual disability (GDD/ID) is among the most common neurodevelopmental disorders, with up to half of cases are attributed to genetic factors. Chromosome microarray (CMA) has traditionally been the primary genetic test for idiopathic GDD/ID. However, whole exome sequencing (WES) and whole genome sequencing (WGS) have recently emerged, substantially increasing diagnostic yields in these populations. METHODS: We conducted a comprehensive literature search of PubMed, Scopus, EMBASE, and the Cochrane Library from inception to April 29, 2025. Studies reporting the diagnostic utility of these tests in children with GDD/ID were included and analyzed. RESULTS: A total of 102 studies, comprising 55,752 children, were reviewed. The pooled diagnostic yield of WES was 0.37 (95% CI: 0.33-0.41; I2 = 93%), significantly higher than that of CMA at 0.19 (95% CI: 0.16-0.21; I2 = 95%). Subgroup analyses showed that WES yielded significantly higher diagnostic rates than CMA in both same-sample comparisons (OR = 2.27, 95% CI: 1.08-4.78) and different-sample comparisons (OR = 1.65, 95% CI: 1.15-2.37). Only one study evaluated WGS, reporting a diagnostic yield of 0.27. Meta-regression revealed a significant association between CMA diagnostic yield and the proportion of male participants (p&#x2009;<&#x2009;0.01), but not with WES. No significant difference in diagnostic utility was observed between isolated GDD/ID and GDD/ID with comorbidities. CONCLUSION: In children with unexplained GDD/ID, WES demonstrates superior diagnostic and clinical utility compared to CMA. Incorporating WES as a first-line investigation in the diagnostic evaluation of GDD/ID may be warranted.

Humans

Whole genome sequencing of unusual Hepatitis C virus subtypes and drug resistance analysis during direct-acting antiviral therapy in India.

INTRODUCTION AND OBJECTIVES: Pangenotypic direct-acting antivirals (DAA) are effective against highly prevalent Hepatitis C virus (HCV) subtypes, but have been clinically validated almost exclusively in high-income countries. Unusual HCV subtypes may carry natural polymorphisms, potentially impacting DAA susceptibility. We conducted full-genome characterization and resistance analysis of unusual HCV subtypes in patients receiving DAA treatment. PATIENTS AND METHODS: In this prospective hospital-based study, eligible patients were screened for anti-HCV antibodies and active infection was confirmed by diagnostic 5'NCR-based HCV RNA detection. Genotyping was performed by core region sequencing, and viral load quantified by real-time PCR. For whole genome sequencing, multiplex primers were designed using alignments of global reference sequences. Sequencing was carried out using the Oxford Nanopore Technology platform. Phylogenetic analysis used multiple sequence alignment and the HCV-GLUE resource for resistance-associated substitution (RAS) analysis. RESULTS: Predominant genotype was genotype 3 in 64.3% (n = 45); genotype 6 in 21.4% (n = 15); and genotype 1 in 14.2% (n = 10). Unusual HCV subtype 6xa was detected in two patients and showed no NS5A resistance mutations. One genotype 3b patient relapsed at 24 weeks post-DAA treatment completion and carried NS5A resistance-associated substitutions 30 K and 31 M both at baseline and at relapse, conferring high-level resistance to NS5A inhibitors. CONCLUSION: This is the first report from India of whole genome sequencing of HCV subtype 6xa. The identification of NS5A resistance mutations in the 3b relapse case underscores challenges for global HCV elimination strategies.

Humans

Comprehensive analysis of mRNA-microRNA-lncRNA expression profiles in post-traumatic elbow heterotopic ossification using RNA sequencing and experimental validation.

BACKGROUND: This study aimed to profile the molecular signatures of post-traumatic elbow heterotopic ossification (HO) to identify key regulators and potential therapeutic targets. METHODS: Total RNA from post-traumatic elbow HO tissues (n=4) and normal bone tissues (n=6) was subjected to high-throughput sequencing to identify differentially expressed mRNAs (DEGs), microRNAs (DEMs), and lncRNAs (DELs). Bioinformatics analyses included Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment, protein-protein interaction network construction, and transcription factor (TF)-microRNA-mRNA network analysis. The expression trends of four most upregulated and four most downregulated DEGs were validated by real-time quantitative reverse transcription polymerase chain reaction (qRT-PCR). RESULTS: We identified 2,138 DEGs, 40 DEMs, and 905 DELs. DEGs were significantly enriched in biological process "bone mineralization," cellular component "plasma membrane," molecular function "integrin binding," and pathways including PI3K-Akt, NF-&#x3ba;B, JAK-STAT, and TNF signaling pathways. Hub genes with high connectivity included MMP9, IL6, MMP3, CTSK, and BGLAP. Integrated network analysis highlighted the transcription factor JUN and key microRNAs (hsa-miR-124-3p, hsa-miR-548c-3p, and hsa-miR-135b). The qRT-PCR results confirmed the expression trends of selected DEGs. CONCLUSIONS: This study, for the first time, profiled the differentially expressed mRNAs, microRNAs, and lncRNAs in post-traumatic elbow HO using high-throughput RNA sequencing. These findings provide valuable insights into the molecular mechanisms of HO following elbow trauma. The identified hub genes (MMP9, IL6, MMP3, CTSK, and BGLAP), key TF (JUN), and key microRNAs (hsa-miR-124-3p, hsa-miR-548c-3p, and hsa-miR-135b) may serve as potential therapeutic targets for preventing and treating post-traumatic elbow HO.

Humans

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (&#x3a8;) represents one of the most abundant and conserved RNA modifications. &#x3a8; provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of &#x3a8; sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel &#x3a8; site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA &#x3a8;-site prediction. The &#x3a8; modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA &#x3a8;-site prediction. Meta-PseU offers a new framework for robust &#x3a8;-site identification by using long sequences.

Pseudouridine

Evaluation of one-step amplicon-based targeted enrichment for SARS-CoV-2 whole-genome sequencing using the Midnight amplicon scheme.

Genomic surveillance proved invaluable during the COVID-19 pandemic for tracking SARS-CoV-2 variants and guiding outbreak responses, underscoring the ongoing need to reduce whole-genome sequencing (WGS) costs and improve workflow efficiency to ensure accessibility in resource limited settings. Here, we evaluated a one-step reverse transcription polymerase chain reaction (RT-PCR) approach using the Midnight V2 primer scheme for targeted amplification of the SARS-CoV-2 genome, assessed its compatibility with Illumina sequencing, and compared its performance to a well-established two-step method. Initially, we determined optimal RT-PCR reaction conditions using the Midnight V2 primer panel for the one-step RT-PCR kit and scaled reaction volumes for both RT-PCR and library preparation. Clinical specimens (n&#x202f;=&#x202f;53) that had undergone routine WGS for surveillance purposes using the established two-step RT-PCR method were compared using the one-step RT-PCR assay. For samples with genome completeness greater than 70%, both methods gave comparable results with similar sequence coverage and 100% concordance for lineage assignment. Further investigation revealed a higher percentage of reads aligning to the SARS-CoV-2 genome with a greater depth of coverage using the one-step method compared to the two-step method. Finally, analysis of scaled one-step and library reaction volumes revealed significant cost savings for samples undergoing WGS. Overall, the results presented here verify the accuracy and reproducibility of one-step targeted amplification and offer an efficient and cost-effective workflow for routine SARS-CoV-2 genomic surveillance.

Humans

Single-cell RNA sequencing provides further insights into the immunostimulatory action of freeze-dried Lactiplantibacillus plantarum on Penaeus vannamei shrimp.

Immunostimulation through dietary interventions opened new avenues in developing disease control and prevention tools for shrimp aquaculture. We have previously shown that feeding with freeze-dried Lactiplantibacillus plantarum (LAB) increased disease resistance of Penaeus vannamei against both Vibrio parahaemolyticus and white spot syndrome virus (WSSV) based on bulk RNA sequencing of shrimp gills. This tissue participates in ion transport and serves as a first line of defense against environmental stressors and pathogenic infections. However, characterization of their cell composition and functions remains limited. Here, we implemented a single-cell RNA sequencing approach to further gather insights into how feeding with freeze-dried LAB modulates host immunity which may not be evident with bulk RNA sequencing approach. A total of five clusters with unique transcriptional signatures were identified, corresponding to pillar cells, septal cells, and sessile hemocytes. Pseudo-bulk analyses at global- and cluster-levels showed differential expression of genes related to host immunity and metabolism. We further revealed how overall transcriptomic changes are not exclusively caused by gene expression changes but may also be driven by cell population dynamics. This study highlighted how single-cell RNA sequencing approach may shed light on the mechanisms of action of immunostimulants which may be masked in bulk transcriptome analyses.

Animals

Integrated exome and mitochondrial genome sequencing reveals the genetic landscape of primary mitochondrial diseases: findings from a large Tunisian cohort.

Primary mitochondrial diseases are a heterogeneous group of neurometabolic disorders recognized as the most common metabolic genetic diseases. They manifest at any age, affecting any tissue or organ, especially those with high energy demands, and are caused by pathogenic variants in both mitochondrial and nuclear genomes. Here, we aimed to describe the genetic spectrum of a Tunisian pediatric cohort with suspected mitochondrial diseases. We recruited 47 unrelated families who underwent exome sequencing as a first-tier test followed by whole mitochondrial genome sequencing for unsolved cases. Dedicated bioinformatic pipelines and prediction tools were used to determine the potential disease-causing variants. Sanger sequencing confirmed the presence and segregation within parents. For the newly identified variants, structural modeling was conducted to study the impact of these variants on protein structure and motions. Dual genome sequencing yielded a molecular diagnosis in 33/47 families (70%) and 18/47 (38%) showed disease-causing variants in genes encoding mitochondrial proteins. Among them, four families disclosed novel variants in FASTKD2, SERAC1 and GATB, which were supported by in-depth in silico and structural analyses demonstrating their deleterious effect. The remaining families (32%, 15/47) disclosed other metabolic and neurological disorders. An exome-first strategy delivers a high diagnostic yield in Tunisia, where consanguinity remains high and simultaneously captures mitochondrial and non-mitochondrial etiologies. Mitochondrial sequencing remains indispensable in the case of an inconclusive exome. Thus, our data expand the clinical and genetic spectrum of primary mitochondrial diseases in Tunisia, an underrepresented and admixed population.

Humans

Diagnostic value of plasma cell-free DNA metagenomic next-generation sequencing in patients with suspected infections and exploration of clinical scenarios-a retrospective study from a single center.

BACKGROUND: Plasma cell-free DNA metagenomic next-generation sequencing (mNGS) is a non-invasive comprehensive method for the etiological diagnosis of various infectious diseases. However, research on the early diagnosis and real-world clinical impact of plasma mNGS in patients with suspected infection are still limited. MATERIALS AND METHODS: This study retrospectively included 140 patients with suspected infections who underwent early plasma mNGS and conventional culture testing. Referring to the clinical diagnosis of infectious diseases, the diagnostic performance of plasma mNGS and culture tests was compared, and the application scenarios and clinical effects of plasma mNGS were evaluated. RESULTS: The positive rate of plasma mNGS was significantly higher than that of culture methods (55.71% vs 25.10%, p&#x2009;<&#x2009;0.001) and blood cultures (55.71% vs 12.86%, p&#x2009;<&#x2009;0.001). Regarding clinical diagnosis, the sensitivity of plasma mNGS was significantly higher than that of culture (58.27% vs 37.80%, p&#x2009;=&#x2009;0.002). The combination of mNGS and culture achieved a higher detection sensitivity (69.29%), especially in patients with multi-site co-infections (73.68%) and blood infections (73.17%). Plasma mNGS demonstrated higher sensitivity in patients with procalcitonin (PCT) index > 5&#x2009;ng/ml or human neutrophil lipocalin (HNL) index > 200&#x2009;ng/ml. In terms of treatment, a total of 69 patients (54.33%) benefited from plasma mNGS. CONCLUSION: This study highlights the significant improvement in pathogen detection performance by combining conventional culture with plasma mNGS detection, especially in patients with multi-site co-infections and blood infections. Early use of plasma mNGS as an adjunct to culture can better guide clinicians to initiate appropriate anti-infective therapy.

Humans

Nanopore-based epigenomic profiling reveals the absence of widespread CpG methylation in the African swine fever virus genome.

DNA methylation is a critical epigenetic mechanism implicated in regulating replication and transcription in DNA viruses. However, the epigenetic landscape of African swine fever virus (ASFV), a large double-stranded DNA virus infecting pigs, remains controversial. Here, we systematically profiled the DNA methylome of the first ASFV strain isolated in Hong Kong (HK_NT_202103) using Oxford Nanopore Technologies (ONT) R10.4.1 sequencing. We employed a paired design: native whole-genome sequencing (WGS) against a methylation-free whole-genome amplification (WGA) control. Using conservative thresholds, we found no evidence of 5-methylcytosine (5mC), especially typical CpG methylation, across the viral genome. Importantly, clear CpG methylation signals were successfully detected in the host genome from WGS data, confirming the functionality of the workflow to detect 5mC at CG sites. While widespread 5mC seems absent, a small number of putative N6-methyladenine (6mA) loci were identified. A specific 6mA candidate exhibited raw ionic current disruptions and gene-level intersection with another ASFV isolate (CAS19-01/2019), although it lacked single-base consensus across different methylation callers or between the two isolates. Although our biological findings are restricted to a single isolate under specific experimental conditions, this study introduces a novel, highly rigorous ONT framework for viral epigenomics research. Furthermore, the absence of ASFV CpG methylation indicates that host CpG-depletion remains a viable strategy for viral metagenomic enrichment. Ultimately, our work offers a critical methodological baseline for ASFV surveillance and highlights the necessity of targeted experimental validation for rare viral modifications.

African Swine Fever Virus

LitCTL1: A novel C-type lectin involved in the mucosal and cellular immunity of the common periwinkle Littorinalittorea.

C-type lectins (CTLs) are vital pattern-recognition receptors (PRRs) that mediate innate immune responses in mollusks, yet their characterization in Caenogastropoda, the largest gastropod group, remains limited. This study characterizes LitCTL1, a novel secreted single-domain C-type lectin from the common periwinkle, Littorina littorea. The 199-amino acid polypeptide contains a conserved carbohydrate recognition domain with canonical QPD and WND motifs and is predicted to form a homodimer. Uniquely, LitCTL1 was localized in both circulating hemocytes and mucus-secreting epithelial cells of the foot, mantle, and hypobranchial gland - the first report of such dual localization for a molluscan lectin, linking systemic and mucosal defense. Expression analysis revealed that LitCTL1 is constitutively expressed in hemocytes. Functional assays with recombinant LitCTL1 demonstrated its role as a potent opsonin with hemagglutinating activity, significantly enhancing hemocyte spreading and the phagocytosis of zymosan. Genomic analysis reveals that LitCTL1 belongs to a rapidly diversifying, genus-specific expansion distinct from conserved perlucin-like lineages. These results identify LitCTL1 as a key effector molecule in both systemic and mucosal innate immunity, likely reflecting an evolutionary adaptation to the microbial challenges of the intertidal environment.

Animals

Routine methods misidentify Serratia spp.: Limitations of MALDI-TOF MS revealed by whole-genome sequencing.

Accurate species-level identification within the genus Serratia remains challenging due to extensive phenotypic overlap and high genomic relatedness among closely related and recently described taxa. This study presents an evaluation of routine and genome-based identification approaches applied to clinical Serratia isolates, integrating phenotypic assays, MALDI-TOF MS (Bruker Daltonics), 16S rRNA gene sequencing, and Whole-Genome Sequencing (WGS). A total of 103 isolates collected from a teaching hospital were analyzed. WGS was performed on a subset of isolates. Conventional biochemical methods classified all isolates as Serratia marcescens, whereas MALDI-TOF MS identified 60.1% as S. marcescens, 11.6% as S. ureilytica, and 28.1% just at the genus level. Peak analysis from MALDI-TOF MS revealed specific peaks associated with S. marcescens and S. ureilytica, but limited discriminatory power. WGS of six isolates initially identified as S. ureilytica by MALDI-TOF MS revealed reclassification as Serratia sarumanii (n = 5) and Serratia montpellierensis (n = 1), supported by Average Nucleotide Identity (ANI), Average Amino Acid Identity (AAI), and Digital DNA-DNA Hybridization (dDDH) thresholds. In contrast, 16S rRNA analysis showed limited species-level resolution. Phylogenomic and SNP-based analyses confirmed these classifications with strong support. Overall, this study underscores the critical role of high-resolution genomic approaches for precise species identification and highlights the need for continuous expansion and curation of MALDI-TOF MS reference databases to support reliable clinical diagnostics and epidemiological surveillance of emerging Serratia species.

Spectrometry, Mass, Matrix-Assisted Laser Desorpti

Whole-transcriptome RNA sequencing and ceRNA network analyses provide novel insights into the antibacterial immune response of Hippocampus abdominalis against Vibrio harveyi.

Long non-coding RNAs (lncRNAs) stand as newly-arisen molecular types that exert regulatory effects, able to operate as competitive endogenous RNAs (ceRNAs) to engage microRNAs (miRNAs) in interaction, resulting in the recovery of target mRNA expression and activity. Increasing evidences indicate that the ceRNA network affects various biological processes in mammals, including development, cellular differentiation, metabolism, immune response, and disease pathogenesis. In teleost fish, the lncRNA-miRNA-mRNA regulatory networks have been reported occasionally. However, up to now, the roles of lncRNAs in the big-belly seahorse (Hippocampus abdominalis) remains unclear. In this study, we reported for the first time, via whole-transcriptome RNA sequencing, the lncRNA mediated ceRNA regulatory network in Vibrio harveyi-infected H. abdominalis. A total of 4197 differentially expressed mRNAs (DE-mRNAs), 1317 DE-lncRNAs, and 183 DE-miRNAs were identified. Furthermore, the crosstalk between miRNAs and lncRNAs as well as between miRNAs and mRNAs was inferred based on the negative correlations between miRNAs and their target lncRNAs/mRNAs. A core immune associated lncRNA-miRNA-mRNA putative regulatory network was thus constructed, comprising 211 lncRNA-miRNA and 224 mRNA-miRNA pairs. In conclusion, our findings provide an integrative overview of the ceRNA regulatory networks on the underlying immune responses to V. harveyi infection in the big-belly seahorse, and offer a solid theoretical foundation for the comparative immunological research of teleost fish.

Animals

A genome-wide coverage-based pipeline for the identification of host-derived candidate DNA biomarkers from cell-free blood.

We have created a new data-analysis pipeline for the discovery of host-specific candidate DNA biomarkers derived from sequencing data of cell-free blood. Unlike approaches that rely on specific molecular or genetic signatures, our method leverages the coverage distribution of cell-free DNA sequences mapped to a reference genome, applying statistical analyses to identify informative short genomic regions for biomarker discovery. The pipeline is applicable to diverse diseases and can be used to analyze cell-free DNA sequences from plasma or serum to identify candidate biomarkers that are characteristic of disease states in mammals. Core functionalities were developed in Java and integrated with open-source software tools for the preprocessing of raw sequencing data, complemented by Python scripts for the machine-learning analysis and statistical validation. The pipeline is designed for HPC use and users can access the pipeline through a Galaxy workflow, which offers a user-friendly web interface for input selection prior to execution and analysis progress monitoring. Performance tests, carried out using duplicate sets of COVID-19 samples and controls, showed linear scalability of execution time with an increasing dataset size, as well as a substantial reduction in execution time through parallelized computation, whereby each HPC node is used to process the data of one chromosome. Further statistical tests confirmed the quality of the pipeline's results by showing that the set of identified candidate biomarkers remained stable across varying dataset sizes.

Biomarkers

Comparison of paralog identification methods and their impact on species tree topologies in target capture phylogenomics within the Sindora clade (Detarioideae: Leguminosae).

Target capture is a common method of generating high throughput DNA sequencing data for phylogenetic reconstruction of species relationships, for which single copy genes are usually most informative. However, a pervasive problem with target capture is that putatively single copy genes may in fact be paralogs resulting from gene duplication, which are problematic for phylogenetic inference because their evolutionary history may differ from the divergence history of species. Here, we use as a case study a target enrichment dataset of 88 species of Detarioideae (Leguminosae) with a focus on the Sindora clade to examine approaches for handling paralogs, including the built-in paralog handling functions in HybPiper and CAPTUS, plus subsequent steps using Putative Paralog Detection and the tree-based Yang & Smith orthology inference approach. We compare the paralogs flagged using these methods and verify their performance with BLAST mapping against a reference genome sequence of Sindora glabra, and then subsequently compare the species tree topologies produced across these methods. Our comparisons of paralogs flagged across the Sindora clade show that the Putative Paralog Detection pipeline was the most accurate in identifying paralogs in terms of its similarity to the BLAST mapping, followed by the built-in paralog identification function of CAPTUS. However, the results we recovered for the Detarioideae subfamily suggest that the largest differences in species tree topology resulted from the use of paralog-filtered alignments (such as with the Putative Paralog Detection pipeline and the Yang & Smith orthology inference approaches) rather than just by removing the sequences of identified paralogous genes. This was the true for HybPiper-assembled datasets but was not seen in CAPTUS-assembled datasets. In all comparisons, the topological differences caused by different paralog handling methods tended to be confined to clades where processes such as hybridisation and introgression are prevalent. Our study provides a roadmap to establish the best approach to identify, eliminate or separate paralogs in the absence of a chromosomally contiguous reference genome for a study group, and highlights the importance of careful data inspection and processing in addition to understanding the extent of paralogy and paralog characteristics (e.g. sequence divergence between copies) for their study group.

Phylogeny

Conserved host-exclusive oligonucleotide motifs enriched in pathogenic genes of human oncogenic viruses.

Comparative viral genomics can reveal sequence-level constraints influencing virus-host interactions. Relative minimal absent words (rMAWs) are short oligonucleotide motifs present in viral genomes but completely absent from the host, potentially reflecting selective pressures related to host adaptation and immune evasion. Using the EAGLE algorithm and the GRCh38 human reference genome, we systematically screened for prevalent rMAWs (prMAWs) across six major human oncogenic viruses: Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human papillomavirus (HPV), human T-cell leukemia virus type 1 (HTLV-1), and human herpesvirus 8/Kaposi's sarcoma-associated herpesvirus (HHV-8/KSHV). highly conserved 11- and 12-bp prMAWs were identified in EBV, HBV, HTLV-1, and HHV-8/KSHV, with sequence prevalences ranging from 91.5% to 97.9%. Conversely, no short prMAWs were detected in HCV or HPV, likely reflecting differences in genome architecture, mutation rates, and long-term host adaptation to the human host. Importantly, the identified host-exclusive motifs exhibited non-random genomic distribution and were preferentially embedded within viral genes central to replication, persistence, immune modulation, and oncogenesis, including EBNA-1 (EBV), HBx (HBV), Tax-associated regions (HTLV-1), and lytic replication genes of HHV-8/KSHV. Notably, all detected prMAWs were enriched in GC nucleotides and exhibited marked CpG over-representation, suggesting sequence constraints associated with epigenetic regulation and viral persistence. Collectively, these highly conserved, host-exclusive signatures offer promising, candidates for sequence-directed approaches in the diagnosis, monitoring, and investigation of virus-associated cancers.

Humans

A conserved distal-tail helical extension defines a tailspike attachment architecture in Gram-negative siphophages.

Rapid growth of bacteriophage genome collections has outpaced functional annotation of tail-tip proteins, limiting comparative analysis of host-recognition structures. Starting from a shared distal-tail gene organization in the Salmonella phages 9NA and Jersey, I developed a morphogenetic bioinformatic framework integrating gene synteny, sequence comparison, profile hidden Markov model (HMM) screening, structural evidence, structure-aware searching, and AlphaFold modeling. Comparison with the experimentally characterized lambda and Sf11 tail assemblies identified a predominantly alpha-helical C-terminal extension of the distal-tail (DT) protein associated with tailspike attachment, termed the distal-tail helical extension (DT-helix). Screening 541,986 proteins from 5167 complete NCBI RefSeq tailed-phage genomes, followed by evidence-based evaluation of sequence, genomic context, and structural architecture, identified 165 curated DT-helical-extension-associated phages. Their DT proteins segregated into six sequence groups. In the four principal multi-member groups, cognate tailspikes showed group-specific conservation in proximal N-terminal regions but substantially greater downstream diversity, consistent with sequence constraint at the DT-tailspike attachment boundary. A complementary ProstT5/Foldseek search supported the established groups but revealed no convincing additional highly divergent family. Together with the experimentally characterized Sf11 attachment interface, these findings define a recurrent morphogenetic architecture linking conserved distal-tail scaffolds to more variable receptor-binding proteins across siphophages infecting Gram-negative bacteria. Although universal exchangeability is not established, the identified scaffold-receptor-binding boundaries provide a framework for molecular characterization and rational phage engineering. Accession-level information for the 165 curated phages is available through PhageTailDB.

Viral Tail Proteins

Mitochondrial DNA diversity in Ecuadorian populations: Recurrence of variant 16136 within haplogroup B2.

The identification of lineage-defining variants, frequently found in the coding region of mitochondrial DNA (mtDNA), is essential for refining haplogroup classification. Most mtDNA studies in South American populations have focused on the control region (CR), which has provided important insights into population structure and maternal lineage origins, although information needed for more robust phylogenetic resolution has been neglected. This study investigates the maternal genetic structure of Ecuadorian populations by combining CR and whole mitogenome analyses. Sequences from the mtDNA CR were obtained from 461 individuals (253 Mestizos and 208 Native Americans), while complete mitogenomes were sequenced for 127 individuals to improve phylogenetic resolution by identifying lineage-defining variants present in coding region. Most mtDNA haplogroups in the two population groups analyzed were of Native American origin (A2, B2, B4, C1, D1, D4), with significant differences in the distribution of specific lineages between them. Among Mestizos, African haplogroups (all within the L branches) and Eurasian haplogroups (H, K, R, U) were detected at low frequencies, whereas no African lineages were observed among Native Americans. The results obtained highlighted a heterogeneity within Ecuadorian populations that must be considered when developing mtDNA haplotype databases for forensic purposes. Whole mitogenome sequences enabled the identification of variants that refined haplogroup classifications, provided a more accurate reconstruction of the maternal genetic diversity, and improve the discrimination between Native American and Asian maternal lineages within haplogroup B4b.

Humans