Search PubMedSearch

SEARCH · Search PubMed

Results for “Genome-wide SNPs”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

271 records · Page 4Linked to original sources

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Artificial Intelligence

Comparative genomic epidemiology of food- and patient-derived diarrheagenic Escherichia coli from sentinel surveillance in Southeast China.

Diarrheagenic Escherichia coli (DEC) remains an important foodborne pathogen, yet long-term comparative genomic surveillance data jointly characterizing food-derived and patient-derived isolates remain limited. This surveillance-based comparative study integrated antimicrobial susceptibility testing and whole-genome sequencing to characterize diarrheagenic Escherichia coli isolates recovered from food and patient sources in Lishui, Southeast China, during 2018-2025, with emphasis on occurrence, resistance profiles, genomic backgrounds, and plasmid replicon-associated features. Antimicrobial susceptibility testing was performed for 258 selected isolates, and whole-genome sequencing was conducted for a curated analytical subset of 204 isolates. The sequenced subset was used for diversity-oriented comparative genomic analysis rather than for unbiased prevalence estimation of the entire DEC collection. EAEC predominated in both sources, although food-associated occurrence was heterogeneous across categories, with the highest recovery rate observed in raw meat. Patient-derived isolates showed a broader overall resistance burden, whereas food-derived isolates retained substantial resistance to tetracycline, chloramphenicol, and florfenicol. Phylogenetic analysis showed partial overlap in genomic backgrounds between food-derived and patient-derived isolates, while representative resistance determinants displayed both broadly distributed and lineage-enriched patterns. Replicon-based plasmid profiling identified 42 plasmid types, including 12 detected in both sources, with IncF-related replicons predominating among these shared profiles. Several food-derived isolates carried multiple plasmid replicon types that were also observed in patient-derived isolates. Overall, food-derived and patient-derived DEC showed partial overlap in genomic backgrounds, resistance determinants, and replicon-defined plasmid profiles within this surveillance setting, while retaining source-associated heterogeneity. These findings should be interpreted as surveillance-based comparative evidence rather than as evidence of direct source attribution or transmission.

Humans

Genomic science and the nurse educator's role: Promoting integration from curriculum to clinical practice.

BACKGROUND: Registered nurses and nurse educators play a critical role in preparing future clinicians to translate genomic discoveries into practice. However, emerging evidence suggests that both groups may lack sufficient knowledge and confidence in genomics, potentially limiting their ability to teach, mentor, and apply genomics in real-world settings. This gap is especially concerning in Aotearoa New Zealand, where the genomic literacy of nurse educators and clinicians remains underexplored. OBJECTIVE: This study aims to: (1) assess nurse educators' genomic literacy and confidence in teaching genomics; and (2) evaluate registered nurses' knowledge and confidence in applying and teaching genomics in clinical practice. DESIGN: Exploratory descriptive qualitative. SETTING: This study was conducted in the greater Auckland area. PARTICIPANTS: A total of 17 participants were recruited using purposive sampling to ensure a diverse range of perspectives across varying levels of teaching experience, disciplinary backgrounds, and exposure to genomic content. METHODS: Data were collected using semi-structured focus group interviews, a method well-suited for generating in-depth discussion and facilitating interaction among participants with shared professional interests. The collected data were analysed using thematic analysis methods. RESULTS: The findings offer insight into the preparedness of New Zealand's nursing workforce to engage with genomic-informed healthcare and inform strategies for integrating genomics into nursing curricula and continuing professional development. Given the interdisciplinary nature of genomic healthcare, these insights may also be relevant to other health professionals-including midwives, pharmacists, and allied health practitioners-who increasingly encounter genomic information in clinical practice and require foundational competencies to support patient care. CONCLUSION: Addressing this educational gap is critical to ensuring that nurses-key facilitators of patient care and public health-are equipped to deliver safe, equitable, and evidence-based genomic healthcare.

Humans

Evaluation of one-step amplicon-based targeted enrichment for SARS-CoV-2 whole-genome sequencing using the Midnight amplicon scheme.

Genomic surveillance proved invaluable during the COVID-19 pandemic for tracking SARS-CoV-2 variants and guiding outbreak responses, underscoring the ongoing need to reduce whole-genome sequencing (WGS) costs and improve workflow efficiency to ensure accessibility in resource limited settings. Here, we evaluated a one-step reverse transcription polymerase chain reaction (RT-PCR) approach using the Midnight V2 primer scheme for targeted amplification of the SARS-CoV-2 genome, assessed its compatibility with Illumina sequencing, and compared its performance to a well-established two-step method. Initially, we determined optimal RT-PCR reaction conditions using the Midnight V2 primer panel for the one-step RT-PCR kit and scaled reaction volumes for both RT-PCR and library preparation. Clinical specimens (n = 53) that had undergone routine WGS for surveillance purposes using the established two-step RT-PCR method were compared using the one-step RT-PCR assay. For samples with genome completeness greater than 70%, both methods gave comparable results with similar sequence coverage and 100% concordance for lineage assignment. Further investigation revealed a higher percentage of reads aligning to the SARS-CoV-2 genome with a greater depth of coverage using the one-step method compared to the two-step method. Finally, analysis of scaled one-step and library reaction volumes revealed significant cost savings for samples undergoing WGS. Overall, the results presented here verify the accuracy and reproducibility of one-step targeted amplification and offer an efficient and cost-effective workflow for routine SARS-CoV-2 genomic surveillance.

Humans

Genome-guided stage- and tissue-resolved transcriptome analysis of Serrodes campana identifies sex-biased antennal expression and candidate chemosensory-related genes.

Serrodes campana is an erebid moth of ecological and forestry relevance; its larvae are mainly associated with the soapberry tree, Sapindus mukorossi, whereas adults exhibit fruit-piercing behavior. However, stage- and tissue-resolved transcriptomic resources for this species remain limited. Here, using a chromosome-level reference genome, we performed a genome-guided transcriptome analysis of S. campana based on 12 RNA-seq libraries representing major developmental stages and key adult tissues. Global transcriptomic analyses revealed pronounced transcriptional differentiation across developmental stages and tissue types. Tissue-enriched gene sets and functional enrichment analyses identified distinct molecular signatures associated with developmental, sensory, and pheromone-associated tissues. Comparative analysis of female and male antennae further revealed sex-biased expression of several candidate chemosensory-related genes. Among 153 curated chemosensory-related candidate genes, most odorant receptor genes showed strong antennal enrichment, whereas other major chemosensory gene families displayed broader but still tissue-preferential expression patterns. In addition, an exploratory comparison of female terminal abdominal gland tissue and male terminal abdominal coremata revealed divergent expression profiles and highlighted candidate genes potentially associated with pheromone-related physiology, reproduction, and tissue-specific signaling. Together, this study provides the first genome-guided stage- and tissue-resolved transcriptomic resource for S. campana and offers a useful foundation for future studies of chemosensory detection, sex-biased gene expression, and pheromone-associated biology in this species.

Animals

Cytonuclear conflict and reticulate evolution in the Morelloid clade (Solanum, Solanaceae): Insights from genome skimming and network Phylogenomics.

The Morelloid clade (black nightshades) is one of the most strongly supported clades within the megadiverse Solanum genus. It comprises 76 globally distributed, non-spiny herbaceous and suffrutescent species. While often erroneously considered poisonous weeds, several species are economically important as orphan crops. The clade is closely related to tomato and potato but, due to a lack of focused breeding efforts, remains a putative reservoir of genetic diversity for crop improvement. Despite this potential, we lack fundamental knowledge on the evolution of the Morelloid clade. The group includes polyploid species with unknown parental origins-likely reflecting reticulate processes such as hybridization, introgression, and associated backcrossing events. Prior analyses have been unable to disentangle these processes, leaving the mechanisms underlying reticulate evolution in the Morelloid clade poorly understood. Here, we use genome skimming to produce a well-supported maximum likelihood plastid phylogeny from complete circularized plastomes and a coalescent-based species tree from combined Angiosperms353 and conserved ortholog set nuclear markers. Our dataset, composed of previously published data and deep genome skimming from herbarium samples, spans 26 Morelloid species. To investigate phylogenetic discordance, we used a nuclear phylogenetic network, multispecies coalescent simulations, a fused rooted nuclear chloroplast tree, and quantification of nuclear gene tree concordance. We show that incongruence between nuclear and plastid trees is pervasive and cannot be explained by incomplete lineage sorting alone. Instead, our results demonstrate that events consistent with repeated chloroplast capture have shaped the reticulate evolutionary history of the clade, especially among African polyploid and Pan-American diploid lineages.

Phylogeny

Ramu stunt virus genome reveals previously unreported segments and nucleocapsid domain duplication in Mechlorovirus.

Ramu stunt virus (RmSV), a member of the genus Mechlorovirus within the family Phenuiviridae, was previously described as a six-segmented RNA virus infecting sugarcane. In this study, we re-examined type material and additional isolates using high-throughput sequencing and RT-PCR validation, revealing that RmSV possesses a nine-segmented genome, making it the largest reported in the Phenuiviridae. This expanded architecture includes duplicated RNA segments (RNA 2a and RNA 2b) encoding nucleocapsid-like proteins and two novel segments (RNA 7 and RNA 8). Comparative analysis showed that RNA 2a and 2b share about 84% amino acid identity, while RNA 5 encodes a third nucleocapsid homolog, indicating unprecedented domain redundancy. Structural modeling confirmed that all three nucleocapsid proteins maintain a conserved fold despite low sequence identity, with electrostatic mapping suggesting differential RNA-binding potential. Additionally, RNA 6 encodes a hypothetical protein structurally similar to the rice stripe virus disease-specific S-protein, implicating a role in symptom development. Transcript abundance analysis revealed RNA 6 as the most highly expressed segment across isolates. These findings revise the genomic composition of RmSV, highlight mechanisms of genome plasticity and adaptive evolution in plant-infecting bunyaviruses, and underscore practical implications for diagnostic assay design, resistance breeding, and biosecurity surveillance.

Genome, Viral

Whole genome sequencing of unusual Hepatitis C virus subtypes and drug resistance analysis during direct-acting antiviral therapy in India.

INTRODUCTION AND OBJECTIVES: Pangenotypic direct-acting antivirals (DAA) are effective against highly prevalent Hepatitis C virus (HCV) subtypes, but have been clinically validated almost exclusively in high-income countries. Unusual HCV subtypes may carry natural polymorphisms, potentially impacting DAA susceptibility. We conducted full-genome characterization and resistance analysis of unusual HCV subtypes in patients receiving DAA treatment. PATIENTS AND METHODS: In this prospective hospital-based study, eligible patients were screened for anti-HCV antibodies and active infection was confirmed by diagnostic 5'NCR-based HCV RNA detection. Genotyping was performed by core region sequencing, and viral load quantified by real-time PCR. For whole genome sequencing, multiplex primers were designed using alignments of global reference sequences. Sequencing was carried out using the Oxford Nanopore Technology platform. Phylogenetic analysis used multiple sequence alignment and the HCV-GLUE resource for resistance-associated substitution (RAS) analysis. RESULTS: Predominant genotype was genotype 3 in 64.3% (n = 45); genotype 6 in 21.4% (n = 15); and genotype 1 in 14.2% (n = 10). Unusual HCV subtype 6xa was detected in two patients and showed no NS5A resistance mutations. One genotype 3b patient relapsed at 24 weeks post-DAA treatment completion and carried NS5A resistance-associated substitutions 30 K and 31 M both at baseline and at relapse, conferring high-level resistance to NS5A inhibitors. CONCLUSION: This is the first report from India of whole genome sequencing of HCV subtype 6xa. The identification of NS5A resistance mutations in the 3b relapse case underscores challenges for global HCV elimination strategies.

Humans

Nanopore-based epigenomic profiling reveals the absence of widespread CpG methylation in the African swine fever virus genome.

DNA methylation is a critical epigenetic mechanism implicated in regulating replication and transcription in DNA viruses. However, the epigenetic landscape of African swine fever virus (ASFV), a large double-stranded DNA virus infecting pigs, remains controversial. Here, we systematically profiled the DNA methylome of the first ASFV strain isolated in Hong Kong (HK_NT_202103) using Oxford Nanopore Technologies (ONT) R10.4.1 sequencing. We employed a paired design: native whole-genome sequencing (WGS) against a methylation-free whole-genome amplification (WGA) control. Using conservative thresholds, we found no evidence of 5-methylcytosine (5mC), especially typical CpG methylation, across the viral genome. Importantly, clear CpG methylation signals were successfully detected in the host genome from WGS data, confirming the functionality of the workflow to detect 5mC at CG sites. While widespread 5mC seems absent, a small number of putative N6-methyladenine (6mA) loci were identified. A specific 6mA candidate exhibited raw ionic current disruptions and gene-level intersection with another ASFV isolate (CAS19-01/2019), although it lacked single-base consensus across different methylation callers or between the two isolates. Although our biological findings are restricted to a single isolate under specific experimental conditions, this study introduces a novel, highly rigorous ONT framework for viral epigenomics research. Furthermore, the absence of ASFV CpG methylation indicates that host CpG-depletion remains a viable strategy for viral metagenomic enrichment. Ultimately, our work offers a critical methodological baseline for ASFV surveillance and highlights the necessity of targeted experimental validation for rare viral modifications.

African Swine Fever Virus

Systematic modular engineering of genome-integrated Escherichia coli MG1655 for high-level 2'-fucosyllactose production.

2'-Fucosyllactose (2'-FL), the most abundant human milk oligosaccharide (HMO), has attracted considerable interest for its prebiotic and immunomodulatory functions, with broad applications in infant nutrition. In this study, we report the development of a high-yield, genome-integrated 2'-FL-producing strain based on Escherichia coli MG1655 through systematic modular optimization. Starting from a single-copy BKHT strain (MGC06), we first optimized the copy number of the α-1,2-fucosyltransferase (α-1,2-FT) gene BKHT. Subsequently, the GDP-L-fucose supply was enhanced through coordinated genomic integration of the gene clusters cpsG-cpsB and gmd-fcl, while the multidrug efflux transporter gene mdfA was integrated to improve product export and strain robustness. BKHT copy number was then re-evaluated in the optimized background, with four copies yielding the highest production. The final engineered strain, harboring all genetic modifications stably integrated into the chromosome, produced 17.18 g/L 2'-FL in shake-flask culture. In fed-batch fermentation using a 5-L bioreactor, this strain achieved a titer of 154.12 g/L after 60 h, with a productivity of 2.57 g/L/h. Notably, throughout the entire fermentation process, no antibiotics or inducers were supplemented, underscoring the genetic stability and regulatory compliance of this plasmid-free system. To our knowledge, this represents the highest 2'-FL titer reported to date, positioning our engineered strain as a promising candidate for commercial 2'-FL production.

Escherichia coli

Integrating genomic distance analyses in the description of a new family, genus, and species of sponge-associated antipatharians (black corals).

Antipatharians (black corals) are among the least studied coral groups, with much of their diversity still undescribed. Here, we present an integrative morphological, phylogenomic and genomic distance study of deep-sea antipatharians sampled in high seas areas of the North Pacific Ocean and from New Zealand's Exclusive Economic Zone. These corals grow on hexactinellid sponges - a unique characteristic in the order Antipatharia. Using a dataset of ultra-conserved elements and exons, combined with morphological analyses, we reconstruct phylogenomic relationships and formally describe a new family (Eidikopathidae fam. nov.), a new genus (Eidikopathesgen. nov.), and two new species (E. korallispongiasp. nov., E. zealandkoralliasp. nov.). Morphologically, the new family is distinguished by a corallum consisting of a network of loose branches that fuse with the sponge skeletal framework. Phylogenomic analyses recovered consistent topologies with strong nodal support, corroborating the distinct evolutionary placement of this sponge-associated lineage. Pairwise genomic distances estimated using the Tamura-Nei model were concordant with patristic genomic distances, identifying Pteridopathidae as the genetically closest family to Eidikopathidae fam. nov., followed by Myriopathidae and Stylopathidae, which were recovered as sister families in the phylogeny. This pattern shows that genomic distance complements, rather than simply mirrors, tree topology by quantifying accumulated sequence divergence among lineages. Together, these results provide the first genomic distance framework for Antipatharia, offering a baseline for future systematic, evolutionary, and biodiversity studies on this fundamental shallow, mesophotic and deep-sea coral group.

Animals

Regional genomic analysis of lineage distribution and transferable multidrug resistance among chicken-associated Salmonella Kentucky isolates in China.

Salmonella enterica serovar Kentucky is an important multidrug-resistant foodborne pathogen in the poultry meat supply chain. Although recent broader genomic studies have elucidated the population structure and epidemiological significance of major lineages in China (e.g., ST198 and ST314), the regional dynamics within local poultry supply chains remain insufficiently characterized. In this study, 31 chicken meat-derived isolates from Shanghai and 39 publicly available genomes from China were analyzed using antimicrobial susceptibility testing, whole-genome sequencing, phylogenetic analysis, conjugation experiments, and complete sequencing of representative plasmids. This enabled a systematic characterization of the molecular epidemiological features of the population and the mechanisms underlying resistance dissemination. Population genomic analysis revealed a lineage composition markedly different from the global epidemiological pattern: ST314 was the predominant sequence type among the Shanghai chicken-derived isolates (74.2%), whereas the internationally recognized high-risk clone ST198 accounted for only 25.8% of the local isolates. However, risk stratification analysis indicated that although ST198 was detected less frequently, it carried a significantly greater burden of acquired resistance genes and therefore represented a higher-risk resistant lineage. Functional and structural validation further elucidated the molecular basis of resistance dissemination within this high-risk lineage. Conjugation experiments confirmed the co-transfer of a multidrug resistance module carrying blaTEM-1 and blaCTX-M-267 to the recipient strain Escherichia coli J53. Complete plasmid analysis revealed that these two β-lactam resistance genes were co-localized on a 242-kb transferable plasmid flanked by Tn1331, Tn3, and multiple transposase-associated elements, thereby providing a structural basis for their horizontal transfer. This study provides important molecular epidemiological evidence for lineage-specific surveillance and risk-stratified control of resistant Salmonella in the poultry meat supply chain and further underscores the need for continuous monitoring of mobile genetic elements within a One Health framework.

Animals

Comparative analyses of olfactory receptor repertoires in Schizothorax fish based on the chromosome-level genomes: Implications for regulatory roles of dietary differentiation and ploidy variation.

The olfactory receptor (OR) genes constitute the molecular basis of fish olfaction, mediating survival behaviors and environmental adaptation while coevolving with habitat-driven evolution. Schizothorax, a cyprinid genus endemic to the Qinghai-Tibetan Plateau, exhibits remarkable dietary divergence and ploidy variation in response to plateau environmental changes, which presumably facilitates the adaptive evolution of OR genes. However, the evolutionary patterns of OR genes associated with trophic divergence and ploidy variation in this genus remain unclear. In this study, three species were selected: the herbivorous diploid S. macropogon, the carnivorous diploid S. lantsangensis, and the herbivorous tetraploid S. curvilabiatus. S. macropogon possessed 142 OR genes (92.25% functional), primarily located on chromosomes 14 and 24, with the fewest sequence clusters. Such compact gene repertoire and highly overlapping chromosomal clusters indicated specialization for a herbivorous olfactory niche. S. lantsangensis contained 127 OR genes (93.70% functional), concentrated on chromosomes 4 and 5, with fewer sequence clusters and a scattered distribution, reflecting evolution of OR genes under carnivorous feeding habits. The herbivorous tetraploid S. curvilabiatus exhibited striking features: 316 OR genes (94.30% functional), the most subfamilies, unique ε and κ OR subfamilies, and species-specific motifs. These characteristics revealed that ploidy, rather than herbivory, dominated OR gene evolution. In conclusion, dietary differentiation and ploidy variation together drove olfactory adaptive evolution in Schizothorax, providing new insights into vertebrate OR gene ecological adaptation.

Animals

Genomic epidemiology of clinically critical antibiotic resistance in Salmonella enterica causing bloodstream infections across six Chinese provinces, 1994-2023.

Clinically critical antibiotic-resistant Salmonella enterica (S. enterica) causing bloodstream infections remains a public health challenge. Here, we aim to reveal the emergence and trends of clinically important antibiotic resistance in S. enterica causing bloodstream infections using 833 isolates from six Chinese provincial-level administrative areas during 1994-2023. We identified 48 serovars and 64 sequence types (STs). Overall, 8.52% of 833 isolates were resistant or had decreased susceptibility to ciprofloxacin, 4.32% and 6.84% reported resistance or decreased susceptibility to third- and fourth-generation cephalosporins (3GCs and 4GCs), 1.80% reported resistance to fosfomycin, and 2.16% reported resistance to azithromycin. Across these six regions, azithromycin and fosfomycin resistance is increasing, as is decreased susceptibility or resistance to ciprofloxacin, 3GCs, and 4GCs, especially among younger children and elderly people. Clinically prioritized antibiotic resistance also varies by region, serovar, and age group. S. Paratyphi A genotype 2.3.3 strains are mainly divided into 2 lineages distributed in Guangxi and Shanghai. Within the scope of this passive surveillance dataset, S. Typhi genotype 4.3.1.2.1 was identified as the earliest documented case among the collected isolates. Our retrospective and longitudinal genomic epidemiology study provides critical data for the formulation of treatment guidelines and policies for bloodstream infections and for the monitoring and control of antimicrobial resistance.

Humans

PGR expression as a pharmacogenomic companion biomarker to GENE70-derived genomic risk in ER-positive/HER2-negative breast cancer.

BACKGROUND: The biology of the estrogen receptor-positive (ER+) and human epidermal growth factor receptor 2-negative (HER2-) breast cancers is heterogeneous even when they are categorized by their risk via genomics. Transcriptomic PGR expression reflects endocrine pathway activity and may provide complementary biological information within established GENE70-derived genomic-risk categories. Whether this molecular marker improves the biological interpretation of genomic-risk stratification beyond conventional clinicopathological assessment remains uncertain. OBJECTIVES: The aim of this study was to determine whether transcriptomic PGR expression provides complementary biological and prognostic information within reconstructed GENE70-derived genomic-risk categories and refines the characterization of endocrine-related tumour biology in ER-positive/HER2-negative breast cancer. METHODS: This study analysed publicly available transcriptomic and clinical data from three cohorts: METABRIC (discovery cohort), GSE96058/SCAN-B cohort (validation cohort) and TCGA-BRCA cohort (molecular validation cohort). The GENE70-derived genomic-risk score was reconstructed for each cohort using matched genes. Cox regression, Kaplan-Meier analysis and subgroup comparisons were used to assess relationships between PGR expression, clinicopathologic variables, molecular features and survival outcomes. RESULTS: Across the three independent cohorts, low transcriptomic PGR expression was consistently associated with higher GENE70-derived genomic risk, increased MKI67 expression, reduced ESR1 expression and enrichment of the Luminal B subtype. Survival findings differed between cohorts. In the discovery METABRIC cohort, transcriptomic PGR expression showed heterogeneous associations with survival, particularly within GENE70-derived high-risk subgroups, whereas the external GSE96058/SCAN-B validation cohort demonstrated consistent associations between low PGR expression and poorer overall survival in both the overall ER-positive/HER2-negative population and GENE70-derived high-risk subgroups. CONCLUSION: These findings suggest that transcriptomic PGR provides complementary biological and prognostic information within GENE70-derived genomic-risk categories. However, because treatment response was not evaluated in the present study, the findings should not be interpreted as evidence of predictive or pharmacogenomic utility and prospective studies incorporating treatment-response analyses are required before such applications can be established.

Humans

Conserved host-exclusive oligonucleotide motifs enriched in pathogenic genes of human oncogenic viruses.

Comparative viral genomics can reveal sequence-level constraints influencing virus-host interactions. Relative minimal absent words (rMAWs) are short oligonucleotide motifs present in viral genomes but completely absent from the host, potentially reflecting selective pressures related to host adaptation and immune evasion. Using the EAGLE algorithm and the GRCh38 human reference genome, we systematically screened for prevalent rMAWs (prMAWs) across six major human oncogenic viruses: Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human papillomavirus (HPV), human T-cell leukemia virus type 1 (HTLV-1), and human herpesvirus 8/Kaposi's sarcoma-associated herpesvirus (HHV-8/KSHV). highly conserved 11- and 12-bp prMAWs were identified in EBV, HBV, HTLV-1, and HHV-8/KSHV, with sequence prevalences ranging from 91.5% to 97.9%. Conversely, no short prMAWs were detected in HCV or HPV, likely reflecting differences in genome architecture, mutation rates, and long-term host adaptation to the human host. Importantly, the identified host-exclusive motifs exhibited non-random genomic distribution and were preferentially embedded within viral genes central to replication, persistence, immune modulation, and oncogenesis, including EBNA-1 (EBV), HBx (HBV), Tax-associated regions (HTLV-1), and lytic replication genes of HHV-8/KSHV. Notably, all detected prMAWs were enriched in GC nucleotides and exhibited marked CpG over-representation, suggesting sequence constraints associated with epigenetic regulation and viral persistence. Collectively, these highly conserved, host-exclusive signatures offer promising, candidates for sequence-directed approaches in the diagnosis, monitoring, and investigation of virus-associated cancers.

Humans