Search PubMedSearch

SEARCH · Search PubMed

Results for “Sequence Analysis, RNA”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,667 records · Page 9Linked to original sources

Biomarker Analysis from Patients with Metastatic PDAC Treated with TGFβ Antibody NIS793 plus Abraxane + Gemcitabine versus Abraxane + Gemcitabine Alone in a Phase II, Open-Label, Randomized Study.

PURPOSE: Transforming growth factor β (TGFβ) plays a dual role in cancer, acting as a tumor suppressor early in the disease but promoting progression and immune evasion when dysregulated. In pancreatic ductal adenocarcinoma (PDAC), TGFβ-driven desmoplasia fosters chemoresistance and immunosuppression, limiting therapeutic efficacy. NIS793, a fully human mAb targeting TGFβ, demonstrated antifibrotic and immunomodulatory activity in preclinical models and early-phase trials. PATIENTS AND METHODS: We conducted a randomized, open-label, phase II study in treatment-naïve patients with metastatic PDAC (mPDAC) to evaluate NIS793 ± spartalizumab (anti-PD-1) combined with nab-paclitaxel (or Abraxane)/gemcitabine (ABRA/GEM) versus ABRA/GEM alone. The primary endpoint was progression-free survival (PFS); secondary endpoints included overall survival (OS), safety, pharmacokinetics, and biomarker analyses. Exploratory assessments included paired tumor RNA sequencing, cell-free DNA profiling, and plasma proteomics. RESULTS: NIS793 demonstrated target engagement and suppression of TGFβ signaling, confirmed by transcriptomic and proteomic analyses. Stromal remodeling was evident, with significant downregulation of cancer-associated fibroblast markers (Acta2, Fap) and collagen-related signatures. Despite proof of mechanism, clinical efficacy was not observed: Median PFS and OS were comparable or numerically worse in the NIS793 arm versus control (HR for OS in NIS793 + ABRA/GEM vs. ABRA/GEM: 1.32; 95% confidence interval, 0.84-2.07). The safety profile was manageable, with no unexpected toxicities. Biomarker data revealed increased expression of neutrophil-related genes after treatment, suggesting potential induction of tumor-promoting inflammation. CONCLUSIONS: NIS793 effectively inhibited TGFβ signaling and led to stromal remodeling but failed to improve outcomes in mPDAC. These findings highlight the complexity of TGFβ biology and caution against its blockade in combination with chemotherapy for PDAC. Future strategies should consider context-dependent effects of TGFβ inhibition.

Humans

The cold case of state transition 7 (stt7) mutants of Chlamydomonas reinhardtii, solved by whole-genome sequencing.

The process of State Transitions (ST) corresponds to an STT7 kinase-driven redistribution of the transmembrane LHCII antenna proteins between Photosystem II (PSII) and Photosystem I (PSI), which results from changes in their phosphorylation state. For the past two decades, two LHCII-kinase mutants, stt7-1 and stt7-9, have been instrumental in the study of STs in Chlamydomonas reinhardtii, the former being a null mutant for the kinase but quasi-sterile in crosses, while the latter, although fertile, has a leaky phenotype. Using long-read sequencing, this study further characterized the genetic lesions of the stt7 mutant strains through whole-genome reconstruction and de novo chromosome assembly. In addition, two new stt7 null mutants were generated, one derived by crosses from the original stt7-1 and one obtained by Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated protein 9 (Cas9) technology. This work provides a comprehensive genomic characterization of the original stt7-1 null mutant, revealing extensive chromosomal rearrangements and high levels of aneuploidy, associated with increased cell size and meiotic dysfunction. Reassessment of their physiology and genetic backgrounds highlights the need for caution in interpreting genetic information. We thus produced more reliable null mutants for the LHCII-kinase, amenable to genetic crosses for the study of STs in a variety of genetic backgrounds.

Chlamydomonas reinhardtii

Complete genome sequence of multidrug-resistant Salmonella enterica subsp. enterica serovar Enteritidis SD191 isolated from chicken liver, harboring a novel imipenem resistance mechanism.

We present the complete genome sequence of Salmonella enterica subsp. enterica serovar Enteritidis SD191 isolated from Gallus gallus liver in China, harboring plasmid pSE191. The genome reveals multiple antibiotic resistance mechanisms and phenotypic imipenem resistance without canonical genes.

antibiotic resistance

Microbial signal profiles and organism-level concordance between plasma metagenomic sequencing and blood culture in suspected bloodstream infection.

Plasma metagenomic next-generation sequencing (mNGS) and blood culture detect different components of the microbial signal and frequently produce discordant organism reports. We characterized microbial signal class, report-derived burden, organism-level concordance, and independent clinical attribution in a retrospective, single-center, episode-level cohort. Among 329 episodes with evaluable plasma mNGS reports, 315 had blood culture performed; 232 were mNGS positive/culture negative and 53 were positive by both methods. In the 232 discordant episodes, the recorded routine-care diagnosis classified 124 as bloodstream infection (BSI) and 108 as non-BSI. Nonviral signals were present in 78.2% and 42.6%, respectively (P&#x2009;<&#x2009;0.001), and median maximum report-derived sequence counts were 98.5 and 11.5 (P&#x2009;<&#x2009;0.001). Two laboratory physicians then independently reviewed source records using structured criteria while masked to the recorded BSI label and mNGS organism and sequence-count information. Initial agreement for the five-category BSI assessment was 97.6% (Cohen's kappa, 0.960). Within the mNGS-positive/culture-negative subgroup, adjudicated BSI likelihood showed a modest ordinal association with report burden (Spearman rho&#x2009;=&#x2009;0.190; P&#x2009;=&#x2009;0.004), while mNGS organisms were considered supported in 1 episode, plausible in 158, unlikely or contaminant in 72, and unresolved in 1. Among 53 dual-positive episodes, 33 (62.3%) shared at least one species, but only 5 (9.4%) had complete species-set concordance. Plasma mNGS and blood culture therefore frequently generated non-equivalent organism sets. Signal class and report burden contributed graded contextual evidence, but organism-level attribution required clinical review and orthogonal microbiology rather than binary positivity alone.

Humans

Cost-Effectiveness and the Economics of Genomic Testing and Molecularly Matched Therapies.

Cost-effectiveness analysis of precision oncology can help guide value-driven care. Next-generation sequencing is increasingly cost-efficient over single gene testing because diagnostic algorithms require multiple individual gene tests to determine biomarker status. Matched targeted therapy is often not cost-effective due to the high cost associated with drug treatment. However, genomic profiling can promote cost-effective care by identifying patients who are unlikely to benefit from therapy. Additional applications of genomic profiling such as universal testing for hereditary cancer syndromes and germline testing in patients with cancer may represent cost-effective approaches compared with traditional history-based diagnostic methods.

Humans

Prenatal exome sequencing of fetuses with central nervous system anomalies based on prenatal ultrasound and magnetic resonance imaging diagnosis: A retrospective cohort study with a systematic review and meta-analysis.

INTRODUCTION: Fetal central nervous system (CNS) abnormalities have diverse etiologies, with genetic factors as a major contributor. Prenatal exome sequencing (ES) is a powerful tool for precise molecular diagnosis of CNS anomalies, but its diagnostic yield varies among studies. This study aimed to evaluate the additional diagnostic yield of prenatal ES compared with chromosomal microarray analysis (CMA) in fetuses with CNS anomalies detected by prenatal imaging. MATERIAL AND METHODS: We collected ES results from fetuses diagnosed with CNS anomalies by prenatal imaging (2019-2024) who had negative results. Subgroup analyses assessed phenotype-specific ES diagnostic yield for associated genes and variants. A systematic review and meta-analysis incorporating our data and published studies further explored the association between phenotype and diagnostic yield. RESULTS: In the cohort study of 219 cases, ES identified pathogenic/likely pathogenic single nucleotide variations in 36 cases (16%). The highest diagnostic yield of ES was in cases with multisystem malformations (25%, 14/55), followed by multiple CNS anomalies (15%, 2/13) and isolated CNS anomalies (13%, 20/151). The most commonly identified isolated CNS anomaly was agenesis of the corpus callosum (31%, 5/16). Neural tube defects with urogenital anomalies were associated with a positive ES finding in 57% (4/7) of cases. The meta-analysis of 989 cases from 22 studies showed a pooled diagnostic yield of ES of 27% (95% CI, 21%-34%). The highest diagnostic yield of ES was in cases of corpus callosum anomalies with facial abnormalities (75%, 8/11) and neural tube defects with urogenital malformations (80%, 12/15). The diagnostic yield of ES for three or more CNS abnormalities was 43% (95% CI, 31%-58%), significantly higher than that for only two abnormalities (10%, 95% CI, 4%-18%). No significant difference in diagnostic yield was found between cases identified by prenatal MRI combined with ultrasound (27%, 95% CI, 20%-36%) and those identified by ultrasound alone (25%, 95% CI, 17%-35%). CONCLUSIONS: ES provided a significantly higher diagnostic yield than CMA for fetal CNS abnormalities, with diagnostic yields varying by phenotype. The systematic review and meta-analysis confirmed that the complexity and combination of malformations are key factors associated with differences in ES diagnostic yield.

Humans

LitCTL1: A novel C-type lectin involved in the mucosal and cellular immunity of the common periwinkle Littorinalittorea.

C-type lectins (CTLs) are vital pattern-recognition receptors (PRRs) that mediate innate immune responses in mollusks, yet their characterization in Caenogastropoda, the largest gastropod group, remains limited. This study characterizes LitCTL1, a novel secreted single-domain C-type lectin from the common periwinkle, Littorina littorea. The 199-amino acid polypeptide contains a conserved carbohydrate recognition domain with canonical QPD and WND motifs and is predicted to form a homodimer. Uniquely, LitCTL1 was localized in both circulating hemocytes and mucus-secreting epithelial cells of the foot, mantle, and hypobranchial gland - the first report of such dual localization for a molluscan lectin, linking systemic and mucosal defense. Expression analysis revealed that LitCTL1 is constitutively expressed in hemocytes. Functional assays with recombinant LitCTL1 demonstrated its role as a potent opsonin with hemagglutinating activity, significantly enhancing hemocyte spreading and the phagocytosis of zymosan. Genomic analysis reveals that LitCTL1 belongs to a rapidly diversifying, genus-specific expansion distinct from conserved perlucin-like lineages. These results identify LitCTL1 as a key effector molecule in both systemic and mucosal innate immunity, likely reflecting an evolutionary adaptation to the microbial challenges of the intertidal environment.

Animals

A genome-wide coverage-based pipeline for the identification of host-derived candidate DNA biomarkers from cell-free blood.

We have created a new data-analysis pipeline for the discovery of host-specific candidate DNA biomarkers derived from sequencing data of cell-free blood. Unlike approaches that rely on specific molecular or genetic signatures, our method leverages the coverage distribution of cell-free DNA sequences mapped to a reference genome, applying statistical analyses to identify informative short genomic regions for biomarker discovery. The pipeline is applicable to diverse diseases and can be used to analyze cell-free DNA sequences from plasma or serum to identify candidate biomarkers that are characteristic of disease states in mammals. Core functionalities were developed in Java and integrated with open-source software tools for the preprocessing of raw sequencing data, complemented by Python scripts for the machine-learning analysis and statistical validation. The pipeline is designed for HPC use and users can access the pipeline through a Galaxy workflow, which offers a user-friendly web interface for input selection prior to execution and analysis progress monitoring. Performance tests, carried out using duplicate sets of COVID-19 samples and controls, showed linear scalability of execution time with an increasing dataset size, as well as a substantial reduction in execution time through parallelized computation, whereby each HPC node is used to process the data of one chromosome. Further statistical tests confirmed the quality of the pipeline's results by showing that the set of identified candidate biomarkers remained stable across varying dataset sizes.

Biomarkers

Diagnostic value of plasma cell-free DNA metagenomic next-generation sequencing in patients with suspected infections and exploration of clinical scenarios-a retrospective study from a single center.

BACKGROUND: Plasma cell-free DNA metagenomic next-generation sequencing (mNGS) is a non-invasive comprehensive method for the etiological diagnosis of various infectious diseases. However, research on the early diagnosis and real-world clinical impact of plasma mNGS in patients with suspected infection are still limited. MATERIALS AND METHODS: This study retrospectively included 140 patients with suspected infections who underwent early plasma mNGS and conventional culture testing. Referring to the clinical diagnosis of infectious diseases, the diagnostic performance of plasma mNGS and culture tests was compared, and the application scenarios and clinical effects of plasma mNGS were evaluated. RESULTS: The positive rate of plasma mNGS was significantly higher than that of culture methods (55.71% vs 25.10%, p&#x2009;<&#x2009;0.001) and blood cultures (55.71% vs 12.86%, p&#x2009;<&#x2009;0.001). Regarding clinical diagnosis, the sensitivity of plasma mNGS was significantly higher than that of culture (58.27% vs 37.80%, p&#x2009;=&#x2009;0.002). The combination of mNGS and culture achieved a higher detection sensitivity (69.29%), especially in patients with multi-site co-infections (73.68%) and blood infections (73.17%). Plasma mNGS demonstrated higher sensitivity in patients with procalcitonin (PCT) index > 5&#x2009;ng/ml or human neutrophil lipocalin (HNL) index > 200&#x2009;ng/ml. In terms of treatment, a total of 69 patients (54.33%) benefited from plasma mNGS. CONCLUSION: This study highlights the significant improvement in pathogen detection performance by combining conventional culture with plasma mNGS detection, especially in patients with multi-site co-infections and blood infections. Early use of plasma mNGS as an adjunct to culture can better guide clinicians to initiate appropriate anti-infective therapy.

Humans

Mapping antibody sequences and effector functions across spatial niches.

Antibodies are fundamental to human health but can also drive pathology. Each antibody has a molecular specificity, encoded by their clonally heritable B cell receptor (BCR). Recent advances in spatial transcriptomics coupled with repertoire sequencing have enabled capturing antibody-secreting cells (ASCs) and their clonal BCR within their tissue microenvironment. However, our understanding of antibody production niches remains limited. Furthermore, where antibodies are produced can be distinct from where antibodies exert their effector function. Here, we propose a conceptual spatial framework to distinguish between 'antibody production niches', defined by the ASC, BCR, and niche composition, versus 'antibody functional niches', composed of the antibody, antigen, and effector landscape. We then examine the possibilities and challenges to map and link antibody-encoding sequences and antibody effector functions using current and emerging technologies. Combined, we argue that integrating spatial sequence data with the antibody functional context is essential to decode the architecture of antibody-mediated immunity.

Humans

Comparison of paralog identification methods and their impact on species tree topologies in target capture phylogenomics within the Sindora clade (Detarioideae: Leguminosae).

Target capture is a common method of generating high throughput DNA sequencing data for phylogenetic reconstruction of species relationships, for which single copy genes are usually most informative. However, a pervasive problem with target capture is that putatively single copy genes may in fact be paralogs resulting from gene duplication, which are problematic for phylogenetic inference because their evolutionary history may differ from the divergence history of species. Here, we use as a case study a target enrichment dataset of 88 species of Detarioideae (Leguminosae) with a focus on the Sindora clade to examine approaches for handling paralogs, including the built-in paralog handling functions in HybPiper and CAPTUS, plus subsequent steps using Putative Paralog Detection and the tree-based Yang & Smith orthology inference approach. We compare the paralogs flagged using these methods and verify their performance with BLAST mapping against a reference genome sequence of Sindora glabra, and then subsequently compare the species tree topologies produced across these methods. Our comparisons of paralogs flagged across the Sindora clade show that the Putative Paralog Detection pipeline was the most accurate in identifying paralogs in terms of its similarity to the BLAST mapping, followed by the built-in paralog identification function of CAPTUS. However, the results we recovered for the Detarioideae subfamily suggest that the largest differences in species tree topology resulted from the use of paralog-filtered alignments (such as with the Putative Paralog Detection pipeline and the Yang & Smith orthology inference approaches) rather than just by removing the sequences of identified paralogous genes. This was the true for HybPiper-assembled datasets but was not seen in CAPTUS-assembled datasets. In all comparisons, the topological differences caused by different paralog handling methods tended to be confined to clades where processes such as hybridisation and introgression are prevalent. Our study provides a roadmap to establish the best approach to identify, eliminate or separate paralogs in the absence of a chromosomally contiguous reference genome for a study group, and highlights the importance of careful data inspection and processing in addition to understanding the extent of paralogy and paralog characteristics (e.g. sequence divergence between copies) for their study group.

Phylogeny

Nanopore-based epigenomic profiling reveals the absence of widespread CpG methylation in the African swine fever virus genome.

DNA methylation is a critical epigenetic mechanism implicated in regulating replication and transcription in DNA viruses. However, the epigenetic landscape of African swine fever virus (ASFV), a large double-stranded DNA virus infecting pigs, remains controversial. Here, we systematically profiled the DNA methylome of the first ASFV strain isolated in Hong Kong (HK_NT_202103) using Oxford Nanopore Technologies (ONT) R10.4.1 sequencing. We employed a paired design: native whole-genome sequencing (WGS) against a methylation-free whole-genome amplification (WGA) control. Using conservative thresholds, we found no evidence of 5-methylcytosine (5mC), especially typical CpG methylation, across the viral genome. Importantly, clear CpG methylation signals were successfully detected in the host genome from WGS data, confirming the functionality of the workflow to detect 5mC at CG sites. While widespread 5mC seems absent, a small number of putative N6-methyladenine (6mA) loci were identified. A specific 6mA candidate exhibited raw ionic current disruptions and gene-level intersection with another ASFV isolate (CAS19-01/2019), although it lacked single-base consensus across different methylation callers or between the two isolates. Although our biological findings are restricted to a single isolate under specific experimental conditions, this study introduces a novel, highly rigorous ONT framework for viral epigenomics research. Furthermore, the absence of ASFV CpG methylation indicates that host CpG-depletion remains a viable strategy for viral metagenomic enrichment. Ultimately, our work offers a critical methodological baseline for ASFV surveillance and highlights the necessity of targeted experimental validation for rare viral modifications.

African Swine Fever Virus

A conserved distal-tail helical extension defines a tailspike attachment architecture in Gram-negative siphophages.

Rapid growth of bacteriophage genome collections has outpaced functional annotation of tail-tip proteins, limiting comparative analysis of host-recognition structures. Starting from a shared distal-tail gene organization in the Salmonella phages 9NA and Jersey, I developed a morphogenetic bioinformatic framework integrating gene synteny, sequence comparison, profile hidden Markov model (HMM) screening, structural evidence, structure-aware searching, and AlphaFold modeling. Comparison with the experimentally characterized lambda and Sf11 tail assemblies identified a predominantly alpha-helical C-terminal extension of the distal-tail (DT) protein associated with tailspike attachment, termed the distal-tail helical extension (DT-helix). Screening 541,986 proteins from 5167 complete NCBI RefSeq tailed-phage genomes, followed by evidence-based evaluation of sequence, genomic context, and structural architecture, identified 165 curated DT-helical-extension-associated phages. Their DT proteins segregated into six sequence groups. In the four principal multi-member groups, cognate tailspikes showed group-specific conservation in proximal N-terminal regions but substantially greater downstream diversity, consistent with sequence constraint at the DT-tailspike attachment boundary. A complementary ProstT5/Foldseek search supported the established groups but revealed no convincing additional highly divergent family. Together with the experimentally characterized Sf11 attachment interface, these findings define a recurrent morphogenetic architecture linking conserved distal-tail scaffolds to more variable receptor-binding proteins across siphophages infecting Gram-negative bacteria. Although universal exchangeability is not established, the identified scaffold-receptor-binding boundaries provide a framework for molecular characterization and rational phage engineering. Accession-level information for the 165 curated phages is available through PhageTailDB.

Viral Tail Proteins

Mitochondrial DNA diversity in Ecuadorian populations: Recurrence of variant 16136 within haplogroup B2.

The identification of lineage-defining variants, frequently found in the coding region of mitochondrial DNA (mtDNA), is essential for refining haplogroup classification. Most mtDNA studies in South American populations have focused on the control region (CR), which has provided important insights into population structure and maternal lineage origins, although information needed for more robust phylogenetic resolution has been neglected. This study investigates the maternal genetic structure of Ecuadorian populations by combining CR and whole mitogenome analyses. Sequences from the mtDNA CR were obtained from 461 individuals (253 Mestizos and 208 Native Americans), while complete mitogenomes were sequenced for 127 individuals to improve phylogenetic resolution by identifying lineage-defining variants present in coding region. Most mtDNA haplogroups in the two population groups analyzed were of Native American origin (A2, B2, B4, C1, D1, D4), with significant differences in the distribution of specific lineages between them. Among Mestizos, African haplogroups (all within the L branches) and Eurasian haplogroups (H, K, R, U) were detected at low frequencies, whereas no African lineages were observed among Native Americans. The results obtained highlighted a heterogeneity within Ecuadorian populations that must be considered when developing mtDNA haplotype databases for forensic purposes. Whole mitogenome sequences enabled the identification of variants that refined haplogroup classifications, provided a more accurate reconstruction of the maternal genetic diversity, and improve the discrimination between Native American and Asian maternal lineages within haplogroup B4b.

Humans

Whole-Exome Sequencing in a Consanguinity-Enriched South Indian Retinitis Pigmentosa Cohort: Diagnostic Yield and Molecular Spectrum.

PURPOSE: To determine the molecular diagnostic yield, variant spectrum, inheritance architecture, and influence of consanguinity on whole-exome sequencing outcomes in a South Indian retinitis pigmentosa (RP) cohort. DESIGN: Prospective, registry-based cohort study. SUBJECTS: A total of 113 affected participants were enrolled through the Aravind Registry for Inherited Diseases of the Eye, including 109 unrelated probands and 4 affected relatives from already represented families. Primary analyses were restricted to the 109 unrelated probands. METHODS: Whole-exome sequencing was performed using a clinical exome workflow. Variants were interpreted using American College of Medical Genetics and Genomics/Association for Molecular Pathology criteria and cases were categorized as solved, possibly solved, inconclusive, or unsolved using prespecified inheritance-aware rules. MAIN OUTCOME MEASURES: Molecular diagnostic yield, distribution of implicated genes and variant classes, inheritance architecture, and diagnostic yield stratified by consanguinity status. RESULTS: Among the 109 unrelated probands, mean age at testing was 39.3 &#xb1; 14.1 years and 58.7% were male. Whole-exome sequencing identified 186 distinct rare variants across 92 inherited retinal disease genes, including 26 pathogenic and 33 likely pathogenic variants. A molecular diagnosis was established in 50 of 109 probands (45.9%), including 42 solved and 8 possibly solved cases; 45 (41.3%) were inconclusive and 14 (12.8%) remained unsolved, including 4 (3.7%) in whom no candidate variant was identified. EYS, USH2A, and ADGRV1 were the most frequently implicated genes. Autosomal recessive (AR) disease predominated (44/50, 88.0%). Consanguineous AR cases were exclusively homozygous (17/17); notably, 68.0% of nonconsanguineous AR cases were also homozygous (P = 0.013). Diagnostic yield was higher in consanguineous probands (51.4% vs. 41.7%), without reaching significance. Recurrent alleles included an established South Asian founder variant (MFSD8 c.1361T>C) and candidate founder alleles in EYS (c.4321C>T) and ADGRV1 (c.14329C>T). CONCLUSIONS: Whole-exome sequencing established a molecular diagnosis in nearly half of this South Indian RP cohort and revealed a predominantly recessive, homozygosity-enriched architecture shaped by consanguinity. These findings define a region-specific variant landscape to support clinical interpretation, genetic counseling, and future trial enrollment in this underrepresented population. FINANCIAL DISCLOSURES: The authors have no proprietary or commercial interest in any materials discussed in this article.

Consanguinity

Exploring the mechanism of aroma production in fermented cherry juice by L. brevis LD1.0600 using flavomics and whole genome analysis.

This study focused on L.brevis LD1.0600 with excellent fermentation traits: it analyzed genome-wide key regulatory genes for micro-metabolites, combined with fermented cherry juice flavor metabolomics data, and used machine learning to explore correlations between gene regulation, metabolite production, and flavor formation. The SVM model screened and verified fermented cherry juice VOCs; through OAV and flavor wheel analysis, LD1.0600 emerged as the top-performing strain, with a sweet, fruity dominant aroma. Key aroma-active components (OAV&#xa0;>&#xa0;100) included 2-methoxy-4-vinylphenol, benzaldehyde, 2-methyl-butanoic acid and hexanoic acid, and 2-methoxy-4-vinylphenol and hexanoic acid elevated by LD1.0600-regulated genes (Chrom1-001884, Chrom1-000925, fabF and Chrom1-000199). At the same time, through research, a "strain screening-SVM screening of DVCs-OAV screening of key aroma components-whole genome sequencing of flavor regulatory genes" system was established. This system can not only be applied to the screen fermentation strains, but also can be extended to the application of other fermentation products.

Fermentation

Conserved host-exclusive oligonucleotide motifs enriched in pathogenic genes of human oncogenic viruses.

Comparative viral genomics can reveal sequence-level constraints influencing virus-host interactions. Relative minimal absent words (rMAWs) are short oligonucleotide motifs present in viral genomes but completely absent from the host, potentially reflecting selective pressures related to host adaptation and immune evasion. Using the EAGLE algorithm and the GRCh38 human reference genome, we systematically screened for prevalent rMAWs (prMAWs) across six major human oncogenic viruses: Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human papillomavirus (HPV), human T-cell leukemia virus type 1 (HTLV-1), and human herpesvirus 8/Kaposi's sarcoma-associated herpesvirus (HHV-8/KSHV). highly conserved 11- and 12-bp prMAWs were identified in EBV, HBV, HTLV-1, and HHV-8/KSHV, with sequence prevalences ranging from 91.5% to 97.9%. Conversely, no short prMAWs were detected in HCV or HPV, likely reflecting differences in genome architecture, mutation rates, and long-term host adaptation to the human host. Importantly, the identified host-exclusive motifs exhibited non-random genomic distribution and were preferentially embedded within viral genes central to replication, persistence, immune modulation, and oncogenesis, including EBNA-1 (EBV), HBx (HBV), Tax-associated regions (HTLV-1), and lytic replication genes of HHV-8/KSHV. Notably, all detected prMAWs were enriched in GC nucleotides and exhibited marked CpG over-representation, suggesting sequence constraints associated with epigenetic regulation and viral persistence. Collectively, these highly conserved, host-exclusive signatures offer promising, candidates for sequence-directed approaches in the diagnosis, monitoring, and investigation of virus-associated cancers.

Humans