Search PubMedSearch

SEARCH · Search PubMed

Results for “Sequence analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,620 records · Page 2Linked to original sources

Cost-Effectiveness and the Economics of Genomic Testing and Molecularly Matched Therapies.

Cost-effectiveness analysis of precision oncology can help guide value-driven care. Next-generation sequencing is increasingly cost-efficient over single gene testing because diagnostic algorithms require multiple individual gene tests to determine biomarker status. Matched targeted therapy is often not cost-effective due to the high cost associated with drug treatment. However, genomic profiling can promote cost-effective care by identifying patients who are unlikely to benefit from therapy. Additional applications of genomic profiling such as universal testing for hereditary cancer syndromes and germline testing in patients with cancer may represent cost-effective approaches compared with traditional history-based diagnostic methods.

Humans

Diving Deeper Into Mechanisms of Acrylamide-Induced Toxicity: RNA Sequencing Reveals Transcriptomic Alteration and Retrotransposon Expression in Drosophila melanogaster.

Given the inevitability of human and animal exposure to acrylamide, there is increasing concern regarding its potential health risks. While a number of molecular mechanisms have been proposed, the complexity of acrylamide toxicological pathways and interactions remains incompletely characterized. In this study, we employed a transcriptomic approach to investigate the transcriptional responses of Drosophila melanogaster following exposure to acrylamide (100 mg/kg). Our analysis identified 634 differentially expressed genes (DEGs), with 362 upregulated and 272 downregulated. Functional analysis revealed these DEGs are enriched in pathways related to reproduction, detoxification, cellular and metabolic processes, signaling, synaptic formation and organization. Notably, acrylamide exposure upregulated the expression of tau and beta-amyloid protein precursor-like genes, both implicated in Alzheimer's disease pathology. An aversive memory test further demonstrated that acrylamide impaired the short-term memory of treated flies. Additionally, acrylamide-induced toxicity altered the expression of nine long terminal repeat retrotransposons, belonging to the gypsy and pao superfamilies. By exploring the potential role of transposable element activity in acrylamide-mediated toxicity, this study provides novel insights into the molecular mechanisms underlying its effects. Collectively, these findings offer a more comprehensive understanding of the mechanisms and pathways associated with the toxic action and detoxification of acrylamide in D. melanogaster.

Animals

Prenatal exome sequencing of fetuses with central nervous system anomalies based on prenatal ultrasound and magnetic resonance imaging diagnosis: A retrospective cohort study with a systematic review and meta-analysis.

INTRODUCTION: Fetal central nervous system (CNS) abnormalities have diverse etiologies, with genetic factors as a major contributor. Prenatal exome sequencing (ES) is a powerful tool for precise molecular diagnosis of CNS anomalies, but its diagnostic yield varies among studies. This study aimed to evaluate the additional diagnostic yield of prenatal ES compared with chromosomal microarray analysis (CMA) in fetuses with CNS anomalies detected by prenatal imaging. MATERIAL AND METHODS: We collected ES results from fetuses diagnosed with CNS anomalies by prenatal imaging (2019-2024) who had negative results. Subgroup analyses assessed phenotype-specific ES diagnostic yield for associated genes and variants. A systematic review and meta-analysis incorporating our data and published studies further explored the association between phenotype and diagnostic yield. RESULTS: In the cohort study of 219 cases, ES identified pathogenic/likely pathogenic single nucleotide variations in 36 cases (16%). The highest diagnostic yield of ES was in cases with multisystem malformations (25%, 14/55), followed by multiple CNS anomalies (15%, 2/13) and isolated CNS anomalies (13%, 20/151). The most commonly identified isolated CNS anomaly was agenesis of the corpus callosum (31%, 5/16). Neural tube defects with urogenital anomalies were associated with a positive ES finding in 57% (4/7) of cases. The meta-analysis of 989 cases from 22 studies showed a pooled diagnostic yield of ES of 27% (95% CI, 21%-34%). The highest diagnostic yield of ES was in cases of corpus callosum anomalies with facial abnormalities (75%, 8/11) and neural tube defects with urogenital malformations (80%, 12/15). The diagnostic yield of ES for three or more CNS abnormalities was 43% (95% CI, 31%-58%), significantly higher than that for only two abnormalities (10%, 95% CI, 4%-18%). No significant difference in diagnostic yield was found between cases identified by prenatal MRI combined with ultrasound (27%, 95% CI, 20%-36%) and those identified by ultrasound alone (25%, 95% CI, 17%-35%). CONCLUSIONS: ES provided a significantly higher diagnostic yield than CMA for fetal CNS abnormalities, with diagnostic yields varying by phenotype. The systematic review and meta-analysis confirmed that the complexity and combination of malformations are key factors associated with differences in ES diagnostic yield.

Humans

LitCTL1: A novel C-type lectin involved in the mucosal and cellular immunity of the common periwinkle Littorinalittorea.

C-type lectins (CTLs) are vital pattern-recognition receptors (PRRs) that mediate innate immune responses in mollusks, yet their characterization in Caenogastropoda, the largest gastropod group, remains limited. This study characterizes LitCTL1, a novel secreted single-domain C-type lectin from the common periwinkle, Littorina littorea. The 199-amino acid polypeptide contains a conserved carbohydrate recognition domain with canonical QPD and WND motifs and is predicted to form a homodimer. Uniquely, LitCTL1 was localized in both circulating hemocytes and mucus-secreting epithelial cells of the foot, mantle, and hypobranchial gland - the first report of such dual localization for a molluscan lectin, linking systemic and mucosal defense. Expression analysis revealed that LitCTL1 is constitutively expressed in hemocytes. Functional assays with recombinant LitCTL1 demonstrated its role as a potent opsonin with hemagglutinating activity, significantly enhancing hemocyte spreading and the phagocytosis of zymosan. Genomic analysis reveals that LitCTL1 belongs to a rapidly diversifying, genus-specific expansion distinct from conserved perlucin-like lineages. These results identify LitCTL1 as a key effector molecule in both systemic and mucosal innate immunity, likely reflecting an evolutionary adaptation to the microbial challenges of the intertidal environment.

Animals

Routine methods misidentify Serratia spp.: Limitations of MALDI-TOF MS revealed by whole-genome sequencing.

Accurate species-level identification within the genus Serratia remains challenging due to extensive phenotypic overlap and high genomic relatedness among closely related and recently described taxa. This study presents an evaluation of routine and genome-based identification approaches applied to clinical Serratia isolates, integrating phenotypic assays, MALDI-TOF MS (Bruker Daltonics), 16S rRNA gene sequencing, and Whole-Genome Sequencing (WGS). A total of 103 isolates collected from a teaching hospital were analyzed. WGS was performed on a subset of isolates. Conventional biochemical methods classified all isolates as Serratia marcescens, whereas MALDI-TOF MS identified 60.1% as S. marcescens, 11.6% as S. ureilytica, and 28.1% just at the genus level. Peak analysis from MALDI-TOF MS revealed specific peaks associated with S. marcescens and S. ureilytica, but limited discriminatory power. WGS of six isolates initially identified as S. ureilytica by MALDI-TOF MS revealed reclassification as Serratia sarumanii (n = 5) and Serratia montpellierensis (n = 1), supported by Average Nucleotide Identity (ANI), Average Amino Acid Identity (AAI), and Digital DNA-DNA Hybridization (dDDH) thresholds. In contrast, 16S rRNA analysis showed limited species-level resolution. Phylogenomic and SNP-based analyses confirmed these classifications with strong support. Overall, this study underscores the critical role of high-resolution genomic approaches for precise species identification and highlights the need for continuous expansion and curation of MALDI-TOF MS reference databases to support reliable clinical diagnostics and epidemiological surveillance of emerging Serratia species.

Spectrometry, Mass, Matrix-Assisted Laser Desorpti

Whole-transcriptome RNA sequencing and ceRNA network analyses provide novel insights into the antibacterial immune response of Hippocampus abdominalis against Vibrio harveyi.

Long non-coding RNAs (lncRNAs) stand as newly-arisen molecular types that exert regulatory effects, able to operate as competitive endogenous RNAs (ceRNAs) to engage microRNAs (miRNAs) in interaction, resulting in the recovery of target mRNA expression and activity. Increasing evidences indicate that the ceRNA network affects various biological processes in mammals, including development, cellular differentiation, metabolism, immune response, and disease pathogenesis. In teleost fish, the lncRNA-miRNA-mRNA regulatory networks have been reported occasionally. However, up to now, the roles of lncRNAs in the big-belly seahorse (Hippocampus abdominalis) remains unclear. In this study, we reported for the first time, via whole-transcriptome RNA sequencing, the lncRNA mediated ceRNA regulatory network in Vibrio harveyi-infected H. abdominalis. A total of 4197 differentially expressed mRNAs (DE-mRNAs), 1317 DE-lncRNAs, and 183 DE-miRNAs were identified. Furthermore, the crosstalk between miRNAs and lncRNAs as well as between miRNAs and mRNAs was inferred based on the negative correlations between miRNAs and their target lncRNAs/mRNAs. A core immune associated lncRNA-miRNA-mRNA putative regulatory network was thus constructed, comprising 211 lncRNA-miRNA and 224 mRNA-miRNA pairs. In conclusion, our findings provide an integrative overview of the ceRNA regulatory networks on the underlying immune responses to V. harveyi infection in the big-belly seahorse, and offer a solid theoretical foundation for the comparative immunological research of teleost fish.

Animals

A genome-wide coverage-based pipeline for the identification of host-derived candidate DNA biomarkers from cell-free blood.

We have created a new data-analysis pipeline for the discovery of host-specific candidate DNA biomarkers derived from sequencing data of cell-free blood. Unlike approaches that rely on specific molecular or genetic signatures, our method leverages the coverage distribution of cell-free DNA sequences mapped to a reference genome, applying statistical analyses to identify informative short genomic regions for biomarker discovery. The pipeline is applicable to diverse diseases and can be used to analyze cell-free DNA sequences from plasma or serum to identify candidate biomarkers that are characteristic of disease states in mammals. Core functionalities were developed in Java and integrated with open-source software tools for the preprocessing of raw sequencing data, complemented by Python scripts for the machine-learning analysis and statistical validation. The pipeline is designed for HPC use and users can access the pipeline through a Galaxy workflow, which offers a user-friendly web interface for input selection prior to execution and analysis progress monitoring. Performance tests, carried out using duplicate sets of COVID-19 samples and controls, showed linear scalability of execution time with an increasing dataset size, as well as a substantial reduction in execution time through parallelized computation, whereby each HPC node is used to process the data of one chromosome. Further statistical tests confirmed the quality of the pipeline's results by showing that the set of identified candidate biomarkers remained stable across varying dataset sizes.

Biomarkers

Diagnostic value of plasma cell-free DNA metagenomic next-generation sequencing in patients with suspected infections and exploration of clinical scenarios-a retrospective study from a single center.

BACKGROUND: Plasma cell-free DNA metagenomic next-generation sequencing (mNGS) is a non-invasive comprehensive method for the etiological diagnosis of various infectious diseases. However, research on the early diagnosis and real-world clinical impact of plasma mNGS in patients with suspected infection are still limited. MATERIALS AND METHODS: This study retrospectively included 140 patients with suspected infections who underwent early plasma mNGS and conventional culture testing. Referring to the clinical diagnosis of infectious diseases, the diagnostic performance of plasma mNGS and culture tests was compared, and the application scenarios and clinical effects of plasma mNGS were evaluated. RESULTS: The positive rate of plasma mNGS was significantly higher than that of culture methods (55.71% vs 25.10%, p&#x2009;<&#x2009;0.001) and blood cultures (55.71% vs 12.86%, p&#x2009;<&#x2009;0.001). Regarding clinical diagnosis, the sensitivity of plasma mNGS was significantly higher than that of culture (58.27% vs 37.80%, p&#x2009;=&#x2009;0.002). The combination of mNGS and culture achieved a higher detection sensitivity (69.29%), especially in patients with multi-site co-infections (73.68%) and blood infections (73.17%). Plasma mNGS demonstrated higher sensitivity in patients with procalcitonin (PCT) index > 5&#x2009;ng/ml or human neutrophil lipocalin (HNL) index > 200&#x2009;ng/ml. In terms of treatment, a total of 69 patients (54.33%) benefited from plasma mNGS. CONCLUSION: This study highlights the significant improvement in pathogen detection performance by combining conventional culture with plasma mNGS detection, especially in patients with multi-site co-infections and blood infections. Early use of plasma mNGS as an adjunct to culture can better guide clinicians to initiate appropriate anti-infective therapy.

Humans

Mapping antibody sequences and effector functions across spatial niches.

Antibodies are fundamental to human health but can also drive pathology. Each antibody has a molecular specificity, encoded by their clonally heritable B cell receptor (BCR). Recent advances in spatial transcriptomics coupled with repertoire sequencing have enabled capturing antibody-secreting cells (ASCs) and their clonal BCR within their tissue microenvironment. However, our understanding of antibody production niches remains limited. Furthermore, where antibodies are produced can be distinct from where antibodies exert their effector function. Here, we propose a conceptual spatial framework to distinguish between 'antibody production niches', defined by the ASC, BCR, and niche composition, versus 'antibody functional niches', composed of the antibody, antigen, and effector landscape. We then examine the possibilities and challenges to map and link antibody-encoding sequences and antibody effector functions using current and emerging technologies. Combined, we argue that integrating spatial sequence data with the antibody functional context is essential to decode the architecture of antibody-mediated immunity.

Humans

Comparison of paralog identification methods and their impact on species tree topologies in target capture phylogenomics within the Sindora clade (Detarioideae: Leguminosae).

Target capture is a common method of generating high throughput DNA sequencing data for phylogenetic reconstruction of species relationships, for which single copy genes are usually most informative. However, a pervasive problem with target capture is that putatively single copy genes may in fact be paralogs resulting from gene duplication, which are problematic for phylogenetic inference because their evolutionary history may differ from the divergence history of species. Here, we use as a case study a target enrichment dataset of 88 species of Detarioideae (Leguminosae) with a focus on the Sindora clade to examine approaches for handling paralogs, including the built-in paralog handling functions in HybPiper and CAPTUS, plus subsequent steps using Putative Paralog Detection and the tree-based Yang & Smith orthology inference approach. We compare the paralogs flagged using these methods and verify their performance with BLAST mapping against a reference genome sequence of Sindora glabra, and then subsequently compare the species tree topologies produced across these methods. Our comparisons of paralogs flagged across the Sindora clade show that the Putative Paralog Detection pipeline was the most accurate in identifying paralogs in terms of its similarity to the BLAST mapping, followed by the built-in paralog identification function of CAPTUS. However, the results we recovered for the Detarioideae subfamily suggest that the largest differences in species tree topology resulted from the use of paralog-filtered alignments (such as with the Putative Paralog Detection pipeline and the Yang & Smith orthology inference approaches) rather than just by removing the sequences of identified paralogous genes. This was the true for HybPiper-assembled datasets but was not seen in CAPTUS-assembled datasets. In all comparisons, the topological differences caused by different paralog handling methods tended to be confined to clades where processes such as hybridisation and introgression are prevalent. Our study provides a roadmap to establish the best approach to identify, eliminate or separate paralogs in the absence of a chromosomally contiguous reference genome for a study group, and highlights the importance of careful data inspection and processing in addition to understanding the extent of paralogy and paralog characteristics (e.g. sequence divergence between copies) for their study group.

Phylogeny

Nanopore-based epigenomic profiling reveals the absence of widespread CpG methylation in the African swine fever virus genome.

DNA methylation is a critical epigenetic mechanism implicated in regulating replication and transcription in DNA viruses. However, the epigenetic landscape of African swine fever virus (ASFV), a large double-stranded DNA virus infecting pigs, remains controversial. Here, we systematically profiled the DNA methylome of the first ASFV strain isolated in Hong Kong (HK_NT_202103) using Oxford Nanopore Technologies (ONT) R10.4.1 sequencing. We employed a paired design: native whole-genome sequencing (WGS) against a methylation-free whole-genome amplification (WGA) control. Using conservative thresholds, we found no evidence of 5-methylcytosine (5mC), especially typical CpG methylation, across the viral genome. Importantly, clear CpG methylation signals were successfully detected in the host genome from WGS data, confirming the functionality of the workflow to detect 5mC at CG sites. While widespread 5mC seems absent, a small number of putative N6-methyladenine (6mA) loci were identified. A specific 6mA candidate exhibited raw ionic current disruptions and gene-level intersection with another ASFV isolate (CAS19-01/2019), although it lacked single-base consensus across different methylation callers or between the two isolates. Although our biological findings are restricted to a single isolate under specific experimental conditions, this study introduces a novel, highly rigorous ONT framework for viral epigenomics research. Furthermore, the absence of ASFV CpG methylation indicates that host CpG-depletion remains a viable strategy for viral metagenomic enrichment. Ultimately, our work offers a critical methodological baseline for ASFV surveillance and highlights the necessity of targeted experimental validation for rare viral modifications.

African Swine Fever Virus

A conserved distal-tail helical extension defines a tailspike attachment architecture in Gram-negative siphophages.

Rapid growth of bacteriophage genome collections has outpaced functional annotation of tail-tip proteins, limiting comparative analysis of host-recognition structures. Starting from a shared distal-tail gene organization in the Salmonella phages 9NA and Jersey, I developed a morphogenetic bioinformatic framework integrating gene synteny, sequence comparison, profile hidden Markov model (HMM) screening, structural evidence, structure-aware searching, and AlphaFold modeling. Comparison with the experimentally characterized lambda and Sf11 tail assemblies identified a predominantly alpha-helical C-terminal extension of the distal-tail (DT) protein associated with tailspike attachment, termed the distal-tail helical extension (DT-helix). Screening 541,986 proteins from 5167 complete NCBI RefSeq tailed-phage genomes, followed by evidence-based evaluation of sequence, genomic context, and structural architecture, identified 165 curated DT-helical-extension-associated phages. Their DT proteins segregated into six sequence groups. In the four principal multi-member groups, cognate tailspikes showed group-specific conservation in proximal N-terminal regions but substantially greater downstream diversity, consistent with sequence constraint at the DT-tailspike attachment boundary. A complementary ProstT5/Foldseek search supported the established groups but revealed no convincing additional highly divergent family. Together with the experimentally characterized Sf11 attachment interface, these findings define a recurrent morphogenetic architecture linking conserved distal-tail scaffolds to more variable receptor-binding proteins across siphophages infecting Gram-negative bacteria. Although universal exchangeability is not established, the identified scaffold-receptor-binding boundaries provide a framework for molecular characterization and rational phage engineering. Accession-level information for the 165 curated phages is available through PhageTailDB.

Viral Tail Proteins

Longitudinal whole-genome analysis of bluetongue virus identifies conserved serotype-specific genomes and distinct genomic constellations within a Colorado sheep flock (2021-2023).

Bluetongue virus (BTV) is a segmented double-stranded RNA virus of ruminants transmitted by Culicoides spp. biting midges. Although the genome consists of ten segments, classification into serotypes is primarily based on genome segment 2. However, reassortment among genomic segments is a major driver of BTV evolution and diversity. This study used longitudinal whole-genome sequencing to characterize BTV genomes collected from 2021 to 2023 within a single sheep flock in Colorado, where multiple serotypes co-circulate. Whole-genome sequences were generated from fourteen blood samples representing four serotypes: BTV-6, -11, -13, and -17. Longitudinal sampling identified multiple BTV serotypes within individual sheep across consecutive years. Tanglegram analysis comparing segment phylogenies to the segment 2 tree demonstrated incongruent topologies across all genomic segments, suggestive of reassortment or the circulation of distinct genomic constellations. Nucleotide-level comparisons revealed high sequence homology among same-serotype samples from the same year, while the greatest genetic divergence was observed among BTV-17 genomes collected in different years. Additionally, all BTV-13 genomes contained a previously undescribed nonsynonymous substitution in segment 10 predicted to extend the encoded protein by three amino acids. Together, these findings demonstrate that highly conserved BTV genomes and distinct genomic constellations can be detected at the flock level across multiple years. This longitudinal whole-genome approach reveals the genetic complexity of endemic BTV populations, including novel variants and genomic patterns consistent with reassortment that are lost with conventional serotyped-based approaches, highlighting the need to integrate whole-genome characterization into endemic BTV monitoring programs.

Animals

Mitochondrial DNA diversity in Ecuadorian populations: Recurrence of variant 16136 within haplogroup B2.

The identification of lineage-defining variants, frequently found in the coding region of mitochondrial DNA (mtDNA), is essential for refining haplogroup classification. Most mtDNA studies in South American populations have focused on the control region (CR), which has provided important insights into population structure and maternal lineage origins, although information needed for more robust phylogenetic resolution has been neglected. This study investigates the maternal genetic structure of Ecuadorian populations by combining CR and whole mitogenome analyses. Sequences from the mtDNA CR were obtained from 461 individuals (253 Mestizos and 208 Native Americans), while complete mitogenomes were sequenced for 127 individuals to improve phylogenetic resolution by identifying lineage-defining variants present in coding region. Most mtDNA haplogroups in the two population groups analyzed were of Native American origin (A2, B2, B4, C1, D1, D4), with significant differences in the distribution of specific lineages between them. Among Mestizos, African haplogroups (all within the L branches) and Eurasian haplogroups (H, K, R, U) were detected at low frequencies, whereas no African lineages were observed among Native Americans. The results obtained highlighted a heterogeneity within Ecuadorian populations that must be considered when developing mtDNA haplotype databases for forensic purposes. Whole mitogenome sequences enabled the identification of variants that refined haplogroup classifications, provided a more accurate reconstruction of the maternal genetic diversity, and improve the discrimination between Native American and Asian maternal lineages within haplogroup B4b.

Humans

Whole-Exome Sequencing in a Consanguinity-Enriched South Indian Retinitis Pigmentosa Cohort: Diagnostic Yield and Molecular Spectrum.

PURPOSE: To determine the molecular diagnostic yield, variant spectrum, inheritance architecture, and influence of consanguinity on whole-exome sequencing outcomes in a South Indian retinitis pigmentosa (RP) cohort. DESIGN: Prospective, registry-based cohort study. SUBJECTS: A total of 113 affected participants were enrolled through the Aravind Registry for Inherited Diseases of the Eye, including 109 unrelated probands and 4 affected relatives from already represented families. Primary analyses were restricted to the 109 unrelated probands. METHODS: Whole-exome sequencing was performed using a clinical exome workflow. Variants were interpreted using American College of Medical Genetics and Genomics/Association for Molecular Pathology criteria and cases were categorized as solved, possibly solved, inconclusive, or unsolved using prespecified inheritance-aware rules. MAIN OUTCOME MEASURES: Molecular diagnostic yield, distribution of implicated genes and variant classes, inheritance architecture, and diagnostic yield stratified by consanguinity status. RESULTS: Among the 109 unrelated probands, mean age at testing was 39.3 &#xb1; 14.1 years and 58.7% were male. Whole-exome sequencing identified 186 distinct rare variants across 92 inherited retinal disease genes, including 26 pathogenic and 33 likely pathogenic variants. A molecular diagnosis was established in 50 of 109 probands (45.9%), including 42 solved and 8 possibly solved cases; 45 (41.3%) were inconclusive and 14 (12.8%) remained unsolved, including 4 (3.7%) in whom no candidate variant was identified. EYS, USH2A, and ADGRV1 were the most frequently implicated genes. Autosomal recessive (AR) disease predominated (44/50, 88.0%). Consanguineous AR cases were exclusively homozygous (17/17); notably, 68.0% of nonconsanguineous AR cases were also homozygous (P = 0.013). Diagnostic yield was higher in consanguineous probands (51.4% vs. 41.7%), without reaching significance. Recurrent alleles included an established South Asian founder variant (MFSD8 c.1361T>C) and candidate founder alleles in EYS (c.4321C>T) and ADGRV1 (c.14329C>T). CONCLUSIONS: Whole-exome sequencing established a molecular diagnosis in nearly half of this South Indian RP cohort and revealed a predominantly recessive, homozygosity-enriched architecture shaped by consanguinity. These findings define a region-specific variant landscape to support clinical interpretation, genetic counseling, and future trial enrollment in this underrepresented population. FINANCIAL DISCLOSURES: The authors have no proprietary or commercial interest in any materials discussed in this article.

Consanguinity

Exploring the mechanism of aroma production in fermented cherry juice by L. brevis LD1.0600 using flavomics and whole genome analysis.

This study focused on L.brevis LD1.0600 with excellent fermentation traits: it analyzed genome-wide key regulatory genes for micro-metabolites, combined with fermented cherry juice flavor metabolomics data, and used machine learning to explore correlations between gene regulation, metabolite production, and flavor formation. The SVM model screened and verified fermented cherry juice VOCs; through OAV and flavor wheel analysis, LD1.0600 emerged as the top-performing strain, with a sweet, fruity dominant aroma. Key aroma-active components (OAV&#xa0;>&#xa0;100) included 2-methoxy-4-vinylphenol, benzaldehyde, 2-methyl-butanoic acid and hexanoic acid, and 2-methoxy-4-vinylphenol and hexanoic acid elevated by LD1.0600-regulated genes (Chrom1-001884, Chrom1-000925, fabF and Chrom1-000199). At the same time, through research, a "strain screening-SVM screening of DVCs-OAV screening of key aroma components-whole genome sequencing of flavor regulatory genes" system was established. This system can not only be applied to the screen fermentation strains, but also can be extended to the application of other fermentation products.

Fermentation

Conserved host-exclusive oligonucleotide motifs enriched in pathogenic genes of human oncogenic viruses.

Comparative viral genomics can reveal sequence-level constraints influencing virus-host interactions. Relative minimal absent words (rMAWs) are short oligonucleotide motifs present in viral genomes but completely absent from the host, potentially reflecting selective pressures related to host adaptation and immune evasion. Using the EAGLE algorithm and the GRCh38 human reference genome, we systematically screened for prevalent rMAWs (prMAWs) across six major human oncogenic viruses: Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human papillomavirus (HPV), human T-cell leukemia virus type 1 (HTLV-1), and human herpesvirus 8/Kaposi's sarcoma-associated herpesvirus (HHV-8/KSHV). highly conserved 11- and 12-bp prMAWs were identified in EBV, HBV, HTLV-1, and HHV-8/KSHV, with sequence prevalences ranging from 91.5% to 97.9%. Conversely, no short prMAWs were detected in HCV or HPV, likely reflecting differences in genome architecture, mutation rates, and long-term host adaptation to the human host. Importantly, the identified host-exclusive motifs exhibited non-random genomic distribution and were preferentially embedded within viral genes central to replication, persistence, immune modulation, and oncogenesis, including EBNA-1 (EBV), HBx (HBV), Tax-associated regions (HTLV-1), and lytic replication genes of HHV-8/KSHV. Notably, all detected prMAWs were enriched in GC nucleotides and exhibited marked CpG over-representation, suggesting sequence constraints associated with epigenetic regulation and viral persistence. Collectively, these highly conserved, host-exclusive signatures offer promising, candidates for sequence-directed approaches in the diagnosis, monitoring, and investigation of virus-associated cancers.

Humans

Translating single-cell RNA sequencing into monocyte direct leukocyte subpopulation-transcript abundance assay ratio-based biomarkers (IFI27/PSAP or IFI27/CTSS) for clinical detection of viral infection.

A rapid method for triaging febrile patients by aetiology (e.g., viral or bacterial infection) using gene expression in peripheral blood (PB) is an intensively researched area. However, gene expression in blood represents a composite sum of gene expression of all the component cell types present in the sample. As a result, numerous genes are measured in most proposed signatures. Herein, we propose a simple ratio-based biomarker (RBB) called direct leukocyte subpopulation-transcript abundance assay (DIRECT LS-TA) that recapitulates gene expressions of a single cell type in PB (i.e., monocytes). Based on single-cell RNA sequencing (scRNAseq) data and bulk expression data, IFI27 and SIGLEC1 are found as interferon-stimulated genes (ISGs) predominantly expressed by monocytes. The DIRECT LS-TA method can use a simple ratio of two genes measured in PB as an RBB to represent the target gene expression in monocytes without the need for monocyte purification. Both scRNAseq and bulk RNA sequencing datasets were used to evaluate the correlation between ISG expression in monocytes and PB, with a particular focus on monocyte expression of IFI27. An iceberg plot of bulk transcriptome data was used to identify genes that were predominantly expressed by monocytes in PB. DIRECT LS-TA RBBs of the three genes (IFI27, IFI44L and SIGLEC1) were evaluated by group-wise comparison, receiver operating characteristic and meta-analysis. In addition, the conventional interferon (IFN) score was evaluated for comparison of diagnostic performance. In viral infection datasets, DIRECT LS-TA of IFI27 (IFI27/PSAP or IFI27/CTSS) was most intensely activated (p value by t test <1e-9) and had the best area under the curve (0.94) among the three potential monocyte ISGs analysed. DIRECT LS-TA SIGLEC1 was also another monocyte biomarker but showed a lower activation (p<9e-5). IFI27/PSAP showed better diagnostic performance than the conventional IFN score. On the other hand, IFI44L was not a predominant monocyte expression gene. DIRECT LS-TA of IFI27 (IFI27/PSAP or IFI27/CTSS) measured in PB was the best biomarker of viral infection and IFN activation among ISGs predominantly expressed by monocytes. It performed even better than the conventional IFN score which required quantification of eight genes. The results suggest that DIRECT LS-TA of IFI27 is a monocyte-informative biomarker which is easy to determine in PB without the need for cell sorting.

Humans