Search PubMedSearch

SEARCH · Search PubMed

Results for “genome sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

630 records · Page 2Linked to original sources

A genome-wide coverage-based pipeline for the identification of host-derived candidate DNA biomarkers from cell-free blood.

We have created a new data-analysis pipeline for the discovery of host-specific candidate DNA biomarkers derived from sequencing data of cell-free blood. Unlike approaches that rely on specific molecular or genetic signatures, our method leverages the coverage distribution of cell-free DNA sequences mapped to a reference genome, applying statistical analyses to identify informative short genomic regions for biomarker discovery. The pipeline is applicable to diverse diseases and can be used to analyze cell-free DNA sequences from plasma or serum to identify candidate biomarkers that are characteristic of disease states in mammals. Core functionalities were developed in Java and integrated with open-source software tools for the preprocessing of raw sequencing data, complemented by Python scripts for the machine-learning analysis and statistical validation. The pipeline is designed for HPC use and users can access the pipeline through a Galaxy workflow, which offers a user-friendly web interface for input selection prior to execution and analysis progress monitoring. Performance tests, carried out using duplicate sets of COVID-19 samples and controls, showed linear scalability of execution time with an increasing dataset size, as well as a substantial reduction in execution time through parallelized computation, whereby each HPC node is used to process the data of one chromosome. Further statistical tests confirmed the quality of the pipeline's results by showing that the set of identified candidate biomarkers remained stable across varying dataset sizes.

Biomarkers

Assembly and Characterization of the First Complete Mitochondrial Genome of Tussilago farfara L.: Insights into Biological Functions and Phylogenetic Relationships within the Asteraceae Family.

Tussilago farfara L., a member of the Asteraceae family, is an economically valuable species due to its edible and medicinal properties. To elucidate the structural characteristics, genetic mechanisms, and evolutionary pathways of the organelle genomes of T. farfara, we sequenced, assembled, and annotated its mitochondrial genome for the first time. The complete mitochondrial genome of T. farfara spans 306,024 bp and contains 33 mitochondrial protein-coding genes (PCGs), 3 rRNAs, and 22 tRNAs. Analysis of the nucleotide substitution rate and genetic diversity revealed that most mitochondrial genome genes may have undergone purifying selection, indicating a slow evolutionary rate and a relatively conserved genomic structure. We further identified 13 fragments of chloroplast-derived DNA integrated into the mitochondrial genome, evidencing intracellular gene transfer. Collinearity analysis showed that Arctium lappa shares the most extensive mitochondrial homologous sequences and the highest sequence similarity with T. farfara. Phylogenetic analysis based on the mitochondrial genome helped to clarify the evolutionary and taxonomic position of T. farfara within the Asteraceae family. The mitochondrial genome sequence of T. farfara provides a valuable genomic resource for species identification and for evolutionary studies within the Asteraceae family.

Genome, Mitochondrial

Exploring the mechanism of aroma production in fermented cherry juice by L. brevis LD1.0600 using flavomics and whole genome analysis.

This study focused on L.brevis LD1.0600 with excellent fermentation traits: it analyzed genome-wide key regulatory genes for micro-metabolites, combined with fermented cherry juice flavor metabolomics data, and used machine learning to explore correlations between gene regulation, metabolite production, and flavor formation. The SVM model screened and verified fermented cherry juice VOCs; through OAV and flavor wheel analysis, LD1.0600 emerged as the top-performing strain, with a sweet, fruity dominant aroma. Key aroma-active components (OAV > 100) included 2-methoxy-4-vinylphenol, benzaldehyde, 2-methyl-butanoic acid and hexanoic acid, and 2-methoxy-4-vinylphenol and hexanoic acid elevated by LD1.0600-regulated genes (Chrom1-001884, Chrom1-000925, fabF and Chrom1-000199). At the same time, through research, a "strain screening-SVM screening of DVCs-OAV screening of key aroma components-whole genome sequencing of flavor regulatory genes" system was established. This system can not only be applied to the screen fermentation strains, but also can be extended to the application of other fermentation products.

Fermentation

Conserved host-exclusive oligonucleotide motifs enriched in pathogenic genes of human oncogenic viruses.

Comparative viral genomics can reveal sequence-level constraints influencing virus-host interactions. Relative minimal absent words (rMAWs) are short oligonucleotide motifs present in viral genomes but completely absent from the host, potentially reflecting selective pressures related to host adaptation and immune evasion. Using the EAGLE algorithm and the GRCh38 human reference genome, we systematically screened for prevalent rMAWs (prMAWs) across six major human oncogenic viruses: Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human papillomavirus (HPV), human T-cell leukemia virus type 1 (HTLV-1), and human herpesvirus 8/Kaposi's sarcoma-associated herpesvirus (HHV-8/KSHV). highly conserved 11- and 12-bp prMAWs were identified in EBV, HBV, HTLV-1, and HHV-8/KSHV, with sequence prevalences ranging from 91.5% to 97.9%. Conversely, no short prMAWs were detected in HCV or HPV, likely reflecting differences in genome architecture, mutation rates, and long-term host adaptation to the human host. Importantly, the identified host-exclusive motifs exhibited non-random genomic distribution and were preferentially embedded within viral genes central to replication, persistence, immune modulation, and oncogenesis, including EBNA-1 (EBV), HBx (HBV), Tax-associated regions (HTLV-1), and lytic replication genes of HHV-8/KSHV. Notably, all detected prMAWs were enriched in GC nucleotides and exhibited marked CpG over-representation, suggesting sequence constraints associated with epigenetic regulation and viral persistence. Collectively, these highly conserved, host-exclusive signatures offer promising, candidates for sequence-directed approaches in the diagnosis, monitoring, and investigation of virus-associated cancers.

Humans

A conserved distal-tail helical extension defines a tailspike attachment architecture in Gram-negative siphophages.

Rapid growth of bacteriophage genome collections has outpaced functional annotation of tail-tip proteins, limiting comparative analysis of host-recognition structures. Starting from a shared distal-tail gene organization in the Salmonella phages 9NA and Jersey, I developed a morphogenetic bioinformatic framework integrating gene synteny, sequence comparison, profile hidden Markov model (HMM) screening, structural evidence, structure-aware searching, and AlphaFold modeling. Comparison with the experimentally characterized lambda and Sf11 tail assemblies identified a predominantly alpha-helical C-terminal extension of the distal-tail (DT) protein associated with tailspike attachment, termed the distal-tail helical extension (DT-helix). Screening 541,986 proteins from 5167 complete NCBI RefSeq tailed-phage genomes, followed by evidence-based evaluation of sequence, genomic context, and structural architecture, identified 165 curated DT-helical-extension-associated phages. Their DT proteins segregated into six sequence groups. In the four principal multi-member groups, cognate tailspikes showed group-specific conservation in proximal N-terminal regions but substantially greater downstream diversity, consistent with sequence constraint at the DT-tailspike attachment boundary. A complementary ProstT5/Foldseek search supported the established groups but revealed no convincing additional highly divergent family. Together with the experimentally characterized Sf11 attachment interface, these findings define a recurrent morphogenetic architecture linking conserved distal-tail scaffolds to more variable receptor-binding proteins across siphophages infecting Gram-negative bacteria. Although universal exchangeability is not established, the identified scaffold-receptor-binding boundaries provide a framework for molecular characterization and rational phage engineering. Accession-level information for the 165 curated phages is available through PhageTailDB.

Viral Tail Proteins

Complete mitochondrial genomes of eight cyclophyllidean tapeworms: genome pattern and phylogenetic analysis.

Cyclophyllidean tapeworms are widespread parasites of significant medical and veterinary importance. However, mitochondrial (mt) genomic resources for cyclophyllideans from China, particularly those recovered from wildlife hosts, remain comparatively limited. In this study, we sequenced and characterized the complete mt genomes of eight cyclophyllidean isolates collected from diverse wild and domestic hosts in China, including two Hymenolepis sp. isolates and two Raillietina sp. isolates from China, and four additional isolates of previously sequenced Taenia species. The circular mt genomes ranged from 13,387 to 14,021 bp in length, encoding 36 typical genes with variable non-coding regions. Comparative analysis revealed highly conserved gene composition and mostly conserved mt architecture, with localized rearrangement patterns detected among the cyclophyllidean lineages examined. In particular, all sampled Taeniidae exhibited a consistent trnL1-trnS2 arrangement, whereas the examined non-Taeniidae families showed the trnS2-trnL1 arrangement, confirming and extending, across additional wildlife-associated isolates, a previously proposed family-associated gene-order marker within Cyclophyllidea. Phylogenetic analyses based on concatenated amino acid sequences of the 12 protein-coding genes placed the eight isolates within their expected families, in topologies broadly consistent with previous mitogenomic studies. These data provide additional Chinese mitogenomic references, especially for underrepresented wildlife-associated isolates, and support family-associated gene-order patterns in Cyclophyllidea.

Animals

Clinical Outcomes and Genomic Epidemiology of Multidrug-Resistant Methicillin-Resistant Staphylococcus aureus Keratitis.

PURPOSE: To characterize the clinical features, management, antimicrobial resistance patterns, and genomic epidemiology of methicillin-resistant Staphylococcus aureus (MRSA) keratitis at two North American centers. DESIGN: Retrospective interventional case series combined with laboratory investigation PARTICIPANTS: Seventy eyes of 67 patients presenting laboratory-confirmed MRSA keratitis were included METHODS: We performed a multicenter retrospective case series of patients with culture-proven MRSA keratitis treated between 2005 and 2022. Demographic and clinical data were collected. Antimicrobial susceptibility testing was conducted, and multidrug resistance (MDR) was defined as resistance to ≥3 antibiotic classes. A subset of isolates underwent whole-genome sequencing with core genome multilocus sequence typing. Vancomycin susceptibility, heteroresistance screening, and tolerance testing were performed on available isolates. MAIN OUTCOME MEASURES: Antimicrobial susceptibility and multidrug resistance rates, vancomycin phenotypic profiles, MRSA genotypic distribution, and final best-corrected visual acuity RESULTS: Median age was 63.5 years, and 61.4% were female. Ocular surface disease (67.7%) and prior ocular surgery (65.2%) were common. Only 25.4% had significant healthcare exposure in the preceding year. Most isolates (85.7%) were MDR. Fluoroquinolone susceptibility was low (moxifloxacin 19.7%). All isolates were susceptible to vancomycin (MIC₉₀ 2 µg/mL), and no vancomycin-intermediate, heteroresistant, or tolerant phenotypes were identified. Whole genome sequencing (n = 41) demonstrated predominance of clonal complexes 5 (68.3%) and 8 (29.2%). Visual outcomes were poor, with most patients (85.2%) having a final visual acuity worse than 20/60 among those with follow-up. CONCLUSIONS: MRSA keratitis is associated with high rates of multidrug resistance and poor visual outcomes despite guideline-based therapy. Infections were predominantly caused by CC5 MDR strains despite limited recent healthcare exposure. These findings highlight the persistence of highly resistant MRSA lineages in community-associated corneal infection and underscore the need for ongoing antimicrobial surveillance and optimized treatment strategies.

Humans

The genomic origins and evolutionary path to a key innovation in the world's most venomous snakes.

Evolutionary innovation is a catalyst for the colonization of new environments and the adaptive radiations of major groups. Novel traits typically evolve through the modification of preexisting characters, but the genetic paths underlying their origin have been challenging to trace, and the general requirements for and relative order of different kinds of gene mutations have been difficult to assess. Here, we trace the genomic origins of four procoagulant venom toxins (factor X, factor V, group I phospholipase A2, and Kunitz-type toxins) that collectively underlie a novel, especially potent blood-clotting venom type in the recently evolved Australian brown snake and taipan clade. We find evidence for a previously unknown fifth toxin, coagulation factor VII, and show that the toxins evolved through two distinct genetic paths. The factor X and factor V toxins evolved through the sequential de novo co-option of ancestral clotting factor proteins that entailed their heterotopic expression in the venom gland, the fixation of segmental duplications containing each locus, and subsequent gain-of-function mutations that rendered factor X and factor V constitutively active. In contrast, the phospholipase A2 and Kunitz-type toxins evolved by modifying the functions of neurotoxins that were part of the venom arsenal. Our findings support models in which innovative mutations in single-copy genes precede gene duplication in the evolution of novel proteins and offer a rare view into the genesis of a complex trait that has played a central role in a major adaptive radiation.

Animals

Comparison of paralog identification methods and their impact on species tree topologies in target capture phylogenomics within the Sindora clade (Detarioideae: Leguminosae).

Target capture is a common method of generating high throughput DNA sequencing data for phylogenetic reconstruction of species relationships, for which single copy genes are usually most informative. However, a pervasive problem with target capture is that putatively single copy genes may in fact be paralogs resulting from gene duplication, which are problematic for phylogenetic inference because their evolutionary history may differ from the divergence history of species. Here, we use as a case study a target enrichment dataset of 88 species of Detarioideae (Leguminosae) with a focus on the Sindora clade to examine approaches for handling paralogs, including the built-in paralog handling functions in HybPiper and CAPTUS, plus subsequent steps using Putative Paralog Detection and the tree-based Yang & Smith orthology inference approach. We compare the paralogs flagged using these methods and verify their performance with BLAST mapping against a reference genome sequence of Sindora glabra, and then subsequently compare the species tree topologies produced across these methods. Our comparisons of paralogs flagged across the Sindora clade show that the Putative Paralog Detection pipeline was the most accurate in identifying paralogs in terms of its similarity to the BLAST mapping, followed by the built-in paralog identification function of CAPTUS. However, the results we recovered for the Detarioideae subfamily suggest that the largest differences in species tree topology resulted from the use of paralog-filtered alignments (such as with the Putative Paralog Detection pipeline and the Yang & Smith orthology inference approaches) rather than just by removing the sequences of identified paralogous genes. This was the true for HybPiper-assembled datasets but was not seen in CAPTUS-assembled datasets. In all comparisons, the topological differences caused by different paralog handling methods tended to be confined to clades where processes such as hybridisation and introgression are prevalent. Our study provides a roadmap to establish the best approach to identify, eliminate or separate paralogs in the absence of a chromosomally contiguous reference genome for a study group, and highlights the importance of careful data inspection and processing in addition to understanding the extent of paralogy and paralog characteristics (e.g. sequence divergence between copies) for their study group.

Phylogeny

Unraveling a Diagnostic Enigma: A TECPR2 Case Solved Through Multi-Omic Genomics.

TECPR2 is a key regulator of autophagy, encoded by the TECPR2 gene. Pathogenic variants in this gene have been linked to a rare hereditary sensory and autonomic neuropathy with intellectual disability (HSAN9). We report a teenage female with a syndromic intellectual disability disorder associated with neuromuscular abnormalities. Multi-omics analysis including genomics, transcriptomics, and proteomics, together with muscle biopsy from the affected individual, were used in this clinical case. Through trio exome sequencing we identified two heterozygous variants in the TECPR2 gene, NM_014844.4: c.480G>A; p.(Gln160=) and c.2846C>A; p.(Ala949Glu). Both were classified as variants of uncertain significance due to the lack of supporting evidence for pathogenicity. Subsequent long-read sequencing phased the variants and confirmed they were in trans. Additional functional studies using RNAseq and proteomics analyses verified the pathogenicity of the variants. This case study demonstrated the value of a multi-omics assisted analysis, which complemented the traditional phenotype-first approach in reaching a definitive clinical diagnosis.

Humans

Ramu stunt virus genome reveals previously unreported segments and nucleocapsid domain duplication in Mechlorovirus.

Ramu stunt virus (RmSV), a member of the genus Mechlorovirus within the family Phenuiviridae, was previously described as a six-segmented RNA virus infecting sugarcane. In this study, we re-examined type material and additional isolates using high-throughput sequencing and RT-PCR validation, revealing that RmSV possesses a nine-segmented genome, making it the largest reported in the Phenuiviridae. This expanded architecture includes duplicated RNA segments (RNA 2a and RNA 2b) encoding nucleocapsid-like proteins and two novel segments (RNA 7 and RNA 8). Comparative analysis showed that RNA 2a and 2b share about 84% amino acid identity, while RNA 5 encodes a third nucleocapsid homolog, indicating unprecedented domain redundancy. Structural modeling confirmed that all three nucleocapsid proteins maintain a conserved fold despite low sequence identity, with electrostatic mapping suggesting differential RNA-binding potential. Additionally, RNA 6 encodes a hypothetical protein structurally similar to the rice stripe virus disease-specific S-protein, implicating a role in symptom development. Transcript abundance analysis revealed RNA 6 as the most highly expressed segment across isolates. These findings revise the genomic composition of RmSV, highlight mechanisms of genome plasticity and adaptive evolution in plant-infecting bunyaviruses, and underscore practical implications for diagnostic assay design, resistance breeding, and biosecurity surveillance.

Genome, Viral

Insights into the fate and dynamics of antibiotic resistance in multidrug-resistant Bacillus cereus during in vitro simulated gastrointestinal digestion.

Bacillus cereus, an important pathogen responsible for causing foodborne diseases worldwide, releases pore-forming enterotoxins, which target host epithelial cells, leading to osmotic lysis and ultimately manifesting as diarrheal syndrome. Moreover, some B. cereus strains carry antimicrobial resistance genes that confer multidrug resistance against a spectrum of antibiotics. Characterizing the survival traits of multidrug-resistant (MDR) B. cereus strains in the intestinal microenvironment is essential for developing targeted strategies to effectively manage diarrheal foodborne diseases caused by this pathogen. This study used whole-genome sequencing (WGS) to evaluate the pre- and post-digestion toxigenic potential, antimicrobial resistance profiles, and genetic diversity of MDR B. cereus strains isolated from food samples in Guangdong Province, China. The four B. cereus isolates investigated in this study exhibited a genetic diversity, as determined by multilocus sequence typing analysis of WGS data. All four isolates produced the diarrheal toxins Hbl, Nhe, and CytK to varying levels, indicative of their potential to cause outbreaks of foodborne diseases. Each of the four isolates exhibited resistance to more than three classes of antibiotics, fulfilling the criterion for multidrug resistance. At an initial concentration of 9 log colony-forming units (CFU)/mL, the intestinal concentration of these four isolates crossed the threshold required to induce widespread diarrhea in the general population. Under rice slurry protection, all tested isolates maintained intestinal concentration beyond the threshold when the initial concentration was increased to ≥8 log CFU/mL. Moreover, the upregulations of genes associated with acid tolerance, bile tolerance and stress response were observed in the surviving MDR B. cereus isolates. Digestion markedly altered the antibiotic resistance profiles of the MDR B. cereus isolates. In the absence of a food matrix, the MDR isolates lost their resistance to imipenem, meropenem, amoxicillin-clavulanic acid, and trimethoprim-sulfamethoxazole post-digestion and was influenced by the initial concentration of the strains. In the presence of food matrix rice slurry, the effects of digestion on the antibiotic resistance of MDR B. cereus isolates can be mitigated, enabling them to maintain their antibiotic resistance to the greatest extent. Most remarkably, after digestion, the isolates Bce055 and Bce166 exhibited newly emergent resistance to cefotetan and trimethoprim-sulfamethoxazole, respectively. Our findings clarify the fate of MDR B. cereus isolates in the gastrointestinal tract and inform the development of prevention and control strategies for foodborne diseases caused by this pathogen.

Drug Resistance, Multiple, Bacterial

Diagnostic and clinical utility of exome sequencing and chromosomal microarray in children with GDD/iD: a meta-analysis.

BACKGROUND: Global developmental delay/intellectual disability (GDD/ID) is among the most common neurodevelopmental disorders, with up to half of cases are attributed to genetic factors. Chromosome microarray (CMA) has traditionally been the primary genetic test for idiopathic GDD/ID. However, whole exome sequencing (WES) and whole genome sequencing (WGS) have recently emerged, substantially increasing diagnostic yields in these populations. METHODS: We conducted a comprehensive literature search of PubMed, Scopus, EMBASE, and the Cochrane Library from inception to April 29, 2025. Studies reporting the diagnostic utility of these tests in children with GDD/ID were included and analyzed. RESULTS: A total of 102 studies, comprising 55,752 children, were reviewed. The pooled diagnostic yield of WES was 0.37 (95% CI: 0.33-0.41; I2 = 93%), significantly higher than that of CMA at 0.19 (95% CI: 0.16-0.21; I2 = 95%). Subgroup analyses showed that WES yielded significantly higher diagnostic rates than CMA in both same-sample comparisons (OR = 2.27, 95% CI: 1.08-4.78) and different-sample comparisons (OR = 1.65, 95% CI: 1.15-2.37). Only one study evaluated WGS, reporting a diagnostic yield of 0.27. Meta-regression revealed a significant association between CMA diagnostic yield and the proportion of male participants (p&#x2009;<&#x2009;0.01), but not with WES. No significant difference in diagnostic utility was observed between isolated GDD/ID and GDD/ID with comorbidities. CONCLUSION: In children with unexplained GDD/ID, WES demonstrates superior diagnostic and clinical utility compared to CMA. Incorporating WES as a first-line investigation in the diagnostic evaluation of GDD/ID may be warranted.

Humans

Genomic and food-safety evaluation of Staphylococcus chromogenes in Chinese dairy milk.

Non-aureus staphylococci and mammaliicocci (NASM) cause mastitis and may contaminate milk and dairy products. Milk samples (n&#xa0;=&#xa0;1916) from cows with subclinical or clinical mastitis (SCM and CM, respectively) were collected from 28 large-scale (> 500 lactating cows) Chinese dairy farms. Overall, 999 NASM isolates representing 19 species were identified by MALDI-TOF MS and cpn60 sequencing, with Staphylococcuschromogenes, Mammaliicoccus sciuri and Staphylococcus haemolyticus being most prevalent. Antimicrobial resistance (AMR) was determined with disc diffusion; non-susceptible to penicillin was most common (SCM, 30% and CM, 29%) whereas cefoxitin non-susceptible NASM accounted for 8-10% of isolates; among these, 12.5% carried mecA but none carried mecC. Galleria mellonella was used to assess virulence of 78 strains of S. chromogenes, a dominant species; subsequently, 32 strains, representing higher- and lower-virulence in the Galleria model, were selected for whole-genome sequencing and comparative genomics. S. chromogenes isolates from CM had higher virulence (p&#xa0;<&#xa0;0.05) than those from SCM. The 32 genomes comprised 20 sequence types, indicating high genetic diversity. No robust genomic marker of Galleria virulence phenotype was identified in this selected WGS subset. Acquired resistance genes (n&#xa0;=&#xa0;5) were detected, including a first report of fusC in S. chromogenes; the fusC-positive isolate had an elevated fusidic acid MIC (8&#xa0;mg/L). Although S. chromogenes persisted in milk at 4&#xa0;&#xb0;C, pasteurization (64&#xa0;&#xb0;C for 30&#xa0;min) reduced viable counts to below detection. This study provided new insights into the prevalence, AMR, genomic diversity, and dairy-chain relevance of milk-derived NASM, particularly S. chromogenes. However, the genomic findings were based on an intentionally selected WGS subset and should be interpreted as hypothesis-generating rather than population-representative.

Animals

Regional genomic analysis of lineage distribution and transferable multidrug resistance among chicken-associated Salmonella Kentucky isolates in China.

Salmonella enterica serovar Kentucky is an important multidrug-resistant foodborne pathogen in the poultry meat supply chain. Although recent broader genomic studies have elucidated the population structure and epidemiological significance of major lineages in China (e.g., ST198 and ST314), the regional dynamics within local poultry supply chains remain insufficiently characterized. In this study, 31 chicken meat-derived isolates from Shanghai and 39 publicly available genomes from China were analyzed using antimicrobial susceptibility testing, whole-genome sequencing, phylogenetic analysis, conjugation experiments, and complete sequencing of representative plasmids. This enabled a systematic characterization of the molecular epidemiological features of the population and the mechanisms underlying resistance dissemination. Population genomic analysis revealed a lineage composition markedly different from the global epidemiological pattern: ST314 was the predominant sequence type among the Shanghai chicken-derived isolates (74.2%), whereas the internationally recognized high-risk clone ST198 accounted for only 25.8% of the local isolates. However, risk stratification analysis indicated that although ST198 was detected less frequently, it carried a significantly greater burden of acquired resistance genes and therefore represented a higher-risk resistant lineage. Functional and structural validation further elucidated the molecular basis of resistance dissemination within this high-risk lineage. Conjugation experiments confirmed the co-transfer of a multidrug resistance module carrying blaTEM-1 and blaCTX-M-267 to the recipient strain Escherichia coli J53. Complete plasmid analysis revealed that these two &#x3b2;-lactam resistance genes were co-localized on a 242-kb transferable plasmid flanked by Tn1331, Tn3, and multiple transposase-associated elements, thereby providing a structural basis for their horizontal transfer. This study provides important molecular epidemiological evidence for lineage-specific surveillance and risk-stratified control of resistant Salmonella in the poultry meat supply chain and further underscores the need for continuous monitoring of mobile genetic elements within a One Health framework.

Animals

Whole-Genome Deep Learning Predicts Chemotherapy Response in Colorectal Cancer.

Chemotherapy response in colorectal cancer (CRC) exhibits significant heterogeneity, with current clinical predictors failing to capture complex genomic determinants of resistance. We developed a hybrid deep learning framework integrating convolutional neural networks (CNNs) and bidirectional long short-term memory (BiLSTM) networks to analyze whole-genome somatic mutations, evolutionary conservation, chromatin accessibility, and 3D genome architecture in 2,546 TCGA patients. An attention mechanism identified predictive genomic regions. The model achieved an AUC of 0.92 (95% CI: 0.89-0.94) in cross-validation and 0.88 (95% CI: 0.85-0.91) in independent validation, outperforming clinical models (&#x394;AUC = +0.18, p < 0.001). Key predictors included non-coding variants in TP53, KRAS, and PIK3CA regulatory regions. Triple-positive patients (mutations in all 3 regions) had significantly worse progression-free survival (HR = 4.7, p < 0.001). Our framework enables accurate chemotherapy response prediction and reveals novel non-coding resistance mechanisms, advancing precision oncology in CRC.

Humans

Urinary Small Extracellular Vesicle DNA as a Biomarker for the Non-Invasive Diagnosis of Bladder Cancer.

Existing diagnostic technologies for bladder cancer (BC) suffer from low sensitivity, low specificity, or a lack of validation. Therefore, validated, non-invasive diagnostic biomarkers with high sensitivity and specificity for early detection of BC are needed to complement and improve upon the limitations of existing diagnostic methods. We used low-pass whole genome sequencing (LP-WGS) technology to detect copy number variations (CNVs) in small extracellular vesicle (sEV) DNA isolated from urine samples of patients. Based on these results, we constructed and validated a diagnostic model to differentiate between benign and malignant bladder lesions. We conducted a receiver operating characteristic analysis and calculated the area under the curve (AUC) to evaluate the performance of the diagnostic model. The urine sEV-DNA LP-WGS data revealed CNV differences between benign and malignant samples. The diagnostic model achieved an AUC of 0.953, a sensitivity of 86.7%, and a specificity of 100% in the training cohort and an AUC of 0.985, a sensitivity of 90%, and a specificity of 100% in the validation cohort. Even at the lowest coverage depth of 0.01X, the performance of the diagnostic model remained relatively robust. Notably, the performance of this diagnostic model surpassed that of the biomarker neuron-specific enolase (sensitivity: 85.7% vs. 64.3%; specificity: 100% vs. 87.5%) and urinary cytology (sensitivity: 100% vs. 66.7%; specificity: 100% vs. 94.1%). Our study demonstrates that urine sEV-DNA exhibits high discriminatory power in distinguishing between benign and malignant bladder lesions, making it a promising tool for auxiliary diagnosis of BC.

Humans

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Artificial Intelligence