Search PubMedSearch

SEARCH · Search PubMed

Results for “simplified genome sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

[Genetic diversity analysis of Forsythia suspensa germplasm resources in Shanxi based on phenotypic traits and SNP molecular markers].

This study aimed to clarify the degree of fruit phenotypic variation and the characteristics of genetic diversity, population structure, and genetic differentiation of Forsythia suspensa resources in Shanxi, providing an important basis for germplasm conservation and breeding of superior varieties. A total of 46 F. suspensa fruits were collected, and 12 agronomic traits were measured and analyzed. The population genetic structure and genetic diversity of F. suspensa germplasm were evaluated using simplified genome sequencing technology. For the five quality traits of the 46 fruits, the Shannon-Wiener index ranged from 0.631 to 1.074, and the Simpson index ranged from 0.379 to 0.560. The seven quantitative traits exhibited abundant genetic variation, with coefficients of variation ranging from 9.764%(fruit shape index) to 45.494%(forsythin content). Principal component analysis reduced the 12 phenotypic traits to four factors, with a cumulative variance contribution of 74.547%. Sequencing data showed mean Q20 and Q30 values of 98.13% and 94.33%, respectively, with an average GC content of 35.95%. After filtering, a total of 12 347 327 high-quality single nucleotide polymorphism(SNP) loci were obtained. Based on these high-quality SNPs, principal component analysis, population structure analysis, and phylogenetic tree construction were carried out. The 46 germplasm resources were divided into four groups; however, grouping showed little relationship with geographic origin, and intermixing occurred among regions. Mantel test revealed a significant but weak positive correlation between phenotypic and genetic distances(r=0.159, P=0.001). At the molecular level, the four groups exhibited moderate genetic diversity overall, and the genetic differentiation index among populations ranged from 0.027 to 0.084, indicating low to moderate differentiation. The rich genetic diversity of the main phenotypic traits provides a solid material basis for screening superior germplasm and genetic breeding of F. suspensa.

Forsythia

Identifying loci under positive selection in complex population histories.

Detailed modeling of a species' history is of prime importance for understanding how natural selection operates over time. Most methods designed to detect positive selection along sequenced genomes, however, use simplified representations of past histories as null models of genetic drift. Here, we present the first method that can detect signatures of strong local adaptation across the genome using arbitrarily complex admixture graphs, which are typically used to describe the history of past divergence and admixture events among any number of populations. The method-called graph-aware retrieval of selective sweeps (GRoSS)-has good power to detect loci in the genome with strong evidence for past selective sweeps and can also identify which branch of the graph was most affected by the sweep. As evidence of its utility, we apply the method to bovine, codfish, and human population genomic data containing panels of multiple populations related in complex ways. We find new candidate genes for important adaptive functions, including immunity and metabolism in understudied human populations, as well as muscle mass, milk production, and tameness in specific bovine breeds. We are also able to pinpoint the emergence of large regions of differentiation owing to inversions in the history of Atlantic codfish.

Animals

Selective restriction endonuclease cleavage of human globin genes.

Double-stranded human globin DNA synthesized in vitro from sickle cell mRNA has been used as a substrate for a series of restriction endonucleases. The double-stranded DNA contained full length transcripts of the alpha- and beta- globin genes. Of the 10 enzymes tested, only 3 (Hpa I, Sal I, and Kpn I) failed to cleave either alpha- or beta-DNA; 2 (Eco RI and Bam HI) cleaved only beta-DNA; 3 (HindIII, Hpa II, and Hha I) cleaved only alpha-DNA; and 2 (Hae III and Alu I) cleaved both alpha- and beta-DNAs. The selective cleavage of human globin genes by restriction endonucleases should provide a strategy for the identification and purification of DNA fragments of genomic DNA containing globin genes plus their flanking sequences, simplify the preparation of pure, chain-specific globin probes, and permit the isolation of DNA probes for specific regions of the globin genes.

DNA

Optimized Amplicon Strategy for Long-Read Sequencing of the Chikungunya Virus Genome.

Chikungunya virus (CHIKV) is a positive-sense RNA alphavirus transmitted to humans primarily by Aedes aegypti and Aedes albopictus mosquitoes. Its global circulation and significant public health impact underscore the need to better understand the molecular mechanisms driving CHIKV pathogenesis and transmission. Although robust molecular biology methods exist for CHIKV genome sequencing, a major limitation for surveillance and research is the inability to determine whether two nucleotide variations co-occur within the same viral genome when they are separated beyond the span of typical short-read designs. Here, we describe an optimized approach for processing CHIKV RNA samples that generates large amplicons suitable for long-read nanopore sequencing. This protocol enables amplification of the complete CHIKV genome in only two or three amplicons and facilitates detection of co-occurring nucleotide variations across 4-7.5 kb within the same molecule, thereby simplifying sequencing workflows and improving resolution in studies of viral evolution.

Chikungunya virus

CRISPRessoSea: streamlined analysis and comparison of pooled amplicon CRISPR screens.

BACKGROUND: CRISPR genome editing enables precise modification of genomic targets but may also induce unintended edits at off-target sites with similar sequences. Pooled amplicon sequencing can assess on- and off-target editing across many samples, yet analyzing, aggregating, and visualizing results from multiple pooled experiments remains challenging. Tools to simplify and standardize these analyses are needed to provide reproducible and comparable interpretation of editing data. RESULTS: We developed CRISPRessoSea, a software package that processes, compares, and visualizes genome editing rates from pooled amplicon sequencing experiments. The tool provides standardized workflows for analyzing editing across multiple targets and samples, supports both nuclease- and base-editing modalities, and generates clear, data-rich summaries suitable for downstream interpretation. CONCLUSIONS: CRISPRessoSea facilitates reproducible, scalable analysis of CRISPR editing outcomes across diverse experimental designs, enabling more efficient and transparent assessment of genome editing specificity. The software is freely available at https://github.com/clementlab/CRISPRessoSea .

Software

Performance of seven carbapenemase detection assays in Pseudomonas aeruginosa across different epidemiological settings: a multicenter cross-sectional study.

The detection of carbapenemases in Pseudomonas aeruginosa remains challenging due to a great variety of other resistance mechanisms, and most laboratories, therefore, do not test for them. This study aimed to comparatively evaluate seven phenotypic carbapenemase detection assays across three epidemiological settings. A total of 320 P. aeruginosa isolates from three German centers with varying carbapenemase prevalences (5.8%-51.4%), including 113 carbapenemase-producing isolates carrying VIM-2 (n = 58), NDM-1 (n = 19), and GIM-1 (n = 16), underwent whole-genome sequencing as reference to assess seven phenotypic carbapenemase-detection tests: modified- and modified-zinc-supplemented carbapenem inactivation method (mCIM and mzCIM), simplified carbapenem inactivation method (sCIM), Carba NP, imipenem-cloxacillin test (IC-4000), and two imipenem-EDTA disk assays. Of all confirmation assays, mzCIM and sCIM showed the best overall performance for carbapenemase detection (sensitivity/specificity: 100%/94.2% and 99.1%/92.8%), followed by mCIM (93.8%/96.1%). Carba NP achieved the highest specificity (99.0%), but the lowest sensitivity (85.8%). EDTA-based assays and IC-4000 were highly sensitive (96.5%-100%) but less specific (79.7%-87.0%). Negative predictive values were consistently high (≥98%-100%) across all assays and prevalence settings, whereas positive predictive values varied (72.5%-98.0%). Both mzCIM and sCIM exhibited robust performance for carbapenemase detection in P. aeruginosa, representing the most suitable approach across diverse epidemiological settings. Their high negative predictive values indicate that these assays are particularly effective for ruling out carbapenemase production. Furthermore, both assays are cost-effective, simple to perform, and can be readily implemented in any routine microbiology laboratory.IMPORTANCEThis study provides comparative diagnostic accuracy data for seven phenotypic carbapenemase detection assays in Pseudomonas aeruginosa (PA) across different prevalence settings. Modified-zinc-supplemented carbapenem inactivation method (mzCIM) and simplified carbapenem inactivation method (sCIM) are the most robust screening tools and show that local carbapenemase-producing P. aeruginosa (CP-PA) prevalence substantially influences the utility of all evaluated assays.

CIM

VDJ-Insights: simplifying the annotation of genomic immunoglobulin and T cell receptor regions.

MOTIVATION: Accurate annotation of germline immunoglobulin (IG) and T cell receptor (TCR) loci is critical for understanding adaptive immunity. RESULTS: VDJ-Insights provides a user-friendly software package for characterizing these complex immune regions. In addition, it assesses gene segment functionality, identifies recombination signal sequences, and annotates complementarity-determining regions 1 and 2. VDJ-Insights achieved over 99% concordance with curated annotations from multiple species, outperforming existing annotation tools. When applied to 95 haplotypes from the Human Pangenome Reference Consortium, VDJ-Insights identified 652 and 275 novel IG and TCR alleles, respectively, highlighting its scalability for large immunogenetic studies. AVAILABILITY AND IMPLEMENTATION: Datasets and software package are available in the VDJ-insights repository, https://github.com/BPRC-Bioinfo and https://doi.org/10.5281/zenodo.17588835. Additional intermediate datasets used and analyzed during the current study are available from the corresponding authors upon reasonable request.

Software

OpenSpliceAI: An efficient, modular implementation of SpliceAI enabling easy retraining on non-human species.

The SpliceAI deep learning system is currently one of the most accurate methods for identifying splicing signals directly from DNA sequences. However, its utility is limited by its reliance on older software frameworks and human-centric training data. Here we introduce OpenSpliceAI, a trainable, open-source version of SpliceAI implemented in PyTorch to address these challenges. OpenSpliceAI supports both training from scratch and transfer learning, enabling seamless retraining on species-specific datasets and mitigating human-centric biases. Our experiments show that it achieves faster processing speeds and lower memory usage than the original SpliceAI code, allowing large-scale analyses of extensive genomic regions on a single GPU. Additionally, OpenSpliceAI's flexible architecture makes for easier integration with established machine learning ecosystems, simplifying the development of custom splicing models for different species and applications. We demonstrate that OpenSpliceAI's output is highly concordant with SpliceAI. In silico mutagenesis (ISM) analyses confirm that both models rely on similar sequence features, and calibration experiments demonstrate similar score probability estimates.

Journal Article

IBDV-SSA, a novel molecular approach for the recovery of infectious bursal disease virus whole genomes from FTA cards.

Infectious bursal disease (IBD), a highly contagious viral disease in young chickens, poses significant economic losses due to high mortality and immunosuppression. While IBD virus (IBDV) virulence is influenced by multiple genes, whole-genome sequencing (WGS) of IBDV is crucial for defining the strain pathotype and clinical profile. Flinders Technology Associates (FTA) cards are convenient for field sample collection, but their filter paper matrix can hinder nucleic acid recovery, impacting sequencing efficiency. This study evaluated two enrichment strategies, single primer amplification (SPA) and IBDV segment-specific amplification (SSA), coupled with short-read (Illumina) and long-read (Oxford Nanopore Technologies, ONT) sequencing platforms, to optimize IBDV whole-genome recovery from FTA cards. Illumina sequencing produced comparable raw read counts for both methods, yet IBDV-SSA samples achieved significantly higher genome mapping rates (76%) than IBDV-SPA (12%). Genome coverage analysis revealed that IBDV-SSA provided uniform read distribution across both genomic segments, ensuring complete coverage, while IBDV-SPA exhibited significant bias, with most reads mapping to segment B, and limited coverage of segment A. Importantly, IBDV-SSA also proved compatible with ONT long-read sequencing, providing complete genome coverage. Notably, IBDV-SSA coupled with short-read sequencing successfully characterized coinfections in two samples. This optimized approach using IBDV-SSA enables efficient and comprehensive WGS of IBDV from FTA cards, facilitating strain characterization, virulence prediction, and epidemiological investigations.IMPORTANCEThis research tackles a significant problem for poultry farmers: a virus called infectious bursal disease virus (IBDV) that harms young chickens, causing high death rates and economic losses. To fight it effectively, scientists need to analyze its complete genetic makeup. Traditionally, collecting and preserving IBDV field samples was challenging. Flinders Technology Associates (FTA) cards have simplified this process, but getting usable genetic material from them has been difficult. This study introduces a new genome enrichment method, IBDV segment-specific amplification (IBDV-SSA), which successfully allows for IBDV complete genome recovery from FTA cards. By using this improved approach, scientists can accurately identify virus strains, assess how harmful they are, and monitor their spread. This, in turn, helps to improve vaccines and protect flocks. IBDV-SSA is a powerful tool for outbreak surveillance, supporting the poultry industry and ensuring a stable food supply.

Infectious bursal disease virus

Identification of lineage-associated polymorphisms in the ROP18 3' flanking region and development of molecular assays for differentiation of Toxoplasma gondii lineages.

BACKGROUND: Toxoplasma gondii (T. gondii) exhibits substantial genetic diversity, and different parasite lineages are associated with distinct epidemiological distributions and biological characteristics. Accurate molecular characterization of T. gondii strains is important for understanding parasite population structure and transmission patterns. However, existing genotyping approaches often require multiple loci, extensive experimental procedures, or complex data analysis. Therefore, simplified and reliable molecular markers for rapid lineage differentiation are still needed. METHODS: In this study, comparative genomic analysis was performed using representative T. gondii strains with well-defined genetic backgrounds and virulence phenotypes. The ROP18 genomic region, including partial genomic sequences, 5' flanking regions, coding sequence (CDS), and 3' flanking regions, was analyzed to identify informative polymorphic signatures. A short conserved sequence region containing lineage-associated polymorphic sites was identified within the ROP18 3' flanking region. Based on these sequence signatures, HRM-PCR and TaqMan MGB probe-based real-time PCR assays were developed and evaluated using plasmid standards and representative T. gondii genomic DNA samples. RESULTS: Phylogenetic analyses based on different ROP18 genomic regions demonstrated distinct clustering patterns among analyzed strains. Although the ROP18 3' flanking region was highly conserved, a short conserved sequence region containing informative polymorphic sites was identified, and the combination of these sites generated three distinct lineage-associated ROP18 patterns. Analysis of publicly available genomic datasets further demonstrated that individual strains contained one of these defined patterns rather than multiple patterns simultaneously. The developed HRM-PCR assay successfully discriminated the three ROP18-associated patterns based on distinct melting profiles with good reproducibility. Furthermore, the TaqMan MGB probe-based assay enabled specific identification of different ROP18-associated patterns through defined probe-recognition combinations and showed good analytical performance. CONCLUSION: This study identifies novel lineage-associated molecular signatures within the ROP18 3' flanking region and establishes complementary HRM-PCR and TaqMan MGB probe-based approaches for rapid molecular differentiation of T. gondii strains. These findings highlight the potential of conserved non-coding regions adjacent to functionally important genes as informative targets for parasite genotyping and provide a practical complementary tool for epidemiological surveillance and strain characterization.

HRM-PCR

Extensive Analysis of Genetic Diversity in HLA-DMA, HLA-DMB, HLA-DOA and HLA-DOB: Characterisation of 236 Novel Alleles.

HLA-DMA, -DMB, -DOA and -DOB are non-classical HLA Class II genes that play a crucial role in the selection of highly stable HLA Class II/peptide complexes on antigen-presenting cells. Although the genes were initially thought to have a limited diversity with less than 13 alleles per gene documented in the IPD-IMGT/HLA Database in 2022, recent studies suggest a potential impact of certain alleles on the outcome of hematopoietic cell transplantation. To gain a deeper understanding of allelic diversity, we sequenced HLA-DMA, -DMB, -DOA and -DOB of 1880 potential stem cell donors from Germany, Poland, Great Britain and Chile, achieving full-gene resolution. Remarkably, we identified 3968 previously undescribed sequences, including 28 distinct novel proteins. The observed allele frequencies were consistent across all studied populations with one dominating protein for each gene: HLA-DMA*01:01 (> 77%), HLA-DMB*01:01 (> 63%), HLA-DOA*01:01 (> 97%) and HLA-DOB*01:01 (> 77%). Notably, a much higher diversity was observed in full-genomic resolution. Finally, we submitted 51 distinct novel sequences for HLA-DMA, 58 for HLA-DMB, 80 for HLA-DOA and 47 for HLA-DOB to the IPD-IMGT/HLA Database. This comprehensive reference database update will not only simplify future genotyping of HLA-DMA, -DMB, -DOA and -DOB but will hopefully also enhance our understanding of the complex process of peptide selection and loading to the HLA Class II proteins.

Humans

Trapped-oligonucleotide nucleotide incorporation (TONI) assay, a simple method for screening point mutations.

We present a simple screening method for detecting a known point mutation, using only one 5'-biotinylated oligonucleotide primer, with its 3' end adjacent to the mutation site. In parallel reactions, an amplified DNA template encompassing the biotinylated oligonucleotide and mutation site undergoes 40 step-cycles of single nucleotide incorporation using Taq thermostable DNA polymerase and only one radioactive [alpha-32P]dNTP, specified by either the normal or mutant sequence. The oligonucleotides, now radioactively labelled at the 3' end according to the template sequence, are then trapped by streptavidin-coated magnetic beads, and the percent of radiolabel incorporated is determined directly by the Cerenkov method in a scintillation counter. The trapped-oligonucleotide nucleotide incorporation (TONI) assay has been used for the screening of a mitochondrial polymorphism, and has also been shown to distinguish the genotypes of hemoglobin A/C, A/A, A/S, and S/S. It is reproducible over at least a 100-fold range of radioisotope and a 10-fold range of oligonucleotide primer. This method is particularly useful for diagnosing mutations which do not produce alterations detectable by restriction enzyme analysis, since optimization of conditions is rarely necessary. In addition, it requires only a single oligonucleotide, and no electrophoretic separation of the allele-specific products. It thus represents an improved and simplified modification of the existing allele-specific primer extension methods (Kuppuswamy et al., Proc Natl Acad Sci USA 88:1143-1147, 1991; Sokolov, Nucl Acids Res 18:3671, 1989; Syvanen et al., Genomics 8:684-692, 1990).

Base Sequence

Diversity at the HYP1 locus in potato cyst nematodes does not result from developmentally-programmed somatic mutations.

Most genetic diversity stems from spontaneous mutations, that is, errors in DNA repair or replication. But for dozens of organisms across the tree of life, mutations at specific loci are not spontaneous but developmentally programmed: effectively, some organisms edit their own DNA sequences. This is perhaps most common among pathogens and parasites, many of which use editing to diversify genes that produce important antigens. Plant-parasitic potato cyst nematodes are damaging agricultural pests that establish a lifelong feeding site inside the root of their host plant. We previously observed extensive diversity of rare alleles at HYP1, the most highly expressed gene that encodes a protein secreted by potato cyst nematodes during parasitism. Importantly, HYP1 alleles differ from each other by complex, in-frame rearrangements of short repeated sequence motifs within a single exon. Combining several lines of evidence, we previously hypothesized that potato cyst nematodes use developmentally-programmed mutations, or editing, to diversify HYP1 alleles in the soma. In the current work, we now test this hypothesis. We employ highly accurate long-read DNA sequencing of a simplified genetic system to identify potential rare edited alleles, we use a transgenic yeast system to describe large de novo mutations at HYP1, and we interpret our findings in light of key population genetic parameters as well as the genetic diversity surrounding HYP1 and across the genome.

Animals

Genomic signatures associated with epidemiologically defined high-risk pathogenic Escherichia coli isolates identified by interpretable machine learning.

Pathogenic Escherichia coli is a major cause of foodborne illness worldwide and includes strains capable of causing severe disease. To establish a genome-informed framework for foodborne outbreak surveillance, we analyzed 1,029 E. coli isolates from clinical, food, livestock, and environmental sources using whole-genome sequencing. Pathogenic isolates obtained from human clinical cases or linked to documented outbreaks were classified as epidemiologically defined high-risk (EpiHR), whereas the remaining pathogenic isolates were classified as non-EpiHR. Virulence-associated genomic features were extracted using a bioinformatics pipeline, and four machine learning (ML) algorithms, including gradient boosting machine, random forest (RF), and support vector machines with linear and radial basis function kernels, were evaluated. Among them, the RF model showed the best performance, achieving an area under the curve (AUC) of 0.98 and accuracy of 0.93 in 10-fold cross-validation. Additional leave-one-group-out validation showed retained discrimination across held-out sequence types and serotypes, although performance was reduced when isolates were grouped by isolation source. Evaluation using an independent test dataset of 1,908 publicly available pathogenic E. coli genomes showed an AUC of 0.97 and a sensitivity of 0.98. Feature importance analysis using Shapley additive explanations identified influential predictive features, including traT, etpB, and enterotoxin-associated genes. A reduced 10-feature model achieved an AUC of 0.79 in the independent test dataset, supporting its exploratory use for future simplified screening approaches. These results indicate that genome-based ML provides a sensitive framework for surveillance-oriented prioritization of EpiHR pathogenic E. coli isolates, with model predictions interpreted together with epidemiological information.

Escherichia coli

IQ-NET: fast and accurate quartet phylogenetic inference using deep learning trained on empirical DNA alignments.

Phylogenetic inference is fundamental to modern biology, with many applications including evolutionary biology, epidemiology, and comparative genomics. While maximum likelihood and Bayesian methods remain the gold standard for phylogenetic analysis, they rely on simplifying assumptions and are computationally intensive. Recent machine learning approaches for phylogenetics offer speed advantages, but have several limitations: exclusive reliance on simulated data for training, inadequate handling of gaps, and sensitivity to input sequence order. Here, we introduce IQ-NET (Intelligent Quartet NETwork), a deep learning framework that solves these limitations to infer four-taxon trees. IQ-NET estimates both tree topology and branch lengths directly from gapped alignments. IQ-NET outperforms existing machine learning methods in terms of accuracy, and obtained a 24-fold speedup compared with the widely used maximum likelihood software, IQ-TREE. We finally introduce a pipeline using IQ-NET and the ASTRAL software to reconstruct a larger species tree, i.e., with more than four taxa.

Empirical data training

Characteristics and assembly mechanisms of tobacco-associated bacteria in typical tobacco-planting regions across China.

INTRODUCTION: Plant-associated microbiota critically modulates host growth and environmental adaptation, yet assembly mechanisms, niche differentiation, and ecological strategies of bacterial communities inhabiting tobacco microhabitats remain poorly elucidated across geographical gradients. METHODS: Here, we systematically characterized bacterial microbiome assembly across five tobacco-associated niches (bulk soil, rhizosphere soil, root, stem, and leaf) from seven typical tobacco-planting regions using 16S rRNA amplicon sequencing, genome annotation, and niche breadth analysis. The independent and interactive effects of geographical location and host compartment on community structure, and further compared genomic traits, functional profiles, and life-history strategies between specialist and generalist bacterial populations were quantified. RESULTS: The results revealed a deterministic soil-plant continuum stratification of bacterial communities and diversity, with progressively simplified communities and decreasing alpha diversity from bulk soil to above-ground tissues, accompanied by progressive dominance of Proteobacteria. Geographical factors predominantly structured soil microbial communities via divergent edaphic properties, while host filtering acted as a universal dominant driver shaping endophytic microbiome assembly. Niche differentiation analysis demonstrated that niche-specialized bacterial ASVs overwhelmingly dominated all microhabitats and geographical sites, whereas generalist taxa only constituted auxiliary populations. Although specialist and generalist microbes exhibited highly conserved core genomic architectures and overall functional repertoires, they displayed distinct niche-specific functional divergence in metabolic pathways, stress resistance, and secondary metabolism across host compartments. Life-history strategy analysis further revealed that Y-strategist represented the core adaptive bacterial population, especially enriched in above-ground tobacco tissues. DISCUSSION: Our study establishes a hierarchical dual-filtering assembly model for tobacco microbiota, clarifies the ecological differentiation and functional adaptation of specialist and generalist bacteria, and provides fundamental insights into the assembly rules and adaptive mechanisms of crop-associated microbiomes for future microbial resource utilization and agricultural microbiome regulation.

biogeography

Association Between cnm-Positive Streptococci and Cerebral Small Vessel Disease: Insights From Oral Health and Microbiome Status.

INTRODUCTION AND AIMS: Cerebral small vessel disease (CSVD) is associated with various severe neurological outcomes; while oral cnm-positive streptococci are suggested to be involved in cerebrovascular lesions, the specific associative features between these bacteria and CSVD have not yet been systematically investigated. This study aims to investigate the prevalence of cnm-positive streptococci in patients with CSVD and explore the correlation between infection and CSVD severity. By integrating oral health indices and microbiome sequencing, we evaluate the oral hygiene status and microbial dysbiosis characteristics of cnm-positive streptococci carriers. Furthermore, cnm-positive streptococci derived from CSVD patients will be isolated, identified, and subjected to whole-genome sequencing to provide a foundation for future research. METHODS: To explore cnm-positive streptococci prevalence and its association with CSVD, we conducted a case-control study comparing their oral detection rates between healthy controls and CSVD patients. We also performed 16S rRNA gene high-throughput sequencing of oral plaque microbiota and assessed oral health, including the simplified oral hygiene index (OHI-S), the decayed, missing, and filled teeth (DMFT) index, oral hygiene practices, gingival status, and saliva scores. RESULTS: cnm-positive streptococci were more prevalent in CSVD patients, correlating with higher OHI-S and microbial dysbiosis. Multivariable regression models (adjusted for demographic/vascular risk factors) linked cnm positivity to periventricular hyperintensities (PVH), deep white matter hyperintensities (DWMH), Fazekas score, and total CSVD burden (not cerebral microbleeds (CMBs)/lacunes). CONCLUSION: Oral cnm-positive streptococci are independently associated with CSVD phenotypes, particularly those characterized by white matter injury. These findings presents a potential oral-cerebrovascular interaction and imply that managing specific virulent oral strains may be a noteworthy consideration in future clinical research.

Humans

Rapid assessment of clinical severity for salmonellosis cases via protein family domain analysis and machine learning.

Salmonella is a common pathogen, infecting more than a million people yearly. Rapid assessment of clinical case severity is essential for improving patient outcomes and optimizing healthcare resources. Advancements in genome sequencing technologies have enabled the analysis of bacterial genomes from many clinical cases, opening up new opportunities for precise and timely diagnosis. This study proposes a genome-based framework for identifying critical Salmonella cases before the onset of critical symptoms and facilitating early medical intervention. By leveraging protein family (Pfam) domains as the representation for genomic data, the complex genetic profiles of Salmonella cases are simplified into interpretable features. The severity levels of cases were investigated through rigorous data analysis, resulting in a set of 70 Pfam domains that could be potentially used as biomarkers. Machine Learning was employed to assess the predictive power of the curated Pfam biomarkers, achieving high accuracy (~93%) in sorting cases into critical, moderate, and mild categories. The results demonstrate the efficacy of the proposed approach. This framework highlights the potential of using bacterial genomic data in clinical decision-making, opening the window for timely personalized interventions for Salmonella infection management.

Domains of unknown function (DUFs)