Search PubMedSearch

SEARCH · Search PubMed

Results for “bioinformatics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Bioinformatic Analysis of Bacillus pacificus B630: Molecular Understanding of Biofilm Production.

The aim of this study was to determine biofilm production and motility in Bacillus pacificus B630 and Bacillus cereus ATCC 14579, and to perform a comparative genome analysis using bioinformatic tools to understand the differences between the two strains. Biofilm production was performed in glass tubes stained with safranin; motility was determined on soft agar. Bioinformatic analysis was performed using genomic information from both strains, including the identification of orthologous genes, the similarity between genes of the eps1 and sipW-tasA-calY operons, and the SipW and TasA model prediction. B. pacificus B630 produces a greater amount of biofilm on glass than B. cereus ATCC 14579 (p < 0.01). Furthermore, B. pacificus B630 shows lower motility than B. cereus ATCC 14579 (p < 0.001). B. pacificus B630 contains 45 unshared genes, whereas B. cereus ATCC 14579 has 27 unshared genes. Differences in similarity were observed between the genes of the eps1 and sipW-tasA-calY operons. These differences between SipW and TasA may affect protein structural predictions. In SipW, the differences may affect the C-terminal region. In TasA, the number of B-sheets differed between the two proteins, and amino acid substitutions were found in regions of high protein aggregation. Genomic differences in genes associated with biofilm production may explain differences in biofilm production between the strains studied.

Biofilms

Identification of a Nonribosomal Peptide Analog With Activity Against Multiple Gram-Positive Bacteria via a Synthetic Bioinformatic Natural Product Discovery Approach.

Nonribosomal peptide (NRP) antibiotics exhibit potent biological activities and are broadly used in clinical therapy. Because most microorganisms are difficult to culture and many antibiotic biosynthetic genes are silent, traditional activity tracking approaches face major limitations in the discovery of novel NRPs. Here, based on a synthetic bioinformatic natural product (syn-BNP) discovery approach that integrates bioinformatics and chemical synthesis, a novel nonribosomal peptide synthetase (NRPS) gene cluster from the genome of Rhodococcus erythropolis D-1 was mined. A putative NRP scaffold synthesized by the NRPS encoded by this cluster was predicted. Through chemical synthesis and four rounds of structure-activity relationship (SAR) studies, 37 NRP analogs were ultimately generated. Among these analogs, ZURJC28 shows activity against multiple Gram-positive bacteria, including two drug-resistant strains. Mechanistic studies and metabolomics analyses revealed that ZURJC28 exerts membrane-disruptive activity associated with interaction with phosphatidylglycerol (PG)-enriched Gram-positive membranes, leading to membrane damage and widespread metabolic dysregulation. ZURJC28 also shows low cytotoxicity and low hemolytic activity, suggesting its preliminary in vitro safety profile.

Gram-Positive Bacteria

PathoSeq-QC: a decision support bioinformatics workflow for robust genomic surveillance.

MOTIVATION: Recommendations on the use of genomics for pathogens surveillance are evidence that high-throughput genomic sequencing plays a key role to fight global health threats. Coupled with bioinformatics and other data types (e.g., epidemiological information), genomics is used to obtain knowledge on health pathogenic threats and insights on their evolution, to monitor pathogens spread, and to evaluate the effectiveness of countermeasures. From a decision-making policy perspective, it is essential to ensure the entire process's quality before relying on analysis results as evidence. Available workflows usually offer quality assessment tools that are primarily focused on the quality of raw NGS reads but often struggle to keep pace with new technologies and threats, and fail to provide a robust consensus on results, necessitating manual evaluation of multiple tool outputs. RESULTS: We present PathoSeq-QC, a bioinformatics decision support workflow developed to improve the trustworthiness of genomic surveillance analyses and conclusions. Designed for SARS-CoV-2, it is suitable for any viral threat. In the specific case of SARS-CoV-2, PathoSeq-QC: (i) evaluates the quality of the raw data; (ii) assesses whether the analysed sample is composed by single or multiple lineages; (iii) produces robust variant calling results via multi-tool comparison; (iv) reports whether the produced data are in support of a recombinant virus, a novel or an already known lineage. The tool is modular, which will allow easy functionalities extension. AVAILABILITY AND IMPLEMENTATION: PathoSeq-QC is a command-line tool written in Python and R. The code is available at https://code.europa.eu/dighealth/pathoseq-qc.

Genomics

Secure bioinformatics: privacy-preserving federated analytics using homomorphic encryption.

MOTIVATION: Large-scale bioinformatics analyses increasingly require collaboration across multiple cohorts and institutions, yet existing workflows often rely on data co-localization, which is slow, difficult to scale, and raises privacy concerns. We present a privacy-preserving federated analytics framework that enables secure statistical analysis across distributed datasets without transferring raw data, by performing all computations on encrypted data via cryptographic methods. RESULTS: We evaluate the framework by validating polygenic risk scores and conducting meta-analyses on two real-world cohorts. The proposed solution achieves over 99.9% accuracy relative to plaintext analyses, while maintaining scalable runtime performance with increasing data size and number of participating sites. These results demonstrate the feasibility of secure federated analytics for practical bioinformatics applications involving sensitive data.

Computational Biology

Exploring the mechanism of Shengmai San in treating lung adenocarcinoma based on bioinformatics and molecular dynamics simulation.

To investigate the mechanism of Shengmai San (SMS) in the treatment of lung adenocarcinoma (LUAD) based on an integrated strategy combining "network pharmacology, bioinformatics, molecular docking, and molecular dynamics simulation," aiming to provide a precise combination therapy strategy and identify potential bioactive compounds. Differentially expressed genes in LUAD were identified from the Gene Expression Omnibus database using R (originally developed at Bell Laboratories and currently managed by Lucent Technologies). SMS components (ginseng, Ophiopogon japonicus, and Schisandra chinensis) were retrieved from encyclopaedia of traditional Chinese medicine, with Lipinski-compliant compounds selected. Compound targets were predicted via SwissTargetPrediction and Similarity Ensemble Approach. Intersecting targets between differentially expressed genes and compound targets were identified for "herbs-compounds-targets-disease" network construction. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes enrichment analyses were performed. Hub targets were identified by analyzing the protein-protein interaction network. High-prognostic relevance targets were screened from The Cancer Genome Atlas. Compounds targeting these were identified through the herbs-compounds-targets-disease network, and absorption, distribution, metabolism, excretion, and toxicity-compliant compounds were selected using SwissADME (a web-based tool provided by the Molecular Modeling Group of the Swiss Institute of Bioinformatics). Core regulatory targets were identified through molecular docking, with complex stability assessed by molecular dynamics simulations. The key bioactive compounds of SMS for treating LUAD were identified as 7-hydroxy-2,5-dimethyl-4H-1-benzopyran-4-one, N-trans-feruloyltyramine, paprazine, and (E)-N-[(2S)-2-hydroxy-2-(4-hydroxyphenyl)ethyl]-3-(4-hydroxyphenyl)prop-2-enamide. Hub targets included AURKA, CCNA2, CCNB1, CDK1, CHEK1, KIF11, NEK2, PLK1, TTK, and TYMS. Among these, CDK1, CHEK1, and PLK1 demonstrated both high-prognostic relevance and strong binding affinity with SMS, emerging as core regulatory targets for SMS in LUAD treatment. Mechanistically, SMS exerts its anticancer effects primarily by modulating the tumor necrosis factor, interleukin-17, cell cycle, and Lipid and atherosclerosis signaling pathways. The active components of SMS, such as paprazine, may exert antitumor effects partly through downregulating CDK1, CHEK1, and PLK1 expression. Although the present study did not examine drug-resistance models or combination regimens, our findings raise the possibility that, in patients with high expression of these genes, combining SMS with standard chemotherapy or targeted therapy could potentially enhance chemosensitivity and mitigate the development of resistance. This hypothesis, however, requires formal testing in appropriate preclinical models and functional validation studies.

Molecular Dynamics Simulation

Triple primary synchronous liver cancer in one patient: the first case report and origin speculation through bioinformatics.

INTRODUCTION: A diagnosis of multiple primary liver tumors is extremely rare. Preoperative diagnosis based on imaging findings is difficult. Moreover, the clinical benefits of treatment strategies for multiple liver cancers remain unclear. Here, we report a case of three synchronous primary liver tumors with three distinct pathological types-hepatocellular carcinoma (HCC), intrahepatic cholangiocarcinoma (ICC), and combined hepatocellular-cholangiocarcinoma (cHCC&#x2011;CCA)-in a single patient. Bioinformatics analysis supported at least two clonal origins, with cHCC&#x2011;CCA and ICC sharing a common lineage based on identical HBV integration sites. CASE PRESENTATION: A 63-year-old female with a history of hepatitis B for several years presented with three lesions in hepatic segment VIII. Multiphase magnetic resonance imaging with gadolinium ethoxybenzyl diethylenetriaminepentaacetic acid revealed a diagnosis of multiple lesions, namely, cHCC&#x2011;CCA, with multiple intrahepatic metastases. The AFP level was normal, while the CA 19&#x2009;-&#x2009;9 level was mildly elevated (normal range&#x2009;&#x2264;&#x2009;30.00 U/ml). Hepatectomy was performed, and postoperative assessment confirmed that the large lesion was cHCC&#x2011;CCA. However, the small lesions close to the large lesion were HCC and ICC. Gene testing revealed distinct mutational profiles among the three tumors. Similar gene mutations were detected in cHCC&#x2011;CCA and ICC. We also found that gene fragments of hepatitis B virus-C (HBV-C) were inserted into the genomes of ICC and cHCC&#x2011;CCA rather than that of HCC. The genomic integration site of HBV-C in cHCC&#x2011;CCA and ICC was the same. CONCLUSION: We report an extremely rare case of three synchronous primary liver tumors with three distinct pathological types (HCC, ICC, and cHCC&#x2011;CCA) in a single patient. Bioinformatics analysis supported at least two clonal origins, with cHCC&#x2011;CCA and ICC sharing a common lineage based on identical HBV integration sites. Hepatectomy represents a potential radical strategy for the treatment of multiple PLCs.

Humans

'PePApipe': A complete bioinformatics analysis pipeline for African Swine Fever Virus genome.

African Swine Fever Virus (ASFV) is of high concern in porcine livestock across the world due to both the high mortality rates and the trade restrictions imposed on affected regions. The viral genome is large and complex, and genomic analysis is essential for tracing its origin and evolution. Although several bioinformatics tools exist for genome assembly and analysis, no single platform integrates all necessary steps in an accessible and systematic way. In this study the authors developed 'PePApipe', a custom-built, user-friendly pipeline that enables rapid, complete, and efficient ASFV genome analysis. It is specifically designed for laboratory professionals with limited bioinformatics experience, requiring only basic command-line knowledge. Starting from raw sequencing data, PePApipe integrates thirteen software tools into one automated workflow, covering quality control and pre-processing of raw reads, de novo genome assembly and variant calling. Programmed in Python, it can be executed locally through bash scripts, or using a Slurm protocol for batch processing of multiple samples. The main outputs are the ASFV consensus genome sequence and a file listing its putative variants compared to the selected reference genome. PePApipe classifies generated files into structured folders and produces intermediate files that can be used as inputs for further or parallel analyses; users can also enable or disable specific steps in each particular case. This pipeline is adaptable and complementary to downstream steps such as viral genome annotation or genome visualization. By consolidating all stages of viral genome analysis into a single automated workflow, PePApipe reduces the likelihood of user error, and enhances reproducibility and efficiency. This user-friendly pipeline facilitates the transition from sequencing to assembly and downstream analysis of viral genomes, ensuring a fast and reliable response to molecular analysis demands. Finally, the pipeline can be easily adapted to the study of other viral species, expanding its application in infectious diseases surveillance.

African Swine Fever Virus

Correlation assessment of SARS-CoV-2 variants and their subvariants present in clinical and wastewater samples in Oregon, USA (February 7, 2021 - February 26, 2022) using the Freyja bioinformatics approach.

BACKGROUND: Wastewater surveillance is a valuable tool for monitoring SARS-CoV-2 at the community level. As the virus diversified into many variants and subvariants that share overlapping mutations, resolving them accurately from wastewater becomes a key bioinformatic challenge. OBJECTIVES AND AIMS: This study evaluated two distinct bioinformatic approaches, multilocus sequence typing (MLST) and Freyja, for identifying SARS-CoV-2 variants and subvariants in Oregon wastewater samples collected from February 2021 to February 2022. METHODS: The MLST approach identified SARS-CoV-2 variants using unique mutations curated from clinical samples. In contrast, the Freyja approach resolved variant and subvariant abundances using genome wide mutation profiles weighted by sequencing depth. In this study, the variant and subvariants relative abundances produced by both approaches were compared against those observed in clinical surveillance data. RESULTS: Both approaches identified SARS-CoV-2 variants at relative abundances that agreed closely with those observed in clinical surveillance data. However, only the Freyja approach identified over 200 Delta subvariants, divided into three clades (21A, 21I and 21J) and two levels (Level 1 and 2) based on Pango subvariants. Delta subvariants showed strong agreement at Level 1 subvariants (rs&#x2009;=&#x2009;0.892-0.944), while agreement at Level 2 subvariants was inconsistent (rs&#x2009;=&#x2009;0.324-0.903). CONCLUSIONS: The Freyja approach provided enhanced resolution of SARS-CoV-2 variants and subvariants in wastewater, at abundances that agreed with clinical surveillance. This added resolution is a critical advantage for public health surveillance as SARS-CoV-2 continues to evolve and share mutations across variants and subvariants.

Oregon

Leveraging bioinformatics approaches for drug repositioning in space radiation protection.

The health effects of space radiation, primarily Galactic Cosmic Rays (GCRs), on humans remain largely unknown, with potential cardiovascular consequences posing a significant threat to astronauts on long-duration spaceflight missions. Currently, there are no established pharmacological countermeasures for GCR exposure. Drug repositioning offers a promising strategy to accelerate pharmaceutical research in space medicine. This study leverages existing bioinformatics techniques to identify and prioritize potential drug candidates associated with proteomic perturbations following simulated GCR exposure using previously published murine cardiac proteomic data. A protein-protein interaction (PPI) network was constructed using the top differentially expressed proteins (DEPs) from murine heart tissue following exposure to 5-ion GCRs as seed nodes, focusing on experimentally supported interactions. Network topology, Markov clustering, and functional enrichment analyses were used to characterize biologically relevant proteins and pathways. Drug-protein interactions were predicted using Drugst.One and mapped to PPI clusters of interest to identify candidate drugs. Selected drug-macromolecule interactions were further explored using CB-Dock2 molecular docking and short-duration molecular dynamics simulations as hypothesis-generating structural assessments. Analysis of a key PPI network cluster consisting of several ATP synthase proteins identified 23 unique drug candidates. These analyses demonstrate a systematic approach for leveraging bioinformatics techniques to identify candidate molecular targets and generate pharmacological hypotheses in the context of space radiation countermeasures. Ultimately, this strategy introduces a hypothesis-generating framework for the prioritization of potential drug candidates for future computational characterization and experimental investigation against spaceflight stressors.

Animals

Functional Annotation Routines Used by ABRF Bioinformatics Core Facilities - Observations, Comparisons, and Considerations.

The functional annotation of gene lists is a common analysis routine required for most genomics experiments, and bioinformatics core facilities must support these analyses. In contrast to methods such as the quantitation of RNA-Seq reads or differential expression analysis, our research group noted a lack of consensus in our preferred approaches to functional annotation. To investigate this observation, we selected 4 experiments that represent a range of experimental designs encountered by our cores and analyzed those data with 6 tools used by members of the Association of Biomolecular Resource Facilities (ABRF) Genomic Bioinformatics Research Group (GBIRG). To facilitate comparisons between tools, we focused on a single biological result for each experiment. These results were represented by a gene set, and we analyzed these gene sets with each tool considered in our study to map the result to the annotation categories presented by each tool. In most cases, each tool produces data that would facilitate identification of the selected biological result for each experiment. For the exceptions, Fisher's exact test parameters could be adjusted to detect the result. Because Fisher's exact test is used by many functional annotation tools, we investigated input parameters and demonstrate that, while background set size is unlikely to have a significant impact on the results, the numbers of differentially expressed genes in an annotation category and the total number of differentially expressed genes under consideration are both critical parameters that may need to be modified during analyses. In addition, we note that differences in the annotation categories tested by each tool, as well as the composition of those categories, can have a significant impact on results.

Computational Biology

Bioinformatics for the Structural Genomics of Poxviruses.

Poxviruses are large, complex viruses, and their host species are widespread across the tree of life. As a result, the bioinformatics analysis of their genomes can be complex. Here we show how a few helpful tools and strategies can be used to inform the analysis, leading to a better understanding of the structural properties of poxvirus genomes and to a more accurate quality control of, or comparison between, assembled sequences.

Poxviridae

A Comprehensive Bioinformatics Approach to Analysis of Variants: Variant Calling, Annotation, and Prioritization.

Next-Generation Sequencing (NGS), also known as high-throughput sequencing technologies, has enabled rapid and efficient sequencing of large amounts of DNA and RNA. These technologies have revolutionized the field of genomics, transcriptomics, and proteomics and have been widely used in cancer research, leading to advances in clinical diagnosis and treatment. Improvements in the NGS technologies enabled millions of fragments to be sequenced simultaneously in a time- and cost-effective manner and resulted in large amount of genomic data which require efficient analysis methods. Analysis of the genomic data requires both efficient computer resources and bioinformatics approaches. This chapter details a comprehensive computational approach and analysis steps for genomic data analysis.

Computational Biology

pSTRminer: integrated bioinformatic software for genome-wide identification and population-scale evaluation of polymorphic short tandem repeats.

Animal forensic genetics plays a critical role in criminal investigations by providing crucial evidence through domestic animal individualization and wildlife species identification. While human forensic genetics benefits from standardized short tandem repeats (STR) genotyping systems, animal forensic applications encounter significant challenges, including the limited availability of validated STR markers, the prevalence of error-prone dinucleotide STRs (di-STRs), and insufficient integration of population data. To address these challenges, we developed pSTRminer, an integrated bioinformatic tool that automates genome-wide STR mining and polymorphism evaluation. By applying pSTRminer to domestic cattle (Bos taurus), we identified 775,444 STRs de novo from the reference genome and genotyped them using whole-genome sequencing data from 60 Chinese and 111 African cattle to evaluate polymorphism across diverse genetic backgrounds. This led to the development of the cattle STR database (CSDB), comprising loci with a genotyping success rate&#x2009;&#x2265;&#x2009;40% and polymorphism information content (PIC)&#x2009;&#x2265;&#x2009;0.5. Experimental validation of 30 randomly selected tetranucleotide STRs (tetra-STRs) and 33 di-STRs via next-generation sequencing in a local Chinese cattle population (n&#x2009;=&#x2009;145) confirmed marker reliability. Although tetra-STRs had lower average polymorphism levels, they exhibited significantly lower stutter ratios (p&#x2009;<&#x2009;0.05), providing a viable path for identifying discriminative markers with fewer artifacts. Systematic screening revealed that certain tetra-STRs could surpass di-STRs in polymorphism. In conclusion, pSTRminer provides a scalable framework for developing standardized STR panels, facilitating the identification of robust and informative markers in forensic applications.

Bioinformatic software

The bioinformatics approach to identifying pathogenic variants for colorectal cancer (CRC).

Colorectal cancer (CRC) is the third most prevalent cancer globally, accounting for 9.6% of newly diagnosed cases and 9.3% of cancer-related deaths. It develops from the uncontrolled proliferation of glandular cells in the colon and rectum and is categorized into three primary types: sporadic, hereditary, and colitis-associated. While genetic susceptibility is a key factor in CRC pathogenesis, identifying high-impact pathogenic variants remains a significant challenge. This study integrates bioinformatics and population genetics approaches to identify CRC-associated single-nucleotide polymorphisms (SNPs) with potential clinical significance. CRC-associated SNPs were extracted from the Genome-Wide Association Studies (GWAS) Catalog, functionally annotated via HaploReg, and validated via Ensembl. In addition, expression quantitative trait locus (eQTL) data from the GTEx database were used to assess the effects of these variants on gene expression across human tissues. Our analysis identified three high-priority SNPs (rs9379084, rs3184504, and rs11557154) associated with the RREB1, ATXN2, SH2B3, and DCAF12 genes, which exhibited marked allele frequency differences among populations. These findings suggest potential biomarkers for CRC risk assessment and highlight the importance of genetic screening across diverse populations.

Bioinformatics

Analysis of differentially expressed genes in schizophrenia based on bioinformatics and corresponding mRNA expression levels.

OBJECTIVE: This study aimed to use bioinformatics analysis to identify differentially expressed genes (DEGs) involved in the pathogenesis of schizophrenia and validate their mRNA expression levels through real-time quantitative PCR (qPCR). MATERIAL/METHODS: Datasets from the publicly available Gene Expression Omnibus (GEO) database were analyzed using R software to identify DEGs. Functional enrichment analyses, including Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways, were conducted. A protein-protein interaction (PPI) network was constructed using Cytoscape software to identify key genes with notable expression changes. The expression levels of these key genes were subsequently validated in schizophrenia patients using qPCR to assess potential susceptibility genes. RESULTS: In total, 813 DEGs were identified, with six key genes highlighted through GO analysis and PPI network screening. Among these, HDAC1, UBA52, and FYN demonstrated statistically significant differences in mRNA expression between schizophrenia patients and healthy controls (P&#xa0;<&#xa0;0.05). CONCLUSIONS: This study identified several DEGs potentially linked to the pathogenesis of schizophrenia, suggesting that HDAC1, UBA52, and FYN could serve as candidate susceptibility genes and diagnostic biomarkers. These findings provide new insights and directions for future schizophrenia research.

Humans

Integrated bioinformatics analysis reveals cross-talking hub genes and therapeutic agents between sepsis and acute myocardial infarction.

BACKGROUND: Sepsis and acute myocardial infarction (AMI) are two significant diseases that may share overlapping etiological mechanisms. This study aims to systematically identify core genes common to both conditions and to explore their potential as therapeutic targets and drug candidates through an integrative analysis of clinical data and bioinformatics. METHODS: The AMI dataset was obtained from the GEO database, and RNA sequencing data were collected from blood samples of patients with sepsis at our hospital. Common genes were identified using differential expression gene analysis (DEG) and weighted gene co-expression network analysis (WGCNA). Functional enrichment analyses, including Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis, were performed. A protein-protein interaction (PPI) network was constructed, and hub genes were identified using the MCC/Degree algorithm. Diagnostic value was assessed via receiver operating characteristic curve analysis. Immune infiltration patterns, single-cell sequencing data, and molecular docking simulations were employed to evaluate immune relevance and identify potential therapeutic compounds. RESULTS: A total of 417 genes were identified between sepsis and AMI, with enrichment analysis revealing significant involvement in inflammatory responses. Three hub genes-JAK2, MYD88, and TIMP1-were selected for further investigation. ROC curves confirmed their strong diagnostic performance for both diseases. Immune infiltration analysis showed that these core genes were significantly correlated with the infiltration levels of various immune cell types. Molecular docking indicated that quercetin exhibited stable binding affinity with the proteins encoded by these genes. qPCR validation further confirmed the upregulation of these three genes, supporting the anti-inflammatory effects of quercetin as a potential targeted therapy. CONCLUSION: JAK2, MYD88, and TIMP1 were identified as shared core genes in sepsis and AMI. These genes not only serve as potential diagnostic biomarkers but also offer novel targets for developing common therapeutic strategies for both conditions. Furthermore, quercetin emerges as a promising candidate for targeted treatment.

Humans

Phage bioinformatics tools: a review of computational approaches for bacteriophage research.

Rising clinical interest in phage therapy and the exponential growth of metagenomic sequence catalogues have driven a rapid expansion of bacteriophage bioinformatics. More than 80 dedicated tools, mostly published since 2020, now span identification, assembly, annotation, taxonomy, lifestyle prediction, defence-system detection, and host prediction. Aimed at experienced practitioners and developers, this review synthesizes the field through the lens of three successive computational paradigms: sequence homology, bounded by database completeness; machine learning, constrained by labelled training data; and foundation models, which now achieve Matthews correlation coefficients above 0.95 in identification tasks and, through structure-informed prediction, raise functional annotation to over half of phage genes. Furthermore, we map the upstream components, namely, gene callers, homology engines, protein language models, and structural search tools, that underpin most downstream pipelines, exposing shared infrastructure and ecosystem-level fragility when dependencies change. To translate this into practice, we propose web-based and command-line reference workflows calibrated to user expertise and sample types. Finally, we set an agenda for the next wave of tool development. Roughly half of phage genes still resist functional annotation despite structural methods; no broadly generalizable strain-level host predictor exists for phage therapy; varying true-positive rates (0%-97%) underscore the absence of standardized community benchmarks analogous to Critical Assessment of Structure Prediction or Critical Assessment of Metagenome Interpretation. As generative genome models begin designing synthetic phages, progress will depend less on producing standalone tools than on rigorous evaluation, interoperable infrastructure, and clinically meaningful prediction targets.

Computational Biology

Comparative performance of portable DNA extraction protocols and bioinformatics workflows for rapid detection of gram-negative bacteria and antimicrobial resistance using Oxford Nanopore sequencing.

Oxford Nanopore Technology (ONT) enables rapid, portable pathogen identification and antimicrobial resistance (AMR) detection, but the reliability of downstream genomic analyses is highly dependent on DNA extraction quality, particularly in resource-limited settings. This study comparatively evaluated four portable bacterial DNA extraction protocols derived from three commercial kits to determine their impact on nanopore sequencing performance, bioinformatics workflow completion, and field deployability. Six gram-negative bacterial isolates (Escherichia coli, n = 4; Pseudomonas sp., n = 1; and Salmonella sp., n = 1) were processed using four extraction protocols: SwiftX DNA, SwiftX DNA with proteinase K (ProtK), SwiftX ParaBact, and NucleoSpin Microbial. Twenty-four resulting DNA extracts were sequenced on a single multiplexed MinION R10.4.1 flow cell. Sequencing data were analyzed using validated Galaxy-based generic and species-specific pipelines. Workflow completion was defined as successful progression through quality control, assembly, virulence, plasmid, and AMR detection modules. DNA purity varied substantially by extraction protocol and was strongly associated with successful workflow completion (Kruskal-Wallis, P = 0.0006). Accordingly, NucleoSpin Microbial achieved 100% workflow completion, and SwiftX ParaBact achieved 83%, while both SwiftX DNA-based protocols failed to complete full workflows. Importantly, key AMR genes required to classify isolates as multidrug-resistant were consistently detected using both NucleoSpin Microbial and SwiftX ParaBact extractions. However, NucleoSpin Microbial assemblies showed significantly higher contiguity and enabled a broader, more complete detection of virulence factors, pathogenicity islands, plasmid replicons, and accessory AMR genes, reflecting enhanced genomic resolution.IMPORTANCERapid whole-genome sequencing is increasingly used to detect antimicrobial resistance and guide public health responses, but its reliability depends strongly on how bacterial DNA is extracted. In this study, we have shown that DNA extraction method choice has a major impact on Oxford Nanopore sequencing performance across clinically relevant gram-negative bacteria. While silica column-based extraction maximized genomic completeness and analytical depth, paramagnetic bead-based reverse purification offered superior portability with sufficient resolution for frontline AMR surveillance. These findings highlight a practical trade-off between field deployability and high-resolution genomic characterization in low-resource settings.

DNA extraction