Search PubMedSearch

SEARCH · Search PubMed

Results for “Health Bioinformatics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Genomic Medicine Sweden: Advancing precision medicine at the national level.

High-throughput sequencing has transformed clinical diagnostics of rare diseases (RD), cancer and infectious diseases by enabling the identification of disease-causing genetic alterations and facilitating individualised treatment and care. In response to these advances, Genomic Medicine Sweden (GMS) was established in 2017 as a national collaborative effort to accelerate implementation of genomics-based precision medicine within Sweden's regionally organized, publicly funded healthcare system. GMS brings together the seven university healthcare regions and their associated medical faculties, in collaboration with healthcare regions across Sweden, Science for Life Laboratory, patient organizations, industry and governmental agencies. Activities are coordinated through national disease-specific expert groups, supported by cross-cutting functions in bioinformatics, health economics, ethics, education and patient engagement. At the operational level, seven Genomic Medicine Centres, embedded at university hospitals, develop and deliver harmonised genomic diagnostics nationwide. The National Genomics Platform provides secure infrastructure for large-scale data storage, analysis, and national and international data sharing. Following initial project-based funding, GMS now receives long-term governmental support. This review describes the national implementation of genomic-based precision diagnostics, discusses challenges and lessons learnt, and highlights key milestones across disease areas, including whole-genome sequencing in RD and paediatric cancer, comprehensive genomic profiling of haematological malignancies and solid tumours, pathogen genomics in microbiology, pharmacogenomic testing and emerging applications of polygenic risk scores in complex diseases. Collectively, these efforts have contributed to more than 500,000 genomic tests being performed within Swedish healthcare between 2017 and 2025. Finally, we outline future diagnostic needs and priority areas to ensure sustainable, scalable and equitable access to precision medicine.

Precision Medicine

Integrated experimental and bioinformatics analysis reveals ECM-integrin and redox signaling associated with PMMA/NiO nanocomposites for craniofacial applications.

BACKGROUND: Poly(methyl methacrylate) (PMMA) is widely used in dental and craniofacial applications; however, its clinical performance is limited by poor surface wettability, moderate mechanical strength, and restricted biological activity. Integrating nanomaterial engineering with computational biology offers an opportunity to better understand biomaterial-cell interactions and support the rational design of functional biomaterials. METHODS: Nickel oxide (NiO) nanoparticles were synthesized via chemical precipitation and incorporated into PMMA to fabricate nanocomposites. Physicochemical characterization included contact angle measurements, Fourier-transform infrared spectroscopy (FTIR), scanning electron microscopy (SEM), energy-dispersive X-ray spectroscopy (EDX), and Vickers hardness testing. Biocompatibility was evaluated using zebrafish embryo developmental assays. To explore biological processes potentially associated with biomaterial-cell interactions, bioinformatics analyses including Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), and STRING protein-protein interaction (PPI) network analyses were performed. RESULTS: Incorporation of NiO nanoparticles improved the surface and mechanical properties of PMMA, reducing the contact angle from 105.35° to 90.46° and increasing Vickers hardness compared with unmodified PMMA. Structural and morphological analyses confirmed successful synthesis and homogeneous nanoparticle incorporation. Zebrafish embryo studies demonstrated minimal developmental toxicity, supporting the biocompatibility of the nanocomposite. Bioinformatics analyses identified significant enrichment of pathways related to extracellular matrix organization, cell adhesion, focal adhesion, PI3K-Akt signaling, and oxidative stress regulation. Protein-protein interaction analysis revealed highly interconnected networks associated with ECM-integrin signaling and redox homeostasis, highlighting biological processes potentially associated with biomaterial-cell communication. CONCLUSIONS: PMMA/NiO nanocomposites exhibited improved physicochemical performance and favorable biocompatibility characteristics. The integration of experimental characterization with bioinformatics and network-based analyses provides a systems-level perspective on biomaterial-associated cellular processes and identifies ECM-integrin signaling and oxidative stress-related pathways as candidate biological processes for future experimental validation. These findings support the continued development of PMMA/NiO nanocomposites for oral and craniofacial biomedical applications.

Nanocomposites

PathoSeq-QC: a decision support bioinformatics workflow for robust genomic surveillance.

MOTIVATION: Recommendations on the use of genomics for pathogens surveillance are evidence that high-throughput genomic sequencing plays a key role to fight global health threats. Coupled with bioinformatics and other data types (e.g., epidemiological information), genomics is used to obtain knowledge on health pathogenic threats and insights on their evolution, to monitor pathogens spread, and to evaluate the effectiveness of countermeasures. From a decision-making policy perspective, it is essential to ensure the entire process's quality before relying on analysis results as evidence. Available workflows usually offer quality assessment tools that are primarily focused on the quality of raw NGS reads but often struggle to keep pace with new technologies and threats, and fail to provide a robust consensus on results, necessitating manual evaluation of multiple tool outputs. RESULTS: We present PathoSeq-QC, a bioinformatics decision support workflow developed to improve the trustworthiness of genomic surveillance analyses and conclusions. Designed for SARS-CoV-2, it is suitable for any viral threat. In the specific case of SARS-CoV-2, PathoSeq-QC: (i) evaluates the quality of the raw data; (ii) assesses whether the analysed sample is composed by single or multiple lineages; (iii) produces robust variant calling results via multi-tool comparison; (iv) reports whether the produced data are in support of a recombinant virus, a novel or an already known lineage. The tool is modular, which will allow easy functionalities extension. AVAILABILITY AND IMPLEMENTATION: PathoSeq-QC is a command-line tool written in Python and R. The code is available at https://code.europa.eu/dighealth/pathoseq-qc.

Genomics

Integrated molecular, epidemiological, and bioinformatics perspectives on the Mpox virus: Implications for surveillance and Global Health preparedness.

Mpox has re-emerged as a significant global zoonotic threat, driven mainly by two large waves the 2022 worldwide Clade IIb outbreak and the 2024 Clade Ib epidemic in Central Africa. This review examines the challenges of interpreting this evolving virus from molecular, epidemiological, and bioinformatics perspectives, with a focus on global health workforce preparedness. Clade IIb largely moved through sexual transmission across countries, but Clade Ib has appeared in a wider population-women, children, and individuals infected through household spread without any sexual contact. Early case series suggest that Clade Ib may cause a more severe disease burden, but more research is needed to directly compare severity and fatality rates with Clade IIb due to the limited number of current studies. The review examines the virus's strategies for evading the host's immune defenses throughout its ∼197 kbp genome, including how it disrupts interferon signaling and creates decoy receptors. This review summarizes the clinical findings of PALM007 and STOMP, noting that neither trial achieved its main efficacy endpoint making routine tecovirimat use less compelling-while leaving open whether it helps particular high-risk groups. A further point is that immunity from the MVA-BN vaccine wanes with time, leading to the growing adoption of booster vaccinations. In conclusion, the review calls for a One Health approach pairing genomic tracking with ecological intelligence and including wastewater surveillance to fill existing gaps in knowledge and enhance the global handling of new orthopoxvirus threats.

Animals

Detection and characterization of antiviral-resistant viruses during the influenza season of 2024-25.

UNLABELLED: During the high severity season of 2024-25, CDC with public health partners sequenced and analyzed genomes of >10,000 influenza viruses for antiviral resistance markers. Available sequence-flagged and representative viruses were tested with antivirals using in vitro assays. In the US, three oseltamivir-resistant A(H3N2) viruses had treatment-emergent neuraminidase (NA) mutations, either E119V or R292K. Oseltamivir-resistant A(H1N1)pdm09 viruses with NA-H275Y were detected in 15 states, albeit at a low frequency (0.53%). They belonged to several phylogenetic groups, with hemagglutinin (HA) subclade D.3.1 combined with either NA subclade D.1 or D.2 being most common. Based on shared sequence data, nearly all H275Y viruses from Australia, Canada, and Chile also belonged to these HA and NA subclades. Conversely, most H275Y viruses (68/81) from China belonged to HA subclade C.1.9 and NA subclade D and shared the permissive mutation R257K. Influenza polymerase acidic (PA) mutations conferring 4- to 92-fold decreased baloxavir susceptibility were detected in nine influenza A viruses. Viruses with PA-I38T showed mild attenuation of replicative fitness in three cell lines. Based on available data, NA-H275Y and PA-I38T viruses were collected from patients with no exposure to antivirals. Baseline susceptibility to all US-approved influenza antivirals remained largely unchanged compared to previous seasons. All swine-origin viruses detected in the US had adamantane resistance-conferring marker, M2-S31N, but remained susceptible to other approved antivirals. Monitoring antiviral susceptibility has substantially improved with increased sequencing capacities and bioinformatic support at public health laboratories. Information gained through influenza surveillance has been used to guide recommendations on antiviral use. IMPORTANCE: Circulation of influenza viruses with reduced susceptibility to antivirals can diminish the usefulness of medications prescribed for influenza. This study informs on the prevalence of drug-resistant influenza viruses in the US during the high severity season of 2024-25. It provides information on susceptibility profile to all approved antiviral medications and on replicative fitness of representative drug-resistant viruses. Most drug-resistant viruses were collected from patients who were not exposed to antivirals indicating their ability to transmit from human to human. Whole-genome sequence (WGS)-based analysis is the cornerstone for surveillance, and numerous laboratories have been utilizing this approach. However, CDC laboratory is the only laboratory in the US conducting phenotypic testing of circulating viruses needed to confirm the outcomes of sequence-based analysis and to identify new molecular markers of resistance. Data gathered through virologic surveillance give much-needed information on drug susceptibility of influenza viruses which are used to guide recommendations on antiviral use.

Antiviral Agents

Toward a unified approach: Considerations for bioinformatic and sequencing activities & data in wastewater surveillance of biologic public health threats.

Genomic technologies such as PCR and next-generation sequencing (NGS) have greatly advanced public health surveillance, especially during COVID-19, by enabling detailed tracking of pathogen spread, origins, and variants. While PCR is vital for targeted detection, falling NGS costs have made large-scale, high-throughput sequencing more feasible, supporting broader pathogen monitoring-including the detection of vaccine escape variants and new strains. Applying NGS to wastewater offers valuable population-level insights but faces challenges such as variable sample complexity, the need for skilled staff, suitable platforms, and robust IT infrastructure. Although there are currently a lot of efforts towards defining guidelines for sampling, analysis, and integrating wastewater data into public health policy, such as the recently published International Cookbook for Wastewater Practitioners, they often lack universal applicability, emphasizing the analytical approaches in favour of the NGS-based approaches. However, standardising protocols for sampling, sequencing, and analysis is crucial to ensure reliable, comparable data across surveillance systems worldwide. Pilot studies and continuous refinement are recommended to overcome implementation hurdles and fully realise the benefits of NGS in wastewater surveillance. This work attempts to outline these challenges and opportunities across the entire wastewater surveillance workflow, from data generation to reporting, and provide some concrete suggestions and considerations across the spectrum of activities. We further highlight that the infrastructure, funding and government-policy context in which surveillance operates acts as an enabling condition for these activities, and that technical standardisation alone is unlikely to deliver durable, comparable surveillance in its absence.

considerations

Robust replication of associations across patient-mediated and provider-sourced EHR data in the All of Us research program.

The All of Us Research Program is assembling a nationwide cohort with electronic health record (EHR) resources through two complementary pathways: healthcare provider organization (HPO)-sourced EHRs and patient-mediated EHR (PME) contributed through patient portal linkages. The comparative research utility of these two data sources has not been systematically evaluated. Here, we compared PME and HPO EHRs with respect to disease prevalence, phenotype-phenotype associations, and replication of established genotype-phenotype associations using data from 19,703 PME and 373,887 HPO participants. We benchmarked disease prevalence against national estimates, conducted phenome-wide association studies for 10 commonly studied diseases, and tested replication of more than 5000 established genotype-phenotype associations across multiple ancestral groups. Disease prevalence was consistently lower in PME than in HPO, although prevalence of most diseases in both cohorts exceeded national estimates. Both data sources reproduced known phenotype-phenotype associations and showed moderate-to-strong concordance in effect sizes across the phenome. The overall genotype-phenotype replication rate was 49.1% (5399/10,999) in HPO and 5.9% (381/6482) in PME across ancestral groups, with effect sizes strongly correlated among well-powered associations (R&#x2009;=&#x2009;0.84, P&#x2009;<&#x2009;0.001). To disentangle the impact of sample size from data quality, we performed 1:1 propensity score matching. After matching, the replication gap in genotype-phenotype associations narrowed from 8.3-fold to 1.3-fold, with equivalent replication rates among adequately powered associations and strongly concordant effect sizes; comorbidity patterns were also consistent across all 10 diseases tested. These findings demonstrate that both data sources are valuable for clinical and genomic research and can inform other cohorts integrating provider-derived and patient-mediated EHRs.

Computational biology and bioinformatics

The cost and cost trajectory of genome sequencing and bioinformatics analysis for Indigenous children with suspected rare diseases.

PURPOSE: Indigenous peoples are underrepresented in reference genome libraries. Consequently, rare disease diagnosis may require bespoke bioinformatics analyses of genome sequences. Establishing diagnostic cost is crucial to support policy development for equitable diagnosis of rare diseases. We estimated the cost and cost trajectory of diagnostic genome sequencing and bioinformatics for Indigenous participants with suspected rare diseases. METHODS: We conducted a microcosting study of Indigenous children and their families receiving genome sequencing through Canada's Silent Genomes Project. Invoice data informed the costs of genome sequencing. We conducted a time-and-motion study for bioinformatics analyses, including labor, computing, and data storage costs. RESULTS: With standard bioinformatics, costs ranged from C$3645 (SD: 455) for singletons to C$7402 (SD: 566) for trios. With advanced, bespoke bioinformatics, costs ranged from C$5344 (SD: 634) for singletons to C$9760 (SD: 822) for trios. Genome sequencing was a primary cost driver; however, sequencing costs decreased by 61% over 4 years. Bioinformatics costs ranged from 21.3% to 58.3% of the total costs. The time required for bioinformatics ranged from 71 hours to 215 hours for standard and advanced analyses, respectively. CONCLUSION: Genome sequencing costs decreased over time. Bioinformatics is a significant cost driver, particularly for bespoke analyses arising from nonrepresentative reference libraries.

Humans

Leveraging bioinformatics approaches for drug repositioning in space radiation protection.

The health effects of space radiation, primarily Galactic Cosmic Rays (GCRs), on humans remain largely unknown, with potential cardiovascular consequences posing a significant threat to astronauts on long-duration spaceflight missions. Currently, there are no established pharmacological countermeasures for GCR exposure. Drug repositioning offers a promising strategy to accelerate pharmaceutical research in space medicine. This study leverages existing bioinformatics techniques to identify and prioritize potential drug candidates associated with proteomic perturbations following simulated GCR exposure using previously published murine cardiac proteomic data. A protein-protein interaction (PPI) network was constructed using the top differentially expressed proteins (DEPs) from murine heart tissue following exposure to 5-ion GCRs as seed nodes, focusing on experimentally supported interactions. Network topology, Markov clustering, and functional enrichment analyses were used to characterize biologically relevant proteins and pathways. Drug-protein interactions were predicted using Drugst.One and mapped to PPI clusters of interest to identify candidate drugs. Selected drug-macromolecule interactions were further explored using CB-Dock2 molecular docking and short-duration molecular dynamics simulations as hypothesis-generating structural assessments. Analysis of a key PPI network cluster consisting of several ATP synthase proteins identified 23 unique drug candidates. These analyses demonstrate a systematic approach for leveraging bioinformatics techniques to identify candidate molecular targets and generate pharmacological hypotheses in the context of space radiation countermeasures. Ultimately, this strategy introduces a hypothesis-generating framework for the prioritization of potential drug candidates for future computational characterization and experimental investigation against spaceflight stressors.

Animals

Identification of NLRP3 and TIPE2 as asthma biomarkers via integrative bioinformatics and Mendelian randomization.

Asthma is a chronic inflammatory airway disease imposing a substantial global health burden. NLRP3 is an immune sensor involved in infection and cellular stress responses. Recent studies suggest that NLRP3 may be involved in the pathogenesis of asthma. We hypothesized that genetic variation in NLRP3 may contribute to asthma susceptibility. However, the causal relationship between NLRP3 and asthma still remains unclear. In this study, bioinformatics analysis using asthma data and R software was performed to identify NLRP3-related genes. We performed weighted gene co-expression network analysis to identify co-expressed genes, resulting in 12 candidate genes. Kyoto Encyclopedia of Genes and Genomes and Gene Ontology enrichment analyses were used to identify the functions of these candidate genes, revealing their involvement in cellular metabolism. Mendelian randomization analysis of the 12 candidate genes identified 2 biomarkers: NLRP3 and TNFAIP8L2 (TIPE2). We validated their diagnostic value for asthma using the GSE182503 dataset, with area under the curve values of 0.83 and 0.66 for NLRP3 and TIPE2, respectively. This project discusses how NLRP3 promotes asthma pathogenesis, whereas TIPE2 may alleviate it, and explores the potential interplay between them. NLRP3 and TIPE2 may serve as diagnostic biomarkers for asthma: NLRP3 may promote, whereas TIPE2 may alleviate asthma development. Both genes represent potential diagnostic biomarkers and therapeutic targets that warrant further functional investigation.

Asthma

Human Wings Apart-Like Protein as a Serum Diagnostic Biomarker in Cervical Cancer: An Integrative Bioinformatics Analysis with Serum Validation.

Cervical cancer remains a major threat to women's health worldwide, and reliable serum biomarkers for early detection and therapeutic stratification remain limited. Human wings-apart-like (hWAPL) protein has been implicated in cervical carcinogenesis, but its diagnostic and clinical value has not been fully elucidated. To address this gap, this study integrated public multi-omics datasets, including The Cancer Genome Atlas, GEPIA2, the Human Protein Atlas, and single-cell transcriptomic data, to characterize hWAPL expression, clinicopathological associations, immune infiltration, co-expression networks, post-translational modifications, and drug sensitivity predictions. These findings were evaluated in an independent single-center serum cohort comprising 89 patients with histologically confirmed cervical squamous cell carcinoma and 89 healthy female controls. Serum hWAPL and squamous cell carcinoma antigen (SCC) levels were measured, and diagnostic performance was assessed by receiver operating characteristic curve analysis. In silico, hWAPL was broadly upregulated across multiple malignancies, particularly cervical cancer, enriched in malignant epithelial cells and monocytes/macrophages, and associated with shorter progression-free interval, predicted reduced sensitivity to cisplatin, paclitaxel, and 5-fluorouracil, and predicted sensitivity to MCL-1 and Wee1 inhibitors. In the serum cohort, hWAPL levels were significantly higher in patients than controls and discriminated cervical cancer with an area under the curve of 0.961, exceeding SCC alone. Combining hWAPL with SCC further improved diagnostic performance (area under the curve, 0.974; sensitivity, 93.3%; specificity, 95.5%). These findings suggest that serum hWAPL is a potential novel diagnostic biomarker for cervical squamous cell carcinoma whose performance is enhanced by SCC, whereas the observed associations with chemoresistance and immune microenvironment remodeling are hypothesis-generating and require experimental confirmation.

Humans

VirDetector: a bioinformatic pipeline for virus surveillance using nanopore sequencing.

SUMMARY: Virus surveillance programmes are designed to counter the growing threat of viral outbreaks to human health. Nanopore sequencing, in particular, has proven to be suitable for this purpose, as it is readily available and provides rapid results. However, as special bioinformatic programs are required to extract the relevant information from the sequencing data, applications are needed that allow users without extensive bioinformatics knowledge to carry out the relevant analysis steps. We present VirDetector, a bioinformatic pipeline for virus surveillance using nanopore sequencing. The pipeline automatically installs all required programs and databases and allows all its steps to be executed with a single console command. After preprocessing the samples, including the possibility for basecalling, the pipeline classifies each sample taxonomically and reconstructs the viral consensus genomes, which are then used in phylogenetic analyses. This streamlined workflow provides a user-friendly and efficient solution for monitoring viral pathogens. AVAILABILITY AND IMPLEMENTATION: VirDetector is freely available at https://github.com/NLKaiser/VirDetector and https://zenodo.org/records/14637302 (10.5281/zenodo.14637302).

Nanopore Sequencing

Longitudinal characterization of mixed-genotype SARS-CoV-2 infections in a military cohort reveals compartmentalized viral populations.

UNLABELLED: Mixed-genotype severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infections are a concern due to the potential generation of novel recombinants that give rise to new variants. To better understand intra-host viral dynamics, we analyzed specimens from 24 participants from the U.S. Military Health System's Epidemiology, Immunology, and Clinical Characteristics of Emerging Infectious Diseases with Pandemic Potential COVID-19 cohort with suspected mixed-genotype SARS-CoV-2 infections. From an initial 24 suspected cases, we confirmed 17 as genuine coinfections and graded them by evidence: 7 were "strong"; 4 were "moderate"; 6 were "weak"; and 7 were deemed unlikely to be true mixed-genotype infections. Access to swabs from multiple body sites across the course of infection allowed us to observe compartmentalization and shifts in variant dominance that would have been missed by a single-timepoint analysis, as well as one recombinant Omicron BA.1/BA.2 genome. By using an evidence-based bioinformatic framework to assess sequencing data from well-characterized clinical cases, we distinguished genuine coinfections from bioinformatic artifacts. Our findings emphasize the importance of both extensive specimen collection and careful bioinformatic approaches in ascertaining dual genotype infections. IMPORTANCE: Novel recombinants of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) arise from coinfections with different lineages, but mixed infections are not screened for despite risk to public health, and most surveillance relies on single swabs. We analyzed a longitudinal data set with specimens from multiple body sites, providing an opportunity to assess intra-host dynamics. To distinguish true coinfection from bioinformatic artifacts with confidence, we applied a framework that grades evidence for mixed genotypes by incorporating lineage and clade with manually validated variant calls. This allowed investigation beyond abundance levels of mixed genotypes within a single specimen, including observations of compartmentalization and a recombinant virus. This work enables further study of evolutionary, immunological, and clinical implications of mixed SARS-CoV-2 genotypes. Detecting dual-genotype infections and discriminating between true dual-genotype infection vs potential bioinformatics-based artifacts support public health and military readiness. These efforts provide evidence to bolster decision-making in molecular epidemiological studies to track transmission and for the choice of effective countermeasures.

SARS-CoV-2

Global Genomic Surveillance.

Global genomic surveillance has emerged as a foundational pillar of public health in the twenty-first century, enabling real-time tracking of pathogen evolution and informing outbreak response. This chapter examines the strategic architecture of global genomic surveillance, focusing on its application to arboviruses such as chikungunya virus (CHIKV). It explores the integration of genomic data with epidemiological, clinical, and environmental information within a One Health framework, while addressing critical challenges in governance, equity, and interoperability. The discussion covers the entire genomic surveillance workflow, from sample collection and sequencing to bioinformatic analysis and phylogenetic inference, and highlights the transformative role of artificial intelligence (AI) in predictive surveillance. By analyzing global initiatives, operational barriers, and emerging technologies, this chapter underscores the necessity of sustainable, equitable, and interoperable genomic systems to proactively address current and future infectious disease threats.

Humans

Large language models in bioinformatics: a comprehensive survey.

The emergence of foundation models with trillion-level parameters has redefined the landscape of artificial intelligence. Various fields are developing their own large-scale models, which can solve many problems within the field and improve work efficiency. Biological large-scale models are a cross-disciplinary research field that combines mathematics, computer science, and biology, aiming to simulate and understand the structure, function, and dynamic changes of biological systems through the establishment of complex computational models. This field covers multiple levels such as biological pathways, population dynamics, protein folding, etc., providing us with tools for deep exploration of the mysteries of life and applications in medicine, ecology, and other fields. This article reviews the background and research status of biological large-scale models, and discusses future directions. Large language models (LLMs) and other large-scale foundation models have rapidly advanced in recent years, enabling powerful representation learning and generation across text, sequences, and multimodal data. In bioinformatics and biomedicine, these models are increasingly used to analyze genomic sequences, infer protein properties and structures, support drug discovery, and integrate heterogeneous biomedical evidence. This survey reviews the basic principles of LLMs and summarizes representative applications in (i) gene and genome sequence analysis, (ii) protein structure and function prediction, and (iii) drug design, including virtual screening and personalized medicine. We also discuss emerging multi-model modeling approaches, as well as key challenges such as data quality and privacy, interpretability, generalization to new organisms and tasks, and responsible deployment in health-related settings. Finally, we outline future directions for developing reliable, scalable, and explainable bioinformatics foundation models.

bioinformatics

Bioactive macromolecules in LAB-fermented cereals: Mechanisms of formation, functional properties, and health benefits.

Cereal and pseudo-cereal based fermented food products represent a substantial segment of global diet, nutrition as well as food security. Fermentation, especially by Lactic Acid Bacteria (LAB) increases the nutritional and functional values of foods by increasing palatability, bioavailability and minimizing antinutritional factors. LAB plays a pivotal role in synthesizing bioactive peptides, vitamins, minerals and reducing anti-nutrients parallelly. This review elucidates the mechanism through which LAB revamping nutritional macromolecules, such as peptides and polysaccharides, during fermentation and their role in the development of traditional as well as modern fermented foods. Additionally, these fermented foods have been associated with several health benefits. Recent advancement in biotechnology such as genome sequencing, functional genomics, and AI-assisted bioinformatics, have significantly enhanced our understanding of the diversity of LAB, the metabolism, and adaptation mechanisms. The combination of in silico and experimental methods has enabled the development of novel food enzymes as well as highly precise fermentation processes. Together with new innovations, growing demands for quality, consistency, safety as well as health benefits point out the significance of continued research. More studies employing both conventional and modern methods are necessary to explore these food groups completely and achieve better food quality, increased nutrition, more health benefits and comprehensive socioeconomic advantages.

Bioactive macromolecules

Expanding vaginal microbiome pangenomes via a custom MIDAS database reveals Lactobacillus crispatus accessory genes associated with cervical dysplasia.

The vaginal microbiome plays a central role in reproductive health. Vaginal microbiome dysbiosis is associated with many adverse reproductive health outcomes, but most studies have focused on associations at the species level. The potential contribution of intraspecies microbial variation, especially gene content differences across bacterial strains, remains underexplored in reproductive health contexts. The Metagenomic Intra-Species Diversity Analysis (MIDAS) framework enables such analyses, but depends on comprehensive reference databases. We constructed a MIDAS-compatible pangenome database from over 18,000 genomes in the Vaginal Microbiome Genome Collection (VMGC). Compared to the Genome Taxonomy Database (GTDB)-derived reference, the VMGC-derived database expanded the pangenomes of prevalent vaginal species, better capturing vaginal-specific intraspecies diversity. Applying this database to vaginal samples from a cervical dysplasia cohort, we identified 13 Lactobacillus crispatus accessory genes significantly associated with cervical dysplasia, including a HicAB toxin-antitoxin system, three transcriptional regulators, and three phage-derived genes. These findings highlight the utility of body site-specific reference resources and shotgun metagenomic sequencing for uncovering intraspecies microbial variation relevant to reproductive health.IMPORTANCEThe vaginal microbiome plays a critical role in reproductive health, and different bacteria from the same species can carry different genes that influence how the strains interact with the host and other microbes. These strain-level differences are often overlooked when microbiomes are analyzed only at the species level. Existing genomic reference databases are heavily biased toward gut and environmental bacteria, leaving the genetic diversity of vaginal microbes understudied. We built a specialized reference database from over 18,000 vaginal bacterial genomes that better reflects this diversity. We then applied this resource to quantify gene-level variation in vaginal samples from a cervical dysplasia cohort. Focusing on Lactobacillus crispatus, a prevalent and often beneficial vaginal species, we identified 13 genes that were more common in women with cervical dysplasia than in controls. This work demonstrates that body site-specific genomic resources are essential for uncovering strain-level bacterial differences relevant to reproductive health.

Lactobacillus crispatus

Integrated Bioinformatics Analysis Revealing that the NSDHL Gene Might Be Associated with the Progression of Western HFD/SW-Induced Hepatocellular Carcinoma.

BACKGROUND AND OBJECTIVE: Hepatocellular carcinoma (HCC) remains a significant global health concern. However, the etiology and pathogenesis of HCC have yet to be fully elucidated. Previous studies have indicated a close association between obesity and the occurrence and progression of HCC. The objective of this study was to employ bioinformatics strategies in order to explore key genes associated with the clinical diagnosis and prognosis of HCC induced by a Western high-fat diet and sugar water (HFD/SW). MATERIALS AND METHODS: We obtained the expression profile chip data GSE197884 from the Gene Expression Omnibus (GEO) database. Subsequently, &#x201c;DESeq&#x201d; and &#x201c;Limma&#x201d; R packages were employed to identify differentially expressed genes (DEGs) while constructing a co-expressed gene network using weighted gene co-expression analysis (WGCNA). Functional enrichment analyses were then carried out, followed by the construction of a protein-protein interaction (PPI) network to uncover core genes. The core genes were confirmed through data retrieved from The Cancer Genome Atlas (TCGA) database in order to determine their status as hub genes. Finally, survival and tumor immune infiltration analyses were performed to unveil the prognostic significance of these hub genes. RESULTS: In total, 126 intersection targets were retrieved through the Venn diagram. Gene ontology (GO) enrichment and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses revealed that the DEGs were primarily related to the proliferation and apoptosis of HCC cells, the digestion and metabolism of liver cells, the HCC tumor microenvironment, and immune response. The PPI network analysis identified 11 core targets, among which seven hub genes, including NSDHL, MVK, SQLW, GCAT, ALAS2, GLDC, and AGXT, were obtained after TCGA database validation. Furthermore, it was found that NSDHL was closely associated with the clinical diagnosis and prognosis of HCC induced by HFD/SW and also affected the cellular immune infiltration in the HCC tumor microenvironment. CONCLUSION: The present study demonstrated a significantly elevated expression of NSDHL in HCC tissues, suggesting its potential as a specific biomarker for precise clinical diagnosis and prognosis assessment of HCC induced by HFD/SW.

Computational Biology