Search PubMedSearch

SEARCH · Search PubMed

Results for “virtual screening”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Genomic mapping of diabetic kidney disease biomarkers and identification of potential inhibitors through virtual screening.

BACKGROUND: Diabetic kidney disease (DKD) is a common and serious complication of diabetes mellitus, marked by a multifactorial pathogenesis and the absence of sensitive diagnostic biomarkers. Identifying novel molecular targets and therapeutic options is essential to improve early diagnosis and treatment outcomes. METHODS: To uncover potential biomarkers and therapeutic candidates, we performed an integrated genomic analysis using microarray and RNA-seq datasets from the Gene Expression Omnibus (GEO) and Sequence Read Archive (SRA) databases. Differentially expressed genes (DEGs) were identified and subjected to protein-protein interaction (PPI) network analysis. Key genes were further explored through virtual screening of an FDA-approved compound library using molecular docking techniques. Drug-likeness was assessed via Lipinski's rule of five. RESULTS: A total of 40 DEGs were identified, among which ISCU (downregulated; involved in iron-sulfur cluster biogenesis) and AP1S2 (upregulated; associated with vesicular trafficking) emerged as potential biomarkers. PPI analysis revealed their involvement in critical DKD-related pathways, such as extracellular matrix remodeling and oxidative stress. Virtual screening identified six FDA-approved compounds with high binding affinity (≤-7.96 kcal/mol) to ISCU, notably ZINC000001576020, all of which complied with Lipinski's rule. CONCLUSIONS: This in-silico study nominates ISCU and AP1S2 as candidate diagnostic biomarkers for DKD and identifies computationally prioritized inhibitors targeting ISCU. These findings require experimental validation but provide a molecular framework for precision diagnosis and therapeutic development. These findings offer new molecular insights that could inform precision diagnosis and personalized treatment strategies for diabetic kidney disease.

Diabetic Nephropathies

Structure-based virtual screening, multi-score docking, and molecular dynamics simulation of novel small molecules targeting the epidermal growth factor receptor for potential management of oral squamous cell carcinoma.

UNLABELLED: Oral squamous cell carcinoma (OSCC) is a major global health burden, with epidermal growth factor receptor (EGFR) serving as an important therapeutic target. However, resistance to currently available EGFR inhibitors limits the efficacy of long-term treatment. In this study, a structure-based virtual screening approach was employed using the Mcule database to identify novel small molecules with potential EGFR-inhibitory activity. The top-ranked compounds were subjected to consensus docking using multiple docking platforms and compared with established EGFR inhibitors. The most promising complexes were further evaluated using 500 ns molecular dynamics simulations to investigate their structural stability, conformational flexibility, and binding persistence. ADMET and pharmacokinetic analyses were performed to assess the drug-like and safety profiles. Five lead compounds (C1-C5) demonstrated significant binding affinities toward EGFR, ranging from - 9.9 to - 9.2 kcal/mol, while satisfying the major drug-likeness criteria. Molecular dynamics simulations suggested that C1 and C4 may form relatively stable EGFR-ligand complexes, as supported by stable RMSD convergence and persistent interactions with key active-site residues throughout the simulation period. Trajectory-based interaction analyses further indicated a sustained binding behavior. ADMET profiling predicted favorable oral bioavailability and low predicted toxicity for most compounds, particularly C3 and C5, although a potential risk of cytochrome P450-mediated drug-drug interactions was observed. Overall, the shortlisted compounds exhibited docking and dynamic stability profiles comparable to those of the reference inhibitor Lapatinib. These findings suggest the potential therapeutic relevance of structurally novel EGFR-targeting scaffolds in OSCC and provide a foundation for future experimental validation through in vitro and in vivo studies. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s40203-026-00722-4.

ADMET

Chromosome scale genomes of two invasive Adelges species enable virtual screening for selective adelgicides.

Two invasive hemipteran adelgids are associated with widespread damage to several North American conifer species. Adelges tsugae, hemlock woolly adelgid, was introduced from Japan and reproduces parthenogenetically in North America, where it has rapidly decimated Tsuga canadensis and Tsuga caroliniana (the eastern and Carolina hemlocks, respectively). Adelges abietis, eastern spruce gall adelgid, introduced from Europe, forms distinctive pineapple-shaped galls on several native spruce species. While not considered a major forest pest, it weakens trees and increases susceptibility to additional stressors. Broad-spectrum insecticides that are often used to control adelgid populations can have off-target impacts on beneficial insects. Whole genome sequencing was performed on both species to aid in development of targeted solutions that may minimize ecological impact. Adelges abietis was sequenced using Illumina Linked-Read technology from 30 pooled individuals, with Hi-C scaffolding performed using data from a single individual collected from the same host plant. Adelges tsugae used Oxford Nanopore long-read sequencing from pooled nymphs. The assembled A. tsugae and A. abietis genomes, pooled from several parthenogenetic females, are 220.75 Mbp and 253.16 Mbp, respectively. Each consists of eight autosomal chromosomes, as well as two sex chromosomes (X1/X2), supporting the XX-XO sex determination system. The genomes are over 96% complete based on BUSCO assessment. Genome annotation identified 11,424 and 12,060 protein-coding genes in A. tsugae and A. abietis, respectively. Comparative analysis of proteins across 29 hemipteran species and 14 arthropod outgroups identified 31,666 putative gene families. Gene family evolution analysis with CAFE revealed lineage-specific expansions in immune-related aminopeptidases (ERAP1) and juvenile hormone binding proteins (JHBP), contractions in juvenile hormone acid methyltransferases (JHAMT), and conservation of nicotinic acetylcholine receptors (nAChR). These genes were explored as candidate families towards a long-term objective of developing adelgid-selective insecticides. Structural comparisons of proteins across seven focal species (Adelges tsugae, Adelges abietis, Adelges cooleyi, Rhopalosiphum maidis, Apis mellifera, Danaus plexippus, and Drosophila melanogaster) revealed high conservation of nAChR and ERAP1, while JHAMT exhibited species-specific structural divergence. The potential of JHAMT as a lineage-specific target for pest control was explored through virtual drug and pesticide screening.

adelgids

Oncogenic EME1 promotes tumor progression and immune modulation in human cancers with therapeutic targeting potential.

BACKGROUND: EME1, a critical DNA repair endonuclease, has emerged as a potential oncogene implicated in genome instability and cancer progression. However, its pan-cancer roles, prognostic significance, immune interactions, and therapeutic targeting remain underexplored. METHODS: We conducted a comprehensive pan-cancer analysis integrating multi-omics data from public databases, including TIMER2.0, GEPIA2, TISIDB, and cBioPortal, to evaluate EME1 expression, genetic alterations, and their association with clinical outcomes, immune infiltration, and molecular pathways. Virtual screening of 3180 FDA-approved drugs and molecular dynamics (MD) simulations were employed to identify and validate potential EME1 inhibitors. RESULTS: EME1 was significantly overexpressed in various human cancers and positively associated with advanced tumor grade and stage. High EME1 expression and mutations were linked to poor overall and disease-free survival. Immunogenomic profiling revealed strong positive correlations between EME1 and myeloid-derived suppressor cells (MDSCs), alongside a negative association with endothelial cell function, suggesting immunosuppressive roles. Machine learning models based on EME1-associated genes demonstrated high predictive accuracy for liver hepatocellular carcinoma (AUC > 0.90). Virtual screening identified eight promising drug candidates, including Everolimus and Dioscin, with strong binding affinities. MD simulations confirmed the stability of these interactions, particularly for Dioscin. CONCLUSION: This study reveals the multifaceted oncogenic roles of EME1 in tumor progression, immune evasion, and prognosis. It proposes EME1 as a promising biomarker and therapeutic target across multiple cancer types. The identified drug candidates warrant further in vitro and in vivo validation for potential repurposing in EME1-targeted cancer therapy.

EME1

Integrative TWAS and multi-omics analyses prioritize HSPE1 as a candidate risk gene for bipolar disorder with immune cell-specific regulatory evidence.

BACKGROUND: Bipolar disorder (BD) is a severe psychiatric disorder associated with substantial disability. Although genome-wide association studies have identified multiple BD-associated loci, the underlying genes and mechanisms remain incompletely understood. METHODS: We integrated a European-ancestry BD genome-wide association dataset with cross-tissue and tissue-specific transcriptome-wide association studies (TWAS) and complementary gene-based analysis. Candidate genes were further evaluated using differential expression analysis, consensus clustering, immune infiltration analysis, machine learning, summary-data-based Mendelian randomization, Mendelian randomization using single-cell expression quantitative trait locus data, single-nucleus transcriptomics, phenome-wide association analysis, and virtual screening. RESULTS: The integrative analyses prioritized 37 candidate genes. Peripheral-blood differential-expression analysis identified 14 genes that remained significant after FDR correction, and their expression profiles separated BD samples into two expression-defined clusters. Machine-learning analysis selected UNC50, LMAN2L, LYG2, HSPE1, and KANSL3 for an exploratory classification nomogram. SMR associated genetically predicted higher HSPE1 expression with increased BD risk in two blood eQTL datasets. Cell-type-specific analyses indicated HSPE1-related associations in T-cell and natural killer cell subsets, while single-nucleus analysis descriptively showed higher HSPE1 expression in medial thalamic T cells from BD samples. PheWAS identified no genome-wide significant associations for HSPE1, whereas virtual screening identified candidate compounds with favorable predicted docking scores against the HSPE1 structure. CONCLUSION: This integrative multi-omics study identified HSPE1 as a candidate BD risk gene with immune-cell-related regulatory evidence, providing insight into BD pathogenesis and supporting functional validation.

Humans

Structure-based discovery of inhibitors of Mac1 domain of nonstructural protein-3 of SARS-CoV-2 by machine learning-augmented screening of chemical space.

Significant efforts have been recently dedicated to the discovery of small molecule inhibitors against the Macrodomain 1 (Mac1) of nonstructural protein 3 (NSP3) as potential antivirals for SARS-CoV-2. Thus, Mac1 has also been selected as the target for the Critical Assessment of Hit-finding Experiments (CACHE) challenge #3. As contestants in that challenge, we developed a computational strategy that ranked on the top among all 23 participants in the competition and resulted in the discovery of a novel chemical series of non-charged Mac1 inhibitors. Those have been identified through the combination of machine learning-accelerated virtual screening of Enamine REAL Diversity Subset of approximately 25 million compounds and consequent hit expansion into the entire Enamine REAL Space library. In particular, the initially identified hit compound CACHE3-HI_1706_56 (KD = 20 μM) was explored by probing 17 close analogues from a library of 44 billion molecules from the Enamine REAL. All those analogues effectively displaced the Mac1-binding ADP-ribose peptide, and 12 were confirmed to engage with Mac1 by the Surface Plasmon Resonance experiments, revealing a new chemical series of compounds for hit-to-lead optimization. The structure of the CACHE3-HI_1706_56-Mac1 complex was further determined at high resolution with crystallography, confirming initial computational predictions. Our results illustrate the effectiveness of ML-accelerated docking to rapidly identify novel chemical series and provide a strong foundation for the development of SARS-CoV-2 NSP3 Mac1 inhibitors.

CACHE challenge

Machine Learning in Hyperlipidaemia Research: Screening and Experimental Insights into Lipid Metabolism Modulators.

Hyperlipidemia, characterized by elevated blood lipid levels, represents a major global health concern due to its strong association with cardiovascular disease, diabetes, and metabolic syndrome. While current therapies - such as statins, fibrates, bile acid sequestrants, and PCSK9 inhibitors - are effective in controlling hyperlipidemia, they are often associated with adverse effects, potential drug resistance, and suboptimal efficacy in certain patient populations. All of the above underscore the urgent need for safer and more effective therapeutic alternatives. Among the major molecular targets involved in the regulation of lipid metabolism are HMG-CoA reductase, PCSK9, peroxisome proliferator-activated receptors (PPARs), cholesteryl ester transfer protein (CETP), and nuclear receptors, including the liver X receptor (LXR) and farnesoid X receptor (FXR), which are also targets for future antihyperlipidemic drug development. Recent advancements in artificial intelligence (AI) and machine learning (ML) have significantly transformed and accelerated drug discovery by enabling the processing of vast amounts of genomic, proteomic, and chemical data. Furthermore, ML tools such as quantitative structure-activity relationship (QSAR) modelling, deep learning, random forest, and support vector machines (SVM) have proven predictive and effective in identifying novel lipid metabolism modulators, thereby enhancing the efficacy and accuracy of virtual screening. Meanwhile, molecular docking has become an integral part of structure-based drug design (SBDD), and software such as AutoDock, Glide, and GOLD have proven effective in generating accurate ligand-target docking models. Molecular docking, together with ML-based approaches, enables the identification of potent and selective drug candidates. Overall, the combination of ML and molecular docking offers an efficient and accurate platform for antihyperlipidemic drug discovery, helping to overcome the limitations of currently available therapeutic strategies.

HMG-CoA reductase

Artificial Intelligence for Natural Products Discovery and Development.

Natural products (NPs) remain a cornerstone of modern drug discovery, offering stereochemical complexity and diverse bioactivities that precisely modulate therapeutic targets, refined through billions of years of evolution. However, their research has long been hindered by inefficient, empirical workflows, high resource consumption, structural complexity, and the "multicomponent, multi-target" nature of their mechanisms. The exponential growth of genomic, metabolomic, and spectral data has overwhelmed conventional analytical methods, exposing critical bottlenecks in handling high-dimensional, heterogeneous datasets that exceed human interpretive capacity. Artificial intelligence (AI) is emerging as a transformative paradigm to address these challenges, integrating multi-omics and chemical data to shift NP research from fragmented empiricism toward mechanism-driven, precision-oriented development. By leveraging deep learning architectures- including graph neural networks, Transformers, and diffusion-based generative models-AI enables systematic decoding of NP biosynthesis, automated structure elucidation, rational target identification, knowledge extraction from vast unstructured scientific literature, and de novo molecular design. This review comprehensively surveys recent advances in AI applications across the full NP discovery and development pipeline, encompassing genome mining, structure-based and ligand-based virtual screening, multimodal structural characterization, lead optimization, and biosynthetic pathway engineering. We further examine the emerging roles of protein-centric, molecule- centric, and multimodal foundation models, as well as large language models, in bridging genotype-to-chemotype gaps and unlocking unstructured scientific knowledge. Finally, we discuss critical challenges including data scarcity, representational limitations for complex stereochemistry, physical plausibility in generative models, and the urgent need for experimental validation, while outlining future directions toward autonomous experimentation, closed-loop optimization, and human-AI collaborative discovery.

Artificial intelligence

Structure-based drug design of small-molecule c-Myc G-quadruplex binders.

The c-Myc oncogene is crucial in tumorigenesis. Although it is a promising therapeutic target, its protein lacks a conventional drug-binding pocket, making it traditionally "undruggable". Recent studies show that the c-Myc promoter can form a G-quadruplex (G4) structure, which suppresses transcription and offers a new strategy for indirect inhibition. In this study, structure-based virtual screening was performed using the c-Myc G4 crystal structure to screen the ChemDiv compound library, aiming to identify small molecules that bind to the G4 structure. Candidate compounds were evaluated in preliminary in vitro assays for biological activity. The results showed that Y502-3888 binds to the c-Myc G4 and downregulates c-Myc expression at both mRNA and protein levels. Collectively, these findings support the potential of Y502-3888 as a c-Myc G4 binder for the treatment of multiple myeloma (MM), providing a foundation for future development of anticancer agents targeting the c-Myc G4.

G-Quadruplexes

Structure-guided discovery of non-catechol dopamine D1 receptor ligands with biased agonism and antagonism.

The catechol L-DOPA, a cornerstone of Parkinson's disease (PD) treatment, has two major drawbacks: poor pharmacokinetics and, more significantly, debilitating dyskinesias from chronic dopamine D1 receptor (D1R) activation. Preclinical rodent studies suggest that D1R antagonism or β-arrestin-biased agonism can alleviate these motor complications, highlighting the need for next-generation non-catechol ligands. Through virtual screening, we identified eight novel chemotypes as D1R ligands, including two G protein-biased agonists, two β-arrestin-biased agonists and four antagonists. Structure-activity relationship (SAR) optimization led to the development of A82R, a non-catechol D1R antagonist (Ki 733 nM) with high D1 family over D2 family selectivity. Additionally, we present A69, a novel non-catechol β-arrestin-biased partial agonist for D1R (Ki 86.9 nM, stronger than representative D1R commercial drugs) with a sustained half-life of 1 h in the mouse brain. We show that the observed selectivity patterns are consistent with structural and information-theoretic limits on dopamine's ability to encode receptor subtype identity. Within these bounds, the non-catechol ligand chemotypes represent promising leads for developing therapies that modulate D1R signaling and reduce L-DOPA-induced dyskinesia in PD.

Receptors, Dopamine D1

AI-Driven Precision Medicine in Alzheimer's Disease: Drug Repurposing, Digital Therapeutics and Clinical Decision Support.

Alzheimer's Disease (AD) is a neurodegenerative disease that causes significant clinical, social, and economic burden worldwide. Despite improvements in understanding its multifaceted pathogenesis, current treatments are mostly symptomatic and ineffective across varied patient populations. To overcome these constraints, AI-driven precision medicine allows tailored risk assessment, treatment selection, and disease monitoring. This review covers AI's role in AD precision medicine, focusing on drug repurposing, digital therapies and clinical decision support systems. Machine and deep learning models are used to predict medication response, integrate heterogeneous data sources such as genomics, transcriptomics, neuroimaging and electronic health records, and uncover pharmacogenomic treatment success factors. The paper covers AIenabled precision pharmacology, including tailored dosing algorithms, adaptive therapeutic monitoring, and adverse drug reaction prediction. Bioinformatics-based target identification, network pharmacology, graphbased AI models, virtual screening, and real-world and clinical data validation are emphasized in AI-driven medication repurposing. AI-powered digital treatments like personalized cognitive training platforms, wearable- derived digital biomarkers, virtual and mixed reality interventions, adherence monitoring, and digital twins for therapy optimization have been discussed. AI-based clinical decision support systems are also thoroughly assessed for clinical value, accuracy, and explainability in disease subtyping, trajectory prediction, and risk stratification in preclinical and prodromal AD. Despite these promises, data heterogeneity, algorithmic bias, legal barriers, and privacy concerns exist. Federated learning enables safe multi-center collaboration and hybrid AI-human approaches, and it represents the future. AI's ability to alter AD care opens the door to precision medicine paradigms that use repurposed medications, digital tools and intelligent decision-making to improve patient outcomes.

Alzheimer’s disease

Discovery of NAT-6-321056 as a novel modulator of VEGFR2 signaling to suppress tumor angiogenesis.

Vascular endothelial growth factor receptor 2 (VEGFR2) is a master regulator of angiogenesis and cancer progression. However, current VEGFR2 modulators face significant challenges, including off-target toxicity and acquired resistance, underscoring the urgent need for novel therapeutic agents with improved efficacy and safety profiles. Here, we reported that virtual screening of 39,442 natural products from the ZINC natural products-derived library, coupled with molecular docking and molecular dynamics (MD) simulations to evaluate the binding stability of candidate compounds, identified NAT-6-321056 as a highly promising modulator of VEGFR2 signaling. Biological evaluations demonstrated that NAT-6-321056 exerted potent inhibition on the growth of a broad spectrum of cancer cells, including both solid tumors and hematological malignancies. In EA.hy 926 endothelial cells and SK-N-DZ neuroblast cells, the compound significantly suppressed proliferation, migration, and invasion. Microscale thermophoresis (MST) confirmed direct binding of NAT-6-321056 to VEGFR2 with favorable affinity. Kinase profiling against a panel of 33 kinases indicated that NAT-6-321056 exhibited a multi-kinase modulation profile. Mechanistic studies revealed that NAT-6-321056 suppressed the expression of hypoxia-inducible factor 1-alpha (HIF-1α) and was associated with reduced VEGFR2 phosphorylation and attenuation of the downstream ERK/JNK/AKT signaling pathways. Moreover, NAT-6-321056 exhibited robust in vivo anti-angiogenic effects in both the chick chorioallantoic membrane (CAM) assay and transgenic zebrafish vascular fluorescence imaging models. Computational absorption, distribution, metabolism, excretion, and toxicity (ADMET) prediction suggested acceptable drug-like properties. Collectively, these findings demonstrated that NAT-6-321056 is a promising modulator of VEGFR2 signaling with potent anti-angiogenic activity and represents a viable candidate for cancer therapy.

Vascular Endothelial Growth Factor Receptor-2

FusionTarget: Computational framework for drug repurposing against modeled fusion protein structures from genomic breakpoints.

Many fusion genes have been recognized as biomarkers and therapeutic targets. However, the lack of knowledge on protein structures and targeting approaches made it challenging to develop effective targeting therapeutics. To fill this, we developed a computational pipeline, FusionTarget, which annotates the genomic DNA breakage to RNA and protein sequences, predicts the 3D structures of fusion proteins, and performs comparative virtual screening, comparative molecular dynamics simulation, and quantitative analyses to identify the fusion protein-selective small molecules by selecting drugs with consistent high-fold binding affinity between fusion and wild-type proteins in multiple isoforms. We applied our pipeline to EWSR1::FLI1 in Ewing sarcoma and KMT2A::AFF1 in infant acute lymphoblastic leukemia. Further cell assay experiments confirmed that cells expressing individual fusion genes were more sensitive to the suggested drugs, and the key downstream genes were affected by our drugs. FusionTarget provides a unique foundation for developing therapeutics targeting fusion proteins.

applied computing in medical science

Gain-of-function PPM1D mutations attenuate ischemic stroke.

Identification of genetic aberrations in stroke, the second leading cause of death worldwide, is of paramount importance for understanding the disease pathogenesis and generating new therapies. Whole-genome sequencing from 10,241 ischemic stroke patients identified eight patients carrying gain-of-function mutations on coding variants in the protein phosphatase magnesium-dependent 1 δ (PPM1D) gene. Patients carrying PPM1D mutations exhibit better stroke-related clinical phenotypes, including improvements in peripheral inflammation, fibrinogen, low-density lipoprotein, cholesterol and plateletcrit level. Experimental brain ischemia in Ppm1d-deficient (Ppm1d-/-) mice resulted in enlarged lesions and pronounced neurological impairments. Spatial transcriptomics revealed a distinct Ppm1d-associated gene expression pattern, indicating disrupted endothelial homeostasis during ischemic brain injury. Proteomic analysis demonstrated that differentially expressed proteins in primary brain endothelial cells from Ppm1d-/- mice were significantly enriched in the peroxisome proliferator-activated receptors (PPARs)-mediated metabolic signaling. Mechanistically, Ppm1d deficiency promoted aberrant fatty acid β-oxidation and increased oxidative stress, which impaired endothelial cell function through the PPARα pathway. A small molecule, T2755, was identified to engage Trp427 and stabilize PPM1D, thereby mitigating ischemic brain injury in mice. Collectively, we find that PPM1D protects against ischemic brain injury and validates its pharmacological stabilizer T2755 as a promising therapy for ischemic stroke. Gain-of-function PPM1D mutations attenuate ischemic cerebral injury. Whole-genome sequencing data of 10,241 ischemic stroke patients from the Third Chinese National Stroke Registry (CNSR-III) identified eight patients with gain-of-function mutations in the protein phosphatase magnesium-dependent 1 δ (PPM1D) gene (17q23.2). These mutation carriers displayed improved peripheral inflammation, decreased fibrinogen, low-density lipoprotein, cholesterol and plateletcrit level. Ppm1d-deficient (Ppm1d-/-) mice exhibited exacerbated stroke outcomes, characterized by enlarged infarct volumes, disrupted cerebrovascular architecture, and enhanced neuro-inflammation. Mechanistically, Ppm1d deficiency induced the disturbance of endothelial fatty acid metabolism involving the PPARα pathway. Through integrated computational modeling, virtual screening, and in vitro validation, T2755 was identified as a small molecule PPM1D stabilizer. Pharmacological PPM1D stabilization with T2755 significantly attenuated ischemic brain injury in murine models.

Aged

Comparative and Subtractive Genomics Analysis of Multidrug-Resistant Klebsiella pneumoniae Strains for Novel Target Identification and Drug Repurposing Strategies.

The rapid rise of multidrug-resistant (MDR) Klebsiella pneumoniae has created a major global health challenge due to the limited availability of conserved therapeutic targets effective across diverse resistant strains. In this study, an integrative computational target-discovery and drug-repurposing framework was applied to six clinically relevant K. pneumoniae strains. Comparative genomic analysis identified 3012 conserved genes, which were subsequently filtered to nine essential, non-host homologous proteins. Among these, three conserved cytoplasmic proteins (accD, cpxR, and mraZ) were prioritized for functional analysis, with acetyl-CoA carboxylase subunit beta (accD) emerging as the most promising therapeutic target based on sequence conservation, predicted essentiality, subcellular localization, and pathway association. Structural assessment supported the reliability of the predicted accD model, whereas consensus binding-site analysis identified key residues suitable for ligand interaction. Virtual screening of FDA-approved drugs followed by molecular docking identified several compounds with favorable binding profiles toward accD. Subsequent molecular dynamics simulations, including root mean square deviation (RMSD), root mean square fluctuation (RMSF), radius of gyration (Rg), hydrogen-bond occupancy, principal component analysis (PCA), and PCA-based free energy landscape (FEL) analyses, consistently identified tenapanor, micafungin, deferoxamine, and cobicistat as the most stable protein-ligand complexes, with tenapanor exhibiting the most favorable overall structural and thermodynamic stability profile. These findings identify accD as a promising therapeutic target in MDR K. pneumoniae and suggest several FDA-approved compounds as potential candidates for drug repurposing. Although experimental validation is needed to confirm their biological activity and therapeutic potential, this study demonstrates the potential of integrating comparative genomics with molecular dynamics analyses to support antimicrobial target identification and drug repurposing against MDR bacterial pathogens.

Klebsiella pneumoniae

Large language models in bioinformatics: a comprehensive survey.

The emergence of foundation models with trillion-level parameters has redefined the landscape of artificial intelligence. Various fields are developing their own large-scale models, which can solve many problems within the field and improve work efficiency. Biological large-scale models are a cross-disciplinary research field that combines mathematics, computer science, and biology, aiming to simulate and understand the structure, function, and dynamic changes of biological systems through the establishment of complex computational models. This field covers multiple levels such as biological pathways, population dynamics, protein folding, etc., providing us with tools for deep exploration of the mysteries of life and applications in medicine, ecology, and other fields. This article reviews the background and research status of biological large-scale models, and discusses future directions. Large language models (LLMs) and other large-scale foundation models have rapidly advanced in recent years, enabling powerful representation learning and generation across text, sequences, and multimodal data. In bioinformatics and biomedicine, these models are increasingly used to analyze genomic sequences, infer protein properties and structures, support drug discovery, and integrate heterogeneous biomedical evidence. This survey reviews the basic principles of LLMs and summarizes representative applications in (i) gene and genome sequence analysis, (ii) protein structure and function prediction, and (iii) drug design, including virtual screening and personalized medicine. We also discuss emerging multi-model modeling approaches, as well as key challenges such as data quality and privacy, interpretability, generalization to new organisms and tasks, and responsible deployment in health-related settings. Finally, we outline future directions for developing reliable, scalable, and explainable bioinformatics foundation models.

bioinformatics

The impact of screening on the incidence of cervical cancer in England and Wales.

Age-specific incidence curves for clinical cancer of the cervix in England and Wales show progressive changes over the period 1963-1978; in particular, a large reduction in incidence is seen in the age group 35-54. Since screening on any scale began in the early 1960s, we have investigated how much of this reduction in incidence in the middle age range can be attributed to detection of pre-invasive disease. Data on registrations of in-situ cancer have been used to estimate the patterns that might have been observed in the absence of screening. The results indicate clear cohort effects on incidence, with rising rates in the generations born 1906-1921 and since 1931, with a fall in the decade between. In addition to this, screening has probably led to a substantial reduction in the number of cases of clinical cancer in women aged 35-54, but has had little effect over the age of 60 where virtually no screening has been performed. Below age 35 the observed increase in incidence may be considerably less than it would have been in the absence of screening.

Adult

Survival experience in the Breast Cancer Detection Demonstration Project.

Very high survival rates have been observed in four through 11 years of follow-up in 4,240 women with a histologically confirmed diagnosis of breast cancer in the Breast Cancer Detection Demonstration Project (BCDDP). The relative five, eight, and 10-year survival rates were 88, 83, and 79 percent, respectively. Allowances were made for lead-time bias among cancers detected through screening, and the validity of the findings was supported by internal analyses, which showed that length-time bias was of little, if any, importance, and that any possible "overdiagnosis" of cancer cases was also of small relevance. In view of current interest in the value of screening women before age 50, intensive analyses were made comparing the BCDDP data for women in their 40s with women in their 50s. In terms of kinds of breast cancers found, modality of finding them, and survival rates once they have been found, the parallel results for the two groups show that screening was virtually as effective in the younger as in the older women. Some authorities are of the opinion that the benefits of mammography after age 50 are well documented, but at younger ages the evidence is still inconclusive. The findings in this study show there is no doubt of the very successful results of screening for breast cancer with mammography in younger as well as older women. In comparing relative five-year and eight-year survival rates for women with invasive breast cancers detected through screening in the BCDDP, with those for cases diagnosed in the National Cancer Institute's Surveillance, Epidemiology and End Results (NCI SEER) program from 1977 to 1982, it is seen that for individual subcategories by tumor size and nodal class, the survival rates are about the same. However, for overall invasive cancers, the five-year and eight-year survival rates were 87 and 81 percent, respectively, for the BCDDP compared with 74 and 65 percent for SEER. Thus the substantial gains in survival followed the large shift toward a high proportion of cancers being diagnosed and treated in more favorable stages through the screening accomplishments. With respect to the relative case fatality rates, the complements of the relative survival rates, the eight-year rate of 19 percent for the BCDDP versus that of 35 percent for SEER connotes 46 percent fewer women dying in the BCDDP group.

Adult