Search PubMedSearch

SEARCH · Search PubMed

Results for “Protein function prediction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

A machine learning-based predictive model for radiosensitivity in nasopharyngeal carcinoma utilizing serum proteomics.

BACKGROUND: Nasopharyngeal carcinoma (NPC) remains highly sensitive to radiotherapy; however, radioresistance in a subset of patients leads to local recurrence and distant metastasis. Serum proteomics provides a minimally invasive approach to capturing dynamic physiological changes, and machine learning enables efficient construction of predictive models. This study aimed to develop and validate a serum proteomics–based machine-learning model for predicting radiotherapy sensitivity in nasopharyngeal carcinoma (NPC). METHODS: Pretreatment serum samples from newly diagnosed NPC patients were analyzed using SELDI-TOF-MS. Differentially expressed proteins between radiosensitive and radioresistant groups were identified using limma. GO and KEGG analyses were performed to explore functional enrichment. Twelve machine-learning algorithms were used to construct predictive models, and the top-performing models were optimized through feature selection. A Random Forest model with seven features was identified as the optimal model. External validation was performed using an independent cohort with ELISA-quantified protein levels. Model performance was assessed using Receiver operating characteristic curve (ROC), calibration analysis, decision curve analysis (DCA), and 10-fold cross-validation. SHapley Additive exPlanations (SHAP) analysis was applied for model interpretability, and the final model was deployed via a ShinyAPP. RESULTS: A total of 96 differentially expressed proteins were identified, which involved multiple function and signaling pathways. The Random Forest model demonstrated the best predictive performance, achieving an area under the curve (AUC) of 0.963 in the training set and 0.975 in the validation set. Cross-validation yielded an average AUC of 0.965. DCA indicated high clinical utility across a broad threshold range, and calibration curves showed good model agreement. Seven proteins (PLXND1, GSR, PGD, PTPRC, OR2T29, ACTG2, CHAD) were selected as final features. SHAP analysis provided global and individual-level interpretability. A web-based tool was developed to facilitate clinical application. CONCLUSION: This study establishes a robust serum proteomics–based machine-learning model capable of accurately predicting radiotherapy sensitivity in NPC. The model offers clinical interpretability and practical implementation, supporting personalized radiotherapy decision-making.

Humans

Genomic insights of first varicella zoster clade9 strain: a potential silent surge in Pakistan.

The study presents the first-time detection of one of the rare clades (clade9 strain) of varicella zoster virus (VZV) from Pakistan. The next-generation sequencing confirmed wild-type clade9 strain through clade-specific markers at C5827A, T33722C, T33725C, T33728C, T38055C, G69424A, C87841T and T95241C and restriction profile of PstI+BgII+SmaI-. The rarely reported SNPs (22/134) were detected along with 12/42 rare amino-acid mutations. However, the mutations at C77Y, Q43H, D613E and A2V were predicted to be not-tolerated hence might affect protein function. The VZV (PV934234) strain clustered with clade9 strains upon phylogenetics. Thus, the first-time detection of clade9 raises concern of limited genomic surveillance of VZV in Pakistan. This necessitates genomic surveillance and continuous clinical vigilance in Pakistan to avoid any potential silent surge in the country.

Clade 9

Systematic Proteome Profiling of Maternal Plasma for Development of Preeclampsia Biomarkers.

Preeclampsia (PE) is a hypertensive disorder of pregnancy with various clinical symptoms. However, traditional markers for the disease including high blood pressure and proteinuria are poor indicators of the related adverse outcomes. Here, we performed systematic proteome profiling of plasma samples obtained from pregnant women with PE to identify clinically effective diagnostic biomarkers. Proteome profiling was performed using TMT-based liquid chromatography-mass spectrometry (LC-MS/MS) followed by subsequent verification by multiple reaction monitoring (MRM) analysis on normal and PE maternal plasma samples. Functional annotations of differentially expressed proteins (DEPs) in PE were predicted using bioinformatic tools. The diagnostic accuracies of the biomarkers for PE were estimated according to the area under the receiver-operating characteristics curve (AUC). A total of 1307 proteins were identified, and 870 proteins of them were quantified from plasma samples. Significant differences were evident in 138 DEPs, including 71 upregulated DEPs and 67 downregulated DEPs in the PE group, compared with those in the control group. Upregulated proteins were significantly associated with biological processes including platelet degranulation, proteolysis, lipoprotein metabolism, and cholesterol efflux. Biological processes including blood coagulation and acute-phase response were enriched for down-regulated proteins. Of these, 40 proteins were subsequently validated in an independent cohort of 26 PE patients and 29 healthy controls. APOM, LCN2, and QSOX1 showed high diagnostic accuracies for PE detection (AUC >0.9 and p&#xa0;<&#xa0;0.001, for all) as validated by MRM and ELISA. Our data demonstrate that three plasma biomarkers, identified by systematic proteomic profiling, present a possibility for the assessment of PE, independent of the clinical characteristics of pregnant women.

Humans

Cloning and partial sequencing of an operon encoding two Pseudomonas putida haloalkanoate dehalogenases of opposite stereospecificity.

We have cloned fragments of DNA (up to 13 kb), from Pseudomonas putida AJ1, that code for two stereospecific haloalkanoate dehalogenases. These enzymes are highly specific for D and L substrates. The two genes, designated hadD and hadL, have been isolated and independently expressed in Escherichia coli and P. putida hosts by using broad-host-range vectors. They are closely adjacent and inducible in what appears to be an operon with an upstream open reading frame of unknown function. Nucleotide sequence determination of hadD predicts a mature, cytoplasmic protein of 300 amino acid residues (molecular weight of 33,601). This has no significant homology with the L-specific haloalkanoate dehalogenases from Pseudomonas sp. strain CBS3 (B. Schneider, R. Muller, R. Frank, and F. Lingens, J. Bacteriol. 173:1530-1535, 1991) nor with any other known DNA or protein sequences.

Amino Acid Sequence

Identification of novel cytoskeleton protein involved in spermatogenic cells and sertoli cells of non-obstructive azoospermia based on microarray and bioinformatics analysis.

BACKGROUND: During mammalian spermatogenesis, the cytoskeleton system plays a significant role in morphological changes. Male infertility such as non-obstructive azoospermia (NOA) might be explained by studies of the cytoskeletal system during spermatogenesis. METHODS: The cytoskeleton, scaffold, and actin-binding genes were analyzed by microarray and bioinformatics (771 spermatogenic cellsgenes and 774 Sertoli cell genes). To validate these findings, we cross-referenced our results with data from a single-cell genomics database. RESULTS: In the microarray analyses of three human cases with different NOA spermatogenic cells, the expression of TBL3, MAGEA8, KRTAP3-2, KRT35, VCAN, MYO19, FBLN2, SH3RF1, ACTR3B, STRC, THBS4, and CTNND2 were upregulated, while expression of NTN1, ITGA1, GJB1, CAPZA1, SEPTIN8, and GOLGA6L6 were downregulated. There was an increase in KIRREL3, TTLL9, GJA1, ASB1, and RGPD5 expression in the Sertoli cells of three human cases with NOA, whereas expression of DES, EPB41L2, KCTD13, KLHL8, TRIOBP, ECM2, DVL3, ARMC10, KIF23, SNX4, KLHL12, PACSIN2, ANLN, WDR90, STMN1, CYTSA, and LTBP3 were downregulated. A combined analysis of Gene Ontology (GO) and STRING, were used to predict proteins' molecular interactions and then to recognize master pathways. Functional enrichment analysis showed that the biological process (BP) mitotic cytokinesis, cytoskeleton-dependent cytokinesis, and positive regulation of cell-substrate adhesion were significantly associated with differentially expressed genes (DEGs) in spermatogenic cells. Moleculare function (MF) of DEGs that were up/down regulated, it was found that tubulin bindings, gap junction channels, and tripeptide transmembrane transport were more significant in our analysis. An analysis of GO enrichment findings of Sertoli cells showed BP and MF to be common DEGs. Cell-cell junction assembly, cell-matrix adhesion, and regulation of SNARE complex assembly were significantly correlated with common DEGs for BP. In the study of MF, U3 snoRNA binding, and cadherin binding were significantly associated with common DEGs. CONCLUSION: Our analysis, leveraging single-cell data, substantiated our findings, demonstrating significant alterations in gene expression patterns.

Male

Significance of autogenously regulated and constitutive synthesis of regulatory proteins in repressible biosynthetic systems.

The functional implications of the different modes of regulation have been examined systematically. The results lead to certain predictions. The regulatory protein in repressor-controlled systems is constitutively synthesised. In activator-controlled systems synthesis of the regulatory protein is autogenously regulated. There is favourable agreement between these predictions and published experimental evidence.

Amino Acids

Molecular mechanisms for proton transport in membranes.

Likely mechanisms for proton transport through biomembranes are explored. The fundamental structural element is assumed to be continuous chains of hydrogen bonds formed from the protein side groups, and a molecular example is presented. From studies in ice, such chains are predicted to have low impedance and can function as proton wires. In addition, conformational changes in the protein may be linked to the proton conduction. If this possibility is allowed, a simple proton pump can be described that can be reversed into a molecular motor driven by an electrochemical potential across the membrane.

Biological Transport, Active

shinyDeepGxP: a user-friendly R shiny app for predicting surface protein abundance from scRNA-seq expression using deep learning in blood cells.

MOTIVATION: Understanding accurate immune cell heterogeneity and function in single-cell datasets requires access to protein-level information, which is often unavailable due to experimental limitations. RESULTS: We present shinyDeepGxP, an interactive web application featuring our deep learning model, DeepGxP, for predicting surface protein abundance from single-cell RNA-sequencing (scRNA-seq) data. This platform makes DeepGxP accessible to researchers without programming skills. Users can upload scRNA-seq count matrices and use "Predict Protein" to predict the abundance of 224 biologically relevant surface proteins. shinyDeepGxP provides visualizations to help identify distinct cell populations based on predicted protein profiles. Moreover, users can choose "Explore Model" to reveal key RNA predictors and their associated biological pathways for each protein. Overall, shinyDeepGxP is a user-friendly, freely available web tool that provides protein-level detail for RNA-only single-cell datasets, enabling multimodal discovery without additional experiments. AVAILABILITY AND IMPLEMENTATION: shinyDeepGxP can be launched on https://shiny.crc.pitt.edu/deepgxp/.

Journal Article

Freely available genomic datasets for atrial fibrillation research: current resources and analytical pipeline.

Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia, characterized by clinical and genetic heterogeneity. Increasing use of genomics and other omics approaches has driven reliance on publicly available AF datasets to advance biological discovery. Thus, this systematic review aimed to identify freely available genomic AF datasets through Mendeley Data and its interconnected repositories, and to characterize the most common analyses performed on these data. The search was conducted in adherence to the PRISMA 2020 guideline. Nineteen freely available genomic AF datasets were identified: Summary statistics for 'Biobank-driven genomic discovery yields new insight into atrial fibrillation biology', hum0014.v8.58qt.v1, AF GWAS in UK Biobank, UK Biobank (Publication 9659), GWAS summary statistics from a 2025 multi-ancestry AF meta-analysis, GSE115574, GSE128188, GSE14975, GSE2240, GSE238242, GSE254133, GSE261170, GSE271748, GSE271839, GSE293813, GSE294456, GSE31821, GSE41177, and GSE79768. The GEO datasets were further examined using differential gene expression, functional enrichment, protein-protein interaction networks, hub gene analysis, microRNA target prediction, and gene clustering, as well as, for the more recently deposited datasets, eQTL colocalization, single-cell/single-nucleus clustering, cell-cell communication analysis, and gene-dosage-dependent transcriptional and electrophysiological profiling. These analyses show some consistency but also considerable heterogeneity in initial conditions, data normalization, and analytical methodological settings. In conclusion, only a limited number of datasets are freely available, so additional, well-characterized and standardized datasets are needed to provide a complete picture of the AF pathology.

Mendeley Data

From pan-life phase insights to PhaseHub: Analyzing protein condensate complexity.

Intracellular biomolecular condensation forms multicomponent signaling hubs that regulate development, stress responses, and environmental adaptation. While the molecular grammar encoded within scaffold proteins defines the basal associative features driving condensation, heterotypic condensates are intrinsically dynamic, multicomponent, and far-from-equilibrium systems. Consequently, how condensates organize component composition, stoichiometry, and functional specificity in space and time under physiological conditions remains poorly understood. Addressing this challenge requires integrative frameworks that combine predictive biophysical features with experimental information on protein abundance, interaction networks, subcellular localization, and evolutionary conservation. In this study, we first analyzed phase separation (PS) proteins across the tree of life in 1106 species, revealing a stark contrast in computationally predicted PS propensity between eukaryotes and prokaryotes, with genome size as a key determinant. Through a broad analysis of amino acid homorepeat-containing proteins (HRPs) across all species, we uncovered how PS evolves via a balance between functional condensation and avoidance of harmful, aggregation-prone sequences. We further identified potential signaling hubs and components across kingdoms by integrating PS-positive proteins with experimentally derived abundance and interactome data from four model eukaryotic species. Using Arabidopsis as a model, we dissected the relationships among PS propensity, condensation hub prediction, HRPs, subcellular localization, and structural conservation. Finally, we developed PhaseHub, a user-friendly interface for exploring scaffold-client dynamics, PS components, sequence signatures within each PS protein, and hubs. Collectively, our work provides an evolutionary framework for understanding multicomponent PS hubs by integrating molecular grammar with physiological context, thereby facilitating hypothesis generation and rational design.

Phase Separation

Expression and characterization of human FKBP52, an immunophilin that associates with the 90-kDa heat shock protein and is a component of steroid receptor complexes.

Using an FK506 affinity column to identify mammalian immunosuppressant-binding proteins, we identified an immunophilin with an apparent M(r) approximately 55,000, which we have named FKBP52. We used chemically determined peptide sequence and a computerized algorithm to search GenPept, the translated GenBank data base, and identified two cDNAs likely to encode the murine FKBP52 homolog. We amplified a murine cDNA fragment, used it to select a human FKBP52 (hFKBP52) cDNA clone, and then used the clone to deduce the hFKBP52 sequence (calculated M(r) 51,810) and to express hFKBP52 in Escherichia coli. Recombinant hFKBP52 has peptidyl-prolyl cis-trans isomerase activity that is inhibited by FK506 and rapamycin and an FKBP12-like consensus sequence that probably defines the immunosuppressant-binding site. FKBP52 is apparently common to several vertebrate species and associates with the 90-kDa heat shock protein (hsp90) in untransformed mammalian steroid receptor complexes. The putative immunosuppressant-binding site is probably distinct from the hsp90-binding site, and we predict that FKBP52 has different structural domains to accommodate these functions. hFKBP52 contains 12 protein kinase phosphorylation-site motifs and a potential calmodulin-binding site, implying that posttranslational phosphorylation could generate multiple isoforms of the protein and that calmodulin and intracellular Ca2+ levels could affect FKBP52 function. FKBP52 transcripts are present in a variety of human tissues and could vary in abundance and/or stability.

Amino Acid Isomerases

The DNA sequence of equine herpesvirus-1.

The complete DNA sequence was determined of a pathogenic British isolate of equine herpesvirus-1, a respiratory virus which can cause abortion and neurological disease. The genome is 150,223 bp in size, has a base composition of 56.7% G + C, and contains 80 open reading frames likely to encode protein. Since four open reading frames are duplicated in the major inverted repeat, two are probably expressed as a spliced mRNA, and one may contain an internal transcriptional promoter, the genome is considered to contain 76 distinct genes. The genes are arranged collinearly with those in the genomes of the two previously sequenced alphaherpesviruses, varicella-zoster virus, and herpes simplex virus type-1, and comparisons of predicted amino acid sequences allowed the functions of many equine herpesvirus 1 proteins to be assigned.

Amino Acid Sequence

The signed two-space proximity model for learning representations in protein-protein interaction networks.

MOTIVATION: Accurately predicting complex protein-protein interactions (PPIs) is crucial for decoding biological processes, from cellular functioning to disease mechanisms. However, experimental methods for determining PPIs are computationally expensive. Thus, attention has been recently drawn to machine learning approaches. Furthermore, insufficient effort has been made toward analyzing signed PPI networks, which capture both activating (positive) and inhibitory (negative) interactions. To accurately represent biological relationships, we present the Signed Two-Space Proximity Model (S2-SPM) for signed PPI networks, which explicitly incorporates both types of interactions, reflecting the complex regulatory mechanisms within biological systems. This is achieved by leveraging two independent latent spaces to differentiate between positive and negative interactions while representing protein similarity through proximity in these spaces. Our approach also enables the identification of archetypes representing extreme protein profiles. RESULTS: S2-SPM's superior performance in predicting the presence and sign of interactions in SPPI networks is demonstrated in link prediction tasks against relevant baseline methods. Additionally, the biological prevalence of the identified archetypes is confirmed by an enrichment analysis of Gene Ontology (GO) terms, which reveals that distinct biological tasks are associated with archetypal groups formed by both interactions. This study is also validated regarding statistical significance and sensitivity analysis, providing insights into the functional roles of different interaction types. Finally, the robustness and consistency of the extracted archetype structures are confirmed using the Bayesian Normalized Mutual Information (BNMI) metric, proving the model's reliability in capturing meaningful SPPI patterns. AVAILABILITY: S2-SPM is implemented and freely available under the MIT license at https://github.com/Nicknakis/S2SPM.

Protein Interaction Mapping

Unravelling the genomic potential of sponge-associated Streptomyces sp. BLC 17-3 from Indonesia for mannooligosaccharide production.

This research aims to show the promising capacity of Streptomyces sp. BLC 17-3 to produce high &#x3b2;-mannanase enzymes and generate mannooligosaccharide (MOS) such as mannobiose, mannotriose, mannotetraose and mannopentaose when exposed to mannan polymers. Streptomyces sp. BLC 17-3 was isolated from the sponge (Rhabdastrella globostellata) Put4 obtained from the marine waters of Putus Island in Bitung, North Sulawesi, Indonesia. The characterization results showed that the peak enzyme activity was achieved at 50&#xa0;mM sodium acetate, 6.0 pH, and 60&#xa0;&#xb0;C temperature on the seventh day of production with a value of 155.77&#xa0;&#xb1;&#xa0;3.21&#xa0;U/mL. The SDS-PAGE and zymograms also showed that the size of the enzyme molecule was approximately &#xb1;34.8-49.1&#xa0;kDa. Moreover, whole-genome sequencing was conducted to identify the genetic basis of MOS-synthesizing capabilities in the selected strain, followed by functional annotation of genes encoding mannan degradation and associated functions. The results showed an 8,248,862&#xa0;Mb complete draft genome of the strain which comprised 111 predicted gene models. Gene annotation also provided important information about the location and function of protein-encoding genes. A total of 6 mannan degradation-related genes encoding mannanase-related metabolism were identified and the three-dimensional structures were predicted using AlphaFold 3. This characterization and modeling further enhanced the bioprospecting and development of this strain which exhibited efficient mannose metabolism. The results showed Streptomyces sp. BLC 17-3 as a promising microorganism for the future bioproduction of MOS which were discovered to have the capability of serving as a potential prebiotic substance to enhance digestion and promote health.

Bioprospecting

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Artificial Intelligence

Safe and Stable Germline Transmission of MSTN Mutations in Cattle.

With the global population expected to reach 10 billion by 2050, sustainable livestock production is critical. Gene editing of the myostatin (MSTN) gene represents a promising strategy to enhance muscle growth in cattle. In this study, MSTN-mutated founder (F0) cows were used to generate F1 offspring via ovum pick-up, in&#xa0;vitro fertilization, and embryo transfer. Four F1 calves were born, all confirmed to be heterozygous for the MSTN mutation. Long-term monitoring showed normal growth and no visible health abnormalities. Whole-genome sequencing identified SNPs, INDELs, and structural variants, most with minimal predicted functional effects. Proteomic profiling of Longissimus dorsi muscle quantified 2947 proteins, revealing only subtle expression differences between MSTN-mutated and wild-type cattle. These results demonstrate stable inheritance and confirm that MSTN editing does not disrupt genome integrity or protein expression. Overall, our findings support the safety and utility of MSTN gene editing to improve livestock productivity for future food security.

Animals

Proteomics combined with single-cell sequencing reveals key genes and computational lead compound related to ligamentum flavum hypertrophy, lactate metabolism and lactate modification.

Ligamentum flavum hypertrophy (LFH) is a hallmark pathological feature of lumbar spinal stenosis; however, its underlying molecular mechanisms remain incompletely understood. Lactate metabolism and related lactylation modifications have emerged as critical links between cellular metabolism and epigenetic regulation, with established roles in various fibrotic and inflammatory diseases. Nevertheless, the specific contribution of lactylation to LFH pathogenesis remains unexplored. In this study, we integrated proteomic profiling of ligamentum flavum tissues with single-cell transcriptomic data to identify differentially expressed proteins associated with LFH. Cross-referencing these genes with genes involved in lactate metabolism and lactylation yielded 16 candidate genes. Through functional enrichment analysis, protein-protein interaction network construction, and GraphBAN model prediction, we identified five hub genes (NDUFS2, HMOX1, SPR, FABP5, and PFKP) and two potential lead compounds (ZINC000014879975 and ZINC000242437513). Molecular docking analysis confirmed favorable binding affinities between these compounds, suggesting that they may serve as potential lead compounds worthy of further experimental investigation. Single-cell analysis further revealed that macrophages occupy a central position in the LFH microenvironment, resulting in pronounced metabolic reprogramming and remodeling of intercellular communication networks, particularly via the MIF-CD74/CD44 axis, under pathological conditions.

Proteomics

Chromosome-level genome assembly with telomeric repeats at scaffold ends for Rhabdosargus sarba.

Rhabdosargus sarba, the goldlined seabream, is a euryhaline marine fish of great aquaculture potential. Genome sequencing and assembly of R. sarba was carried utilizing a multi-platform sequencing strategy that included long-read sequencing (PacBio HiFi), short-read sequencing (Illumina), and chromatin interaction mapping (Hi-C). The final genome assembly size after scaffolding was 764.59&#x2009;Mb in 31 scaffolds with an N50 length of 33.98&#x2009;Mb. Repeat profiling of primary assembly showed that 28.71% of the genome comprises of repeat elements. Gene prediction utilising the evidence from ab initio prediction and transcriptome data revealed 26,913 protein encoding genes and functional annotation and pathway analysis showed their participation in 332 pathways. This genome is an excellent resource for future research on genetic improvement and molecular breeding programmes for R. sarba.

Animals