Search PubMedSearch

SEARCH · Search PubMed

Results for “Protein function prediction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa

On the state of protein function prediction: a report on the fourth CAFA challenge.

BACKGROUND: The Critical Assessment of Functional Annotation (CAFA) is a community effort held to understand the field of computational protein function prediction. Every three years, since 2010, the organizers initiate an experiment to collect function predictions on a large set of proteins and then evaluate the performance of predicting methods on a subset of proteins that have accumulated experimental annotations between the submission deadline and the evaluation time. CAFA provides an independent and rigorous assessment of the current state of the art, thus leveling the playing field, highlighting successes, revealing bottlenecks, and offering a forum for the exchange of ideas in protein science. Here, we report the results of the fourth CAFA experiment (CAFA4). RESULTS: CAFA4 featured the participation of 148 methods from 70 research groups on a total of 46,205 unique proteins over a 5-year annotation accumulation phase, the longest in any CAFA. In a comparison across CAFA2-CAFA4 methods, the prediction of Gene Ontology (GO) terms has clearly improved across all three GO aspects and traditional evaluation settings. While not achieving the first rank, several CAFA2 and CAFA3 methods featured in the top ten methods in many evaluations, suggesting that earlier methods still hold relevance. The performance is weaker in the newly introduced "partial knowledge" evaluation category (proteins with experimental annotations before submission deadline that gained additional annotations in the same GO aspect during the annotation accumulation phase), highlighting the need for a new class of methods. The rankings of the methods were stable over the years in traditional evaluation settings, but less so in the new partial knowledge evaluation. Overall, the field continues to progress with some influx of new participants. Sustained efforts will be necessary to substantially advance it.

Journal Article

GOtcha: a new method for prediction of protein function assessed by the annotation of seven genomes.

BACKGROUND: The function of a novel gene product is typically predicted by transitive assignment of annotation from similar sequences. We describe a novel method, GOtcha, for predicting gene product function by annotation with Gene Ontology (GO) terms. GOtcha predicts GO term associations with term-specific probability (P-score) measures of confidence. Term-specific probabilities are a novel feature of GOtcha and allow the identification of conflicts or uncertainty in annotation. RESULTS: The GOtcha method was applied to the recently sequenced genome for Plasmodium falciparum and six other genomes. GOtcha was compared quantitatively for retrieval of assigned GO terms against direct transitive assignment from the highest scoring annotated BLAST search hit (TOPBLAST). GOtcha exploits information deep into the 'twilight zone' of similarity search matches, making use of much information that is otherwise discarded by more simplistic approaches. At a P-score cutoff of 50%, GOtcha provided 60% better recovery of annotation terms and 20% higher selectivity than annotation with TOPBLAST at an E-value cutoff of 10(-4). CONCLUSIONS: The GOtcha method is a useful tool for genome annotators. It has identified both errors and omissions in the original Plasmodium falciparum annotation and is being adopted by many other genome sequencing projects.

Animals

Decoding the Functional Interactome of Non-Model Organisms with PHILHARMONIC.

Despite the widespread availability of genome sequencing pipelines, many genes remain part of the genome's "dark matter," where existing inference tools cannot even begin to guess the biological function of their proteins from sequence alone. This challenge is especially pronounced in organisms that are highly evolutionarily distant from well-studied models, where homology-based methods break down. Here, we describe PHILHARMONIC, a computational method that combines deep learning-based de novo protein interaction network inference with robust unsupervised spectral clustering and remote homology to illuminate functional organization in any non-model organism. From only a sequenced proteome, we show PHILHARMONIC predicts protein functions, functional communities, and higher-order network structure with high accuracy. We validate its performance using experimental gene expression and pathway data in D. melanogaster, and we demonstrate its broad utility by analyzing temperature sensing and stress response pathways in the reef-building coral P. damicornis and its algal symbiont C. goreaui. PHILHARMONIC provides a general-purpose engine for functional discovery and biological hypothesis generation in non-model organisms, enabling systems-level insights across the full diversity of life.

Journal Article

mettannotator: a comprehensive and scalable Nextflow annotation pipeline for prokaryotic assemblies.

SUMMARY: In recent years, there has been a surge in prokaryotic genome assemblies, coming from both isolated organisms and environmental samples. These assemblies often include novel species that are poorly represented in reference databases creating a need for a tool that can annotate both well-described and novel taxa, and can run at scale. Here, we present mettannotator-a comprehensive, scalable Nextflow pipeline for prokaryotic genome annotation that identifies coding and noncoding regions, predicts protein functions, including antimicrobial resistance, and delineates gene clusters. The pipeline summarizes these results in a GFF (General Feature Format) file that can be easily utilized in downstream analysis or visualized using common genome browsers. Here, we show how it works on 200 genomes from 29 prokaryotic phyla, including isolate genomes and known and novel metagenome-assembled genomes, and present metrics on its performance in comparison to other tools. AVAILABILITY AND IMPLEMENTATION: The pipeline is written in Nextflow and Python and published under an open source Apache 2.0 licence. Instructions and source code can be accessed at https://github.com/EBI-Metagenomics/mettannotator. The pipeline is also available on WorkflowHub: https://workflowhub.eu/workflows/1069.

Software

Challenges and Opportunities in Analyzing Cancer-Associated Microbiomes.

The study of cancer-associated microbiomes has gained significant attention in recent years, spurred by advances in high-throughput sequencing and metagenomic analysis. Microbiome research holds promise for identifying noninvasive biomarkers and possibly new paradigms for cancer treatment. In this review, we explore the key computational challenges and opportunities in analyzing cancer-associated microbiomes (in tumor/normal tissues and other body sites, e.g., gut, oral, and skin), focusing on sequencing-driven strategies and associated considerations for taxonomic and functional characterization. The discussion covers the strengths and limitations of current analysis tools for identifying contamination, determining compositional bias, and resolving species and strains, as well as the statistical, metabolic, and network inferences that are essential to uncover host-microbiome interactions. Several key considerations are required to guide the choice of databases used for metagenomic analysis in such studies. Recent advances in spatial and single-cell technologies have provided insights into cancer-associated microbiomes, and Artificial Intelligence-driven protein function prediction might enable rapid advances in this field. Finally, we provide a perspective on how the field can evolve to manage the ever-growing size of datasets and generate robust and testable hypotheses. This article is part of a special series: Driving Cancer Discoveries with Computational Research, Data Science, and Machine Learning/AI .

Humans

DescribePROT Database of Residue-Level Protein Structure and Function Annotations.

DescribePROT is a freely available online database of structural and functional descriptors of proteins at the amino acid level. It provides access to 13 diverse descriptors that include sequence conservation, putative secondary structure, solvent accessibility, intrinsic disorder, and signal peptides, and putative annotations of residues that interact with proteins, peptides and nucleic acids. These data can be used to elucidate protein functions, to support efforts to develop therapeutics, and to develop and evaluate future predictors of protein structure and function. DescribePROT includes 7.8 billion predictions for 1.4 million proteins from 83 complete proteomes of popular model organisms. This information can be downloaded at multiple levels of scope (entire database, specific organisms, and individual proteins) and can be interacted with using a graphical interface that simultaneously displays data on multiple descriptors. We describe the contents of this resource, provide directions on how to use its interface, and offer instructions on how to obtain and interact with the underlying data. Moreover, we briefly discuss plans for a future expansion of this database. DescribePROT is available at http://biomine.cs.vcu.edu/servers/DESCRIBEPROT/ .

Databases, Protein

Complete genome sequence of a novel alternavirus infecting Fusarium falciforme.

We present the complete genome sequence of a novel alternavirus, tentatively named "Fusarium falciforme alternavirus 1 (FfAV1)", isolated from Fusarium falciforme. The host, F. falciforme strain Fod375, was isolated from a soil sample in Spain in 2012 and was found to be infected with a virus containing a tetra-segmented double-stranded (ds) RNA genome. The genome segments, designated as dsRNA1 (3529 bp), dsRNA2 (2641 bp), dsRNA3 (2459 bp), and dsRNA4 (1471 bp), each possess a single open reading frame (ORF). The protein predicted from dsRNA1 contains the typical domains of an RNA-dependent RNA polymerase (RdRP) homologous to those of previously reported alternaviruses, while the protein predicted from dsRNA3 shows homology to alternavirus capsid proteins. The proteins encoded by dsRNA2 and dsRNA4 are of unknown function. All predicted proteins exhibited the highest sequence identity with their counterparts in Hebei alternavirus and Marquandomyces marquandii alternavirus 1. Phylogenetic analysis supported the placement of this FfAV1 isolate within the genus Alternavirus. Considering these results, we propose that FfAV1, along with the two closely related unassigned alternaviruses, represents a new species within the genus.

Genome, Viral

How Not to be Seen: Predicting Unseen Enzyme Functions using Contrastive Learning.

MOTIVATION: Predicting enzyme function from its sequence is still an unsolved problem in the life sciences. Moreover, with the explosion of annotated genome data, we are inundated with potential enzymatic sequences that have not yet been biochemically characterized. While it is not possible to assign a not-yet-existing label to such a sequence, there is high value in placing the sequence as accurately as possible in known function space. Doing so can help provide more accurate falsifiable hypotheses for experimentalists wishing to characterize enzymes from specific functional families. RESULTS: Here we present a contrastive learning algorithm for predicting enzyme function from sequence. Our method, EnzPlacer, predicts the third, second, and first EC numbers for a protein whose fourth EC number is not in the training corpus. This novel prediction mechanism accurately places a protein sequence within a narrowed-down functional context, even if the precise function remains unknown. AVAILABILITY: EnzPlacer is available from https://github.com/drxiangma/EnzPlacer under a GPL3 license.

Contrastive learning

Phenotypic pleiotropy of missense variants in human B cell confinement receptor P2RY8.

Missense variants can have pleiotropic effects on protein function, and predicting these effects can be difficult. We performed near-saturation deep mutational scanning of P2RY8, a G protein-coupled receptor that promotes germinal center B cell confinement. We assayed the effect of each variant on surface expression, migration, and proliferation. We delineated variants that affected both expression and function, affected function independently of expression, and discrepantly affected migration and proliferation. We also used cryo-electron microscopy to determine the structure of activated, ligand-bound P2RY8, providing structural insights into the effects of variants on ligand binding and signal transmission. We applied the deep mutational scanning results to both improve computational variant effect predictions and to characterize the phenotype of germline variants and lymphoma-associated variants. Together, our results demonstrate the power of integrating deep mutational scanning, structure determination, and in silico prediction to advance the understanding of a receptor important in human health.

Humans

Deciphering the genetic background of an industrial 2-ketogluconic acid-producing strain Pseudomonas plecoglossicida JUIM01 using whole-genome sequencing.

2-Ketogluconic acid (2KGA) is an important precursor for the food antioxidant erythorbic acid, currently produced via microbial fermentation using Pseudomonas species. To facilitate the genetic improvement of production strains, the complete genome of an industrial 2KGA producer P. plecoglossicida JUIM01 was sequenced and analyzed. The genome consists of a 5.13-Mb circular chromosome with a GC content of 63.58%, encoding 4,517 predicted proteins. Comprehensive functional annotation identified a putative global regulatory network comprising 75 core regulators, which were classified into six functionally cooperative modules, potentially governing the strain's metabolism and environmental adaptability. We further delineated the genetic determinants hypothetically linked to efficient 2KGA synthesis, including glucose metabolism, fatty acid metabolism, and the oxidative phosphorylation system. These outputs could provide the genomic resource for elucidating high productivity and robustness, and rationally engineering the high-performance chassis cells toward robust 2KGA production.

P. plecoglossicida

Gastrointestinal digestion governs insect protein hydrolysis and predicted bioactive peptide release: Species-dependent implications for functional food applications.

This study investigates the digestion of insect proteins and the release of predicted bioactive peptides during human gastrointestinal digestion. Using the Infogest in vitro model, mealworm, cricket, and black soldier fly larvae (BSFL) proteins were digested and analyzed through discovery proteomics and bioinformatics to identify predicted bioactive peptides. Sequential windowed acquisition of all theoretical fragment ion mass spectra (SWATH-MS) quantified insect proteins including predicted bioactive peptide precursor proteins, the precursors of predicted bioactive peptides. Results indicated that gastrointestinal digestion strongly influences peptide release, with the gastric phase exhibiting a richer predicted bioactive peptide profile than the small intestinal phase. Many predicted bioactive peptides were rapidly hydrolysed under small intestine conditions, which may lead to reduced stability or diminished activity in vivo, potentially explaining why certain peptides show strong bioactivity in vitro but limited effects in vivo. Additionally, predicted bioactive peptide release varied by insect species, influenced by genetic factors and peptide abundance. These findings highlight the importance of species selection and consideration of proteolytic digestion patterns in optimizing insect-derived bioactive peptides for functional foods and nutraceutical applications.

Animals

Prophage landscapes in clinical MRSA: safety profiling and discovery of Lys81, a broad-spectrum bacteriolytic enzyme.

INTRODUCTION: Methicillin-resistant Staphylococcus aureus (MRSA) poses a significant threat to global healthcare, requiring novel therapeutic strategies. Prophages, latent phage genomes integrated into bacterial chromosomes, are important resources for antimicrobial development due to their genomic stability and genetic engineering potential. METHODS: In this study, we performed genomewide sequencing on 329 MRSA isolates to predict prophage sequences, followed by analyses of these prophages-including examinations of virulence genes, antibiotic resistance genes, homologous proteins of pathogenic MRSA phages, and functional predictions of these homologous proteins-to evaluate their safety and value as genetic engineering scaffolds and to screen for novel broadspectrum bacteriolytic enzymes. RESULTS: Our data indicate that 85.7% (282/329) of strains carried complete prophage sequences; 64 strains lacked virulence factors or genes, meeting the core criteria for safe vectors. Resistance screening found only 6 prophages carried msrA, confirming the biosafety of the remaining strains. A significant correlation existed between prophage virulence gene capacity and genomic structure (R2 = 0.99986684, p = 3.64e-69). High-virulence clusters (>10 factors) showed high structural similarity; 10 characteristic sequences linked to S. aureus phages and their prevalence patterns were identified via conserved motif analysis. Collinearity analysis with reference to virulent MRSA phages and 3D structural predictions of orthologous proteins identified two lysozymes and a host-recognition device. Notably, Lys81, an N-acetylmuramoyl-L-alanine amidase ortholog, was prioritized and characterized as a broad-spectrum lytic enzyme. Our data show Lys81 has key properties: (1) Broad-spectrum antibacterial activity, lysing 52.3% (23/44) of clinical S. aureus strains and cross-acting against Gram-positive bacteria such as Pseudomonas aeruginosa and Listeria; (2) Excellent environmental adaptability, maintaining activity at pH 5.0 and 0°C, with 25 mM Na+ and Ca2 + enhancing function; (3) Potent biofilm clearance, achieving 83% MRSA biofilm reduction at 50 μg/mL; and (4) Favorable in vivo safety/efficacy, eradicating MRSA infections in lung organoid models with minimal cytotoxicity. DISCUSSION: This study establishes a theoretical foundation for the clinical translation of MRSA prophages, positioning Lys81 as a novel candidate for treating drug-resistant bacterial infections.

Lys81

Searching for New Genes That Cause Usher Syndrome.

PURPOSE: The purpose of this project was to identify novel Usher syndrome (USH) candidate genes from phenotyping data of 9139 knockout (KO) mouse lines. METHODS: We evaluated phenotype data for concurrent retinopathy and hearing abnormalities in single-gene KO mice generated by the International Mouse Phenotyping Consortium (IMPC). A search was performed to determine whether each gene had been previously associated with retinopathy and/or deafness in humans. Bioinformatic tools were used to predict protein interactions, molecular functions, signaling pathways, and the expression of human orthologues of candidate genes in the retina and inner ear. RESULTS: We identified 18 single-gene KO lines exhibiting hearing abnormality and retinopathy after ear and eye examinations, respectively, and/or by histopathology. The molecular functions and signaling pathways of the human orthologues of the 18 candidate genes partially overlapped with those of USH genes. Particularly, FER and DYRK1B proteins were predicted to interact with proteins encoded by known ciliopathy genes. ADIPOR1, ATP8B1, and MPDZ were associated with retinal degeneration in humans. CHSY1 and IDUA may be pathogenic causes of hearing impairment in people. Furthermore, CHSY1, CSTB, and SPRED1 were located adjacent to unsolved genetic loci related to USH. CONCLUSIONS: A screen of 9139 KO mouse lines revealed 18 candidate genes exhibiting both retinal and inner ear abnormalities consistent with the principal clinical features associated with USH. As the observed phenotypes are attributed to gene deletion in mice, these genes warrant further study to determine the causation of retinal degeneration and hearing loss in patients.

Animals

Conserved protein folds underpin the diversification of secreted proteins in a fungal pathogen.

BACKGROUND: During host colonization, fungal plant pathogens secrete effector-like proteins that alter host cell physiology and target plant-associated microbes. However, rapid evolution and low sequence conservation hinder the study and characterization of these proteins. The fungus Zymoseptoria passerinii infects Hordeum spp. and includes lineages adapted to wild and domesticated barley. To date, the evolution of effector-like proteins in this species has not been addressed. RESULTS: We combined multiple structure-based and network analyses to unravel the secretome of Z. passerinii. We first compared AlphaFold2 and ESMFold predictions to establish the baseline for structural analyses. We identified 72 structural clusters in the secretome, revealing fold-level relationships across divergent sequences. We showed that effector-like proteins with predicted host immune-interfering functions evolved from a limited group of protein folds, whereas proteins with predicted antimicrobial properties were distributed across fold groups. Physicochemical comparisons indicate that putative antimicrobial effectors predominantly emerged through amino acid replacements on common effector-enriched scaffolds in Z. passerinii, reconfiguring surface charge and electrostatics. We analyzed intra- and interspecific variation in selected effector-enriched families by comparing Z. passerinii proteins and homologs across the genus Zymoseptoria. We describe constrained core folds, with local variation in loop and surface-exposed regions, consistent with fold stability while still enabling protein diversification. We further report that putative antimicrobial effector homologs are broadly distributed across the genus despite sequence divergence. CONCLUSIONS: The secretome of Z. passerinii is organized around common structural folds that support diverse biological roles, including host manipulation and host-associated microbial interactions. Conserved scaffolds combined with surface and physicochemical variation likely contribute to rapid adaptive evolution of effector-like proteins in Z. passerinii.

Fungal Proteins

Large language models in bioinformatics: a comprehensive survey.

The emergence of foundation models with trillion-level parameters has redefined the landscape of artificial intelligence. Various fields are developing their own large-scale models, which can solve many problems within the field and improve work efficiency. Biological large-scale models are a cross-disciplinary research field that combines mathematics, computer science, and biology, aiming to simulate and understand the structure, function, and dynamic changes of biological systems through the establishment of complex computational models. This field covers multiple levels such as biological pathways, population dynamics, protein folding, etc., providing us with tools for deep exploration of the mysteries of life and applications in medicine, ecology, and other fields. This article reviews the background and research status of biological large-scale models, and discusses future directions. Large language models (LLMs) and other large-scale foundation models have rapidly advanced in recent years, enabling powerful representation learning and generation across text, sequences, and multimodal data. In bioinformatics and biomedicine, these models are increasingly used to analyze genomic sequences, infer protein properties and structures, support drug discovery, and integrate heterogeneous biomedical evidence. This survey reviews the basic principles of LLMs and summarizes representative applications in (i) gene and genome sequence analysis, (ii) protein structure and function prediction, and (iii) drug design, including virtual screening and personalized medicine. We also discuss emerging multi-model modeling approaches, as well as key challenges such as data quality and privacy, interpretability, generalization to new organisms and tasks, and responsible deployment in health-related settings. Finally, we outline future directions for developing reliable, scalable, and explainable bioinformatics foundation models.

bioinformatics

Two new cases of fragile X syndrome without CGG triplet expansion. Clinical-molecular characterization and review of the literature.

Fragile X syndrome (FXS) is classically caused by CGG repeat expansion in the FMR1 gene leading to gene silencing. We describe two unrelated patients with clinical features consistent with FXS but harbouring distinct molecular mechanisms. The first patient had a pathogenic intronic variant in FMR1, predicted to impair protein function, and no repeat expansion. The second patient carried a complete deletion of the FMR1 gene, resulting in loss of gene expression. Both individuals presented with developmental delay, intellectual disability, and behavioural manifestations typical of FXS. These cases expand the causal heterogeneity underlying a clinically recognizable phenotype. They reinforce the concept that comprehensive molecular testing beyond repeat expansion analysis is mandatory in individuals with a strong clinical suspicion of FXS, in absence of a typical CGG triplet expansion.

Child, Preschool

Genomic and Structural Analysis of Gamete Recognition Proteins in a Broadcast Spawning Echinoderm Mesocentrotus franciscanus.

Gamete recognition proteins are expressed on the surfaces of sperm and eggs, where they mediate interactions between gametes. The genetic basis for gamete recognition proteins, as well as their structure and interactions, have yet to be fully resolved. Using a new high-quality de novo genome assembly for the sea urchin Mesocentrotus franciscanus, we investigated the genomic structure, expression, and protein forms of several gamete recognition proteins: sperm bindin, egg receptor for sperm (HSP110), and egg bindin receptor (EBR1), as well as the receptor for egg jelly (REJ) and its paralogs. To inform future population genetic and evolutionary studies, we resolve the genomic structure of the large EBR1 protein, identifying fewer tandem CUB-TSP1 repeats in EBR1 compared to the initial characterization of this protein. As expected for an egg receptor for sperm, EBR1 is highly expressed in female reproductive tissues (eggs and female gonad), compared to other tissues. In contrast, HSP110 shows similar levels of expression across male and female reproductive tissues, as well as across non-reproductive tissues and development stages. HSP110 might be a pleiotropic gene that in part influences fertilization. Using protein structural modeling and functional domain predictions, we propose hypotheses about potential interactions among EBR1, bindin, and HSP110 proteins that may provide insight into sperm-egg interactions in sea urchins. Resolving the genomic structure of genes encoding gamete recognition proteins, in combination with functional annotations and protein structural modeling, enables deeper investigation into the consequences of variation in gamete recognition proteins and the evolution of reproductive isolation.

Mesocentrotus franciscanus