Search PubMedSearch

SEARCH · Search PubMed

Results for “Protein function”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa

Modular Photoswitchable Molecular Glues for Chemo-Optogenetic Control of Protein Function in Living Cells.

Optogenetic systems using photosensitive proteins and chemically induced dimerization/proximity (CID/CIP) approaches enabled by chemical dimerizers (also termed molecular glues), are powerful tools to elucidate the dynamics of biological systems and to dissect complex biological regulatory networks. Here, we report a versatile chemo-optogenetic system using modular, photoswitchable molecular glues (sMGs) that can undergo repeated cycles of optical control to switch protein function on and off. We use molecular dynamics (MD) simulations to rationally design the sMGs and further expand their scope by incorporating different photoswitches, resulting in sMGs with customizable properties. We demonstrate that this system can be used to reversibly control protein localization, organelle positioning, protein-fragment complementation as well as posttranslational protein levels by light with high spatiotemporal precision. This system enables sophisticated optical manipulation of cellular processes and thus opens up a new avenue for chemo-optogenetics.

Optogenetics

On the state of protein function prediction: a report on the fourth CAFA challenge.

BACKGROUND: The Critical Assessment of Functional Annotation (CAFA) is a community effort held to understand the field of computational protein function prediction. Every three years, since 2010, the organizers initiate an experiment to collect function predictions on a large set of proteins and then evaluate the performance of predicting methods on a subset of proteins that have accumulated experimental annotations between the submission deadline and the evaluation time. CAFA provides an independent and rigorous assessment of the current state of the art, thus leveling the playing field, highlighting successes, revealing bottlenecks, and offering a forum for the exchange of ideas in protein science. Here, we report the results of the fourth CAFA experiment (CAFA4). RESULTS: CAFA4 featured the participation of 148 methods from 70 research groups on a total of 46,205 unique proteins over a 5-year annotation accumulation phase, the longest in any CAFA. In a comparison across CAFA2-CAFA4 methods, the prediction of Gene Ontology (GO) terms has clearly improved across all three GO aspects and traditional evaluation settings. While not achieving the first rank, several CAFA2 and CAFA3 methods featured in the top ten methods in many evaluations, suggesting that earlier methods still hold relevance. The performance is weaker in the newly introduced "partial knowledge" evaluation category (proteins with experimental annotations before submission deadline that gained additional annotations in the same GO aspect during the annotation accumulation phase), highlighting the need for a new class of methods. The rankings of the methods were stable over the years in traditional evaluation settings, but less so in the new partial knowledge evaluation. Overall, the field continues to progress with some influx of new participants. Sustained efforts will be necessary to substantially advance it.

Journal Article

Structural genomics sheds light on protein functions and remote homologs across the insect tree of life.

Protein structure bridges the sequence-function relationship, enabling deep exploration of biological processes across diverse organisms. Insects, the most diverse animal lineage, accounting for over 50% of all described animal species, provide an exceptional system for exploring sequence-structure-function relationships. Here, we reconstructed a comprehensive and well-resolved phylogeny of 4854 insects, spanning all orders. Leveraging this framework, we created an atlas of 13.29 million predicted protein structures from 824 representative species, including 11.63 million newly predicted structures. Structural clustering revealed that proteins with divergent sequences but similar structures could be effectively grouped together. Structural similarity searches against proteins with well-characterized functions yielded annotations for 7.61 million insect proteins, including up to 14% of previously unannotated proteins. We further identified 750 million remote homologs between insect proteins, many of which trace back to ancient branches of the insect phylogeny. Remarkably, despite extensive sequence divergence, cGAS-like receptors (cGLRs) were structurally conserved across all 824 insects. Experimental assays demonstrated that these structurally identified cGLRs play a crucial role in antiviral defense in the yellow fever mosquito. Our findings highlight the significance of structural genomics for understanding protein function and evolution across the tree of life.

Animals

Properties Governing Native State Entanglements and Relationships to Protein Function.

Non-covalent lasso entanglements are structural motifs found in a majority of globular proteins, and their misfolding has been linked to a range of biological consequences. Here, we characterize these motifs' structural and physicochemical properties, sequence biases, functional site correlations, and universal features across E. coli, S. cerevisiae, and H. sapiens. We find that the crossing residues, which pierce the plane of the entanglement loop, are 11-times more likely to be a β-strand than an α-helix or random coil, and that around this position the protein sequence is 2.5-times more likely to be composed of a stretch of all hydrophobic residues (most often Val, Ile, or Phe) compared to other sequence motifs. Functionally, crossing residues are enriched at enzyme active sites in S. cerevisiae and small molecule binding residues across all species to degrees greater than expected by random chance. Metal binding residues are enriched in these entanglements in H. sapiens. Increasing statistical power by pooling together these species data, we find RNA-binding residues are enriched in these entanglement components. On the other hand, there is a spatial depletion of crossing residues at sites involved in protein binding. Using machine learning, we identified eight robust features predictive of these entanglements, achieving AUROC scores of 0.8 across species. These results are significant because they suggest a direct role for components of native entanglements in particular protein functions, as well as identifying strong secondary structure and sequence preferences in native entanglements.

Humans

Peptide-phosphorodiamidate morpholino oligomer therapy for dysferlinopathy induces pseudoexon skipping and restoration of functional protein.

The dysferlinopathies are a spectrum of autosomal recessive muscle diseases caused by mutations in the dysferlin gene (DYSF). Clinical manifestations vary from asymptomatic hyperCKemia to severe muscle pathology and loss of muscle function. These are designated as limb-girdle muscular dystrophy type 2R (LGMDR2; formerly LGMD2B or Miyoshi myopathy). Among other functions, dysferlin is crucial for plasma membrane repair and maintenance of intracellular calcium homeostasis. In previous studies, we identified 2 independent point mutations deep within introns that cause aberrant DYSF mRNA splicing and the inclusion of pseudoexons within transcripts that diminish protein expression. In this study, we generated and characterized a mouse model for 1 of these mutations (within DYSF intron 44). In these mice, a segment of human DYSF DNA containing the mutant intronic sequence flanked by surrounding human exon sequences replaced the normal homologous mouse DNA. These mice exhibited aberrant Dysf pre-mRNA splicing, pseudoexon inclusion, loss of DYSF protein expression, and muscle pathology similar to that observed in patients. Using this model, we identified antisense oligonucleotides and a peptide-phosphorodiamidate morpholino oligomer that blocks the mouse Dysf pre-mRNA splicing complexes from binding the mutant pre-mRNA, thereby restoring nearly normal muscle pathology and function.

Animals

An Evolutionary Framework Exploiting Virologs and Their Host Origins to Inform Poxvirus Protein Functions.

Poxviruses represent evolutionary successful infectious agents. As a family, poxviruses can infect a wide variety of species including humans, fish, and insects. While many other viruses are species-specific, an individual poxvirus species is often capable of infecting diverse hosts and cell types. For example, the prototypical poxvirus, vaccinia, is well known to infect numerous human cell types but can also infect cells from divergent hosts like frog neurons. Notably, poxvirus infections result in both detrimental human and animal diseases. The most infamous disease linked to a poxvirus is smallpox caused by variola virus. Poxviruses are large double-stranded DNA viruses, which uniquely replicate in the cytoplasm of cells. The model poxvirus genome encodes ~200 nonoverlapping protein-coding open reading frames (ORFs). Poxvirus gene products impact various biological processes like the production of virus particles, the host range of infectivity, and disease pathogenesis. In addition, poxviruses and their gene products have biomedical application with several species commonly engineered for use as vaccines and oncolytic virotherapy. Nevertheless, we still have an incomplete understanding of the functions associated with many poxvirus genes. In this chapter, we outline evolutionary insights that can complement ongoing studies of poxvirus gene functions and biology, which may serve to elucidate new molecular activities linked to this biomedically relevant class of viruses.

Animals

GOtcha: a new method for prediction of protein function assessed by the annotation of seven genomes.

BACKGROUND: The function of a novel gene product is typically predicted by transitive assignment of annotation from similar sequences. We describe a novel method, GOtcha, for predicting gene product function by annotation with Gene Ontology (GO) terms. GOtcha predicts GO term associations with term-specific probability (P-score) measures of confidence. Term-specific probabilities are a novel feature of GOtcha and allow the identification of conflicts or uncertainty in annotation. RESULTS: The GOtcha method was applied to the recently sequenced genome for Plasmodium falciparum and six other genomes. GOtcha was compared quantitatively for retrieval of assigned GO terms against direct transitive assignment from the highest scoring annotated BLAST search hit (TOPBLAST). GOtcha exploits information deep into the 'twilight zone' of similarity search matches, making use of much information that is otherwise discarded by more simplistic approaches. At a P-score cutoff of 50%, GOtcha provided 60% better recovery of annotation terms and 20% higher selectivity than annotation with TOPBLAST at an E-value cutoff of 10(-4). CONCLUSIONS: The GOtcha method is a useful tool for genome annotators. It has identified both errors and omissions in the original Plasmodium falciparum annotation and is being adopted by many other genome sequencing projects.

Animals

Search for millimeter microwave effects on enzyme or protein functions.

Recent observations of nonthermal, resonant biological responses to weak millimeter microwave irradiation have led us to investigate whether similar influences exist on enzymatic functions in vitro. We chose (i) the reduction of ethanol in the presence of alcohol dehydrogenase and (ii) the cooperative binding of oxygen on hemoglobin. Using an irradiation intensity near 10 mW/cm2 the frequency was continuously varied from 40 to 115 GHz with a resolution of a few MHz. No microwave influences were detectable within our experimental sensitivity of about 0.1% of the reaction rate in (i), or of the amount of bound oxygen at half saturation in (ii).

Alcohol Oxidoreductases

Thermodynamic investigations of proteins. I. Standard functions for proteins with lysozyme as an example.

A direct method is proposed for obtaining thermodynamic standard functions for native and denatured proteins using experimental data from scanning calorimetry, isothermal calorimetry and potentiometric titrations. The possibility of this approach is demonstrated on the example of lysozyme in the range of pH 1.5-7.0 and temperature 0-100 degrees C. Tests for the validity of the obtained functions of enthalpy and entropy are presented in the form of cyclic processes using experimental data obtained from thermodynamically different pathways. The Gibbs function is checked by comparison with results of an independent method. The methodic problems in determining and checking standard functions for proteins are discussed in detail.

Calorimetry

FANTASIA suite: a reproducible and configurable framework for embedding-based functional annotation of proteins.

Embedding-based annotation transfer is increasingly used for protein function inference due to protein language models capture sequence, structural, and functional signals that may extend beyond conventional pairwise similarity. However, systematic application of these approaches requires control over model choice, reference composition, lookup parameters, evidence traceability, and output formats. We developed the FANTASIA suite, a configurable framework for embedding-based functional annotation of proteins. The suite combines a database-backed implementation for reproducible and extensible analyses with a portable flat-file implementation for rapid local annotation and pipeline integration. Using non-model and model-organism proteomes, we show that larger neighbourhood sizes remain practical for proteome-scale analyses and that taxonomy and sequence-identity filtering support leakage-aware benchmarking. We also compare the supported models with baseline methods through external CAFA5 evaluation and provide practical guidance based on empirical evidence variables. FANTASIA provides a controlled, scalable, and reproducible framework for extending functional annotation across the rapidly expanding diversity of sequenced organisms.

Software

Yeast proteins: recovery, nutritional and functional properties.

The future need for supplementary sources of food grade functional proteins is emphasized and the potential role of yeasts as a source of proteins is discussed. The problems of proteolysis and nucleic acid contamination can be circumvented by succinylation or derivatization of the yeast protein during extraction. Succinylated yeast proteins demonstrate improved functional properties though their nutritional value needs to be investigated.

Candida

Spacer-engineered donor DNA enhances CRISPR-Cas9-mediated knockin to establish a chemical knockdown platform for endogenous proteins.

Precise installation of functional protein domains at endogenous loci is a powerful approach for interrogating protein functions, but its broad application is limited by the low efficiency of homology-directed repair (HDR)-mediated knockin during Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-Cas9 gene editing. Here, we investigated a simple donor DNA engineering strategy that enhances HDR-mediated gene knockin by appending additional gRNA-recognizable spacer sequences to donor templates. Systematic analysis of linear dsDNA and plasmid donors showed that spacer position, length, and orientation influenced HDR efficiency, and that spacer-containing donors improved knockin across multiple genomic loci, insertion sizes, cell types and delivery modalities. Mechanistic analyses revealed that spacer-containing donors formed stable complexes with Cas9/gRNA and showed increased nuclear localization, supporting nuclear delivery as a key contributor to improved editing outcomes. We then applied this gene-editing strategy to establish a chemical knockdown platform by installing drug-responsive degrons at endogenous loci, generating cell lines in which GSK3β or Lin28A protein could be rapidly, potently and reversibly depleted by drug treatment. These platforms enable selective modulation of endogenous proteins and reveal cellular responses that may differ from those obtained using conventional genetic perturbation. Together, this work establishes a readily implementable framework that integrates improved gene editing with on-demand chemical knockdown of endogenous proteins.

CRISPR-Cas9

DescribePROT Database of Residue-Level Protein Structure and Function Annotations.

DescribePROT is a freely available online database of structural and functional descriptors of proteins at the amino acid level. It provides access to 13 diverse descriptors that include sequence conservation, putative secondary structure, solvent accessibility, intrinsic disorder, and signal peptides, and putative annotations of residues that interact with proteins, peptides and nucleic acids. These data can be used to elucidate protein functions, to support efforts to develop therapeutics, and to develop and evaluate future predictors of protein structure and function. DescribePROT includes 7.8 billion predictions for 1.4 million proteins from 83 complete proteomes of popular model organisms. This information can be downloaded at multiple levels of scope (entire database, specific organisms, and individual proteins) and can be interacted with using a graphical interface that simultaneously displays data on multiple descriptors. We describe the contents of this resource, provide directions on how to use its interface, and offer instructions on how to obtain and interact with the underlying data. Moreover, we briefly discuss plans for a future expansion of this database. DescribePROT is available at http://biomine.cs.vcu.edu/servers/DESCRIBEPROT/ .

Databases, Protein

An Arabidopsis Protein-Flavonoid Interactome Identifies Peroxiredoxin A as a Candidate for Flavonoid Action in Chloroplasts.

The ability of phytochemicals to act as small molecule effectors of protein function is a largely overlooked dimension of plant biochemistry. This is particularly true for the ubiquitous flavonoids where, despite abundant examples of functional interactions with human proteins, biological activities in plants are primarily attributed to ROS scavenging. We used affinity capture to explore the protein interactome of the flavonoid glycoside, rutin, in Arabidopsis seedlings. Unexpectedly, the 397 high-confidence candidates included numerous proteins associated with chloroplasts, where flavonoids are present at exceedingly low levels. Intriguingly, several identified targets are conserved with known flavonoid-interacting proteins in mammals, where the bioavailability of flavonoids is similarly low. Using one of these, the Arabidopsis plastidial 2-cys peroxiredoxin A, as a test case, this study substantiated the potential of affinity proteomics for identifying novel protein targets of phytochemicals and suggests that flavonoids modulate protein function in plants to a larger extent than previously suspected.

Arabidopsis