Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

PotatoRTD and TomatoRTD: Comprehensive Reference Transcript Datasets for Accurate Transcriptome Analysis and Isoform Discovery.

Transcriptome annotations provide essential information on transcript locations, sequences and structures, including transcription start, end sites and splice junctions. They underpin key biological analyses such as gene and transcript quantification, and the study of transcriptional and post-transcriptional regulation, including alternative transcription initiation, polyadenylation and splicing. Accurate characterisation of transcript isoforms is critical for understanding how gene expression relates to functional protein products. However, for many species-including Solanaceae crops such as potato and tomato-current annotations suffer from limited isoform coverage, with tens or hundreds of thousands of splice junctions and transcript isoforms missing. This undermines the completeness and accuracy of transcript-level analyses. Here, by generating Iso-seq and RNA-seq on a range of tissues and samples, we have produced transcriptome annotations for both potato and tomato with improved coverage, diversity, accurate splice junctions, and transcript start and end sites. We have also made these high-quality resources accessible through genome browsers. These enhanced annotations will enable more accurate transcriptome analyses, supporting higher-resolution and novel biological discoveries.

Solanum tuberosum↗

Electron transport in the pathway of acetate conversion to methane in the marine archaeon Methanosarcina acetivorans.

A liquid chromatography-hybrid linear ion trap-Fourier transform ion cyclotron resonance mass spectrometry approach was used to determine the differential abundance of proteins in acetate-grown cells compared to that of proteins in methanol-grown cells of the marine isolate Methanosarcina acetivorans metabolically labeled with 14N versus 15N. The 246 differentially abundant proteins in M. acetivorans were compared with the previously reported 240 differentially expressed genes of the freshwater isolate Methanosarcina mazei determined by transcriptional profiling of acetate-grown cells compared to methanol-grown cells. Profound differences were revealed for proteins involved in electron transport and energy conservation. Compared to methanol-grown cells, acetate-grown M. acetivorans synthesized greater amounts of subunits encoded in an eight-gene transcriptional unit homologous to operons encoding the ion-translocating Rnf electron transport complex previously characterized from the Bacteria domain. Combined with sequence and physiological analyses, these results suggest that M. acetivorans replaces the H2-evolving Ech hydrogenase complex of freshwater Methanosarcina species with the Rnf complex, which generates a transmembrane ion gradient for ATP synthesis. Compared to methanol-grown cells, acetate-grown M. acetivorans synthesized a greater abundance of proteins encoded in a seven-gene transcriptional unit annotated for the Mrp complex previously reported to function as a sodium/proton antiporter in the Bacteria domain. The differences reported here between M. acetivorans and M. mazei can be attributed to an adaptation of M. acetivorans to the marine environment.

Acetates↗

New local potential useful for genome annotation and 3D modeling.

A new potential energy function representing the conformational preferences of sequentially local regions of a protein backbone is presented. This potential is derived from secondary structure probabilities such as those produced by neural network-based prediction methods. The potential is applied to the problem of remote homolog identification, in combination with a distance-dependent inter-residue potential and position-based scoring matrices. This fold recognition jury is implemented in a Java application called JThread. These methods are benchmarked on several test sets, including one released entirely after development and parameterization of JThread. In benchmark tests to identify known folds structurally similar to (but not identical with) the native structure of a sequence, JThread performs significantly better than PSI-BLAST, with 10% more structures identified correctly as the most likely structural match in a fold library, and 20% more structures correctly narrowed down to a set of five possible candidates. JThread also improves the average sequence alignment accuracy significantly, from 53% to 62% of residues aligned correctly. Reliable fold assignments and alignments are identified, making the method useful for genome annotation. JThread is applied to predicted open reading frames (ORFs) from the genomes of Mycoplasma genitalium and Drosophila melanogaster, identifying 20 new structural annotations in the former and 801 in the latter.

Animals↗

Annotating proteins from endoplasmic reticulum and Golgi apparatus in eukaryotic proteomes.

The sub-cellular localization of a native protein constitutes one coarse-grained aspect of its function. Transport between compartments is often regulated through short sequence motifs. Here, we analyzed experimentally characterized endoplasmic reticulum (ER)/ Golgi retrieval motifs and investigated the accuracy of homology-transfer. Only the C-terminal ER retrieval motifs KDEL, HDEL and AIAKE were sufficiently specific. However, even unspecific motifs may help, provided we know the probability for localization given the motif. We provided such estimates. We also rigorously estimated the accuracy and coverage for inferring ER and Golgi localization through homology-transfer by sequence similarity. In entire proteomes, we could thereby annotate 3304 ER (3182 membrane) and 1853 Golgi (759 membrane) proteins. We identified another putative 5157 globular and 3941 membrane ER or Golgi proteins. Each experimental annotation yielded, on average, one to three high-accuracy and five to six low-accuracy homology-transfers in the six proteomes. These numbers will increase with each new experimental annotation.

Amino Acid Motifs↗

A human protein-protein interaction network: a resource for annotating the proteome.

Protein-protein interaction maps provide a valuable framework for a better understanding of the functional organization of the proteome. To detect interacting pairs of human proteins systematically, a protein matrix of 4456 baits and 5632 preys was screened by automated yeast two-hybrid (Y2H) interaction mating. We identified 3186 mostly novel interactions among 1705 proteins, resulting in a large, highly connected network. Independent pull-down and co-immunoprecipitation assays validated the overall quality of the Y2H interactions. Using topological and GO criteria, a scoring system was developed to define 911 high-confidence interactions among 401 proteins. Furthermore, the network was searched for interactions linking uncharacterized gene products and human disease proteins to regulatory cellular pathways. Two novel Axin-1 interactions were validated experimentally, characterizing ANP32A and CRMP1 as modulators of Wnt signaling. Systematic human protein interaction screens can lead to a more comprehensive understanding of protein function and cellular processes.

Axin Protein↗

UTRdb and UTRsite: a collection of sequences and regulatory motifs of the untranslated regions of eukaryotic mRNAs.

The 5' and 3' untranslated regions of eukaryotic mRNAs play crucial roles in the post-transcriptional regulation of gene expression through the modulation of nucleo-cytoplasmic mRNA transport, translation efficiency, subcellular localization and message stability. UTRdb is a curated database of 5' and 3' untranslated sequences of eukaryotic mRNAs, derived from several sources of primary data. Experimentally validated functional motifs are annotated (and also collated as the UTRsite database) and cross-links to genomic and protein data are provided. The integration of UTRdb with genomic and protein data has allowed the implementation of a powerful retrieval resource for the selection and extraction of UTR subsets based on their genomic coordinates and/or features of the protein encoded by the relevant mRNA (e.g. GO term, PFAM domain, etc.). All internet resources implemented for retrieval and functional analysis of 5' and 3' untranslated regions of eukaryotic mRNAs are accessible at http://www.ba.itb.cnr.it/UTR/.

3' Untranslated Regions↗

Redefining ALS: Large-scale proteomic profiling reveals a prolonged pre-diagnostic phase with immune, muscular, metabolic, and brain involvement.

BACKGROUND: Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disorder with a largely unknown duration and pathophysiology of the pre-diagnostic phase, especially for the common non-monogenic form. METHODS: We leveraged the European Prospective Investigation into Cancer and Nutrition (EPIC) cohort with up to 30 years of follow-up to identify incident ALS cases across five European countries. Pre-diagnostic plasma samples from initially healthy participants underwent high-throughput proteomic profiling (7,285 protein markers, SomaScan). Cox proportional hazards models based on 4,567 participants (including 172 incident ALS cases) were used to identify protein biomarkers associated with future ALS diagnosis. Top results were indirectly validated in two independent case-control studies of prevalent ALS (n=417 ALS, 852 controls). Functional annotation included cross-disease comparisons, gene set and tissue enrichment testing, organ-specific proteomic clocks, and the application of large-language models (LLM). FINDINGS: Five proteins (SECTM1, CA3, THAP4, KLHL41, SLC26A7) were identified as significant pre-diagnostic ALS biomarkers (FDR=0.05), detectable approximately two decades before diagnosis. Of these, all except SECTM1 were indirectly validated in independent cohorts of prevalent ALS cases, supporting their clinical significance. Additionally, 22 nominally significant (p<0.05) pre-diagnostic biomarkers were FDR-significant in prevalent ALS with consistent effect directions. Cross-disease comparisons with pre-diagnostic Parkinson's and Alzheimer's disease suggested a largely specific pre-diagnostic ALS biomarker signature. Gene ontology and tissue enrichment highlighted early involvement of immune, muscle, metabolic, and digestive processes. Furthermore, analyses of proteomic clocks revealed accelerated aging in brain-cognition, immune, and muscle tissues before clinical diagnosis. Druggability and LLM analyses revealed possible therapeutic targets and novel strategies, emphasizing translational relevance. INTERPRETATION: Our study provides first evidence of ultra-early molecular changes in common ALS up to two decades prior to clinical onset, mainly affecting immune, muscle, metabolic, digestive, and cognitive systems. Our study nominates several compelling candidates for risk stratification studies and novel therapeutic targets for early intervention. FUNDING: Clinical Research in ALS and Related Disorders for Therapeutic Development (CreATe) Consortium, Cure Alzheimer's Fund, Michael J Fox Foundation, Interdisciplinary Centre for Clinical Research, University M&#xfc;nster.

Journal Article↗

Redefining ALS: Large-scale proteomic profiling reveals a prolonged pre-diagnostic phase with immune, muscular, metabolic, and brain involvement.

BACKGROUND: Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disorder with a largely unknown duration and pathophysiology of the pre-diagnostic phase, especially for the common non-monogenic form. METHODS: We leveraged the European Prospective Investigation into Cancer and Nutrition (EPIC) cohort with up to 30 years of follow-up to identify incident ALS cases across five European countries. Pre-diagnostic plasma samples from initially healthy participants underwent high-throughput proteomic profiling (7,285 protein markers, SomaScan). Cox proportional hazards models based on 4,567 participants (including 172 incident ALS cases) were used to identify protein biomarkers associated with future ALS diagnosis. Top results were indirectly validated in two independent case-control studies of prevalent ALS (n=417 ALS, 852 controls). Functional annotation included cross-disease comparisons, gene set and tissue enrichment testing, organ-specific proteomic clocks, and the application of large-language models (LLM). FINDINGS: Five proteins (SECTM1, CA3, THAP4, KLHL41, SLC26A7) were identified as significant pre-diagnostic ALS biomarkers (FDR=0.05), detectable approximately two decades before diagnosis. Of these, all except SECTM1 were indirectly validated in independent cohorts of prevalent ALS cases, supporting their clinical significance. Additionally, 22 nominally significant (p<0.05) pre-diagnostic biomarkers were FDR-significant in prevalent ALS with consistent effect directions. Cross-disease comparisons with pre-diagnostic Parkinson's and Alzheimer's disease suggested a largely specific pre-diagnostic ALS biomarker signature. Gene ontology and tissue enrichment highlighted early involvement of immune, muscle, metabolic, and digestive processes. Furthermore, analyses of proteomic clocks revealed accelerated aging in brain-cognition, immune, and muscle tissues before clinical diagnosis. Druggability and LLM analyses revealed possible therapeutic targets and novel strategies, emphasizing translational relevance. INTERPRETATION: Our study provides first evidence of ultra-early molecular changes in common ALS up to two decades prior to clinical onset, mainly affecting immune, muscle, metabolic, digestive, and cognitive systems. Our study nominates several compelling candidates for risk stratification studies and novel therapeutic targets for early intervention. FUNDING: Clinical Research in ALS and Related Disorders for Therapeutic Development (CreATe) Consortium, Cure Alzheimer's Fund, Michael J Fox Foundation, Interdisciplinary Centre for Clinical Research, University M&#xfc;nster.

Journal Article↗

Ontology annotation: mapping genomic regions to biological function.

With numerous whole genomes now in hand, and experimental data about genes and biological pathways on the increase, a systems approach to biological research is becoming essential. Ontologies provide a formal representation of knowledge that is amenable to computational as well as human analysis, an obvious underpinning of systems biology. Mapping function to gene products in the genome consists of two, somewhat intertwined enterprises: ontology building and ontology annotation. Ontology building is the formal representation of a domain of knowledge; ontology annotation is association of specific genomic regions (which we refer to simply as 'genes', including genes and their regulatory elements and products such as proteins and functional RNAs) to parts of the ontology. We consider two complementary representations of gene function: the Gene Ontology (GO) and pathway ontologies. GO represents function from the gene's eye view, in relation to a large and growing context of biological knowledge at all levels. Pathway ontologies represent function from the point of view of biochemical reactions and interactions, which are ordered into networks and causal cascades. The more mature GO provides an example of ontology annotation: how conclusions from the scientific literature and from evolutionary relationships are converted into formal statements about gene function. Annotations are made using a variety of different types of evidence, which can be used to estimate the relative reliability of different annotations.

Animals↗

Annotations and functional analyses of the rice WRKY gene superfamily reveal positive and negative regulators of abscisic acid signaling in aleurone cells.

The WRKY proteins are a superfamily of regulators that control diverse developmental and physiological processes. This family was believed to be plant specific until the recent identification of WRKY genes in nonphotosynthetic eukaryotes. We have undertaken a comprehensive computational analysis of the rice (Oryza sativa) genomic sequences and predicted the structures of 81 OsWRKY genes, 48 of which are supported by full-length cDNA sequences. Eleven OsWRKY proteins contain two conserved WRKY domains, while the rest have only one. Phylogenetic analyses of the WRKY domain sequences provide support for the hypothesis that gene duplication of single- and two-domain WRKY genes, and loss of the WRKY domain, occurred in the evolutionary history of this gene family in rice. The phylogeny deduced from the WRKY domain peptide sequences is further supported by the position and phase of the intron in the regions encoding the WRKY domains. Analyses for chromosomal distributions reveal that 26% of the predicted OsWRKY genes are located on chromosome 1. Among the dozen genes tested, OsWRKY24, -51, -71, and -72 are induced by abscisic acid (ABA) in aleurone cells. Using a transient expression system, we have demonstrated that OsWRKY24 and -45 repress ABA induction of the HVA22 promoter-beta-glucuronidase construct, while OsWRKY72 and -77 synergistically interact with ABA to activate this reporter construct. This study provides a solid base for functional genomics studies of this important superfamily of regulatory genes in monocotyledonous plants and reveals a novel function for WRKY genes, i.e. mediating plant responses to ABA.

Abscisic Acid↗

Biological fingerprinting analysis of the interactome of a kinase inhibitor in human plasma by a chemiproteomic approach.

In this study, a gel free chemiproteomic method based on chromatography was developed and applied for the biological fingerprinting analysis of complex biological system. p-Aminobenzamidine (ABA), an inhibitor of trypsin-like serine proteases, was immobilized for characterizing their interacting proteins in human plasma. By the proteomic analysis method, 214 proteins were identified with obvious affinity to the immobilized ABA. By searching the sequences of above proteins with consensus patterns of the two active sites, seven proteins belong to trypsin-like serine protease group were found. Based on the Gene Ontology annotation, the identified trypsin-like serine proteases have the function of catalytic activity and calcium ion binding, and are mainly involved in the biological process of blood coagulation. Eight more other proteins related to calcium ion binding and blood coagulation were found. Nearly all of these proteins cannot be identified by directly analyzing the plasma sample demonstrating the chemiproteomics a useful approach to characterize interacting proteins in the low abundance range.

Adult↗

Translational polymorphism as a potential source of plant proteins variety in Arabidopsis thaliana.

MOTIVATION: According to scanning model, 40S ribosomal subunits can either initiate translation at start AUG codon in suboptimal context or miss it and initiate translation at downstream AUG(s), thereby producing several proteins. Functional significance of such a protein translational polymorphism is still unknown. RESULTS: We compared predicted subcellular localizations of annotated Arabidopsis thaliana proteins and their potential N-terminally truncated forms started from the nearest downstream in-frame AUG codons. It was found that localizations of full and N-truncated proteins differ in many cases: 12.2% of N-truncated proteins acquired sorting signals de novo and 5.7% changed their predicted subcellular locations (mitochodria, chloroplast or secretory pathway). It is likely that the in-frame downstream AUGs may be frequently utilized to synthesize proteins possessing new functional properties and such a translational polymorphism may serve as an important source of cellular and organelle proteomes.

Arabidopsis↗

Database searching by flexible protein structure alignment.

We have recently developed a flexible protein structure alignment program (FATCAT) that identifies structural similarity, at the same time accounting for flexibility of protein structures. One of the most important applications of a structure alignment method is to aid in functional annotations by identifying similar structures in large structural databases. However, none of the flexible structure alignment methods were applied in this task because of a lack of significance estimation of flexible alignments. In this paper, we developed an estimate of the statistical significance of FATCAT alignment score, allowing us to use it as a database-searching tool. The results reported here show that (1) the distribution of the similarity score of FATCAT alignment between two unrelated protein structures follows the extreme value distribution (EVD), adding one more example to the current collection of EVDs of sequence and structure similarities; (2) introducing flexibility into structure comparison only slightly influences the sensitivity and specificity of identifying similar structures; and (3) the overall performance of FATCAT as a database searching tool is comparable to that of the widely used rigid-body structure comparison programs DALI and CE. Two examples illustrating the advantages of using flexible structure alignments in database searching are also presented. The conformational flexibilities that were detected in the first example may be involved with substrate specificity, and the conformational flexibilities detected in the second example may reflect the evolution of structures by block building.

Databases, Protein↗

Plasmodium permeomics: membrane transport proteins in the malaria parasite.

Membrane transport proteins are integral membrane proteins that mediate the passage across the membrane bilayer of specific molecules and/or ions. Such proteins serve a diverse range of physiological roles, mediating the uptake of nutrients into cells, the removal of metabolic wastes and xenobiotics (including drugs), and the generation and maintenance of transmembrane electrochemical gradients. In this chapter we review the present state of knowledge of the membrane transport mechanisms underlying the cell physiology of the intraerythrocytic malaria parasite and its host cell, considering in particular physiological measurements on the parasite and parasitized erythrocyte, the annotation of transport proteins in the Plasmodium genome, and molecular methods used to analyze transport protein function.

Animals↗

The dual nature of human extracellular superoxide dismutase: one sequence and two structures.

Human extracellular superoxide dismutase (EC-SOD; EC 1.15.1.1) is a scavenger of superoxide anions in the extracellular space. The amino acid sequence is homologous to the intracellular counterpart, Cu/Zn superoxide dismutase (Cu/Zn-SOD), apart from N- and C-terminal extensions. Cu/Zn-SOD is a homodimer containing four cysteine residues within each subunit, and EC-SOD is a tetramer composed of two disulfide-bonded dimers in which each subunit contains six cysteines. The amino acid sequences of all EC-SOD subunits are identical. It is known that Cys-219 is involved in an interchain disulfide. To account for the remaining five cysteine residues we purified human EC-SOD and determined the disulfide bridge pattern. The results show that human EC-SOD exists in two forms, each with a unique disulfide bridge pattern. One form (active EC-SOD) is enzymatically active and contains a disulfide bridge pattern similar to Cu/Zn-SOD. The other form (inactive EC-SOD) has a different disulfide bridge pattern and is enzymatically inactive. The EC-SOD polypeptide chain apparently folds in two different ways, most likely resulting in different three-dimensional structures. Our study shows that one gene may produce proteins with different disulfide bridge arrangements and, thus, by definition, different primary structures. This observation adds another dimension to the functional annotation of the proteome.

Aorta↗

Identification of candidate regulators of embryonic stem cell differentiation by comparative phosphoprotein affinity profiling.

Embryonic stem cells are a unique cell population capable both of self-renewal and of differentiation into all tissues in the adult organism. Despite the central importance of these cells, little information is available regarding the intracellular signaling pathways that govern self-renewal or early steps in the differentiation program. Embryonic stem cell growth and differentiation correlates with kinase activities, but with the exception of the JAK/STAT3 pathway, the relevant substrates are unknown. To identify candidate phosphoproteins with potential relevance to embryonic stem cell differentiation, a systems biology approach was used. Proteins were purified using phosphoprotein affinity columns, then separated by two-dimensional gel electrophoresis, and detected by silver stain before being identified by tandem mass spectrometry. By comparing preparations from undifferentiated and differentiating mouse embryonic stem cells, a set of proteins was identified that exhibited altered post-translational modifications that correlated with differentiation state. Evidence for altered post-translational modification included altered gel mobility, altered recovery after affinity purification, and direct mass spectra evidence. Affymetrix microarray analysis indicated that gene expression levels of these same proteins had minimal variability over the same differentiation period. Bioinformatic annotations indicated that this set of proteins is enriched with chromatin remodeling, catabolic, and chaperone functions. This set of candidate phosphoprotein regulators of stem cell differentiation includes products of genes previously noted to be enriched in embryonic stem cells at the mRNA expression level as well as proteins not associated previously with stem cell differentiation status.

Animals↗

Evolutionary trace report_maker: a new type of service for comparative analysis of proteins.

: Evolutionary trace report_maker offers a new type of service for researchers investigating the function of novel proteins. It pools, from different sources, information about protein sequence, structure and elementary annotation, and to that background superimposes inference about the evolutionary behavior of individual residues, using real-valued evolutionary trace method. As its only input it takes a Protein Data Bank identifier or UniProt accession number, and returns a human-readable document in PDF format, supplemented by the original data needed to reproduce the results quoted in the report.

Algorithms↗

Argonaute--a database for gene regulation by mammalian microRNAs.

MicroRNAs (miRNAs) constitute a recently discovered class of small non-coding RNAs that regulate expression of target genes either by decreasing the stability of the target mRNA or by translational inhibition. They are involved in diverse processes, including cellular differentiation, proliferation and apoptosis. Recent evidence also suggests their importance for cancerogenesis. By far the most important model systems in cancer research are mammalian organisms. Thus, we decided to compile comprehensive information on mammalian miRNAs, their origin and regulated target genes in an exhaustive, curated database called Argonaute (http://www.ma.uni-heidelberg.de/apps/zmf/argonaute/interface). Argonaute collects latest information from both literature and other databases. In contrast to current databases on miRNAs like miRBase::Sequences, NONCODE or RNAdb, Argonaute hosts additional information on the origin of an miRNA, i.e. in which host gene it is encoded, its expression in different tissues and its known or proposed function, its potential target genes including Gene Ontology annotation, as well as miRNA families and proteins known to be involved in miRNA processing. Additionally, target genes are linked to an information retrieval system that provides comprehensive information from sequence databases and a simultaneous search of MEDLINE with all synonyms of a given gene. The web interface allows the user to get information for a single or multiple miRNAs, either selected or uploaded through a text file. Argonaute currently has information on 839 miRNAs from human, mouse and rat.

Animals↗