Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Large-scale identification of proteins expressed in mouse embryonic stem cells.

A protein subset expressed in the mouse embryonic stem (ES) cell line, E14-1, was characterized by mass spectrometry-based protein identification technology and data analysis. In total, 1790 proteins including 365 potential nuclear and 260 membrane proteins were identified from tryptic digests of total cell lysates. The subset contained a variety of proteins in terms of physicochemical characteristics, subcellular localization, and biological function as defined by Gene Ontology annotation groups. In addition to many housekeeping proteins found in common with other cell types, the subset contained a group of regulatory proteins that may determine unique ES cell functions. We identified 39 transcription factors including Oct-3/4, Sox-2, and undifferentiated embryonic cell transcription factor I, which are characteristic of ES cells, 88 plasma membrane proteins including cell surface markers such as CD9 and CD81, 44 potential proteinaceous ligands for cell surface receptors including growth factors, cytokines, and hormones, and 100 cell signaling molecules. The subset also contained the products of 60 ES-specific and 41 stemness genes defined previously by the DNA microarray analysis of Ramalho-Santos et al. (Ramalho-Santos et al., Science 2002, 298, 597-600), as well as a number of components characteristic of differentiated cell types such as hematopoietic and neural cells. We also identified potential post-translational modifications in a number of ES cell proteins including five Lys acetylation sites and a single phosphorylation site. To our knowledge, this study provides the largest proteomic dataset characterized to date for a single mammalian cell species, and serves as a basic catalogue of a major proteomic subset that is expressed in mouse ES cells.

Animals↗

The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1998.

SWISS-PROT (http://www.expasy.ch/) is a curated protein sequence database which strives to provide a high level of annotations (such as the description of the function of a protein, its domains structure, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to two additional databases; a variety of new documentation files and improvements to TrEMBL, a computer annotated supplement to SWISS-PROT. TrEMBL consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except the CDS already included in SWISS-PROT.

Amino Acid Sequence↗

ESLpred: SVM-based method for subcellular localization of eukaryotic proteins using dipeptide composition and PSI-BLAST.

Automated prediction of subcellular localization of proteins is an important step in the functional annotation of genomes. The existing subcellular localization prediction methods are based on either amino acid composition or N-terminal characteristics of the proteins. In this paper, support vector machine (SVM) has been used to predict the subcellular location of eukaryotic proteins from their different features such as amino acid composition, dipeptide composition and physico-chemical properties. The SVM module based on dipeptide composition performed better than the SVM modules based on amino acid composition or physico-chemical properties. In addition, PSI-BLAST was also used to search the query sequence against the dataset of proteins (experimentally annotated proteins) to predict its subcellular location. In order to improve the prediction accuracy, we developed a hybrid module using all features of a protein, which consisted of an input vector of 458 dimensions (400 dipeptide compositions, 33 properties, 20 amino acid compositions of the protein and 5 from PSI-BLAST output). Using this hybrid approach, the prediction accuracies of nuclear, cytoplasmic, mitochondrial and extracellular proteins reached 95.3, 85.2, 68.2 and 88.9%, respectively. The overall prediction accuracy of SVM modules based on amino acid composition, physico-chemical properties, dipeptide composition and the hybrid approach was 78.1, 77.8, 82.9 and 88.0%, respectively. The accuracy of all the modules was evaluated using a 5-fold cross-validation technique. Assigning a reliability index (reliability index > or =3), 73.5% of prediction can be made with an accuracy of 96.4%. Based on the above approach, an online web server ESLpred was developed, which is available at http://www.imtech.res.in/raghava/eslpred/.

Artificial Intelligence↗

The SWISS-PROT protein sequence data bank and its supplement TrEMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotations (such as the description of the function of a protein, structure of its domains, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to two additional databases; a variety of new documentation files and the creation of TrEMBL, a computer annotated supplement to SWISS-PROT. This supplement consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except the CDS already included in SWISS-PROT.

Academies and Institutes↗

Visualizing the genome: techniques for presenting human genome data and annotations.

BACKGROUND: In order to take full advantage of the newly available public human genome sequence data and associated annotations, biologists require visualization tools ("genome browsers") that can accommodate the high frequency of alternative splicing in human genes and other complexities. RESULTS: In this article, we describe visualization techniques for presenting human genomic sequence data and annotations in an interactive, graphical format. These techniques include: one-dimensional, semantic zooming to show sequence data alongside gene structures; color-coding exons to indicate frame of translation; adjustable, moveable tiers to permit easier inspection of a genomic scene; and display of protein annotations alongside gene structures to show how alternative splicing impacts protein structure and function. These techniques are illustrated using examples from two genome browser applications: the Neomorphic GeneViewer annotation tool and ProtAnnot, a prototype viewer which shows protein annotations in the context of genomic sequence. CONCLUSION: By presenting techniques for visualizing genomic data, we hope to provide interested software developers with a guide to what features are most likely to meet the needs of biologists as they seek to make sense of the rapidly expanding body of public genomic data and annotations.

Alternative Splicing↗

A high-throughput approach for subcellular proteome: identification of rat liver proteins using subcellular fractionation coupled with two-dimensional liquid chromatography tandem mass spectrometry and bioinformatic analysis.

Four fractions from rat liver (a crude mitochondria (CM) and cytosol (C) fraction obtained with differential centrifugation, a purified mitochondrial (PM) fraction obtained with nycodenz density gradient centrifugation, and a total liver (TL) fraction) were analyzed with two-dimensional liquid chromatography tandem mass spectrometry analysis. A total of 564 rat proteins were identified and were bioinformatically annotated according to their physicochemical characteristics and functions. While most extreme alkaline ribosomal proteins were identified in the TL fraction, the C fraction mainly included neutral enzymes and the PM fraction enriched alkaline proteins and proteins with electron transfer activity or oxygen binding activity. Such characteristics were more apparent in proteins identified only in the TL, C, or PM fraction. The Swiss-Prot annotation and the bioinformatic prediction results proved that the C and PM fractions had enriched cytoplasmic or mitochondrial proteins, respectively. Combination usage of subcellular fractionation with two-dimensional liquid chromatography tandem mass spectrometry was proved to be a high-throughput, sensitive, and effective analytical approach for subcellular proteomics research. Using such a strategy, we have constructed the largest proteome database to date for rat liver (564 rat proteins) and its cytosol (222 rat proteins) and mitochondrial fractions (227 rat proteins). Moreover, the 352 proteins with Swiss-Prot subcellular location annotation in the 564 identified proteins were used as an actual subcellular proteome dataset to evaluate the widely used bioinformatics tools such as PSORT, TargetP, TMHMM, and GRAVY.

Animals↗

ECLIPSE: exploring the dark proteome of ESKAPE pathogens through the sequence similarity network of the Protein Universe Atlas.

MOTIVATION: The accelerating crisis of antimicrobial resistance among the critical so-called ESKAPE pathogens demands the urgent identification of novel molecular targets. However, a substantial fraction of ESKAPE proteomes remains functionally uncharacterized, with many genes annotated as encoding hypothetical proteins. These protein sequences often lack significant similarity to known protein families when conventional homology-based annotation methods are used and thus remain "dark". This limits our ability to explore their roles in pathogenicity, and it is thus crucial to bridge this substantial gap in pathogen biology by developing new strategies to illuminate these "dark" regions of the ESKAPE pan-proteome. RESULTS: We introduce ECLIPSE (ESKAPE Connectome Linkage and Inference for Proteome Sequence Exploration), a network-based computational framework that systematically identifies and prioritizes functionally dark protein families in ESKAPE pan-proteomes. ECLIPSE embeds target ESKAPE pathogen proteomes within the global sequence similarity network of the Protein Universe Atlas. It detects connected components composed entirely of unannotated proteins, called the "dark proteome." As a case study, we applied ECLIPSE to a pan-proteome of 3 460 657 protein sequences from 635 strains of Pseudomonas aeruginosa (PA). ECLIPSE identified 120 985 proteins (4%) residing in completely dark connected components. Furthermore, we have performed a taxonomic diversity analysis using normalized Shannon indices to characterize each dark component by its enrichment in ESKAPE pathogens. The analysis utilized the evenness (E) value (see Methods 2.1), which distinguishes Pseudomonas-specific (target-specific) from ESKAPE-enriched dark components. We then developed the Dark Proteome Prioritization Score (DPPS), a composite multidimensional scoring framework (see Methods 2.5). It ranks these dark components by biological relevance across four orthogonal axes: (i) functional darkness, (ii) P. aeruginosa proportion in the Atlas, (iii) AMR-clade taxonomic restriction, and (iv) conservation across the 635 P. aeruginosa strains. This framework outputs a robust four-tier scoring system; the prioritized Tier I components were validated by weight sensitivity analysis and remained stable across 500 Monte Carlo weight perturbations. Structural characterization of one of the top-ranked ESKAPE-enriched dark components revealed that it belongs to the beta-barrel fold DUF1302 (PF06980) family, for which no experimentally solved three-dimensional structure exists in the PDB. The genomic context analysis indicates that it is co-localized with a LuxR-type transcriptional regulator. Collectively, ECLIPSE identifies evolutionarily conserved, structurally defined, and functionally dark proteins enriched across ESKAPE pathogens; these dark proteins can further be utilized as alternative antimicrobial targets for experimental characterization. AVAILABILITY AND IMPLEMENTATION: The source code and dataset are available for free at: Github: https://github.com/surabhilata/ECLIPSE.git, Zenodo: DOI: 10.5281/zenodo.21064323.

Proteome↗

Type III effector proteins: doppelgangers of bacterial virulence.

Bacterial pathogens have co-evolved with their hosts in their ongoing quest for advantage in the resulting interaction. These intimate associations have resulted in remarkable adaptations of prokaryotic virulence proteins and their eukaryotic molecular targets. An important strategy used by microbial pathogens of animals to manipulate host cellular functions is structural mimicry of eukaryotic proteins. Recent evidence demonstrates that plant pathogens also use structural mimicry of host factors as a virulence strategy. Nearly all virulence proteins from phytopathogenic bacteria have eluded functional annotation on the basis of primary amino-acid sequence. Recent efforts to determine their three-dimensional structures are, however, revealing important clues about the mechanisms of bacterial virulence in plants.

Bacteria↗

Structural insights into adeno-associated virus serotype 5.

The adeno-associated viruses (AAVs) display differential cell binding, transduction, and antigenic characteristics specified by their capsid viral protein (VP) composition. Toward structure-function annotation, the crystal structure of AAV5, one of the most sequence diverse AAV serotypes, was determined to 3.45-Å resolution. The AAV5 VP and capsid conserve topological features previously described for other AAVs but uniquely differ in the surface-exposed HI loop between βH and βI of the core β-barrel motif and have pronounced conformational differences in two of the AAV surface variable regions (VRs), VR-IV and VR-VII. The HI loop is structurally conserved in other AAVs despite amino acid differences but is smaller in AAV5 due to an amino acid deletion. This HI loop is adjacent to VR-VII, which is largest in AAV5. The VR-IV, which forms the larger outermost finger-like loop contributing to the protrusions surrounding the icosahedral 3-fold axes of the AAVs, is shorter in AAV5, creating a smoother capsid surface topology. The HI loop plays a role in AAV capsid assembly and genome packaging, and VR-IV and VR-VII are associated with transduction and antigenic differences, respectively, between the AAVs. A comparison of interior capsid surface charge and volume of AAV5 to AAV2 and AAV4 showed a higher propensity of acidic residues but similar volumes, consistent with comparable DNA packaging capacities. This structure provided a three-dimensional (3D) template for functional annotation of the AAV5 capsid with respect to regions that confer assembly efficiency, dictate cellular transduction phenotypes, and control antigenicity.

Capsid Proteins↗

Protein expression dynamics during replicative senescence of endothelial cells studied by 2-D difference in-gel electrophoresis.

Endothelial senescence contributes to endothelium dysfunctionality and is thereby linked to vascular aging. A dynamic proteomic study on human umbilical vein endothelial cells, isolated from three umbilical cords, was performed. The cells were cultured towards replicative senescence and whole cell lysates were subjected to 2-D difference gel electrophoresis (DIGE). Despite the biological variability of the three independent isolations, a set of proteins was found that showed senescence-dependent expression patterns in all isolations. We focused on those proteins that showed significant changes, with a paired analysis of variance (RM-ANOVA) p-value of < or =0.05. Thirty-five proteins were identified with LC-Fourier transform MS, and functional annotation revealed that endothelial replicative senescence is accompanied by increased cellular stress, protein biosynthesis and reduction in DNA repair and maintenance. Nuclear integrity becomes affected and cytoskeletal structure is also changed. Such important changes in the cell infrastructure might accelerate endothelium dysfunctionality. This study provides biological information that will initiate studies to further unravel endothelial senescence and gain more knowledge about the consequences of this process in the in vivo situation.

Cell Proliferation↗

Transcriptomic and proteomic characterization of the Fur modulon in the metal-reducing bacterium Shewanella oneidensis.

The availability of the complete genome sequence for Shewanella oneidensis MR-1 has permitted a comprehensive characterization of the ferric uptake regulator (Fur) modulon in this dissimilatory metal-reducing bacterium. We have employed targeted gene mutagenesis, DNA microarrays, proteomic analysis using liquid chromatography-mass spectrometry, and computational motif discovery tools to define the S. oneidensis Fur regulon. Using this integrated approach, we identified nine probable operons (containing 24 genes) and 15 individual open reading frames (ORFs), either with unknown functions or encoding products annotated as transport or binding proteins, that are predicted to be direct targets of Fur-mediated repression. This study suggested, for the first time, possible roles for four operons and eight ORFs with unknown functions in iron metabolism or iron transport-related functions. Proteomic analysis clearly identified a number of transporters, binding proteins, and receptors related to iron uptake that were up-regulated in response to a fur deletion and verified the expression of nine genes originally annotated as pseudogenes. Comparison of the transcriptome and proteome data revealed strong correlation for genes shown to be undergoing large changes at the transcript level. A number of genes encoding components of the electron transport system were also differentially expressed in a fur deletion mutant. The gene omcA (SO1779), which encodes a decaheme cytochrome c, exhibited significant decreases in both mRNA and protein abundance in the fur mutant and possessed a strong candidate Fur-binding site in its upstream region, thus suggesting that omcA may be a direct target of Fur activation.

Bacterial Proteins↗

Chromosome-level genome assembly of Manglietia pachyphylla.

Manglietia pachyphylla, an endangered evergreen tree within the Magnoliaceae family, is renowned for its exceptional ornamental value in landscape horticulture. Despite its classification as a Category II nationally protected plant species in China, the genetic basis of its adaptive traits and conservation priorities remains poorly understood. To address this, we present the first chromosome-scale genome assembly of M. pachyphylla utilizing an integrated approach combining PacBio HiFi long-read and Hi-C chromosome conformation capture sequencing technologies. The assembled genome spans 2.15&#x2009;Gb (contig N50&#x2009;=&#x2009;43.57&#x2009;Mb), exhibiting a heterozygosity rate of 0.78% and repeat content of 78.64%, predominantly comprising long terminal repeat (LTR) retrotransposons (52.86%). Hi-C scaffolding anchored 99.57% of the assembly to 19 pseudochromosomes, achieving a BUSCO completeness score of 96.4%. Annotation revealed 42,505 putative protein-coding genes, with 84.46% of predicted genes were functionally annotated. Phylogenomic analysis positioned M. pachyphylla and Oyama sieboldii clustered together in a well-supported group. This high-contiguity genome assembly enables future investigations into adaptive evolution, functional genomics, and evidence-based conservation strategies for this endangered species.

Chromosomes, Plant↗

LC-MS/MS based proteomic analysis and functional inference of hypothetical proteins in Desulfovibrio vulgaris.

High efficiency capillary liquid chromatography-tandem mass spectrometry (LC-MS/MS) was used to examine the proteins extracted from Desulfovibrio vulgaris cells across six treatment conditions. While our previous study provided a proteomic overview of the cellular metabolism based on proteins with known functions [W. Zhang, M.A. Gritsenko, R.J. Moore, D.E. Culley, L. Nie, K. Petritis, E.F. Strittmatter, D.G. Camp II, R.D. Smith, F.J. Brockman, A proteomic view of the metabolism in Desulfovibrio vulgaris determined by liquid chromatography coupled with tandem mass spectrometry, Proteomics 6 (2006) 4286-4299], this study describes the global detection and functional inference for hypothetical D. vulgaris proteins. Using criteria that a given peptide of a protein is identified from at least two out of three independent LC-MS/MS measurements and that for any protein at least two different peptides are identified among the three measurements, 129 open reading frames (ORFs) originally annotated as hypothetical proteins were found to encode expressed proteins. Functional inference for the conserved hypothetical proteins was performed by a combination of several non-homology based methods: genomic context analysis, phylogenomic profiling, and analysis of a combination of experimental information, including peptide detection in cells grown under specific culture conditions and cellular location of the proteins. Using this approach we were able to assign possible functions to 20 conserved hypothetical proteins. This study demonstrated that a combination of proteomics and bioinformatics methodologies can provide verification of the expression of hypothetical proteins and improve genome annotation.

Amino Acid Sequence↗

Enzyme function less conserved than anticipated.

The level of sequence similarity that implies similarity in protein structure is well established. Recently, many groups proposed thresholds for similarity in sequence implying similarity in enzymatic function. All previous results suggest the strong conservation of enzymatic function above levels of 50% pairwise sequence identity. Here, I argue that all groups substantially overestimated the conservation of enzyme function because their data sets were either too biased, or too small. An unbiased analysis suggested that less than 30% of the pair fragments above 50% sequence identity have entirely identical EC numbers. Another surprising finding was that even BLAST E-values below 10(-50) did not suffice to automatically transfer enzyme function without errors. As expected, most misclassifications originated from similarities in relatively short regions and/or from transferring annotations for different domains. Both problems cannot be corrected easily by adjusting the thresholds for automatic transfer of genome annotations. A score relating sequence identity to alignment length (distance from HSSP-threshold) outperformed statistical BLAST scores for high sequence similarity. In particular, the distance score allowed error-free transfer of enzyme function for the 10% most similar enzyme pairs. The results illustrated how difficult it is to assess the conservation of protein function and to guarantee error-free genome annotations, in general: sets with millions of pair comparisons might not suffice to arrive at statistically significant conclusions. In practice, the revised detailed estimates for the sequence conservation of enzyme function may provide important benchmarks for everyday sequence analysis and for more cautious automatic genome annotations.

Amino Acid Sequence↗

An expression and bioinformatics analysis of the Arabidopsis serine carboxypeptidase-like gene family.

The Arabidopsis (Arabidopsis thaliana) genome encodes a family of 51 proteins that are homologous to known serine carboxypeptidases. Based on their sequences, these serine carboxypeptidase-like (SCPL) proteins can be divided into several major clades. The first group consists of 21 proteins which, despite the function implied by their annotation, includes two that have been shown to function as acyltransferases in plant secondary metabolism: sinapoylglucose:malate sinapoyltransferase and sinapoylglucose:choline sinapoyltransferase. A second group comprises 25 SCPL proteins whose biochemical functions have not been clearly defined. Genes encoding representatives from both of these clades can be found in many plants, but have not yet been identified in other phyla. In contrast, the remaining SCPL proteins include five members that are similar to serine carboxypeptidases from a variety of organisms, including fungi and animals. Reverse transcription PCR results suggest that some SCPL genes are expressed in a highly tissue-specific fashion, whereas others are transcribed in a wide range of tissue types. Taken together, these data suggest that the Arabidopsis SCPL gene family encodes a diverse group of enzymes whose functions are likely to extend beyond protein degradation and processing to include activities such as the production of secondary metabolites.

Arabidopsis↗

The PAS fold. A redefinition of the PAS domain based upon structural prediction.

In the postgenomic era it is essential that protein sequences are annotated correctly in order to help in the assignment of their putative functions. Over 1300 proteins in current protein sequence databases are predicted to contain a PAS domain based upon amino acid sequence alignments. One of the problems with the current annotation of the PAS domain is that this domain exhibits limited similarity at the amino acid sequence level. It is therefore essential, when using proteins with low-sequence similarities, to apply profile hidden Markov model searches for the PAS domain-containing proteins, as for the PFAM database. From recent 3D X-ray and NMR structures, however, PAS domains appear to have a conserved 3D fold as shown here by structural alignment of the six representative 3D-structures from the PDB database. Large-scale modelling of the PAS sequences from the PFAM database against the 3D-structures of these six structural prototypes was performed. All 3D models generated (> 5700) were evaluated using prosaii. We conclude from our large-scale modelling studies that the PAS and PAC motifs (which are separately defined in the PFAM database) are directly linked and that these two motifs form the PAS fold. The existing subdivision in PAS and PAC motifs, as used by the PFAM and SMART databases, appears to be caused by major differences in sequences in the region connecting these two motifs. This region, as has been shown by Gardner and coworkers for human PAS kinase (Amezcua, C.A., Harper, S.M., Rutter, J. & Gardner, K.H. (2002) Structure 10, 1349-1361, [1]), is very flexible and adopts different conformations depending on the bound ligand. Some PAS sequences present in the PFAM database did not produce a good structural model, even after realignment using a structure-based alignment method, suggesting that these representatives are unlikely to have a fold resembling any of the structural prototypes of the PAS domain superfamily.

Amino Acid Sequence↗

Helicobacter pylori FlhB function: the FlhB C-terminal homologue HP1575 acts as a "spare part" to permit flagellar export when the HP0770 FlhBCC domain is deleted.

In Helicobacter pylori 26695, a gene annotated HP1575 encodes a putative protein of unknown function which shows significant similarity to part of the C-terminal domain of the flagellar export protein FlhB. In Salmonella enterica, this part (FlhB(CC)) is proteolytically cleaved from the full-length FlhB, a processing event that is required for flagellar protein export and, thus, motility. The role of FlhB (HP0770) and its C-terminal homologue HP1575 was studied in H. pylori using a range of nonpolar deletion mutants defective in HP1575, HP0770, and the CC domain of HP0770 (HP0770(CC)). Deletion of HP0770 abolished swimming motility, whereas mutants carrying a deletion of either HP1575 or HP0770(CC) retained their ability to swim. An H. pylori strain containing deletions in both HP1575 and HP0770(CC) was nonmotile and did not produce flagella, suggesting that at least one of the two proteins had to be present for flagellar assembly to occur. Indeed, motility was restored when HP1575 was reintroduced into this strain immediately downstream of, but not fused to, the truncated HP0770 gene. Thus, HP1575 can functionally replace HP0770(CC) in this background. Like FlhB in S. enterica, HP0770 appeared to be proteolytically processed at a conserved NPTH processing site. However, mutation of the proline contained within the NPTH site of HP0770 did not affect motility and flagellar assembly, although it clearly interfered with processing when the protein was heterologously produced in Escherichia coli.

Amino Acid Sequence↗

pdbFun: mass selection and fast comparison of annotated PDB residues.

pdbFun (http://pdbfun.uniroma2.it) is a web server for structural and functional analysis of proteins at the residue level. pdbFun gives fast access to the whole Protein Data Bank (PDB) organized as a database of annotated residues. The available data (features) range from solvent exposure to ligand binding ability, location in a protein cavity, secondary structure, residue type, sequence functional pattern, protein domain and catalytic activity. Users can select any residue subset (even including any number of PDB structures) by combining the available features. Selections can be used as probe and target in multiple structure comparison searches. For example a search could involve, as a query, all solvent-exposed, hydrophylic residues that are not in alpha-helices and are involved in nucleotide binding. Possible examples of targets are represented by another selection, a single structure or a dataset composed of many structures. The output is a list of aligned structural matches offered in tabular and also graphical format.

Algorithms↗