Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

A new family of CoA-transferases.

CoA-transferases are found in organisms from all lines of descent. Most of these enzymes belong to two well-known enzyme families, but recent work on unusual biochemical pathways of anaerobic bacteria has revealed the existence of a third family of CoA-transferases. The members of this enzyme family differ in sequence and reaction mechanism from CoA-transferases of the other families. Currently known enzymes of the new family are a formyl-CoA: oxalate CoA-transferase, a succinyl-CoA: (R)-benzylsuccinate CoA-transferase, an (E)-cinnamoyl-CoA: (R)-phenyllactate CoA-transferase, and a butyrobetainyl-CoA: (R)-carnitine CoA-transferase. In addition, a large number of proteins of unknown or differently annotated function from Bacteria, Archaea and Eukarya apparently belong to this enzyme family. Properties and reaction mechanisms of the CoA-transferases of family III are described and compared to those of the previously known CoA-transferases.

Bacteria, Anaerobic↗

Transcript annotation in FANTOM3: mouse gene catalog based on physical cDNAs.

The international FANTOM consortium aims to produce a comprehensive picture of the mammalian transcriptome, based upon an extensive cDNA collection and functional annotation of full-length enriched cDNAs. The previous dataset, FANTOM2, comprised 60,770 full-length enriched cDNAs. Functional annotation revealed that this cDNA dataset contained only about half of the estimated number of mouse protein-coding genes, indicating that a number of cDNAs still remained to be collected and identified. To pursue the complete gene catalog that covers all predicted mouse genes, cloning and sequencing of full-length enriched cDNAs has been continued since FANTOM2. In FANTOM3, 42,031 newly isolated cDNAs were subjected to functional annotation, and the annotation of 4,347 FANTOM2 cDNAs was updated. To accomplish accurate functional annotation, we improved our automated annotation pipeline by introducing new coding sequence prediction programs and developed a Web-based annotation interface for simplifying the annotation procedures to reduce manual annotation errors. Automated coding sequence and function prediction was followed with manual curation and review by expert curators. A total of 102,801 full-length enriched mouse cDNAs were annotated. Out of 102,801 transcripts, 56,722 were functionally annotated as protein coding (including partial or truncated transcripts), providing to our knowledge the greatest current coverage of the mouse proteome by full-length cDNAs. The total number of distinct non-protein-coding transcripts increased to 34,030. The FANTOM3 annotation system, consisting of automated computational prediction, manual curation, and final expert curation, facilitated the comprehensive characterization of the mouse transcriptome, and could be applied to the transcriptomes of other species.

Animals↗

Structural genomics of minimal organisms and protein fold space.

The initial aim of the Berkeley Structural Genomics Center is to obtain a near-complete structural complement of two minimal organisms, closely related pathogens Mycoplasma genitalium and M. pneumoniae. The former has fewer than 500 genes and the latter fewer than 700 genes. To achieve this goal, the current protein targets have been selected starting with those predicted to be most tractable and likely to yield new structural and functional information. During the past 3 years, the semi-automated structural genomics pipeline has been set up from cloning, expression, purification, and ultimately to structural determination. The results from the pipeline substantially increased the coverage of the protein fold space of M. pneumoniae and M. genitalium. Furthermore, about 1/2 of the structures of 'unique' protein sequences revealed new and novel folds, and over 2/3 of the structures of previously annotated 'hypothetical proteins' inferred their molecular functions.

Bacterial Proteins↗

Pfarao: a web application for protein family analysis customized for cytoskeletal and motor proteins (CyMoBase).

BACKGROUND: Annotation of protein sequences of eukaryotic organisms is crucial for the understanding of their function in the cell. Manual annotation is still by far the most accurate way to correctly predict genes. The classification of protein sequences, their phylogenetic relation and the assignment of function involves information from various sources. This often leads to a collection of heterogeneous data, which is hard to track. Cytoskeletal and motor proteins consist of large and diverse superfamilies comprising up to several dozen members per organism. Up to date there is no integrated tool available to assist in the manual large-scale comparative genomic analysis of protein families. DESCRIPTION: Pfarao (Protein Family Application for Retrieval, Analysis and Organisation) is a database driven online working environment for the analysis of manually annotated protein sequences and their relationship. Currently, the system can store and interrelate a wide range of information about protein sequences, species, phylogenetic relations and sequencing projects as well as links to literature and domain predictions. Sequences can be imported from multiple sequence alignments that are generated during the annotation process. A web interface allows to conveniently browse the database and to compile tabular and graphical summaries of its content. CONCLUSION: We implemented a protein sequence-centric web application to store, organize, interrelate, and present heterogeneous data that is generated in manual genome annotation and comparative genomics. The application has been developed for the analysis of cytoskeletal and motor proteins (CyMoBase) but can easily be adapted for any protein.

Amino Acid Sequence↗

Proteomic analysis of low-abundant integral plasma membrane proteins based on gels.

To characterize low-copy integral membrane proteins and offer some methods for human liver proteome projects, we fractionated highly purified rat liver plasma membrane (PM). PM was purified through two sucrose density gradient centrifugations, and treated with 0.1 M Na(2)CO(3), chloroform/methanol and Triton X-100. Proteins were separated by electrophoresis and submitted to mass spectrometry analysis. Four hundred and fifty-seven non-redundant membrane proteins were identified, of which 23% (105) were integral membrane proteins with one or more transmembrane domains. One hundred and fifty-three (33.5%) had no location annotation and 68 were unknown-function proteins. The proteins from different fractions were complementory. A database search for all identified proteins revealed that 53 proteins were involved in the cell communication pathway. More interestingly, more than 50% of the proteins had a protein abundance index concentration of less than 0.1 mol/l, and 12% proteins a concentration 100 times less than that of arginase 1 and actin.

Animals↗

PicSNP: a browsable catalog of nonsynonymous single nucleotide polymorphisms in the human genome.

Recent progress in identification and mapping of single nucleotide polymorphisms (SNPs) in the human genome generates an unprecedented opportunity to explore cause-effect relationships between genetic variations and susceptibility to common diseases. For this purpose, one promising strategy would be to select a set of SNPs that potentially alter the function of proteins involved in the pathogenesis of the diseases and compare their frequencies in the affected individuals and the healthy population. In this respect, SNPs that change amino acid sequences (nonsynonymous SNPs; nsSNPs) are of particular interest, since they are more likely to affect protein functions. In this study, we have constructed a catalog of nsSNPs (PicSNP), whose unique features are (i) nsSNPs are classified according to the functions of the affected genes and are searchable under the guidance of hierarchical lists of protein functions and (ii) nsSNPs that lead to amino acid changes in the known functional sites and domains of proteins are highlighted. Out of 1,190,295 SNPs extracted from public database, we identified 3793 nsSNPs and classified them in 1247 categories of protein functions. 495 sites and domains annotated in the Swiss-Prot database were found to include nsSNPs, including 2 nsSNPs in disulfide-binding sites and 38 nsSNPs in transmembrane regions. PicSNP is available via the World Wide Web (http://picsnp.org) and would support research questing for SNPs involved in common diseases.

Databases, Factual↗

Molecular dynamics of the Shewanella oneidensis response to chromate stress.

Temporal genomic profiling and whole-cell proteomic analyses were performed to characterize the dynamic molecular response of the metal-reducing bacterium Shewanella oneidensis MR-1 to an acute chromate shock. The complex dynamics of cellular processes demand the integration of methodologies that describe biological systems at the levels of regulation, gene and protein expression, and metabolite production. Genomic microarray analysis of the transcriptome dynamics of midexponential phase cells subjected to 1 mm potassium chromate (K(2)CrO(4)) at exposure time intervals of 5, 30, 60, and 90 min revealed 910 genes that were differentially expressed at one or more time points. Strongly induced genes included those encoding components of a TonB1 iron transport system (tonB1-exbB1-exbD1), hemin ATP-binding cassette transporters (hmuTUV), TonB-dependent receptors as well as sulfate transporters (cysP, cysW-2, and cysA-2), and enzymes involved in assimilative sulfur metabolism (cysC, cysN, cysD, cysH, cysI, and cysJ). Transcript levels for genes with annotated functions in DNA repair (lexA, recX, recA, recN, dinP, and umuD), cellular detoxification (so1756, so3585, and so3586), and two-component signal transduction systems (so2426) were also significantly up-regulated (p < 0.05) in Cr(VI)-exposed cells relative to untreated cells. By contrast, genes with functions linked to energy metabolism, particularly electron transport (e.g. so0902-03-04, mtrA, omcA, and omcB), showed dramatic temporal alterations in expression with the majority exhibiting repression. Differential proteomics based on multidimensional HPLC-MS/MS was used to complement the transcriptome data, resulting in comparable induction and repression patterns for a subset of corresponding proteins. In total, expression of 2,370 proteins were confidently verified with 624 (26%) of these annotated as hypothetical or conserved hypothetical proteins. The initial response of S. oneidensis to chromate shock appears to require a combination of different regulatory networks that involve genes with annotated functions in oxidative stress protection, detoxification, protein stress protection, iron and sulfur acquisition, and SOS-controlled DNA repair mechanisms.

Anion Transport Proteins↗

Serum Proteomic Profiling Reveals Renin-Associated Immune and Cytoskeletal Dysregulation in Post-COVID-19 Condition Patients with Secondary Adrenal Insufficiency.

Post-COVID-19 condition (PCC) with secondary adrenal insufficiency (SAI) involves multiorgan dysfunction, potentially linked to renin-angiotensin-aldosterone system dysregulation. The molecular basis of renin-associated pathology remains unclear. Here, PCC+SAI patients were stratified by upright renin into low- (<38.8&#x202f;pg/mL) and high-renin (&#x2265;38.8&#x202f;pg/mL) groups. Clinical, endocrine, and proteomic analyses were performed. We found that high-renin patients showed increased BMI, lipids, renin, and aldosterone, but reduced aldosterone-to-renin ratio. Proteomic annalysis identified 20 differentially expressed proteins (DEPs), including 17 upregulated and 3 downregulated proteins in Ren-H patients. Functional annotation revealed that 15 DEPs were immune-related (e.g., APOC4, APOE, C4BPA, CFAH, CFHR3, PF4V, PLF4), while FLNA and COF1 represented cytoskeletal proteins. These DEPs were primarily involved in immune response, complement and coagulation cascades, and MAPK signaling pathways. Correlation analyses indicated that upright renin was positively correlated with complement-related proteins and platelet-derived immune factors, while cytoskeletal proteins (FLNA, COF1) showed positive associations with serum Na+ levels. Additionally, white blood cell and platelet counts were positively correlated with the majority of DEPs. In conclusion, exploratory proteomic analyses suggest that elevated upright renin in PCC+SAI may be associated with immune dysregulation, complement activation, and cytoskeletal remodeling, offering novel insights into the endocrine-immune interactions driving postviral sequelae.

Humans↗

The SBASE domain sequence resource, release 12: prediction of protein domain-architecture using support vector machines.

SBASE (http://www.icgeb.trieste.it/sbase) is an online resource designed to facilitate the detection of domain homologies based on sequence database search. The present release of the SBASE A library of protein domain sequences contains 972,397 protein sequence segments annotated by structure, function, ligand-binding or cellular topology, clustered into 8547 domain groups. SBASE B contains 169,916 domain sequences clustered into 2526 less well-characterized groups. Domain prediction is based on an evaluation of database search results in comparison with a 'similarity network' of inter-sequence similarity scores, using support vector machines trained on similarity search results of known domains.

Artificial Intelligence↗

Support vector machine learning from heterogeneous data: an empirical analysis using protein sequence and structure.

MOTIVATION: Drawing inferences from large, heterogeneous sets of biological data requires a theoretical framework that is capable of representing, e.g. DNA and protein sequences, protein structures, microarray expression data, various types of interaction networks, etc. Recently, a class of algorithms known as kernel methods has emerged as a powerful framework for combining diverse types of data. The support vector machine (SVM) algorithm is the most popular kernel method, due to its theoretical underpinnings and strong empirical performance on a wide variety of classification tasks. Furthermore, several recently described extensions allow the SVM to assign relative weights to various datasets, depending upon their utilities in performing a given classification task. RESULTS: In this work, we empirically investigate the performance of the SVM on the task of inferring gene functional annotations from a combination of protein sequence and structure data. Our results suggest that the SVM is quite robust to noise in the input datasets. Consequently, in the presence of only two types of data, an SVM trained from an unweighted combination of datasets performs as well or better than a more sophisticated algorithm that assigns weights to individual data types. Indeed, for this simple case, we can demonstrate empirically that no solution is significantly better than the naive, unweighted average of the two datasets. On the other hand, when multiple noisy datasets are included in the experiment, then the naive approach fares worse than the weighted approach. Our results suggest that for many applications, a naive unweighted sum of kernels may be sufficient. AVAILABILITY: http://noble.gs.washington.edu/proj/seqstruct

Algorithms↗

Chromosome-Level Genome Assembly and Annotation of the Chinese Lizard Gudgeon (Saurogobio dabryi).

The Chinese lizard gudgeon (Saurogobio dabryi) is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in China, the lack of genomic resources has rendered the genetic breeding and conservation research. Here, we present the first chromosome-level genome assembly of S. dabryi using PacBio HiFi long reads, short reads, and Hi-C sequencing data. The final assembly reaches a total size of 1.09 Gb and Hi-C scaffolding anchors 99.55% of the assembled contigs onto 25 chromosomes, with a scaffold N50 reaching 43.15 Mb. The final genome assembly shows a BUSCO completeness of 98.39%. We annotated 659.55 Mb repetitive sequences and 26,036 protein-coding genes, 99.47% of which are functionally annotated. Comparative phylogenomic analysis clarifies the phylogenetic position of Saurogobio within Gobioninae. This high-quality genome provides a critical genetic basis for exploring cyprinid phylogeny, benthic adaptive evolution, genetic improvement, and conservation efforts of S. dabryi.

Saurogobio dabryi↗

The SBASE domain sequence library, release 10: domain architecture prediction.

SBASE (http://www.icgeb.trieste.it/sbase) is an on-line collection of protein domain sequences and related computational tools designed to facilitate detection of domain homologies based on simple database search. The 10th 'jubilee release' of the SBASE library of protein domain sequences contains 1 052 904 protein sequence segments annotated by structure, function, ligand-binding or cellular topology, clustered into over 6000 domain groups. Domain identification and functional prediction are based on a comparison of BLAST search outputs with a knowledge base of biologically significant similarities extracted from known domain groups. The knowledge base is generated automatically for each domain group from the comparison of within-group ('self') and out-of-group ('non-self') similarities. This is a memory-based approach wherein group-specific similarity functions are automatically learned from the database.

Animals↗

MutDB: annotating human variation with functionally relevant data.

SUMMARY: We have developed a resource, MutDB (http://mutdb.org/), to aid in determining which single nucleotide polymorphisms (SNPs) are likely to alter the function of their associated protein product. MutDB contains protein structure annotations and comparative genomic annotations for 8000 disease-associated mutations and SNPs found in the UCSC Annotated Genome and the human RefSeq gene set. MutDB provides interactive mutation maps at the gene and protein levels, and allows for ranking of their predicted functional consequences based on conservation in multiple sequence alignments. AVAILABILITY: http://mutdb.org/ SUPPLEMENTARY INFORMATION: http://mutdb.org/about/about.html

Database Management Systems↗

A description scheme of biological processes based on elementary bricks of action.

With the fast growth of high-throughput strategies in Biology, there is a strong need to accelerate knowledge acquisition and organization of molecular functions. Unfortunately, although we know that there is a correlation between protein molecules and their functions, we are unable to clearly identify this link. Here, we revisit the current views of protein functions as well as their annotation, and we show that they are incompatible with unambiguous interpretations and the use of this knowledge. We describe herein a description scheme for biological processes based on elementary bricks of action that may be associated with biological molecules. To retrieve the descriptive quality found in annotations of other kinds of biological data, it was decided to develop a scheme involving four levels of abstraction: Basic Elements of Action, Biological Activities, Biological Functionalities and Biological Roles. This multi-level organization is a generic method; it allows for a description of biological processes by using a limited number of elementary bricks of action. Moreover, by using this description of biological processes, it should now be possible to clearly identify unambiguous relationships between the organization of biological processes and the structural or functional organizations of biological molecules.

Algorithms↗

Improved prediction of protein-protein binding sites using a support vector machines approach.

MOTIVATION: Structural genomics projects are beginning to produce protein structures with unknown function, therefore, accurate, automated predictors of protein function are required if all these structures are to be properly annotated in reasonable time. Identifying the interface between two interacting proteins provides important clues to the function of a protein and can reduce the search space required by docking algorithms to predict the structures of complexes. RESULTS: We have combined a support vector machine (SVM) approach with surface patch analysis to predict protein-protein binding sites. Using a leave-one-out cross-validation procedure, we were able to successfully predict the location of the binding site on 76% of our dataset made up of proteins with both transient and obligate interfaces. With heterogeneous cross-validation, where we trained the SVM on transient complexes to predict on obligate complexes (and vice versa), we still achieved comparable success rates to the leave-one-out cross-validation suggesting that sufficient properties are shared between transient and obligate interfaces. AVAILABILITY: A web application based on the method can be found at http://www.bioinformatics.leeds.ac.uk/ppi_pred. The dataset of 180 proteins used in this study is also available via the same web site. CONTACT: westhead@bmb.leeds.ac.uk SUPPLEMENTARY INFORMATION: http://www.bioinformatics.leeds.ac.uk/ppi-pred/supp-material.

Algorithms↗

Assessing semantic similarity measures for the characterization of human regulatory pathways.

MOTIVATION: Pathway modeling requires the integration of multiple data including prior knowledge. In this study, we quantitatively assess the application of Gene Ontology (GO)-derived similarity measures for the characterization of direct and indirect interactions within human regulatory pathways. The characterization would help the integration of prior pathway knowledge for the modeling. RESULTS: Our analysis indicates information content-based measures outperform graph structure-based measures for stratifying protein interactions. Measures in terms of GO biological process and molecular function annotations can be used alone or together for the validation of protein interactions involved in the pathways. However, GO cellular component-derived measures may not have the ability to separate true positives from noise. Furthermore, we demonstrate that the functional similarity of proteins within known regulatory pathways decays rapidly as the path length between two proteins increases. Several logistic regression models are built to estimate the confidence of both direct and indirect interactions within a pathway, which may be used to score putative pathways inferred from a scaffold of molecular interactions.

Databases, Protein↗

A comparison of position-specific score matrices based on sequence and structure alignments.

Sequence comparison methods based on position-specific score matrices (PSSMs) have proven a useful tool for recognition of the divergent members of a protein family and for annotation of functional sites. Here we investigate one of the factors that affects overall performance of PSSMs in a PSI-BLAST search, the algorithm used to construct the seed alignment upon which the PSSM is based. We compare PSSMs based on alignments constructed by global sequence similarity (ClustalW and ClustalW-pairwise), local sequence similarity (BLAST), and local structure similarity (VAST). To assess performance with respect to identification of conserved functional or structural sites, we examine the accuracy of the three-dimensional molecular models predicted by PSSM-sequence alignments. Using the known structures of those sequences as the standard of truth, we find that model accuracy varies with the algorithm used for seed alignment construction in the pattern local-structure (VAST) > local-sequence (BLAST) > global-sequence (ClustalW). Using structural similarity of query and database proteins as the standard of truth, we find that PSSM recognition sensitivity depends primarily on the diversity of the sequences included in the alignment, with an optimum around 30-50% average pairwise identity. We discuss these observations, and suggest a strategy for constructing seed alignments that optimize PSSM-sequence alignment accuracy and recognition sensitivity.

Algorithms↗

Random sequencing of cDNA library derived from partially-fed adult female Haemaphysalis longicornis salivary gland.

A cDNA library was constructed from salivary glands of partially-fed adult female Haemaphysalis longicornis (hard tick). Randomly selected clones were sequenced and a total of 633 sequences were analyzed by bioinformatic programs. The sequences were grouped into 213 clusters, with each cluster being considered to be composed of mRNAs derived from the same gene or closely related genes. About 36% of the mRNA sequences showed significant similarity to known proteins in the non-redundant protein database by the NCBI blastx program and appeared to be coding for functional predicted proteins, whereas the remaining 64% had no similar sequences. Two thirds of the predicted proteins were annotated as basic cellular proteins (housekeeping proteins). Among the functional predicted protein sequences, other than the housekeeping proteins, several protease inhibitors including anticoagulants, two metalloproteases and a potential immunosuppressive protein could be identified. These proteins may play important roles during tick feeding and could be novel anti-tick vaccine candidates.

Amino Acid Sequence↗