Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

PEP: Predictions for Entire Proteomes.

PEP is a database of Predictions for Entire Proteomes. The database contains summaries of analyses of protein sequences from a range of organisms representing all three major kingdoms of life: eukaryotes, prokaryotes and archaea. All proteins publicly available for organisms were aligned against SWISS-PROT, TrEMBL and PDB. Additionally, the following annotations are provided: secondary structure, transmembrane helices, coiled coils, regions of low complexity, signal peptides, PROSITE motifs, nuclear localization signals and classes of cellular function. Proteins that contain long regions without regular secondary structure are also identified. We have produced a related database of structural domain-like fragments derived from PEP and clusters based on homology between all fragments. The PEP database, fragments and clusters are distributed freely as a set of flat files and have been integrated into SRS. The PEP group of databases can be accessed from: http://cubic.bioc.columbia.edu/pep.

Animals↗

A classification of disulfide patterns and its relationship to protein structure and function.

We report a detailed classification of disulfide patterns to further understand the role of disulfides in protein structure and function. The classification is applied to a unique searchable database of disulfide patterns derived from the SwissProt and Pfam databases. The disulfide database contains seven times the number of publicly available disulfide annotations. Each disulfide pattern in the database captures the topology and cysteine spacing of a protein domain. We have clustered the domains by their disulfide patterns and visualized the results using a novel representation termed the "classification wheel." The classification is applied to 40,620 protein domains with 2-10 disulfides. The effectiveness of the classification is evaluated by determining the extent to which proteins of similar structure and function are grouped together through comparison with the SCOP and Pfam databases, respectively. In general, proteins with similar disulfide patterns have similar structure and function, even in cases of low sequence similarity, and we illustrate this with specific examples. Using a measure of disulfide topology complexity, we find that there is a predominance of less complex topologies. We also explored the importance of loss or addition of disulfides to protein structure and function by linking classification wheels through disulfide subpattern comparisons. This classification, when coupled with our disulfide database, will serve as a useful resource for searching and comparing disulfide patterns, and understanding their role in protein structure, folding, and stability. Proteins in the disulfide clusters that do not contain structural information are prime candidates for structural genomics initiatives, because they may correspond to novel structures.

Animals↗

The Feasibility of Using Proteome Expression Profile for Genome Annotation.

By investigating into the expression data from ECO2DBASE (Edition 6),the feasibility of using proteome expression profile for genome annotation was tested. Based on our newly developed CRC (cellular role cluster) method,79 proteins extracted from ECO2DBASE were clustered into 4 CRCs. Function related proteins tend to be clustered into same CRC. Total 9 aminoacyl-tRNA synthetases were clustered into CRC2, whereas 4 heat-shock proteins into CRC3. These results indicate with enough proteome expression data and the efficient algorithm, proteome expression profile can provide very important information for genome annotation, while this kind of information is sequence-independent.

Journal Article↗

Functional cloning, sorting, and expression profiling of nucleic acid-binding proteins.

A major challenge in the post-sequencing era is to elucidate the activity and biological function of genes that reside in the human genome. An important subset includes genes that encode proteins that regulate gene expression or maintain the structural integrity of the genome. Using a novel oligonucleotide-binding substrate as bait, we show the feasibility of a modified functional expression-cloning strategy to identify human cDNAs that encode a spectrum of nucleic acid-binding proteins (NBPs). Approximately 170 cDNAs were identified from screening phage libraries derived from a human colorectal adenocarcinoma cell line and from noncancerous fetal lung tissue. Sequence analysis confirmed that virtually every clone contained a known DNA- or RNA-binding motif. We also report on a complementary sorting strategy that, in the absence of subcloning and protein purification, can distinguish different classes of NBPs according to their particular binding properties. To extend our functional annotation of NBPs, we have used GeneChip expression profiling of 14 different breast-derived cell lines to examine the relative transcriptional activity of genes identified in our screen and cluster analysis to discover other genes that have similar expression patterns. Finally, we present strategies to analyze the upstream regulatory region of each gene within a cluster group and select unique combinations of transcription factor binding sites that may be responsible for dictating the observed synexpression.

Adenocarcinoma↗

Slr2013 is a novel protein regulating functional assembly of photosystem II in Synechocystis sp. strain PCC 6803.

The Synechocystis sp. strain PCC 6803, which has a T192H mutation in the D2 protein of photosystem II, is an obligate photoheterotroph due to the lack of assembled photosystem II complexes. A secondary mutant, Rg2, has been selected that retains the T192H mutation but is able to grow photoautotrophically. Restoration of photoautotrophic growth in this mutant was caused by early termination at position 294 in the Slr2013 protein. The T192H mutant with truncated Slr2013 forms fully functional photosystem II reaction centers that differ from wild-type reaction centers only by a 30% higher rate of charge recombination between the primary electron acceptor, QA-, and the donor side and by a reduced stability of the oxidized form of the redox-active Tyr residue, YD, in the D2 protein. This suggests that the T192H mutation itself did not directly affect electron transfer components, but rather affected protein folding and/or stable assembly of photosystem II, and that Slr2013 is involved in the folding of the D2 protein and the assembly of photosystem II. Besides participation in photosystem II assembly, Slr2013 plays a critical role in the cell, because the corresponding gene cannot be deleted completely under conditions in which photosystem II is dispensable. Truncation of Slr2013 by itself does not affect photosynthetic activity of Synechocystis sp. strain PCC 6803. Slr2013 is annotated in CyanoBase as a hypothetical protein and shares a DUF58 family signature with other hypothetical proteins of unknown function. Genes for close homologues of Slr2013 are found in other cyanobacteria (Nostoc punctiforme, Anabaena sp. strain PCC 7120, and Thermosynechococcus elongatus BP-1), and apparent orthologs of this protein are found in Eubacteria and Archaea, but not in eukaryotes. We suggest that Slr2013 regulates functional assembly of photosystem II and has at least one other important function in the cell.

Amino Acid Sequence↗

Integrative annotation of 21,037 human genes validated by full-length cDNA clones.

The human genome sequence defines our inherent biological potential; the realization of the biology encoded therein requires knowledge of the function of each gene. Currently, our knowledge in this area is still limited. Several lines of investigation have been used to elucidate the structure and function of the genes in the human genome. Even so, gene prediction remains a difficult task, as the varieties of transcripts of a gene may vary to a great extent. We thus performed an exhaustive integrative characterization of 41,118 full-length cDNAs that capture the gene transcripts as complete functional cassettes, providing an unequivocal report of structural and functional diversity at the gene level. Our international collaboration has validated 21,037 human gene candidates by analysis of high-quality full-length cDNA clones through curation using unified criteria. This led to the identification of 5,155 new gene candidates. It also manifested the most reliable way to control the quality of the cDNA clones. We have developed a human gene database, called the H-Invitational Database (H-InvDB; http://www.h-invitational.jp/). It provides the following: integrative annotation of human genes, description of gene structures, details of novel alternative splicing isoforms, non-protein-coding RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein three-dimensional structure, mapping of known single nucleotide polymorphisms (SNPs), identification of polymorphic microsatellite repeats within human genes, and comparative results with mouse full-length cDNAs. The H-InvDB analysis has shown that up to 4% of the human genome sequence (National Center for Biotechnology Information build 34 assembly) may contain misassembled or missing regions. We found that 6.5% of the human gene candidates (1,377 loci) did not have a good protein-coding open reading frame, of which 296 loci are strong candidates for non-protein-coding RNA genes. In addition, among 72,027 uniquely mapped SNPs and insertions/deletions localized within human genes, 13,215 nonsynonymous SNPs, 315 nonsense SNPs, and 452 indels occurred in coding regions. Together with 25 polymorphic microsatellite repeats present in coding regions, they may alter protein structure, causing phenotypic effects or resulting in disease. The H-InvDB platform represents a substantial contribution to resources needed for the exploration of human biology and pathology.

Alternative Splicing↗

Functional annotation of proteomic sequences based on consensus of sequence and structural analysis.

To maximise the assignment of function of the proteins encoded by a genome and to aid the search for novel drug targets, there is an emerging need for sensitive methods of predicting protein function on a genome-wide basis. GeneAtlas is an automated, high-throughput pipeline for the prediction of protein structure and function using sequence similarity detection, homology modelling and fold recognition methods. GeneAtlas is described in detail here. To test GeneAtlas, a 'virtual' genome was used, a subset of PDB structures from the SCOP database, in which the functional relationships are known. GeneAtlas detects additional relationships by building 3D models in comparison with the sequence searching method PSI-BLAST. Functionally related proteins with sequence identity below the twilight zone can be recognised correctly.

Consensus Sequence↗

PASS2: a semi-automated database of protein alignments organised as structural superfamilies.

PASS2 is a nearly automated version of CAMPASS and contains sequence alignments of proteins grouped at the level of superfamilies. This database has been created to fall in correspondence with SCOP database (1.53 release) and currently consists of 110 multi-member superfamilies and 613 superfamilies corresponding to single members. In multi-member superfamilies, protein chains with no more than 25% sequence identity have been considered for the alignment and hence the database aims to address sequence alignments which represent 26 219 protein domains under the SCOP 1.53 release. Structure-based sequence alignments have been obtained by COMPARER and the initial equivalences are provided automatically from a MALIGN alignment and subsequently augmented using STAMP4.0. The final sequence alignments have been annotated for the structural features using JOY4.0. Several interesting links are provided to other related databases and genome sequence relatives. Availability of reliable sequence alignments of distantly related proteins, despite poor sequence identity and single-member superfamilies, permit better sampling of structures in libraries for fold recognition of new sequences and for the understanding of protein structure-function relationships of individual superfamilies. The database can be queried by keywords and also by sequence search, interfaced by PSI-BLAST methods. Structure-annotated sequence alignments and several structural accessory files can be retrieved for all the superfamilies including the user-input sequence. The database can be accessed from http://www.ncbs.res.in/%7Efaculty/mini/campass/pass.html.

Amino Acid Sequence↗

MIPS: a database for genomes and protein sequences.

The Munich Information Center for Protein Sequences (MIPS-GSF), Martinsried near Munich, Germany, develops and maintains genome oriented databases. It is commonplace that the amount of sequence data available increases rapidly, but not the capacity of qualified manual annotation at the sequence databases. Therefore, our strategy aims to cope with the data stream by the comprehensive application of analysis tools to sequences of complete genomes, the systematic classification of protein sequences and the active support of sequence analysis and functional genomics projects. This report describes the systematic and up-to-date analysis of genomes (PEDANT), a comprehensive database of the yeast genome (MYGD), a database reflecting the progress in sequencing the Arabidopsis thaliana genome (MATD), the database of assembled, annotated human EST clusters (MEST), and the collection of protein sequence data within the framework of the PIR-International Protein Sequence Database (described elsewhere in this volume). MIPS provides access through its WWW server (http://www.mips.biochem.mpg.de) to a spectrum of generic databases, including the above mentioned as well as a database of protein families (PROTFAM), the MITOP database, and the all-against-all FASTA database.

Amino Acid Sequence↗

AutoFACT: an automatic functional annotation and classification tool.

BACKGROUND: Assignment of function to new molecular sequence data is an essential step in genomics projects. The usual process involves similarity searches of a given sequence against one or more databases, an arduous process for large datasets. RESULTS: We present AutoFACT, a fully automated and customizable annotation tool that assigns biologically informative functions to a sequence. Key features of this tool are that it (1) analyzes nucleotide and protein sequence data; (2) determines the most informative functional description by combining multiple BLAST reports from several user-selected databases; (3) assigns putative metabolic pathways, functional classes, enzyme classes, GeneOntology terms and locus names; and (4) generates output in HTML, text and GFF formats for the user's convenience. We have compared AutoFACT to four well-established annotation pipelines. The error rate of functional annotation is estimated to be only between 1-2%. Comparison of AutoFACT to the traditional top-BLAST-hit annotation method shows that our procedure increases the number of functionally informative annotations by approximately 50%. CONCLUSION: AutoFACT will serve as a useful annotation tool for smaller sequencing groups lacking dedicated bioinformatics staff. It is implemented in PERL and runs on LINUX/UNIX platforms. AutoFACT is available at http://megasun.bch.umontreal.ca/Software/AutoFACT.htm.

Acanthamoeba castellanii↗

Organelle DB: a cross-species database of protein localization and function.

To efficiently utilize the growing body of available protein localization data, we have developed Organelle DB, a web-accessible database cataloging more than 25,000 proteins from nearly 60 organelles, subcellular structures and protein complexes in 154 organisms spanning the eukaryotic kingdom. Organelle DB is the first on-line resource devoted to the identification and presentation of eukaryotic proteins localized to organelles and subcellular structures. As such, Organelle DB is a strong resource of data from the human proteome as well as from the major model organisms Saccharomyces cerevisiae, Arabidopsis thaliana, Drosophila melanogaster, Caenorhabditis elegans and Mus musculus. In particular, Organelle DB is a central repository of yeast data, incorporating results--and actual fluorescent imagesfrom ongoing large-scale studies of protein localization in S.cerevisiae. Each protein in Organelle DB is presented with its sequence and, as available, a detailed description of its function; functions were extracted from relevant model organism databases, and links to these databases are provided within Organelle DB. To facilitate data interoperability, we have annotated all protein localizations using vocabulary from the Gene Ontology consortium. We also welcome new data for inclusion in Organelle DB, which may be freely accessed at http://organelledb.lsi.umich.edu.

Animals↗

The SCAN domain family of zinc finger transcription factors.

Zinc finger transcription factor genes represent a significant portion of the genes in the vertebrate genome. Some Cys2His2 type zinc fingers are associated with conserved protein domains that help to define these regulators. A novel domain of this type, the SCAN domain, is a highly conserved 84-residue motif that is found near the N-terminus of a subfamily of C2H2 zinc finger proteins. The SCAN domain, which is also known as the leucine rich region, functions as a protein interaction domain, mediating self-association or selective association with other proteins. Here we define the mouse SCAN domain and annotate the mouse SCAN family members. In addition to a single SCAN domain, some of the members of the mouse SCAN family members have a conserved N-terminal motif, a KRAB domain, SANT domains and a variable number of C2H2 type zinc fingers (3-14). The genes encoding mouse SCAN domains are clustered, often in tandem arrays, and are capable of generating isoforms that may affect the function of family members. Although the function of most of the family members is not known, an overview of selected members of this group of transcription factors suggests that some of the mouse SCAN domain family members play roles in cell survival and differentiation.

Amino Acid Sequence↗

Role of context in the relationship between form and function: structural plasticity of some PROSITE patterns.

True positive hits of PROSITE sequence pattern are expected to have a characteristic three-dimensional structure. The combined sequence-structure attributes of PROSITE patterns can be used for function prediction of an uncharacterized protein with known primary and 3D structure, a situation that might arise in structural genomics projects. We have found specific examples of true hits of PROSITE patterns displaying structural plasticity by assuming significantly different local conformation, depending upon the context. Our work highlights the importance of taking into account all the known distinct conformations of PROSITE patterns, while creating a sensitive 3D template for the pattern, for use in functional annotation.

Amino Acid Motifs↗

Large-scale testing of bibliome informatics using Pfam protein families.

Literature mining is expected to help not only with automatically sifting through huge biomedical literature and annotation databases, but also with linking bio-chemical entities to appropriate functional hypotheses. However, there has been very limited success in testing literature mining methods due to the lack of large, objectively validated test sets or "gold standards". To improve this situation we created a large-scale test of literature mining methods and resources. We report on a specific implementation of this test: how well can the Pfam protein family classification be replicated from independently mining different literature/annotation resources? We test and compare different keyterm sets as well as different algorithms for issuing protein family predictions. We find that protein families can indeed be automatically predicted from the literature. Using words from PubMed abstracts, of 3663 proteins tested, over 75% were correctly assigned to one of 618 Pfam families. For 90% of proteins the correct Pfam family was among the top 5 ranked families. We found that protein family prediction is far superior with keywords extracted from PubMed abstracts than with GO annotations or MeSH keyterms, suggesting that the text itself (in combination with the vector space model) is superior to GO and MeSH as a literature mining resources, at least for detecting protein family membership. Finally, we show that Shannon's entropy can be exploited to improve prediction by facilitating the integration of the different literature sources tested.

Algorithms↗

A cytosolic Arabidopsis D-xylulose kinase catalyzes the phosphorylation of 1-deoxy-D-xylulose into a precursor of the plastidial isoprenoid pathway.

Plants are able to integrate exogenous 1-deoxy-D-xylulose (DX) into the 2C-methyl-D-erythritol 4-phosphate pathway, implicated in the biosynthesis of plastidial isoprenoids. Thus, the carbohydrate needs to be phosphorylated into 1-deoxy-D-xylulose 5-phosphate and translocated into plastids, or vice versa. An enzyme capable of phosphorylating DX was partially purified from a cell-free Arabidopsis (Arabidopsis thaliana) protein extract. It was identified by mass spectrometry as a cytosolic protein bearing D-xylulose kinase (XK) signatures, already suggesting that DX is phosphorylated within the cytosol prior to translocation into the plastids. The corresponding cDNA was isolated and enzymatic properties of a recombinant protein were determined. In Arabidopsis, xylulose kinases are encoded by a small gene family, in which only two genes are putatively annotated. The additional gene is coding for a protein targeted to plastids, as was proved by colocalization experiments using green fluorescent protein fusion constructs. Functional complementation assays in an Escherichia coli strain deleted in xk revealed that the cytosolic enzyme could exclusively phosphorylate xylulose in vivo, not the enzyme that is targeted to plastids. xk activities could not be detected in chloroplast protein extracts or in proteins isolated from its ancestral relative Synechocystis sp. PCC 6803. The gene encoding the plastidic protein annotated as "xylulose kinase" might in fact yield an enzyme having different phosphorylation specificities. The biochemical characterization and complementation experiments with DX of specific Arabidopsis knockout mutants seedlings treated with oxo-clomazone, an inhibitor of 1-deoxy-D-xylulose 5-phosphate synthase, further confirmed that the cytosolic protein is responsible for the phosphorylation of DX in planta.

Arabidopsis↗

MMDB: Entrez's 3D-structure database.

Three-dimensional structures are now known within most protein families and it is likely, when searching a sequence database, that one will identify a homolog of known structure. The goal of Entrez's 3D-structure database is to make structure information and the functional annotation it can provide easily accessible to molecular biologists. To this end, Entrez's search engine provides several powerful features: (i) links between databases, for example between a protein's sequence and structure; (ii) pre-computed sequence and structure neighbors; and (iii) structure and sequence/structure alignment visualization. Here, we focus on a new feature of Entrez's Molecular Modeling Database (MMDB): Graphical summaries of the biological annotation available for each 3D structure, based on the results of automated comparative analysis. MMDB is available at: http://www.ncbi.nlm.nih.gov/Entrez/structure.html.

Animals↗

Identifying fundamental gaps in functional metagenomics: a step towards unlocking microbiome research potential.

Incomplete functional annotation limits biological interpretation in microbiome studies and their translational potential. Poor annotation arises from multiple causes, with incomplete gene-protein-reaction mapping being one tractable yet under-examined contributor. We address this gap by developing a comprehensive hierarchical framework that systematically integrates gene families in UniRef, proteins in UniProt, and metabolic reactions in MetaCyc and BioCyc through UniProtKB accession, EC number, and Pfam-domain matching. Applied to a human gut metagenome dataset via HUMAnN3, our MetaCyc-based mapping recovers up to 2.3-fold more unique reaction identifiers than the default pipeline and increases reaction prevalence across samples from ≈32% to 52% core reactions, addressing the data sparsity that limits statistical and machine-learning applications in microbiome research. Biological plausibility for the tested functions was supported by positive and negative controls: gut-microbial hormone-metabolism reactions previously linked to this dataset were recovered, while vertebrate-specific hormone-metabolism reactions remained correctly undetected. These gains derive from systematic database integration alone, without predictive algorithms, indicating that a tractable, mapping-related component of functional dark matter and data sparsity in microbiome studies is directly addressable. Because Pfam- and BioCyc-derived mappings trade specificity for coverage, confidence in any individual reaction assignment depends on the supporting evidence tier and source database.

Humans↗

Counting the zinc-proteins encoded in the human genome.

Metalloproteins are proteins capable of binding one or more metal ions, which may be required for their biological function, or for regulation of their activities or for structural purposes. Genome sequencing projects have provided a huge number of protein primary sequences, but, even though several different elaborate analyses and annotations have been enabled by a rich and ever-increasing portfolio of bioinformatic tools, metal-binding properties remain difficult to predict as well as to investigate experimentally. Consequently, the present knowledge about metalloproteins is only partial. The present bioinformatic research proposes a strategy to answer the question of how many and which proteins encoded in the human genome may require zinc for their physiological function. This is achieved by a combination of approaches, which include: (i) searching in the proteome for the zinc-binding patterns that, on their turn, are obtained from all available X-ray data; (ii) using libraries of metal-binding protein domains based on multiple sequence alignments of known metalloproteins obtained from the Pfam database; and (iii) mining the annotations of human gene sequences, which are based on any type of information available. It is found that 1684 proteins in the human proteome are independently identified by all three approaches as zinc-proteins, 746 are identified by two, and 777 are identified by only one method. By assuming that all proteins identified by at least two approaches are truly zinc-binding and inspecting the proteins identified by a single method, it can be proposed that ca. 2800 human proteins are potentially zinc-binding in vivo, corresponding to 10% of the human proteome, with an uncertainty of 400 sequences. Available functional information suggests that the large majority of human zinc-binding proteins are involved in the regulation of gene expression. The most abundant class of zinc-binding proteins in humans is that of zinc-fingers, with Cys4 and Cys2His2 being the most common types of coordination environment.

Computational Biology↗