Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

iProClass: an integrated, comprehensive and annotated protein classification database.

The iProClass database is an integrated resource that provides comprehensive family relationships and structural and functional features of proteins, with rich links to various databases. It is extended from ProClass, a protein family database that integrates PIR superfamilies and PROSITE motifs. The iProClass currently consists of more than 200,000 non-redundant PIR and SWISS-PROT proteins organized with more than 28,000 superfamilies, 2600 domains, 1300 motifs, 280 post-translational modification sites and links to more than 30 databases of protein families, structures, functions, genes, genomes, literature and taxonomy. Protein and family summary reports provide rich annotations, including membership information with length, taxonomy and keyword statistics, full family relationships, comprehensive enzyme and PDB cross-references and graphical feature display. The database facilitates classification-driven annotation for protein sequence databases and complete genomes, and supports structural and functional genomic research. The iProClass is implemented in Oracle 8i object-relational system and available for sequence search and report retrieval at http://pir.georgetown.edu/iproclass/.

Databases, Factual↗

The DNA sequence and biological annotation of human chromosome 1.

The reference sequence for each human chromosome provides the framework for understanding genome function, variation and evolution. Here we report the finished sequence and biological annotation of human chromosome 1. Chromosome 1 is gene-dense, with 3,141 genes and 991 pseudogenes, and many coding sequences overlap. Rearrangements and mutations of chromosome 1 are prevalent in cancer and many other diseases. Patterns of sequence variation reveal signals of recent selection in specific genes that may contribute to human fitness, and also in regions where no function is evident. Fine-scale recombination occurs in hotspots of varying intensity along the sequence, and is enriched near genes. These and other studies of human biology and disease encoded within chromosome 1 are made possible with the highly accurate annotated sequence, as part of the completed set of chromosome sequences that comprise the reference human genome.

Base Sequence↗

Helicobacter pylori FlhB function: the FlhB C-terminal homologue HP1575 acts as a "spare part" to permit flagellar export when the HP0770 FlhBCC domain is deleted.

In Helicobacter pylori 26695, a gene annotated HP1575 encodes a putative protein of unknown function which shows significant similarity to part of the C-terminal domain of the flagellar export protein FlhB. In Salmonella enterica, this part (FlhB(CC)) is proteolytically cleaved from the full-length FlhB, a processing event that is required for flagellar protein export and, thus, motility. The role of FlhB (HP0770) and its C-terminal homologue HP1575 was studied in H. pylori using a range of nonpolar deletion mutants defective in HP1575, HP0770, and the CC domain of HP0770 (HP0770(CC)). Deletion of HP0770 abolished swimming motility, whereas mutants carrying a deletion of either HP1575 or HP0770(CC) retained their ability to swim. An H. pylori strain containing deletions in both HP1575 and HP0770(CC) was nonmotile and did not produce flagella, suggesting that at least one of the two proteins had to be present for flagellar assembly to occur. Indeed, motility was restored when HP1575 was reintroduced into this strain immediately downstream of, but not fused to, the truncated HP0770 gene. Thus, HP1575 can functionally replace HP0770(CC) in this background. Like FlhB in S. enterica, HP0770 appeared to be proteolytically processed at a conserved NPTH processing site. However, mutation of the proline contained within the NPTH site of HP0770 did not affect motility and flagellar assembly, although it clearly interfered with processing when the protein was heterologously produced in Escherichia coli.

Amino Acid Sequence↗

A focused microarray approach to functional glycomics: transcriptional regulation of the glycome.

Glycosylation is the most common posttranslational modification of proteins, yet genes relevant to the synthesis of glycan structures and function are incompletely represented and poorly annotated on the commercially available arrays. To fill the need for expression analysis of such genes, we employed the Affymetrix technology to develop a focused and highly annotated glycogene-chip representing human and murine glycogenes, including glycosyltransferases, nucleotide sugar transporters, glycosidases, proteoglycans, and glycan-binding proteins. In this report, the array has been used to generate glycogene-expression profiles of nine murine tissues. Global analysis with a hierarchical clustering algorithm reveals that expression profiles in immune tissues (thymus [THY], spleen [SPL], lymph node, and bone marrow [BM]) are more closely related, relative to those of nonimmune tissues (kidney [KID], liver [LIV], brain [BRN], and testes [TES]). Of the biosynthetic enzymes, those responsible for synthesis of the core regions of N- and O-linked oligosaccharides are ubiquitously expressed, whereas glycosyltransferases that elaborate terminal structures are expressed in a highly tissue-specific manner, accounting for tissue and ultimately cell-type-specific glycosylation. Comparison of gene expression profiles with matrix-assisted laser desorption ionization-time of flight (MALDI-TOF) profiling of N-linked oligosaccharides suggested that the alpha1-3 fucosyltransferase 9, Fut9, is the enzyme responsible for terminal fucosylation in KID and BRN, a finding validated by analysis of Fut9 knockout mice. Two families of glycan-binding proteins, C-type lectins and Siglecs, are predominately expressed in the immune tissues, consistent with their emerging functions in both innate and acquired immunity. The glycogene chip reported in this study is available to the scientific community through the Consortium for Functional Glycomics (CFG) (http://www.functionalglycomics.org).

Animals↗

Annotating the genome of Medicago truncatula.

Medicago truncatula will be among the first plant species to benefit from the completion of a whole-genome sequencing project. For each of these species, Arabidopsis, rice and now poplar and Medicago, annotation, the process of identifying gene structures and defining their functions, is essential for the research community to benefit from the sequence data generated. Annotation of the Arabidopsis genome involved gene-by-gene curation of the entire genome, but the larger genomes of rice, Medicago and other species necessitate the automation of the annotation process. Profiting from the experience gained from previous whole-genome efforts, a uniform set of Medicago gene annotations has been generated by coordinated international effort and, along with other views of the genome data, has been provided to the research community at several websites.

Automation↗

oPOSSUM: identification of over-represented transcription factor binding sites in co-expressed genes.

Targeted transcript profiling studies can identify sets of co-expressed genes; however, identification of the underlying functional mechanism(s) is a significant challenge. Established methods for the analysis of gene annotations, particularly those based on the Gene Ontology, can identify functional linkages between genes. Similar methods for the identification of over-represented transcription factor binding sites (TFBSs) have been successful in yeast, but extension to human genomics has largely proved ineffective. Creation of a system for the efficient identification of common regulatory mechanisms in a subset of co-expressed human genes promises to break a roadblock in functional genomics research. We have developed an integrated system that searches for evidence of co-regulation by one or more transcription factors (TFs). oPOSSUM combines a pre-computed database of conserved TFBSs in human and mouse promoters with statistical methods for identification of sites over-represented in a set of co-expressed genes. The algorithm successfully identified mediating TFs in control sets of tissue-specific genes and in sets of co-expressed genes from three transcript profiling studies. Simulation studies indicate that oPOSSUM produces few false positives using empirically defined thresholds and can tolerate up to 50% noise in a set of co-expressed genes.

Algorithms↗

pdbFun: mass selection and fast comparison of annotated PDB residues.

pdbFun (http://pdbfun.uniroma2.it) is a web server for structural and functional analysis of proteins at the residue level. pdbFun gives fast access to the whole Protein Data Bank (PDB) organized as a database of annotated residues. The available data (features) range from solvent exposure to ligand binding ability, location in a protein cavity, secondary structure, residue type, sequence functional pattern, protein domain and catalytic activity. Users can select any residue subset (even including any number of PDB structures) by combining the available features. Selections can be used as probe and target in multiple structure comparison searches. For example a search could involve, as a query, all solvent-exposed, hydrophylic residues that are not in alpha-helices and are involved in nucleotide binding. Possible examples of targets are represented by another selection, a single structure or a dataset composed of many structures. The output is a list of aligned structural matches offered in tabular and also graphical format.

Algorithms↗

Analysis and prediction of functional sub-types from protein sequence alignments.

The increasing number and diversity of protein sequence families requires new methods to define and predict details regarding function. Here, we present a method for analysis and prediction of functional sub-types from multiple protein sequence alignments. Given an alignment and set of proteins grouped into sub-types according to some definition of function, such as enzymatic specificity, the method identifies positions that are indicative of functional differences by comparison of sub-type specific sequence profiles, and analysis of positional entropy in the alignment. Alignment positions with significantly high positional relative entropy correlate with those known to be involved in defining sub-types for nucleotidyl cyclases, protein kinases, lactate/malate dehydrogenases and trypsin-like serine proteases. We highlight new positions for these proteins that suggest additional experiments to elucidate the basis of specificity. The method is also able to predict sub-type for unclassified sequences. We assess several variations on a prediction method, and compare them to simple sequence comparisons. For assessment, we remove close homologues to the sequence for which a prediction is to be made (by a sequence identity above a threshold). This simulates situations where a protein is known to belong to a protein family, but is not a close relative of another protein of known sub-type. Considering the four families above, and a sequence identity threshold of 30 %, our best method gives an accuracy of 96 % compared to 80 % obtained for sequence similarity and 74 % for BLAST. We describe the derivation of a set of sub-type groupings derived from an automated parsing of alignments from PFAM and the SWISSPROT database, and use this to perform a large-scale assessment. The best method gives an average accuracy of 94 % compared to 68 % for sequence similarity and 79 % for BLAST. We discuss implications for experimental design, genome annotation and the prediction of protein function and protein intra-residue distances.

Adenylyl Cyclases↗

Clustering proteins from interaction networks for the prediction of cellular functions.

BACKGROUND: Developing reliable and efficient strategies allowing to infer a function to yet uncharacterized proteins based on interaction networks is of crucial interest in the current context of high-throughput data generation. In this paper, we develop a new algorithm for clustering vertices of a protein-protein interaction network using a density function, providing disjoint classes. RESULTS: Applied to the yeast interaction network, the classes obtained appear to be biological significant. The partitions are then used to make functional predictions for uncharacterized yeast proteins, using an annotation procedure that takes into account the binary interactions between proteins inside the classes. We show that this procedure is able to enhance the performances with respect to previous approaches. Finally, we propose a new annotation for 37 previously uncharacterized yeast proteins. CONCLUSION: We believe that our results represent a significant improvement for the inference of cellular functions, that can be applied to other organism as well as to other type of interaction graph, such as genetic interactions.

Cluster Analysis↗

Gene3D: modelling protein structure, function and evolution.

The Gene3D release 4 database and web portal (http://cathwww.biochem.ucl.ac.uk:8080/Gene3D) provide a combined structural, functional and evolutionary view of the protein world. It is focussed on providing structural annotation for protein sequences without structural representatives--including the complete proteome sets of over 240 different species. The protein sequences have also been clustered into whole-chain families so as to aid functional prediction. The structural annotation is generated using HMM models based on the CATH domain families; CATH is a repository for manually deduced protein domains. Amongst the changes from the last publication are: the addition of over 100 genomes and the UniProt sequence database, domain data from Pfam, metabolic pathway and functional data from COGs, KEGG and GO, and protein-protein interaction data from MINT and BIND. The website has been rebuilt to allow more sophisticated querying and the data returned is presented in a clearer format with greater functionality. Furthermore, all data can be downloaded in a simple XML format, allowing users to carry out complex investigations at their own computers.

Databases, Protein↗

"A system biology" approach to bioinformatics and functional genomics in complex human diseases: arthritis.

Human and other annotated genome sequences have facilitated generation of vast amounts of correlative data, from human/animal genetics, normal and disease-affected tissues from complex diseases such as arthritis using gene/protein chips and SNP analysis. These data sets include genes/proteins whose functions are partially known at the cellular level or may be completely unknown (e.g. ESTs). Thus, genomic research has transformed molecular biology from "data poor" to "data rich" science, allowing further division into subpopulations of subcellular fractions, which are often given an "-omic" suffix. These disciplines have to converge at a systemic level to examine the structure and dynamics of cellular and organismal function. The challenge of characterizing ESTs linked to complex diseases is like interpreting sharp images on a blurred background and therefore requires a multidimensional screen for functional genomics ("functionomics") in tissues, mice and zebra fish model, which intertwines various approaches and readouts to study development and homeostasis of a system. In summary, the post-genomic era of functionomics will facilitate to narrow the bridge between correlative data and causative data by quaint hypothesis-driven research using a system approach integrating "intercoms" of interacting and interdependent disciplines forming a unified whole as described in this review for Arthritis.

Animals↗

Supra-domains: evolutionary units larger than single protein domains.

Domains are the evolutionary units that comprise proteins, and most proteins are built from more than one domain. Domains can be shuffled by recombination to create proteins with new arrangements of domains. Using structural domain assignments, we examined the combinations of domains in the proteins of 131 completely sequenced organisms. We found two-domain and three-domain combinations that recur in different protein contexts with different partner domains. The domains within these combinations have a particular functional and spatial relationship. These units are larger than individual domains and we term them "supra-domains". Amongst the supra-domains, we identified some 1400 (1203 two-domain and 166 three-domain) combinations that are statistically significantly over-represented relative to the occurrence and versatility of the individual component domains. Over one-third of all structurally assigned multi-domain proteins contain these over-represented supra-domains. This means that investigation of the structural and functional relationships of the domains forming these popular combinations would be particularly useful for an understanding of multi-domain protein function and evolution as well as for genome annotation. These and other supra-domains were analysed for their versatility, duplication, their distribution across the three kingdoms of life and their functional classes. By examining the three-dimensional structures of several examples of supra-domains in different biological processes, we identify two basic types of spatial relationships between the component domains: the combined function of the two domains is such that either the geometry of the two domains is crucial and there is a tight constraint on the interface, or the precise orientation of the domains is less important and they are spatially separate. Frequently, the role of the supra-domain becomes clear only once the three-dimensional structure is known. Since this is the case for only a quarter of the supra-domains, we provide a list of the most important unknown supra-domains as potential targets for structural genomics projects.

Animals↗

A novel database of disulfide patterns and its application to the discovery of distantly related homologs.

Disulfide bonds are conserved strongly among proteins of related structure and function. Despite the explosive growth of protein sequence databases and the vast numbers of sequence search tools, no tool exists to draw relations between the disulfide patterns of homologous proteins. We present a comprehensive database of disulfide bonding patterns and a search method to find proteins with similar disulfide patterns. The disulfide database was constructed using disulfide annotations extracted from SwissProt, and was expanded significantly from 16,736 to 94,499 disulfide-containing domains by an inference method that combines SwissProt annotations with Pfam multiple alignments. To search the database, we define a disulfide description, called the disulfide signature, which encodes both spacings between cysteine residues and cysteine connectivity. A web tool was developed that allows users to search for related disulfide patterns and for subpatterns resulting from the removal of one or more disulfides from the pattern. We explore the possibility of using disulfide pattern conservation to identify protein homologs that are undetectable by PSI-BLAST. Examples include the homology between a sea anemone antihypertensive/antiviral protein and a sea anemone neurotoxin, and the homology between tick anticoagulant peptide and bovine trypsin inhibitor. In both examples, there is a clear structural similarity and a functional relationship. We used the database to find structural homologs for the Cripto CFC domain. The identification of a von Willebrand Factor C (VWFC)-like domain agrees with its functional role and explains mutation data. We believe that the rapid increase in structure determinations arising from structural genomics efforts and advances in mass spectrometry techniques will greatly increase the number of disulfide annotations. This information will become a valuable resource for structural and functional annotations of proteins. The availability of a searchable disulfide pattern database will thus provide a powerful new addition to existing homolog discovery methods.

Amino Acid Sequence↗

Clustering the annotation space of proteins.

BACKGROUND: Current protein clustering methods rely on either sequence or functional similarities between proteins, thereby limiting inferences to one of these areas. RESULTS: Here we report a new approach, named CLAN, which clusters proteins according to both annotation and sequence similarity. This approach is extremely fast, clustering the complete SwissProt database within minutes. It is also accurate, recovering consistent protein families agreeing on average in more than 97% with sequence-based protein families from Pfam. Discrepancies between sequence- and annotation-based clusters were scrutinized and the reasons reported. We demonstrate examples for each of these cases, and thoroughly discuss an example of a propagated error in SwissProt: a vacuolar ATPase subunit M9.2 erroneously annotated as vacuolar ATP synthase subunit H. CLAN algorithm is available from the authors and the CLAN database is accessible at http://maine.ebi.ac.uk:8000/cgi-bin/clan/ClanSearch.pl CONCLUSIONS: CLAN creates refined function-and-sequence specific protein families that can be used for identification and annotation of unknown family members. It also allows easy identification of erroneous annotations by spotting inconsistencies between similarities on annotation and sequence levels.

Adenosine Triphosphatases↗

Gene and protein profiling of the response of MA-10 Leydig tumor cells to human chorionic gonadotropin.

Activation of the steroidogenic machinery by peptide hormones involves a number of steps for transmitting signals from the plasma membrane to mitochondria in a spatially and temporally coordinated manner. Although key proteins mediating the hormonal signal have been identified, recent data suggest that the pathway might involve more complex protein-protein and protein-lipid interactions. Genomic and proteomic methods of analysis, namely the Affymetrix Murine Genome U74A v2 GeneChip and the BD PowerBlot Western Array, were used to identify human chorionic gonadotropin (hCG)-induced changes in mRNA and protein of MA-10 Leydig tumor cells that parallel the increase seen in progesterone synthesis. To analyze the massive amount of data that was generated, a comprehensive protein information matrix summarizing the features of each gene or protein, including its known properties, as well as annotations derived by homology-based functional inference, was developed. Of the genes examined by Affymetrix array, approximately 79 were differentially expressed and of gene products examined by PowerBlot, 9 were differentially expressed (above twofold). Changes in the expression of selected transcripts of interest were confirmed using real-time quantitative polymerase chain reaction and immunoblot analyses. Collectively, these results indicate that hormonal regulation of steroidogenesis is a complex phenomenon, involving proteins that participate in various known and novel pathways, which are implicated in transmitting signals from the plasma membrane to mitochondria and nucleus.

Blotting, Western↗

Mapping the surface properties of macromolecules.

Methods are presented for the rapid computation of schematic projections of the surfaces of macromolecules, similar to the "roadmaps" used to illustrate the surfaces of viruses (Rossmann, M.G. & Palmenberg, A.C., 1988, Virology 164, 373-382). Several types of projections are described, extending the application of "roadmaps" to the external surfaces of all macromolecules and their interior binding pockets and pores. The surface projections, showing the positions of residues, can be colored, shaded, contoured, and annotated to show physical, sequence, or functional properties such as surface topology, hydrophobicity, or sequence conservation, for example. The automated procedures are useful for surveys of the surface features of proteins sharing similar functional properties.

Computer Graphics↗

Evaluation of protein fold comparison servers.

When a new protein structure has been determined, comparison with the database of known structures enables classification of its fold as new or belonging to a known class of proteins. This in turn may provide clues about the function of the protein. A large number of fold comparison programs have been developed, but they have never been subjected to a comprehensive and critical comparative analysis. Here we describe an evaluation of 11 publicly available, Web-based servers for automatic fold comparison. Both their functionality (e.g., user interface, presentation, and annotation of results) and their performance (i.e., how well established structural similarities are recognized) were assessed. The servers were subjected to a battery of performance tests covering a broad spectrum of folds as well as special cases, such as multidomain proteins, Calpha-only models, new folds, and NMR-based models. The CATH structural classification system was used as a reference. These tests revealed the strong and weak sides of each server. On the whole, CE, DALI, MATRAS, and VAST showed the best performance, but none of the servers achieved a 100% success rate. Where no structurally similar proteins are found by any individual server, it is recommended to try one or two other servers before any conclusions concerning the novelty of a fold are put on paper.

Computational Biology↗

Structural characterization of Salmonella typhimurium YeaZ, an M22 O-sialoglycoprotein endopeptidase homolog.

The Salmonella typhimurium "yeaZ" gene (StyeaZ) encodes an essential protein of unknown function (StYeaZ), which has previously been annotated as a putative homolog of the Pasteurella haemolytica M22 O-sialoglycoprotein endopeptidase Gcp. YeaZ has also recently been reported as the first example of an RPF from a gram-negative bacterial species. To further characterize the properties of StYeaZ and the widely occurring MK-M22 family, we describe the purification, biochemical analysis, crystallization, and structure determination of StYeaZ. The crystal structure of StYeaZ reveals a classic two-lobed actin-like fold with structural features consistent with nucleotide binding. However, microcalorimetry experiments indicated that StYeaZ neither binds polyphosphates nor a wide range of nucleotides. Additionally, biochemical assays show that YeaZ is not an active O-sialoglycoprotein endopeptidase, consistent with the lack of the critical zinc binding motif. We present a detailed comparison of YeaZ with available structural homologs, the first reported structural analysis of an MK-M22 family member. The analysis indicates that StYeaZ has an unusual orientation of the A and B lobes which may require substantial relative movement or interaction with a partner protein in order to bind ligands. Comparison of the fold of YeaZ with that of a known RPF domain from a gram-positive species shows significant structural differences and therefore potentially distinctive RPF mechanisms for these two bacterial classes.

Amino Acid Sequence↗