Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

A focused microarray approach to functional glycomics: transcriptional regulation of the glycome.

Glycosylation is the most common posttranslational modification of proteins, yet genes relevant to the synthesis of glycan structures and function are incompletely represented and poorly annotated on the commercially available arrays. To fill the need for expression analysis of such genes, we employed the Affymetrix technology to develop a focused and highly annotated glycogene-chip representing human and murine glycogenes, including glycosyltransferases, nucleotide sugar transporters, glycosidases, proteoglycans, and glycan-binding proteins. In this report, the array has been used to generate glycogene-expression profiles of nine murine tissues. Global analysis with a hierarchical clustering algorithm reveals that expression profiles in immune tissues (thymus [THY], spleen [SPL], lymph node, and bone marrow [BM]) are more closely related, relative to those of nonimmune tissues (kidney [KID], liver [LIV], brain [BRN], and testes [TES]). Of the biosynthetic enzymes, those responsible for synthesis of the core regions of N- and O-linked oligosaccharides are ubiquitously expressed, whereas glycosyltransferases that elaborate terminal structures are expressed in a highly tissue-specific manner, accounting for tissue and ultimately cell-type-specific glycosylation. Comparison of gene expression profiles with matrix-assisted laser desorption ionization-time of flight (MALDI-TOF) profiling of N-linked oligosaccharides suggested that the alpha1-3 fucosyltransferase 9, Fut9, is the enzyme responsible for terminal fucosylation in KID and BRN, a finding validated by analysis of Fut9 knockout mice. Two families of glycan-binding proteins, C-type lectins and Siglecs, are predominately expressed in the immune tissues, consistent with their emerging functions in both innate and acquired immunity. The glycogene chip reported in this study is available to the scientific community through the Consortium for Functional Glycomics (CFG) (http://www.functionalglycomics.org).

Animals↗

Synergistic computational and experimental proteomics approaches for more accurate detection of active serine hydrolases in yeast.

An analysis of the structurally and catalytically diverse serine hydrolase protein family in the Saccharomyces cerevisiae proteome was undertaken using two independent but complementary, large-scale approaches. The first approach is based on computational analysis of serine hydrolase active site structures; the second utilizes the chemical reactivity of the serine hydrolase active site in complex mixtures. These proteomics approaches share the ability to fractionate the complex proteome into functional subsets. Each method identified a significant number of sequences, but 15 proteins were identified by both methods. Eight of these were unannotated in the Saccharomyces Genome Database at the time of this study and are thus novel serine hydrolase identifications. Three of the previously uncharacterized proteins are members of a eukaryotic serine hydrolase family, designated as Fsh (family of serine hydrolase), identified here for the first time. OVCA2, a potential human tumor suppressor, and DYR-SCHPO, a dihydrofolate reductase from Schizosaccharomyces pombe, are members of this family. Comparing the combined results to results of other proteomic methods showed that only four of the 15 proteins were identified in a recent large-scale, "shotgun" proteomic analysis and eight were identified using a related, but similar, approach (neither identifies function). Only 10 of the 15 were annotated using alternate motif-based computational tools. The results demonstrate the precision derived from combining complementary, function-based approaches to extract biological information from complex proteomes. The chemical proteomics technology indicates that a functional protein is being expressed in the cell, while the computational proteomics technology adds details about the specific type of function and residue that is likely being labeled. The combination of synergistic methods facilitates analysis, enriches true positive results, and increases confidence in novel identifications. This work also highlights the risks inherent in annotation transfer and the use of scoring functions for determination of correct annotations.

Amino Acid Sequence↗

GOAnno: GO annotation based on multiple alignment.

UNLABELLED: GOAnno is a web tool that automatically annotates proteins according to the Gene Ontology (GO) using evolutionary information available in hierarchized multiple alignments. GO terms present in the aligned functional subfamily can be cross-validated and propagated to obtain highly reliable predicted GO annotation based on the GOAnno algorithm. AVAILABILITY: The web tool and a reduced version for local installation are freely available at http://igbmc.u-strasbg.fr/GOAnno/GOAnno.html SUPPLEMENTARY INFORMATION: The website supplies a detailed explanation and illustration of the algorithm at http://igbmc.u-strasbg.fr/GOAnno/GOAnnoHelp.html.

Algorithms↗

Fragnostic: walking through protein structure space.

The Fragnostic (http://ffas.burnham.org/Fragnostic) web tool implements a novel and useful view of protein structure space. We mined a non-redundant subset of the PDB for common fragments shared between proteins inhabiting different SCOP folds. Subsequently, we formulated an inter-fold similarity measure based on fragment sharing. Fold space is described as a graph whose nodes are folds between which the edges are drawn depending on the extent of fragment sharing. In this fashion, Fragnostic helps discover meaningful relationships between proteins belonging to different folds, based on sharing similar fragments in the proteins comprising those folds. Distant fold similarity information is supplemented by annotations taken from Gene Ontology, SCOP and CATH. Overall, Fragnostic is a tool which helps discover structural and functional relationships between proteins which are distantly related or seemingly unrelated.

Computer Graphics↗

INTEGRATOR: interactive graphical search of large protein interactomes over the Web.

BACKGROUND: The rapid growth of protein interactome data has elevated the necessity and importance of network analysis tools. However, unlike pure text data, network search spaces are of exponential complexity. This poses special challenges for storing, searching, and navigating this data efficiently. Moreover, development of effective web interfaces has been difficult. RESULTS: We present Integrator, a web-integrated graphical search tool for protein-protein interaction networks across 50+ genomes. CONCLUSION: Integrator provides single and multiple protein searches of the Bioverse database containing experimentally-derived and predicted protein-protein interactions. The interface provides animated local network views, rapid subgraph manipulation, and cross-referencing of functional annotations. Integrator is available at http://bioverse.compbio.washington.edu/integrator.

Algorithms↗

The PEDANT genome database in 2005.

The PEDANT genome database (http://pedant.gsf.de) contains pre-computed bioinformatics analyses of publicly available genomes. Its main mission is to provide robust automatic annotation of the vast majority of amino acid sequences, which have not been subjected to in-depth manual curation by human experts in high-quality protein sequence databases. By design PEDANT annotation is genome-oriented, making it possible to explore genomic context of gene products, and evaluate functional and structural content of genomes using a category-based query mechanism. At present, the PEDANT database contains exhaustive annotation of over 1,240,000 proteins from 270 eubacterial, 23 archeal and 41 eukaryotic genomes.

Computational Biology↗

Cloning, expression and characterization of the murine Efemp1, a gene mutated in Doyne-Honeycomb retinal dystrophy.

Development of the bone and cartilage structures is one of the best-studied systems for epithelial-mesenchymal interaction as well as proliferation and differentiation. In a screen for genes differentially expressed in mice deficient for transcription factor AP-2alpha, we have identified a gene which, based on its homology to the human EFEMP-1 gene was designated Efemp1. It encodes for six repeats similar to the domain of the epidermal growth factor. Sequence comparison with EFEMP1 genes of human and rat revealed that the three proteins share a high amino acid identity (92%), suggesting a conserved function during vertebrate development. However, there is no EFEMP1ortholog annotated in sequence databases of other non-mammalian species indicating that it might have evolved in higher vertebrates only. Analysis of the murine genomic locus revealed that the gene is encoded by 11 exons, which are spread over 80 kb of distance on murine chromosome 11A4. The multidomain protein structure may indicate that Efemp1 protein interacts with extracellular matrix components and serves to connect and integrate the function of multiple partner molecules. The gene is expressed in the embryo proper starting from day 9.5 to day 18.5 of murine development. In situ analyses showed that Efemp1 is found in condensing mesenchyme, giving rise to bone and cartilage as well as in developing bone structures of the cranial and the axial skeleton. These results will help in further defining the role of Efemp1 during murine embryogenesis.

Amino Acid Sequence↗

Annotation of bacterial genomes using improved phylogenomic profiles.

MOTIVATION: Phylogenomic profiling is a large-scale comparative genomic method used to infer protein function from evolutionary information first described in a binary form by Pellegrini et al. (1999). Here, we propose improvements of this approach including the use of normalized Blastp bit scores, a normalization of the matrix of profiles to take into account the evolutionary distances between bacteria, the definition of a phylogenomic neighborhood based on continuous pairwise distances between genes and an original annotation procedure including the computation of a p-value for each functional assignment. RESULTS: The method presented here increases the number of Ecocyc enzymes identified as being evolutionarily related by about 25% with respect to the original binary form (absent/present) method. The fraction of 'false' positives is shown to be smaller than 20%. Based on their phylogenomic relationships, genes of unknown function can then be automatically related to annotated genes. Each gene annotation predicted is associated with a p-value, i.e. its probability to be obtained by chance. The validity of this method was extensively tested on a large set of genes of known function using the MultiFun database. We find that 50% of 3122 function attributions that can be made at a p-value level of 10(-11) correspond to the actual gene annotation. The method can be readily applied to any newly sequenced microbial genome. In contrast to earlier work on the same topic, our approach avoids the use of arbitrary cut-off values, and provides a reliability estimate of the functional predictions in form of p-values.

Algorithms↗

Differential detergent fractionation for non-electrophoretic eukaryote cell proteomics.

Differential detergent fractionation (DDF), which relies on detergents to sequentially extract proteins from eukaryotic cells, has been used to increase proteome coverage of 2D-PAGE. Here, we used DDF extraction in conjunction with the nonelectrophoretic proteomics method of liquid chromatography and electrospray ionization tandem mass spectrometry. We demonstrate that DDF can be used with 2D-LC ESI MS2 for comprehensive cellular proteomics, including a large proportion of membrane proteins. Compared to some published methods designed to isolate membrane proteins specifically, DDF extraction yields comprehensive proteomes which include twice as many membrane proteins. Two-thirds of these membrane proteins have more than one trans-membrane domain. Since DDF separates proteins based upon their physicochemistry and subcellular localization, this method also provides data useful for functional genome annotation. As more genome sequences are completed, methods which can aid in functional annotation will become increasingly important.

Animals↗

Analysis of the Thermotoga maritima genome combining a variety of sequence similarity and genome context tools.

The proliferation of genome sequence data has led to the development of a number of tools and strategies that facilitate computational analysis. These methods include the identification of motif patterns, membership of the query sequences in family databases, metabolic pathway involvement and gene proximity. We re-examined the completely sequenced genome of Thermotoga maritima by employing the combined use of the above methods. By analyzing all 1877 proteins encoded in this genome, we identified 193 cases of conflicting annotations (10%), of which 164 are new function predictions and 29 are amendments of previously proposed assignments. These results suggest that the combined use of existing computational tools can resolve inconclusive sequence similarities and significantly improve the prediction of protein function from genome sequence.

Computational Biology↗

Phage bioinformatics tools: a review of computational approaches for bacteriophage research.

Rising clinical interest in phage therapy and the exponential growth of metagenomic sequence catalogues have driven a rapid expansion of bacteriophage bioinformatics. More than 80 dedicated tools, mostly published since 2020, now span identification, assembly, annotation, taxonomy, lifestyle prediction, defence-system detection, and host prediction. Aimed at experienced practitioners and developers, this review synthesizes the field through the lens of three successive computational paradigms: sequence homology, bounded by database completeness; machine learning, constrained by labelled training data; and foundation models, which now achieve Matthews correlation coefficients above 0.95 in identification tasks and, through structure-informed prediction, raise functional annotation to over half of phage genes. Furthermore, we map the upstream components, namely, gene callers, homology engines, protein language models, and structural search tools, that underpin most downstream pipelines, exposing shared infrastructure and ecosystem-level fragility when dependencies change. To translate this into practice, we propose web-based and command-line reference workflows calibrated to user expertise and sample types. Finally, we set an agenda for the next wave of tool development. Roughly half of phage genes still resist functional annotation despite structural methods; no broadly generalizable strain-level host predictor exists for phage therapy; varying true-positive rates (0%-97%) underscore the absence of standardized community benchmarks analogous to Critical Assessment of Structure Prediction or Critical Assessment of Metagenome Interpretation. As generative genome models begin designing synthetic phages, progress will depend less on producing standalone tools than on rigorous evaluation, interoperable infrastructure, and clinically meaningful prediction targets.

Computational Biology↗

Identification of protein-coding genes in the genome of Vibrio cholerae with more than 98% accuracy using occurrence frequencies of single nucleotides.

The published sequence of the Vibrio cholerae genome indicates that, in addition to the genes that encode proteins of known and unknown function, there are 1577 ORFs identified as conserved hypothetical or hypothetical gene candidates. Because the annotation is not 100% accurate, it is not known which of the 1577 ORFs are true protein-coding genes. In this paper, an algorithm based on the Z curve method, with sensitivity, specificity and accuracy greater than 98%, is used to solve this problem. Twenty-fold cross-validation tests show that the accuracy of the algorithm is 98.8%. A detailed discussion of the mechanism of the algorithm is also presented. It was found that 172 of the 1577 ORFs are unlikely to be protein-coding genes. The number of protein-coding genes in the V. cholerae genome was re-estimated and found to be approximately 3716. This result should be of use in microarray analysis of gene expression in the genome, because the cost of preparing chips may be somewhat decreased. A computer program was written to calculate a coding score called VCZ for gene identification in the genome. Coding/noncoding is simply determined by VCZ > 0/VCZ < 0. The program is freely available on request for academic use.

Algorithms↗

Fuzzy measures on the Gene Ontology for gene product similarity.

One of the most important objects in bioinformatics is a gene product (protein or RNA). For many gene products, functional information is summarized in a set of Gene Ontology (GO) annotations. For these genes, it is reasonable to include similarity measures based on the terms found in the GO or other taxonomy. In this paper, we introduce several novel measures for computing the similarity of two gene products annotated with GO terms. The fuzzy measure similarity (FMS) has the advantage that it takes into consideration the context of both complete sets of annotation terms when computing the similarity between two gene products. When the two gene products are not annotated by common taxonomy terms, we propose a method that avoids a zero similarity result. To account for the variations in the annotation reliability, we propose a similarity measure based on the Choquet integral. These similarity measures provide extra tools for the biologist in search of functional information for gene products. The initial testing on a group of 194 sequences representing three proteins families shows a higher correlation of the FMS and Choquet similarities to the BLAST sequence similarities than the traditional similarity measures such as pairwise average or pairwise maximum.

Algorithms↗

Characteristics of the Lotus japonicus gene repertoire deduced from large-scale expressed sequence tag (EST) analysis.

To perform a comprehensive analysis of genes expressed in a model legume, Lotus japonicus, a total of 74472 3'-end expressed sequence tags (EST) were generated from cDNA libraries produced from six different organs. Clustering of sequences was performed with an identity criterion of 95% for 50 bases, and a total of 20457 non-redundant sequences, 8503 contigs and 11954 singletons were generated. EST sequence coverage was analyzed by using the annotated L. japonicus genomic sequence and 1093 of the 1889 predicted protein-encoding genes (57.9%) were hit by the EST sequence(s). Gene content was compared to several plant species. Among the 8503 contigs, 471 were identified as sequences conserved only in leguminous species and these included several disease resistance-related genes. This suggested that in legumes, these genes may have evolved specifically to resist pathogen attack. The rate of gene sequence divergence was assessed by comparing similarity level and functional category based on the Gene Ontology (GO) annotation of Arabidopsis genes. This revealed that genes encoding ribosomal proteins, as well as those related to translation, photosynthesis, and cellular structure were more abundantly represented in the highly conserved class, and that genes encoding transcription factors and receptor protein kinases were abundantly represented in the less conserved class. To make the sequence information and the cDNA clones available to the research community, a Web database with useful services was created at http://www.kazusa.or.jp/en/plant/lotus/EST/.

DNA, Complementary↗

The Riken mouse genome encyclopedia project.

The Riken mouse genome encyclopedia a comprehensive full-length cDNA collection and sequence database. High-level functional annotation is based on sequence homology search, expression profiling, mapping and protein-protein interactions. More than 1000000 clones prepared from 163 tissues were end-sequenced and classified into 128000 clusters, and 60000 representative clones were fully sequenced representing 24000 clear protein-encoding genes. The application of the mouse genome database for positional cloning and gene network regulation analysis is reported.

Animals↗

Moonlighting vacuolar protease: multiple jobs for a busy protein.

In this Genomics Era with a wealth of annotated sequence data, it is easy to pigeonhole a protein into a particular function. However, Noa Matarasso et al. recently found a vacuolar protease that can also function as a transcription factor. This work illustrates that a protein can serve multiple roles in a cell, raising intriguing questions as to the extent that genomic information can be deciphered de novo.

Ethylenes↗

Genome-wide transcription profiling of Corynebacterium glutamicum after heat shock and during growth on acetate and glucose.

To monitor the global gene expression of Corynebacterium glutamicum we established two formats of DNA-arrays on nylon membranes. We produced an ordered DNA-array of PCR fragments from a shotgun library of C. glutamicum representing a threefold coverage of the genome. With this format we studied genome-wide transcriptional changes after heat shock. Sequence and subsequent BLAST analysis of PCR fragments with elevated expression after heat shock revealed PCR fragments harboring genes that encode several proteins of the heat shock family, proteins of the oxidative stress response and proteins with unknown function. DNA-arrays based on PCR fragments representing 2804 annotated ORFs of C. glutamicum were used to monitor the transcript levels during growth on acetate and glucose. We determined minimal detectable ratios and compared labeling approaches with random hexamers and ORF-specific primers. ORF-based DNA-array analysis with different labeling approaches showed similar results: e.g. increased mRNA levels of the pta-ack operon, aceA, aceB and genes encoding phosphoenolpyruvate carboxykinase and enzymes of the citric acid cycle during growth on acetate and elevated mRNA levels of some enzymes of the glycolytic pathway and lactate dehydrogenase upon growth on glucose. These results demonstrate that shotgun DNA-arrays and ORF-based DNA-arrays are appropriate tools to study physiology of microorganism.

Acetates↗

Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis.

Suckermouth catfishes, with their evolved powerful features, have become notorious invasive species, causing significant damage to aquatic ecosystems. However, the lack of high-quality genomes severely restricts research on this group within the field. In this study, we de novo assembled the chromosome-level genome assembly of Pterygoplichthys pardalis using multiple platforms of sequencing data, including Illumina short reads, Nanopore long reads, and Hi-C sequencing reads, resulting in a 1.51&#x2009;Gb genome assembly. Multiple evaluations, including read mapping ratio (98.52%), transcript mapping ratio (99.61%), conserved BUSCO gene set (98.8%), and N50 score (49.47&#x2009;Mb), indicated the high continuity and accuracy of the genome assembly we generated. Genome annotation found that 0.97&#x2009;Gb of genome sequences are repetitive sequences, accounting for 64.47% of the genome assembly. Further, 23,859 protein-coding genes were successfully predicted, 92.92% of which could be annotated in functional databases. This high-quality genome assembly of P. pardalis provides a valuable resource for understanding the genetic underpinnings of P. pardalis's invasive success and offers critical data for future fisheries research and management.

Animals↗