Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Practical limits of function prediction.

The widening gap between known protein sequences and their functions has led to the practice of assigning a potential function to a protein on the basis of sequence similarity to proteins whose function has been experimentally investigated. We present here a critical view of the theoretical and practical bases for this approach. The results obtained by analyzing a significant number of true sequence similarities, derived directly from structural alignments, point to the complexity of function prediction. Different aspects of protein function, including (i) enzymatic function classification, (ii) functional annotations in the form of key words, (iii) classes of cellular function, and (iv) conservation of binding sites can only be reliably transferred between similar sequences to a modest degree. The reason for this difficulty is a combination of the unavoidable database inaccuracies and the plasticity of protein function. In addition, analysis of the relationship between sequence and functional descriptions defines an empirical limit for pairwise-based functional annotations, namely, the three first digits of the six numbers used as descriptors of protein folds in the FSSP database can be predicted at an average level as low as 7.5% sequence identity, two of the four EC digits at 15% identity, half of the SWISS-PROT key words related to protein function would require 20% identity, and the prediction of half of the residues in the binding site can be made at the 30% sequence identity level.

Amino Acid Sequence↗

Primary and secondary metabolism, and post-translational protein modifications, as portrayed by proteomic analysis of Streptomyces coelicolor.

The newly sequenced genome of Streptomyces coelicolor is estimated to encode 7825 theoretical proteins. We have mapped approximately 10% of the theoretical proteome experimentally using two-dimensional gel electrophoresis and matrix-assisted laser desorption ionization time-of-flight (MALDI-TOF) mass spectrometry. Products from 770 different genes were identified, and the types of proteins represented are discussed in terms of their annotated functional classes. An average of 1.2 proteins per gene was observed, indicating extensive post-translational regulation. Examples of modification by N-acetylation, adenylylation and proteolytic processing were characterized using mass spectrometry. Proteins from both primary and certain secondary metabolic pathways are strongly represented on the map, and a number of these enzymes were identified at more than one two-dimensional gel location. Post-translational modification mechanisms may therefore play a significant role in the regulation of these pathways. Unexpectedly, one of the enzymes for synthesis of the actinorhodin polyketide antibiotic appears to be located outside the cytoplasmic compartment, within the cell wall matrix. Of 20 gene clusters encoding enzymes characteristic of secondary metabolism, eight are represented on the proteome map, including three that specify the production of novel metabolites. This information will be valuable in the characterization of the new metabolites.

Acetylation↗

Chromosome-Level Genome Assembly of Solanum carolinense.

Horsenettle (Solanum carolinense L.) is a noxious weed widely distributed across North America and increasingly invasive in other regions. Its strong environmental adaptability, complex defense strategies, and distinctive reproductive traits make it an important model for studying plant-herbivore coevolution. However, the absence of high-quality genomic resources has limited deeper investigation into its adaptive evolutionary mechanisms. In this study, we generated a chromosome-level reference genome assembly for S. carolinense using an integrated approach combining PacBio HiFi long-read sequencing, Illumina second-generation sequencing, and Hi-C chromatin interaction scaffolding. The final genome assembly had a total length of 915.40 Mb, with a contig N50 of 51.06 Mb and a scaffold N50 of 73.17 Mb; 96.05% of the sequences were successfully anchored onto 12 pseudochromosomes. The genome was characterized by a high proportion of repetitive sequences (73.64%) and substantial heterozygosity (1.13%), consistent with a highly repetitive and moderately high heterozygous genome. BUSCO analysis indicated that the chromosome-level genome assembly of S. carolinense reached a completeness score of 94.8%. A total of 32,206 protein-coding genes were annotated, of which 97.95% received functional annotations. The evaluation of the annotated protein-coding gene set returned a completeness value of 94.9%. This reference genome provides a valuable resource for advancing research on the adaptive evolution of weedy Solanaceae species, supports the development of more effective management strategies for this troublesome species, and offers a technical reference for assembling other highly heterozygous weed genomes.

Solanum carolinense↗

Clustering the annotation space of proteins.

BACKGROUND: Current protein clustering methods rely on either sequence or functional similarities between proteins, thereby limiting inferences to one of these areas. RESULTS: Here we report a new approach, named CLAN, which clusters proteins according to both annotation and sequence similarity. This approach is extremely fast, clustering the complete SwissProt database within minutes. It is also accurate, recovering consistent protein families agreeing on average in more than 97% with sequence-based protein families from Pfam. Discrepancies between sequence- and annotation-based clusters were scrutinized and the reasons reported. We demonstrate examples for each of these cases, and thoroughly discuss an example of a propagated error in SwissProt: a vacuolar ATPase subunit M9.2 erroneously annotated as vacuolar ATP synthase subunit H. CLAN algorithm is available from the authors and the CLAN database is accessible at http://maine.ebi.ac.uk:8000/cgi-bin/clan/ClanSearch.pl CONCLUSIONS: CLAN creates refined function-and-sequence specific protein families that can be used for identification and annotation of unknown family members. It also allows easy identification of erroneous annotations by spotting inconsistencies between similarities on annotation and sequence levels.

Adenosine Triphosphatases↗

C-terminal WxL domain mediates cell wall binding in Enterococcus faecalis and other gram-positive bacteria.

Analysis of the genome sequence of Enterococcus faecalis clinical isolate V583 revealed novel genes encoding surface proteins. Twenty-seven of these proteins, annotated as having unknown functions, possess a putative N-terminal signal peptide and a conserved C-terminal region characterized by a novel conserved domain designated WxL. Proteins having similar characteristics were also detected in other low-G+C-content gram-positive bacteria. We hypothesized that the WxL region might be a determinant of bacterial cell location. This hypothesis was tested by generating protein fusions between the C-terminal regions of two WxL proteins in E. faecalis and a nuclease reporter protein. We demonstrated that the C-terminal regions of both proteins conferred a cell surface localization to the reporter fusions in E. faecalis. This localization was eliminated by introducing specific deletions into the domains. Interestingly, exogenously added protein fusions displayed binding to whole cells of various gram-positive bacteria. We also showed that the peptidoglycan was a binding ligand for WxL domain attachment to the cell surface and that neither proteins nor carbohydrates were necessary for binding. Based on our findings, we propose that the WxL region is a novel cell wall binding domain in E. faecalis and other gram-positive bacteria.

Amino Acid Sequence↗

Mitoproteome: human heart mitochondrial protein sequence database.

The human mitochondrial proteome database has been developed by deriving data from a combination of public repositories and experimental and computational prediction methods. The experimental data is derived from highly purified mitochondria from human heart tissue, whereas predictions have been performed by MITOPRED, a genome-scale method for the prediction of nucleus-encoded mitochondrial proteins. Mitochondrial protein sequences from different sources have been clustered to generate a nonredundant dataset. Annotations related to the protein function, structure, disease association, pathways, and so on are collected from a number of public databases using commonly used UNIX and Perl scripts. This chapter provides a detailed description of various data sources and methods used to download, curate, parse, and generate meaningful annotations from primary as well as derived databases.

Computational Biology↗

Proteome annotations and identifications of the human pulmonary fibroblast.

We hereby report on a three year project initiative undertaken by our research team encompassing large-scale protein expression profiling and annotations of human primary lung fibroblast cells. An overview is given of proteomic studies of the fibroblast target cell involved in several diseases such as asthma, idiopatic pulmonary disease, and COPD. It has been the objective within our research team to map and identify the protein expressions occurring in both activated-, as well as resting cell states. The JGGL database www.2DDB.org has been built around these data, allowing advanced hypothesis building using the interactive query bioinformatic tools developed. Gene ontology has been applied to these annotations, classifying and correlating protein expressions to function. The localization as well as the biological processes involved for the annotations are being presented including an annotation-, and sequence-identification strategy, resulting in close to 2000 protein identities. Both gel based, high resolution 2D-gels, and liquid-phase separation (three-dimensional HPLC), as well as the combination of gel- and LC-based approaches (1D-gels and nano-capillary LC, reversed-phase) were utilized. Protein sequencing and structure identities were acquired by a combination of MALDI-, and electrospray-mass spectrometry techniques. Phenotypical and morphological characterizations were also made for this human disease target cell in both stimulated- and resting-cell states. The use of functional assays that demonstrate the key regulating role of growth factors and cytokine stimuli such as PDGF, TGF-beta, and EGF and the effect of ECM molecules such as Biglycan, are also presented and discussed.

Amino Acid Sequence↗

Gene fusions and gene duplications: relevance to genomic annotation and functional analysis.

BACKGROUND: Escherichia coli a model organism provides information for annotation of other genomes. Our analysis of its genome has shown that proteins encoded by fused genes need special attention. Such composite (multimodular) proteins consist of two or more components (modules) encoding distinct functions. Multimodular proteins have been found to complicate both annotation and generation of sequence similar groups. Previous work overstated the number of multimodular proteins in E. coli. This work corrects the identification of modules by including sequence information from proteins in 50 sequenced microbial genomes. RESULTS: Multimodular E. coli K-12 proteins were identified from sequence similarities between their component modules and non-fused proteins in 50 genomes and from the literature. We found 109 multimodular proteins in E. coli containing either two or three modules. Most modules had standalone sequence relatives in other genomes. The separated modules together with all the single (un-fused) proteins constitute the sum of all unimodular proteins of E. coli. Pairwise sequence relationships among all E. coli unimodular proteins generated 490 sequence similar, paralogous groups. Groups ranged in size from 92 to 2 members and had varying degrees of relatedness among their members. Some E. coli enzyme groups were compared to homologs in other bacterial genomes. CONCLUSION: The deleterious effects of multimodular proteins on annotation and on the formation of groups of paralogs are emphasized. To improve annotation results, all multimodular proteins in an organism should be detected and when known each function should be connected with its location in the sequence of the protein. When transferring functions by sequence similarity, alignment locations must be noted, particularly when alignments cover only part of the sequences, in order to enable transfer of the correct function. Separating multimodular proteins into module units makes it possible to generate protein groups related by both sequence and function, avoiding mixing of unrelated sequences. Organisms differ in sizes of groups of sequence-related proteins. A sample comparison of orthologs to selected E. coli paralogous groups correlates with known physiological and taxonomic relationships between the organisms.

Computational Biology↗

Enhanced identification of key bacterial motility genes via a cross-species genomic hybrid feature machine learning approach.

Efficient and accurate identification of functional genes is critical to biological research, yet traditional single-species approaches are often limited by low efficiency. Previously, we established a novel method for identifying key genes using cross-species protein domain features and machine learning. However, the high multiplicity of gene members associated with specific domains creates a substantial workload for subsequent experimental validation. To address this, this study proposes an enhanced approach that integrates EggNOG-based protein sequence annotation with domain analysis. Unannotated sequences are subsequently analyzed for protein domains, generating a comprehensive "direct gene annotation plus domain" hybrid feature matrix. While the hybrid matrix model yielded comparable predictive accuracy, it significantly enhanced feature resolution: the top 50 predicted features were all known motility-related genes or domains. Furthermore, among the top 100 ranked features, 58 are confirmed to be directly related to motility based on experimental evidence. Although strict genus-level control still yielded 51 confirmed features, excessive taxonomic restriction drastically reduces the number of training genomes, which may paradoxically impair identification efficiency. These results demonstrate that the new method effectively reduces the subsequent experimental workload and enables high-throughput identification of functional genes in a single analysis. With accuracy and efficiency far exceeding those of existing single-species identification methods, it provides a highly efficient solution for mining key genes underlying other complex bacterial phenotypes.

Machine Learning↗

Comparative analysis of chloroplast genomes: functional annotation, genome-based phylogeny, and deduced evolutionary patterns.

All protein sequences from 19 complete chloroplast genomes (cpDNA) have been studied using a new computational method able to analyze functional correlations among series of protein sequences contained in complete proteomes. First, all open reading frames (ORFs) from the cpDNAs, comprising a total of 2266 protein sequences, were compared against the 3168 proteins from Synechocystis PCC6803 complete genome to find functionally related orthologous proteins. Additionally, all cpDNA genomes were pairwise compared to find orthologous groups not present in cyanobacteria. Annotations in the cluster of othologous proteins database and CyanoBase were used as reference for the functional assignments. Following this protocol, new functional assignments were made for ORFs of unknown function and for ycfs (hypothetical chloroplast frames), which still lack a functional assignment. Using this information, a matrix of functional relationships was derived from profiles of the presence and/or absence of orthologous proteins; the matrix included 1837 proteins in 277 orthologous clusters. A factor analysis study of this matrix, followed by cluster analysis, allowed us to obtain accurate phylogenetic reconstructions and the detection of genes probably involved in speciation as phylogenetic correlates. Finally, by grouping common evolutionary patterns, we show that it is possible to determine functionally linked protein networks. This has allowed us to suggest putative associations for some unknown ORFs.

Bacterial Proteins↗

Predicted role for the archease protein family based on structural and sequence analysis of TM1083 and MTH1598, two proteins structurally characterized through structural genomics efforts.

Recently, the structures of two proteins belonging to the archease family, TM1083 from Thermotoga maritima and MTH1598 from Methanobacterium thermoautotrophicum, have been solved independently by two Protein Structure Initiative structural genomics pilot centers using X-ray crystallography and NMR, respectively. The archease protein family is a good example of one of the paradoxes of structural genomics: Approximately one third of protein structures produced by structural genomics centers have no known function and are still annotated as "hypothetical proteins" in the Protein Data Bank. In the case of archeases, despite the existence of two protein structures and abundant sequence information, there is still no function assigned to this protein family. Here, our group predicts, based on structural similarity, sequence conservation, and gene context analyses, that members of this protein family might function as chaperones or modulators of proteins involved in DNA/RNA processing. The conservation of genomic context for this protein family is constant from Archaea and Bacteria to humans, and suggests that unannotated open reading frames contiguous to them could be novel RNA/DNA binding proteins.

Amino Acid Motifs↗

Sequence conserved for subcellular localization.

The more proteins diverged in sequence, the more difficult it becomes for bioinformatics to infer similarities of protein function and structure from sequence. The precise thresholds used in automated genome annotations depend on the particular aspect of protein function transferred by homology. Here, we presented the first large-scale analysis of the relation between sequence similarity and identity in subcellular localization. Three results stood out: (1) The subcellular compartment is generally more conserved than what might have been expected given that short sequence motifs like nuclear localization signals can alter the native compartment; (2) the sequence conservation of localization is similar between different compartments; and (3) it is similar to the conservation of structure and enzymatic activity. In particular, we found the transition between the regions of conserved and nonconserved localization to be very sharp, although the thresholds for conservation were less well defined than for structure and enzymatic activity. We found that a simple measure for sequence similarity accounting for pairwise sequence identity and alignment length, the HSSP distance, distinguished accurately between protein pairs of identical and different localizations. In fact, BLAST expectation values outperformed the HSSP distance only for alignments in the subtwilight zone. We succeeded in slightly improving the accuracy of inferring localization through homology by fine tuning the thresholds. Finally, we applied our results to the entire SWISS-PROT database and five entirely sequenced eukaryotes.

Amino Acid Sequence↗

A Bayesian method for identifying missing enzymes in predicted metabolic pathway databases.

BACKGROUND: The PathoLogic program constructs Pathway/Genome databases by using a genome's annotation to predict the set of metabolic pathways present in an organism. PathoLogic determines the set of reactions composing those pathways from the enzymes annotated in the organism's genome. Most annotation efforts fail to assign function to 40-60% of sequences. In addition, large numbers of sequences may have non-specific annotations (e.g., thiolase family protein). Pathway holes occur when a genome appears to lack the enzymes needed to catalyze reactions in a pathway. If a protein has not been assigned a specific function during the annotation process, any reaction catalyzed by that protein will appear as a missing enzyme or pathway hole in a Pathway/Genome database. RESULTS: We have developed a method that efficiently combines homology and pathway-based evidence to identify candidates for filling pathway holes in Pathway/Genome databases. Our program not only identifies potential candidate sequences for pathway holes, but combines data from multiple, heterogeneous sources to assess the likelihood that a candidate has the required function. Our algorithm emulates the manual sequence annotation process, considering not only evidence from homology searches, but also considering evidence from genomic context (i.e., is the gene part of an operon?) and functional context (e.g., are there functionally-related genes nearby in the genome?) to determine the posterior belief that a candidate has the required function. The method can be applied across an entire metabolic pathway network and is generally applicable to any pathway database. The program uses a set of sequences encoding the required activity in other genomes to identify candidate proteins in the genome of interest, and then evaluates each candidate by using a simple Bayes classifier to determine the probability that the candidate has the desired function. We achieved 71% precision at a probability threshold of 0.9 during cross-validation using known reactions in computationally-predicted pathway databases. After applying our method to 513 pathway holes in 333 pathways from three Pathway/Genome databases, we increased the number of complete pathways by 42%. We made putative assignments to 46% of the holes, including annotation of 17 sequences of previously unknown function. CONCLUSIONS: Our pathway hole filler can be used not only to increase the utility of Pathway/Genome databases to both experimental and computational researchers, but also to improve predictions of protein function.

Amino Acid Oxidoreductases↗

Protein-protein interactions of the hyperthermophilic archaeon Pyrococcus horikoshii OT3.

BACKGROUND: Although 2,061 proteins of Pyrococcus horikoshii OT3, a hyperthermophilic archaeon, have been predicted from the recently completed genome sequence, the majority of proteins show no similarity to those from other organisms and are thus hypothetical proteins of unknown function. Because most proteins operate as parts of complexes to regulate biological processes, we systematically analyzed protein-protein interactions in Pyrococcus using the mammalian two-hybrid system to determine the function of the hypothetical proteins. RESULTS: We examined 960 soluble proteins from Pyrococcus and selected 107 interactions based on luciferase reporter activity, which was then evaluated using a computational approach to assess the reliability of the interactions. We also analyzed the expression of the assay samples by western blot, and a few interactions by in vitro pull-down assays. We identified 11 hetero-interactions that we considered to be located at the same operon, as observed in Helicobacter pylori. We annotated and classified proteins in the selected interactions according to their orthologous proteins. Many enzyme proteins showed self-interactions, similar to those seen in other organisms. CONCLUSION: We found 13 unannotated proteins that interacted with annotated proteins; this information is useful for predicting the functions of the hypothetical Pyrococcus proteins from the annotations of their interacting partners. Among the heterogeneous interactions, proteins were more likely to interact with proteins within the same ortholog class than with proteins of different classes. The analysis described here can provide global insights into the biological features of the protein-protein interactions in P. horikoshii.

Genes, Archaeal↗

Re-evaluation and in silico annotation of the Tupaia herpesvirus proteins.

Herpesviruses represent an exceptionally suitable model to analyze evolutionary old pathogens, their competency to adapt to existing and changing molecular niches in host species, and the modulation of the gene content and function to comply with the requirements of life. The basis for numerous studies dealing with these questions are reliable statements about the gene content of herpesviral genomes and the functions of viral proteins. The recent determination of the coding strategy of the chimpanzee cytomegalovirus genome and the re-evaluation of the gene content of the human cytomegalovirus genome made it also necessary to restructure the putative transcription map of the Tupaia herpesvirus (THV) genome. Twenty-three THV-specific ORFs formerly predicted to be coding for viral proteins were deleted from the THV transcription map resulting in a gene layout that is now characterized by the presence of conserved genes in the genome center, that probably reflect the genome structure of common herpesviral ancestors, and species-specific genes at the termini. The conserved regions in the THV genome are characterized by high G + C contents between 60% and 80%, a high CpG dinucleotide frequency, and the presence of densely packed putative CpG islands. The genome termini seem to provide the requirements of large scale rearrangements and complements of the gene content to adapt to new environmental demands. With the help of the recently designed method of dictionary-driven, pattern-based protein annotation it was possible to assign putative functions to almost all potential THV proteins, e.g. 123 were found to be putative membrane or secreted proteins, putative signal domains were identified in 69, and 29 proteins were predicted to be glycosylated. The present study adds new aspects to the knowledge about the precise gene composition of herpesvirus genomes and viral protein functions that are of exceptional importance for studies dealing with the phylogeny, the evolution, vaccine vector development, virus-host interactions, pathogenesis and the determination of protein functions of herpesviruses.

Animals↗

Benchmarking PSI-BLAST in genome annotation.

The recognition of remote protein homologies is a major aspect of the structural and functional annotation of newly determined genomes. Here we benchmark the coverage and error rate of genome annotation using the widely used homology-searching program PSI-BLAST (position-specific iterated basic local alignment search tool). This study evaluates the one-to-many success rate for recognition, as often there are several homologues in the database and only one needs to be identified for annotating the sequence. In contrast, previous benchmarks considered one-to-one recognition in which a single query was required to find a particular target. The benchmark constructs a model genome from the full sequences of the structural classification of protein (SCOP) database and searches against a target library of remote homologous domains (<20 % identity). The structural benchmark provides a reliable list of correct and false homology assignments. PSI-BLAST successfully annotated 40 % of the domains in the model genome that had at least one homologue in the target library. This coverage is more than three times that if one-to-one recognition is evaluated (11 % coverage of domains). Although a structural benchmark was used, the results equally apply to just sequence homology searches. Accordingly, structural and sequence assignments were made to the sequences of Mycoplasma genitalium and Mycobacterium tuberculosis (see http://www.bmm.icnet. uk). The extent of missed assignments and of new superfamilies can be estimated for these genomes for both structural and functional annotations.

Algorithms↗

Differential evolutionary conservation of motif modes in the yeast protein interaction network.

BACKGROUND: The importance of a network motif (a recurring interconnected pattern of special topology which is over-represented in a biological network) lies in its position in the hierarchy between the protein molecule and the module in a protein-protein interaction network. Until now, however, the methods available have greatly restricted the scope of research. While they have focused on the analysis in the resolution of a motif topology, they have not been able to distinguish particular motifs of the same topology in a protein-protein interaction network. RESULTS: We have been able to assign the molecular function annotations of Gene Ontology to each protein in the protein-protein interactions of Saccharomyces cerevisiae. For various motif topologies, we have developed an algorithm, enabling us to unveil one million "motif modes", each of which features a unique topological combination of molecular functions. To our surprise, the conservation ratio, i.e., the extent of the evolutionary constraints upon the motif modes of the same motif topology, varies significantly, clearly indicative of distinct differences in the evolutionary constraints upon motifs of the same motif topology. Equally important, for all motif modes, we have found a power-law distribution of the motif counts on each motif mode. We postulate that motif modes may very well represent the evolutionary-conserved topological units of a protein interaction network. CONCLUSION: For the first time, the motifs of a protein interaction network have been investigated beyond the scope of motif topology. The motif modes determined in this study have not only enabled us to differentiate among different evolutionary constraints on motifs of the same topology but have also opened up new avenues through which protein interaction networks can be analyzed.

Algorithms↗

Aetiology-specific patterns in end-stage heart failure patients identified by functional annotation and classification of microarray data.

BACKGROUND: The objective of the present study was to use gene expression profiling, functional annotations and classification to identify aetiology-specific biological processes and potential molecular markers for different aetiologies of end-stage heart failure. METHODS AND RESULTS: Individual left ventricular myocardial samples from eleven coronary artery disease and nine dilated cardiomyopathy transplant patients were co-hybridized with pooled RNA from four non-failing hearts on custom-made arrays of 7000 human genes. Significance analysis identified differential expression of 153 and 147 genes, respectively, in coronary artery disease or dilated cardiomyopathy versus non-failing hearts. Analysis of Gene Ontology biological process annotations indicated aetiology-specific patterns, primarily related to genes involved in catabolism and in regulation of protein kinase activity. Gene expression classifiers were obtained and used for class prediction of random samples of coronary artery diseased and dilated cardiomyopathic hearts. Best classifiers frequently included matrix metalloproteinase 3, fibulin 1, ATP-binding cassette, sub-family B member 1 and iroquois homeobox protein 5. CONCLUSION: Combining functional annotation from microarray data and classification analysis constitutes a potent strategy to identify disease-specific biological processes and gene expression markers in e.g. end-stage coronary artery disease and dilated cardiomyopathy.

Adult↗