Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

SMART: a web-based tool for the study of genetically mobile domains.

SMART (a Simple Modular Architecture Research Tool) allows the identification and annotation of genetically mobile domains and the analysis of domain architectures (http://SMART.embl-heidelberg.de ). More than 400 domain families found in signalling, extra-cellular and chromatin-associated proteins are detectable. These domains are extensively annotated with respect to phyletic distributions, functional class, tertiary structures and functionally important residues. Each domain found in a non-redundant protein database as well as search parameters and taxonomic information are stored in a relational database system. User interfaces to this database allow searches for proteins containing specific combinations of domains in defined taxa.

Database Management Systems↗

Engineered viruses to select genes encoding secreted and membrane-bound proteins in mammalian cells.

We have developed a functional genomics tool to identify the subset of cDNAs encoding secreted and membrane-bound proteins within a library (the 'secretome'). A Sindbis virus replicon was engineered such that the envelope protein precursor no longer enters the secretory pathway. cDNA fragments were fused to the mutant precursor and expression screened for their ability to restore membrane localization of envelope proteins. In this way, recombinant replicons were released within infectious viral particles only if the cDNA fragment they contain encodes a secretory signal. By using engineered viral replicons to selectively export cDNAs of interest in the culture medium, the methodology reported here efficiently filters genetic information in mammalian cells without the need to select individual clones. This adaptation of the 'signal trap' strategy is highly sensitive (1/200 000) and efficient. Indeed, of the 2546 inserts that were retrieved after screening various libraries, more than 97% contained a putative signal peptide. These 2473 clones encoded 419 unique cDNAs, of which 77% were previously annotated. Of the 94 cDNAs encoding proteins of unknown function, 24% either had no match in databases or contained a secretory signal that could not be predicted from electronic data.

Animals↗

mettannotator: a comprehensive and scalable Nextflow annotation pipeline for prokaryotic assemblies.

SUMMARY: In recent years, there has been a surge in prokaryotic genome assemblies, coming from both isolated organisms and environmental samples. These assemblies often include novel species that are poorly represented in reference databases creating a need for a tool that can annotate both well-described and novel taxa, and can run at scale. Here, we present mettannotator-a comprehensive, scalable Nextflow pipeline for prokaryotic genome annotation that identifies coding and noncoding regions, predicts protein functions, including antimicrobial resistance, and delineates gene clusters. The pipeline summarizes these results in a GFF (General Feature Format) file that can be easily utilized in downstream analysis or visualized using common genome browsers. Here, we show how it works on 200 genomes from 29 prokaryotic phyla, including isolate genomes and known and novel metagenome-assembled genomes, and present metrics on its performance in comparison to other tools. AVAILABILITY AND IMPLEMENTATION: The pipeline is written in Nextflow and Python and published under an open source Apache 2.0 licence. Instructions and source code can be accessed at https://github.com/EBI-Metagenomics/mettannotator. The pipeline is also available on WorkflowHub: https://workflowhub.eu/workflows/1069.

Software↗

Structural characterization of Salmonella typhimurium YeaZ, an M22 O-sialoglycoprotein endopeptidase homolog.

The Salmonella typhimurium "yeaZ" gene (StyeaZ) encodes an essential protein of unknown function (StYeaZ), which has previously been annotated as a putative homolog of the Pasteurella haemolytica M22 O-sialoglycoprotein endopeptidase Gcp. YeaZ has also recently been reported as the first example of an RPF from a gram-negative bacterial species. To further characterize the properties of StYeaZ and the widely occurring MK-M22 family, we describe the purification, biochemical analysis, crystallization, and structure determination of StYeaZ. The crystal structure of StYeaZ reveals a classic two-lobed actin-like fold with structural features consistent with nucleotide binding. However, microcalorimetry experiments indicated that StYeaZ neither binds polyphosphates nor a wide range of nucleotides. Additionally, biochemical assays show that YeaZ is not an active O-sialoglycoprotein endopeptidase, consistent with the lack of the critical zinc binding motif. We present a detailed comparison of YeaZ with available structural homologs, the first reported structural analysis of an MK-M22 family member. The analysis indicates that StYeaZ has an unusual orientation of the A and B lobes which may require substantial relative movement or interaction with a partner protein in order to bind ligands. Comparison of the fold of YeaZ with that of a known RPF domain from a gram-positive species shows significant structural differences and therefore potentially distinctive RPF mechanisms for these two bacterial classes.

Amino Acid Sequence↗

Supra-domains: evolutionary units larger than single protein domains.

Domains are the evolutionary units that comprise proteins, and most proteins are built from more than one domain. Domains can be shuffled by recombination to create proteins with new arrangements of domains. Using structural domain assignments, we examined the combinations of domains in the proteins of 131 completely sequenced organisms. We found two-domain and three-domain combinations that recur in different protein contexts with different partner domains. The domains within these combinations have a particular functional and spatial relationship. These units are larger than individual domains and we term them "supra-domains". Amongst the supra-domains, we identified some 1400 (1203 two-domain and 166 three-domain) combinations that are statistically significantly over-represented relative to the occurrence and versatility of the individual component domains. Over one-third of all structurally assigned multi-domain proteins contain these over-represented supra-domains. This means that investigation of the structural and functional relationships of the domains forming these popular combinations would be particularly useful for an understanding of multi-domain protein function and evolution as well as for genome annotation. These and other supra-domains were analysed for their versatility, duplication, their distribution across the three kingdoms of life and their functional classes. By examining the three-dimensional structures of several examples of supra-domains in different biological processes, we identify two basic types of spatial relationships between the component domains: the combined function of the two domains is such that either the geometry of the two domains is crucial and there is a tight constraint on the interface, or the precise orientation of the domains is less important and they are spatially separate. Frequently, the role of the supra-domain becomes clear only once the three-dimensional structure is known. Since this is the case for only a quarter of the supra-domains, we provide a list of the most important unknown supra-domains as potential targets for structural genomics projects.

Animals↗

Cyanidioschyzon merolae genome. A tool for facilitating comparable studies on organelle biogenesis in photosynthetic eukaryotes.

The ultrasmall unicellular red alga Cyanidioschyzon merolae lives in the extreme environment of acidic hot springs and is thought to retain primitive features of cellular and genome organization. We determined the 16.5-Mb nuclear genome sequence of C. merolae 10D as the first complete algal genome. BLASTs and annotation results showed that C. merolae has a mixed gene repertoire of plants and animals, also implying a relationship with prokaryotes, although its photosynthetic components were comparable to other phototrophs. The unicellular green alga Chlamydomonas reinhardtii has been used as a model system for molecular biology research on, for example, photosynthesis, motility, and sexual reproduction. Though both algae are unicellular, the genome size, number of organelles, and surface structures are remarkably different. Here, we report the characteristics of double membrane- and single membrane-bound organelles and their related genes in C. merolae and conduct comparative analyses of predicted protein sequences encoded by the genomes of C. merolae and C. reinhardtii. We examine the predicted proteins of both algae by reciprocal BLASTP analysis, KOG assignment, and gene annotation. The results suggest that most core biological functions are carried out by orthologous proteins that occur in comparable numbers. Although the fundamental gene organizations resembled each other, the genes for organization of chromatin, cytoskeletal components, and flagellar movement remarkably increased in C. reinhardtii. Molecular phylogenetic analyses suggested that the tubulin is close to plant tubulin rather than that of animals and fungi. These results reflect the increase in genome size, the acquisition of complicated cellular structures, and kinematic devices in C. reinhardtii.

Algal Proteins↗

BEAUTY: an enhanced BLAST-based search tool that integrates multiple biological information resources into sequence similarity search results.

BEAUTY (BLAST enhanced alignment utility) is an enhanced version of the NCBI's BLAST data base search tool that facilitates identification of the functions of matched sequences. We have created new data bases of conserved regions and functional domains for protein sequences in NCBI's Entrez data base, and BEAUTY allows this information to be incorporated directly into BLAST search results. A Conserved Regions Data Base, containing the locations of conserved regions within Entrez protein sequences, was constructed by (1) clustering the entire data base into families, (2) aligning each family using our PIMA multiple sequence alignment program, and (3) scanning the multiple alignments to locate the conserved regions within each aligned sequence. A separate Annotated Domains Data Base was constructed by extracting the locations of all annotated domains and sites from sequences represented in the Entrez, PROSITE, BLOCKS, and PRINTS data bases. BEAUTY performs a BLAST search of those Entrez sequences with conserved regions and/or annotated domains. BEAUTY then uses the information from the Conserved Regions and Annotated Domains data bases to generate, for each matched sequence, a schematic display that allows one to directly compare the relative locations of (1) the conserved regions, (2) annotated domains and sites, and (3) the locally aligned regions matched in the BLAST search. In addition, BEAUTY search results include World-Wide Web hypertext links to a number of external data bases that provide a variety of additional types of information on the function of matched sequences. This convenient integration of protein families, conserved regions, annotated domains, alignment displays, and World-Wide Web resources greatly enhances the biological informativeness of sequence similarity searches. BEAUTY searches can be performed remotely on our system using the "BCM Search Launcher" World-Wide Web pages (URL is < http:/ /gc.bcm.tmc.edu:8088/ search-launcher/launcher.html > ).

Amino Acid Sequence↗

Delineation of modular proteins: domain boundary prediction from sequence information.

The delineation of domain boundaries of a given sequence in the absence of known 3D structures or detectable sequence homology to known domains benefits many areas in protein science, such as protein engineering, protein 3D structure determination and protein structure prediction. With the exponential growth of newly determined sequences, our ability to predict domain boundaries rapidly and accurately from sequence information alone is both essential and critical from the viewpoint of gene function annotation. Anyone attempting to predict domain boundaries for a single protein sequence is invariably confronted with a plethora of databases that contain boundary information available from the internet and a variety of methods for domain boundary prediction. How are these derived and how well do they work? What definition of 'domain' do they use? We will first clarify the different definitions of protein domains, and then describe the available public databases with domain boundary information. Finally, we will review existing domain boundary prediction methods and discuss their strengths and weaknesses.

Algorithms↗

Differential label-free quantitative proteomic analysis of Shewanella oneidensis cultured under aerobic and suboxic conditions by accurate mass and time tag approach.

We describe the application of LC-MS without the use of stable isotope labeling for differential quantitative proteomic analysis of whole cell lysates of Shewanella oneidensis MR-1 cultured under aerobic and suboxic conditions. LC-MS/MS was used to initially identify peptide sequences, and LC-FTICR was used to confirm these identifications as well as measure relative peptide abundances. 2343 peptides covering 668 proteins were identified with high confidence and quantified. Among these proteins, a subset of 56 changed significantly using statistical approaches such as statistical analysis of microarrays, whereas another subset of 56 that were annotated as performing housekeeping functions remained essentially unchanged in relative abundance. Numerous proteins involved in anaerobic energy metabolism exhibited up to a 10-fold increase in relative abundance when S. oneidensis was transitioned from aerobic to suboxic conditions.

Amino Acid Sequence↗

Enzyme genomics: Application of general enzymatic screens to discover new enzymes.

In all sequenced genomes, a large fraction of predicted genes encodes proteins of unknown biochemical function and up to 15% of the genes with "known" function are mis-annotated. Several global approaches are routinely employed to predict function, including sophisticated sequence analysis, gene expression, protein interaction, and protein structure. In the first coupling of genomics and enzymology, Phizicky and colleagues undertook a screen for specific enzymes using large pools of partially purified proteins and specific enzymatic assays. Here we present an overview of the further developments of this approach, which involve the use of general enzymatic assays to screen individually purified proteins for enzymatic activity. The assays have relaxed substrate specificity and are designed to identify the subclass or sub-subclasses of enzymes (phosphatase, phosphodiesterase/nuclease, protease, esterase, dehydrogenase, and oxidase) to which the unknown protein belongs. Further biochemical characterization of proteins can be facilitated by the application of secondary screens with natural substrates (substrate profiling). We demonstrate here the feasibility and merits of this approach for hydrolases and oxidoreductases, two very broad and important classes of enzymes. Application of general enzymatic screens and substrate profiling can greatly speed up the identification of biochemical function of unknown proteins and the experimental verification of functional predictions produced by other functional genomics approaches.

Enzymes↗

Overview of BioCreAtIvE: critical assessment of information extraction for biology.

BACKGROUND: The goal of the first BioCreAtIvE challenge (Critical Assessment of Information Extraction in Biology) was to provide a set of common evaluation tasks to assess the state of the art for text mining applied to biological problems. The results were presented in a workshop held in Granada, Spain March 28-31, 2004. The articles collected in this BMC Bioinformatics supplement entitled "A critical assessment of text mining methods in molecular biology" describe the BioCreAtIvE tasks, systems, results and their independent evaluation. RESULTS: BioCreAtIvE focused on two tasks. The first dealt with extraction of gene or protein names from text, and their mapping into standardized gene identifiers for three model organism databases (fly, mouse, yeast). The second task addressed issues of functional annotation, requiring systems to identify specific text passages that supported Gene Ontology annotations for specific proteins, given full text articles. CONCLUSION: The first BioCreAtIvE assessment achieved a high level of international participation (27 groups from 10 countries). The assessment provided state-of-the-art performance results for a basic task (gene name finding and normalization), where the best systems achieved a balanced 80% precision / recall or better, which potentially makes them suitable for real applications in biology. The results for the advanced task (functional annotation from free text) were significantly lower, demonstrating the current limitations of text-mining approaches where knowledge extrapolation and interpretation are required. In addition, an important contribution of BioCreAtIvE has been the creation and release of training and test data sets for both tasks. There are 22 articles in this special issue, including six that provide analyses of results or data quality for the data sets, including a novel inter-annotator consistency assessment for the test set used in task 2.

Computational Biology↗

QPath: a method for querying pathways in a protein-protein interaction network.

BACKGROUND: Sequence comparison is one of the most prominent tools in biological research, and is instrumental in studying gene function and evolution. The rapid development of high-throughput technologies for measuring protein interactions calls for extending this fundamental operation to the level of pathways in protein networks. RESULTS: We present a comprehensive framework for protein network searches using pathway queries. Given a linear query pathway and a network of interest, our algorithm, QPath, efficiently searches the network for homologous pathways, allowing both insertions and deletions of proteins in the identified pathways. Matched pathways are automatically scored according to their variation from the query pathway in terms of the protein insertions and deletions they employ, the sequence similarity of their constituent proteins to the query proteins, and the reliability of their constituent interactions. We applied QPath to systematically infer protein pathways in fly using an extensive collection of 271 putative pathways from yeast. QPath identified 69 conserved pathways whose members were both functionally enriched and coherently expressed. The resulting pathways tended to preserve the function of the original query pathways, allowing us to derive a first annotated map of conserved protein pathways in fly. CONCLUSION: Pathway homology searches using QPath provide a powerful approach for identifying biologically significant pathways and inferring their function. The growing amounts of protein interactions in public databases underscore the importance of our network querying framework for mining protein network data.

Algorithms↗

Identification and comparative analysis of components from the signal recognition particle in protozoa and fungi.

BACKGROUND: The signal recognition particle (SRP) is a ribonucleoprotein complex responsible for targeting proteins to the ER membrane. The SRP of metazoans is well characterized and composed of an RNA molecule and six polypeptides. The particle is organized into the S and Alu domains. The Alu domain has a translational arrest function and consists of the SRP9 and SRP14 proteins bound to the terminal regions of the SRP RNA. So far, our understanding of the SRP and its evolution in lower eukaryotes such as protozoa and yeasts has been limited. However, genome sequences of such organisms have recently become available, and we have now analyzed this information with respect to genes encoding SRP components. RESULTS: A number of SRP RNA and SRP protein genes were identified by an analysis of genomes of protozoa and fungi. The sequences and secondary structures of the Alu portion of the RNA were found to be highly variable. Furthermore, proteins SRP9/14 appeared to be absent in certain species. Comparative analysis of the SRP RNAs from different Saccharomyces species resulted in models which contain features shared between all SRP RNAs, but also a new secondary structure element in SRP RNA helix 5. Protein SRP21, previously thought to be present only in Saccharomyces, was shown to be a constituent of additional fungal genomes. Furthermore, SRP21 was found to be related to metazoan and plant SRP9, suggesting that the two proteins are functionally related. CONCLUSIONS: Analysis of a number of not previously annotated SRP components show that the SRP Alu domain is subject to a more rapid evolution than the other parts of the molecule. For instance, the RNA portion is highly variable and the protein SRP9 seems to have evolved into the SRP21 protein in fungi. In addition, we identified a secondary structure element in the Saccharomyces RNA that has been inserted close to the Alu region. Together, these results provide important clues as to the structure, function and evolution of SRP.

Amino Acid Sequence↗

The SUPERFAMILY database in 2007: families and functions.

The SUPERFAMILY database provides protein domain assignments, at the SCOP 'superfamily' level, for the predicted protein sequences in over 400 completed genomes. A superfamily groups together domains of different families which have a common evolutionary ancestor based on structural, functional and sequence data. SUPERFAMILY domain assignments are generated using an expert curated set of profile hidden Markov models. All models and structural assignments are available for browsing and download from http://supfam.org. The web interface includes services such as domain architectures and alignment details for all protein assignments, searchable domain combinations, domain occurrence network visualization, detection of over- or under-represented superfamilies for a given genome by comparison with other genomes, assignment of manually submitted sequences and keyword searches. In this update we describe the SUPERFAMILY database and outline two major developments: (i) incorporation of family level assignments and (ii) a superfamily-level functional annotation. The SUPERFAMILY database can be used for general protein evolution and superfamily-specific studies, genomic annotation, and structural genomics target suggestion and assessment.

Databases, Protein↗

A chromosome-level reference genome assembly of the Small snakehead (Channa asiatica).

The Small snakehead (Channa asiatica) is an economically important species in both aquaculture and ornamental trade, mainly distributed in South China and Southeast Asia. Despite its significance, limited genomic resources have impeded in-depth genetic studies and breeding programs. In this study, we used PacBio HiFi long-read sequencing, Illumina short-read sequencing, and Hi-C technologies to generate a high-quality chromosome-level genome of the C. asiatica. The final genome spans 659.44&#x2009;Mb, with an impressive 98.18% anchored to 23 chromosomes. Notably, the contig N50 and scaffold N50 are 23.92&#x2009;Mb and 29.61&#x2009;Mb, validated by a BUSCO completeness score of 98.93%. Genome annotation identified 26,603 protein-coding genes, 99.29% of which were confirmed by BUSCO analysis, and 93.68% were functionally annotated. Approximately 27.72% of the genome sequences were classified as repeat elements. This high-fidelity genome assembly provides a robust foundation for advancing molecular breeding, comparative genomics, and evolutionary studies of C. asiatica and related species.

Animals↗

Expressed sequence tags (ESTs) and simple sequence repeat (SSR) markers from octoploid strawberry (Fragaria x ananassa).

BACKGROUND: Cultivated strawberry (Fragaria x ananassa) represents one of the most valued fruit crops in the United States. Despite its economic importance, the octoploid genome presents a formidable barrier to efficient study of genome structure and molecular mechanisms that underlie agriculturally-relevant traits. Many potentially fruitful research avenues, especially large-scale gene expression surveys and development of molecular genetic markers have been limited by a lack of sequence information in public databases. As a first step to remedy this discrepancy a cDNA library has been developed from salicylate-treated, whole-plant tissues and over 1800 expressed sequence tags (EST's) have been sequenced and analyzed. RESULTS: A putative unigene set of 1304 sequences--133 contigs and 1171 singlets--has been developed, and the transcripts have been functionally annotated. Homology searches indicate that 89.5% of sequences share significant similarity to known/putative proteins or Rosaceae ESTs. The ESTs have been functionally characterized and genes relevant to specific physiological processes of economic importance have been identified. A set of tools useful for SSR development and mapping is presented. CONCLUSION: Sequences derived from this effort may be used to speed gene discovery efforts in Fragaria and the Rosaceae in general and also open avenues of comparative mapping. This report represents a first step in expanding molecular-genetic analyses in strawberry and demonstrates how computational tools can be used to optimally mine a large body of useful information from a relatively small data set.

Chromosome Mapping↗

The past, present and future of genome-wide re-annotation.

Annotation, the process by which structural or functional information is inferred for genes or proteins, is crucial for obtaining value from genome sequences. We define the process of annotating a previously annotated genome sequence as 're-annotation', and examine the strengths and weaknesses of current manual and automatic genome-wide re-annotation approaches.

Computational Biology↗

N-terminal N-myristoylation of proteins: refinement of the sequence motif and its taxon-specific differences.

N-terminal N-myristoylation is a lipid anchor modification of eukaryotic and viral proteins targeting them to membrane locations, thus changing the cellular function of modified proteins. Protein myristoylation is critical in many pathways; e.g. in signal transduction, apoptosis, or alternative extracellular protein export. The myristoyl-CoA:protein N-myristoyltransferase (NMT) recognizes the sequence motif of appropriate substrate proteins at the N terminus and attaches the lipid moiety to the absolutely required N-terminal glycine residue. Reliable recognition of capacity for N-terminal myristoylation from the substrate protein sequence alone is desirable for proteome-wide function annotation projects but the existing PROSITE motif is not practical, since it produces huge numbers of false positive and even some false negative predictions. As a first step towards a new prediction method, it is necessary to refine the sequence motif coding for N-terminal N-myristoylation. Relying on the in-depth study of the amino acid sequence variability of substrate proteins, on binding site analyses in X-ray structures or 3D homology models for NMTs from various taxa, and on consideration of biochemical data extracted from the scientific literature, we found indications that, at least within a complete substrate protein, the N-terminal 17 protein residues experience different types of variability restrictions. We identified three motif regions: region 1 (positions 1-6) fitting the binding pocket; region 2 (positions 7-10) interacting with the NMT's surface at the mouth of the catalytic cavity; and region 3 (positions 11-17) comprising a hydrophilic linker. Each region was characterized by physical requirements to single sequence positions or groups of positions regarding volume, polarity, backbone flexibility and other typical properties of amino acids (http://mendel.imp.univie.ac.at/myristate/). These specificity differences are confined partly to taxonomic ranges and are proposed for the design of NMT inhibitors in pathogenic fungal and protozoan systems including Aspergillus fumigatus, Leishmania major, Trypanosoma cruzi, Trypanosoma brucei, Giardia intestinalis, Entamoeba histolytica, Pneumocystis carinii, Strongyloides stercoralis and Schistosoma mansoni. An exhaustive search for NMT-homologues led to the discovery of two putative entomopoxviral NMTs.

Acyltransferases↗