Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

The sea urchin kinome: a first look.

This paper reports a preliminary in silico analysis of the sea urchin kinome. The predicted protein kinases in the sea urchin genome were identified, annotated and classified, according to both function and kinase domain taxonomy. The results show that the sea urchin kinome, consisting of 353 protein kinases, is closer to the Drosophila kinome (239) than the human kinome (518) with respect to total kinase number. However, the diversity of sea urchin kinases is surprisingly similar to humans, since the urchin kinome is missing only 4 of 186 human subfamilies, while Drosophila lacks 24. Thus, the sea urchin kinome combines the simplicity of a non-duplicated genome with the diversity of function and signaling previously considered to be vertebrate-specific. More than half of the sea urchin kinases are involved with signal transduction, and approximately 88% of the signaling kinases are expressed in the developing embryo. These results support the strength of this nonchordate deuterostome as a pivotal developmental and evolutionary model organism.

Animals↗

High-Resolution Chromosome-Level Genome Assembly and Annotation of Triplophysa stewarti, an Endemic Plateau Loach from the Qinghai-Tibet Plateau.

The bottom-dwelling fish Triplophysa stewarti, endemic to the Qinghai-Tibet Plateau, is a valuable model for studying high-altitude adaptation in aquatic ecosystems. However, the lack of a high-quality reference genome has hindered comparative genomic and evolutionary studies within this genus. Here, we present a chromosome-level genome assembly for T. stewarti, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding. The 697.9 Mb assembly is highly continuous (scaffold N50 of 253.58 Mb) and encompasses 25 chromosomes, representing 92.65% of the genome. BUSCO analysis indicated a 98.4% completeness, supporting the high quality of the assembly. We annotated 28,009 protein-coding genes, with 97.04% being functionally assigned across multiple databases (NR, UniProt, KEGG, GO, Pfam and InterPro). Additionally, repetitive elements constituted 42.47% of the genome, and we identified 52,709 non-coding RNAs. This high-quality reference genome provides a fundamental resource for exploring the adaptive evolution, population structure, and conservation genetics of T. stewarti and related species on the Qinghai-Tibet Plateau.

Animals↗

The proteome: structure, function and evolution.

This paper reports two studies to model the inter-relationships between protein sequence, structure and function. First, an automated pipeline to provide a structural annotation of proteomes in the major genomes is described. The results are stored in a database at Imperial College, London (3D-GENOMICS) that can be accessed at www.sbg.bio.ic.ac.uk. Analysis of the assignments to structural superfamilies provides evolutionary insights. 3D-GENOMICS is being integrated with related proteome annotation data at University College London and the European Bioinformatics Institute in a project known as e-protein (http://www.e-protein.org/). The second topic is motivated by the developments in structural genomics projects in which the structure of a protein is determined prior to knowledge of its function. We have developed a new approach PHUNCTIONER that uses the gene ontology (GO) classification to supervise the extraction of the sequence signal responsible for protein function from a structure-based sequence alignment. Using GO we can obtain profiles for a range of specificities described in the ontology. In the region of low sequence similarity (around 15%), our method is more accurate than assignment from the closest structural homologue. The method is also able to identify the specific residues associated with the function of the protein family.

Computational Biology↗

Complementing genomics with proteomics: the membrane subproteome of Pseudomonas aeruginosa PAO1.

With the completion of many genome projects, a shift is now occurring from the acquisition of gene sequence to understanding the role and context of gene products within the genome. The opportunistic pathogen Pseudomonas aeruginosa is one organism for which a genome sequence is now available, including the annotation of open reading frames (ORFs). However, approximately one third of the ORFs are as yet undefined in function. Proteomics can complement genomics, by characterising gene products and their response to a variety of biological and environmental influences. In this study we have established the first two-dimensional gel electrophoresis reference map of proteins from the membrane fraction of P. aeruginosa strain PA01. A total of 189 proteins have been identified and correlated with 104 genes from the P. aeruginosa genome. Annotated membrane proteins could be grouped into three distinct categories: (i) those with functions previously characterised in P. aeruginosa (38%); (ii) those with significant sequence similarity to proteins with assigned function or hypothetical proteins in other organisms (46%); and (iii) those with unknown function (16%). Transmembrane prediction algorithms showed that each identified protein sequence contained at least one membrane-spanning region. Furthermore, the current methodology used to isolate the membrane fraction was shown to be highly specific since no contaminating cytosolic proteins were characterised. Preliminary analysis showed that at least 15 gel spots may be glycosylated in vivo, including three proteins that have not previously been functionally characterised. The reference map of membrane proteins from this organism is now the basis for determining surface molecules associated with antibiotic resistance and efflux, cell-cell signalling and pathogen-host interactions in a variety of P. aeruginosa strains.

Bacterial Proteins↗

TomatEST database: in silico exploitation of EST data to explore expression patterns in tomato species.

TomatEST is a secondary database integrating expressed sequence tag (EST)/cDNA sequence information from different libraries of multiple tomato species. Redundant EST collections from each species are organized into clusters (gene indices). A cluster consists of one or multiple contigs. Multiple contigs in a cluster represent alternatively transcribed forms of a gene. The set of stand-alone EST sequences (singletons) and contigs, representing all the computationally defined 'Transcript Indices', are annotated according to similarity versus protein and RNA family databases. Sequence function description is integrated with the Gene Ontologies and the Enzyme Commission identifiers for a standard classification of gene products and for the mapping of the expressed sequences onto metabolic pathways. Information on the origin of the ESTs, on their structural features, on clusters and contigs, as well as on functional annotations are accessible via a user-friendly web interface. Specific facilities in the database allow Transcript Indices from a query be automatically classified in Enzyme classes and in metabolic pathways. The 'on the fly' mapping onto the metabolic maps is integrated in the analytical tools. The TomatEST database website is freely available at http://biosrv.cab.unina.it/tomatestdb.

Computational Biology↗

Identification of functional modules in a PPI network by clique percolation clustering.

Large-scale experiments and data integration have provided the opportunity to systematically analyze and comprehensively understand the topology of biological networks and biochemical processes in cells. Modular architecture which encompasses groups of genes/proteins involved in elementary biological functional units is a basic form of the organization of interacting proteins. Here we apply a graph clustering algorithm based on clique percolation clustering to detect overlapping network modules of a protein-protein interaction (PPI) network. Our analysis of the yeast Sacchromyces cerevisiae suggests that most of the detected modules correspond to one or more experimentally functional modules and half of these annotated modules match well with experimentally determined protein complexes. Our method of analysis can of course be applied to protein-protein interaction data for any species and even other biological networks.

Algorithms↗

OntoBlast function: From sequence similarities directly to potential functional annotations by ontology terms.

OntoBlast allows one to find information about potential functions of proteins by presenting a weighted list of ontology entries associated with similar sequences from completely sequenced genomes identified in a BLAST search. It combines, in a single analysis step, the search for sequence similarities in several species with the association of information stored in ontologies. From each identified ontology term a list of genes, which share the functional annotation, can be retrieved. The OntoBlast function is an integral part of the 'Ontologies TO GenomeMatrix' tool which provides an alternative entry point from ontology terms to the Genome-Matrix database. OntoBlast's web interface is accessible on the 'Ontologies TO GenomeMatrix Gate' page at http://functionalgenomics.de/ontogate/.

Animals↗

HCVDB: hepatitis C virus sequences database.

UNLABELLED: To date, more than 30 000 hepatitis C virus (HCV) sequences have been deposited in the generalist databases DNA Data Bank of Japan (DDBJ), EMBL Nucleotide Sequence Database (EMBL) and GenBank. The main difficulties with HCV sequences in these databases are their retrieval, annotation and analyses. To help HCV researchers face the increasing needs of HCV sequence analyses, we developed a specialised database of computer-annotated HCV sequences, called HCVDB. HCVDB is re-built every month from an up-to-date EMBL database by an automated process. HCVDB provides key data about the HCV sequences (e.g. genotype, genomic region, protein names and functions, known 3-dimensional structures) and ensures consistency of the annotations, which enables reliable keyword queries. The database is highly integrated with sequence and structure analysis tools and the SRS (LION bioscience) keywords query system. Thus, any user can extract subsets of sequences matching particular criteria or enter their own sequences and analyse them with various bioinformatics programs available on the same server. AVAILABILITY: HCVDB is available from http://hepatitis.ibcp.fr.

Amino Acid Sequence↗

The role of alternative translation start sites in the generation of human protein diversity.

According to the scanning model, 40S ribosomal subunits initiate translation at the first (5' proximal) AUG codon they encounter. However, if the first AUG is in a suboptimal context, it may not be recognized, and translation can then initiate at downstream AUG(s). In this way, a single RNA can produce several variant products. Earlier experiments suggested that some of these additional protein variants might be functionally important. We have analysed human mRNAs that have AUG triplets in 5' untranslated regions and mRNAs in which the annotated translational start codon is located in a suboptimal context. It was found that 3% of human mRNAs have the potential to encode N-terminally extended variants of the annotated proteins and 12% could code for N-truncated variants. The predicted subcellular localizations of these protein variants were compared: 31% of the N-extended proteins and 30% of the N-truncated proteins were predicted to localize to subcellular compartments that differed from those targeted by the annotated protein forms. These results suggest that additional AUGs may frequently be exploited for the synthesis of proteins that possess novel functional properties.

5' Untranslated Regions↗

Characterization of 954 bovine full-CDS cDNA sequences.

BACKGROUND: Genome assemblies rely on the existence of transcript sequence to stitch together contigs, verify assembly of whole genome shotgun reads, and annotate genes. Functional genomics studies also rely on transcript sequence to create expression microarrays or interpret digital tag data produced by methods such as Serial Analysis of Gene Expression (SAGE). Transcript sequence can be predicted based on reconstruction from overlapping expressed sequence tags (EST) that are obtained by single-pass sequencing of random cDNA clones, but these reconstructions are prone to errors caused by alternative splice forms, transcripts from gene families with related sequences, and expressed pseudogenes. These errors confound genome assembly and annotation. The most useful transcript sequences are derived by complete insert sequencing of clones containing the entire length, or at least the full protein coding sequence (CDS) portion, of the source mRNA. While the bovine genome sequencing initiative is nearing completion, there is currently a paucity of bovine full-CDS mRNA and protein sequence data to support bovine genome assembly and functional genomics studies. Consequently, the production of high-quality bovine full-CDS cDNA sequences will enhance the bovine genome assembly and functional studies of bovine genes and gene products. The goal of this investigation was to identify and characterize the full-CDS sequences of bovine transcripts from clones identified in non-full-length enriched cDNA libraries. In contrast to several recent full-length cDNA investigations, these full-CDS cDNAs were selected, sequenced, and annotated without the benefit of the target organism's genomic sequence, by using comparison of bovine EST sequence to existing human mRNA to identify likely full-CDS clones for full-length insert cDNA (FLIC) sequencing. RESULTS: The predicted bovine protein lengths, 5' UTR lengths, and Kozak consensus sequences from 954 bovine FLIC sequences (bFLICs; average length 1713 nt, representing 762 distinct loci) are all consistent with previously sequenced mammalian full-length transcripts. CONCLUSION: In most cases, the bFLICs span the entire CDS of the genes, providing the basis for creating predicted bovine protein sequences to support proteomics and comparative evolutionary research as well as functional genomics and genome annotation. The results demonstrate the utility of the comparative approach in obtaining predicted protein sequences in other species.

5' Untranslated Regions↗

Protective effects of seminal exosomes on cryopreserved sperm via inhibiting oxidative damage.

This study aimed to explore the protective effect of seminal plasma exosomes (SPEs) on human sperm structure and function during cryopreservation and its potential mechanism. The samples were divided into two groups: the control group was treated solely with sperm cryoprotectant before freezing, while the exosome group was supplemented with SPEs. After cryopreservation and thawing, sperm progressive motility, normal morphological rate, and survival rate were evaluated. Furthermore, PKH67 labeling experiments were performed, and oxidative stress markers as well as energy metabolism indicators in sperm were detected. Subsequent mechanism exploration was conducted via proteomic analysis and protein validation assays. This work reveals that adding SPEs at a concentration of 1 or 2 mg/ml effectively improves sperm progressive motility after cryopreservation. After supplementing with SPEs, sperm glucose levels are reduced and mitochondrial membrane potential is enhanced. Simultaneously, SPEs alleviate oxidative stress by decreasing reactive oxygen species (ROS) and DNA fragment index (DFI) while increasing superoxide dismutase (SOD) activity. Functional annotation of proteomics reveals that 14 of the differentially expressed proteins (DEPs) are associated with sperm motility. Enriched metabolic pathways related to sperm motility and sperm protein validation experiments indicate that the expression of MAPK, p-MAPK, and p-JNK proteins in sperm is higher in the Exosome group than in the Control group. This study provides important theoretical support for the application of SPEs in mitigating cryopreservation damage to sperm by enhancing antioxidant capacity. The specific mechanism may be mediated by the MAPK/p-JNK pathway.

Male↗

ESTHER, the database of the alpha/beta-hydrolase fold superfamily of proteins.

The alpha/beta-hydrolase fold is characterized by a beta-sheet core of five to eight strands connected by alpha-helices to form a alpha/beta/alpha sandwich. In most of the family members the beta-strands are parallels, but some show an inversion in the order of the first strands, resulting in antiparallel orientation. The members of the superfamily diverged from a common ancestor into a number of hydrolytic enzymes with a wide range of substrate specificities, together with other proteins with no recognized catalytic activity. In the enzymes the catalytic triad residues are presented on loops, of which one, the nucleophile elbow, is the most conserved feature of the fold. Of the other proteins, which all lack from one to all of the catalytic residues, some may simply be 'inactive' enzymes while others are known to be involved in surface recognition functions. The ESTHER database (http://bioweb.ensam.inra.fr/esther) gathers and annotates all the published information related to gene and protein sequences of this superfamily, as well as biochemical, pharmacological and structural data, and connects them so as to provide the bases for studying structure-function relationships within the family. The most recent developments of the database, which include a section on human diseases related to members of the family, are described.

Animals↗

A sequence alignment-independent method for protein classification.

Annotation of the rapidly accumulating body of sequence data relies heavily on the detection of remote homologues and functional motifs in protein families. The most popular methods rely on sequence alignment. These include programs that use a scoring matrix to compare the probability of a potential alignment with random chance and programs that use curated multiple alignments to train profile hidden Markov models (HMMs). Related approaches depend on bootstrapping multiple alignments from a single sequence. However, alignment-based programs have limitations. They make the assumption that contiguity is conserved between homologous segments, which may not be true in genetic recombination or horizontal transfer. Alignments also become ambiguous when sequence similarity drops below 40%. This has kindled interest in classification methods that do not rely on alignment. An approach to classification without alignment based on the distribution of contiguous sequences of four amino acids (4-grams) was developed. Interest in 4-grams stemmed from the observation that almost all theoretically possible 4-grams (20(4)) occur in natural sequences and the majority of 4-grams are uniformly distributed. This implies that the probability of finding identical 4-grams by random chance in unrelated sequences is low. A Bayesian probabilistic model was developed to test this hypothesis. For each protein family in Pfam-A and PIR-PSD, a feature vector called a probe was constructed from the set of 4-grams that best characterised the family. In rigorous jackknife tests, unknown sequences from Pfam-A and PIR-PSD were compared with the probes for each family. A classification result was deemed a true positive if the probe match with the highest probability was in first place in a rank-ordered list. This was achieved in 70% of cases. Analysis of false positives suggested that the precision might approach 85% if selected families were clustered into subsets. Case studies indicated that the 4-grams in common between an unknown and the best matching probe correlated with functional motifs from PRINTS. The results showed that remote homologues and functional motifs could be identified from an analysis of 4-gram patterns.

Algorithms↗

The COG database: a tool for genome-scale analysis of protein functions and evolution.

Rational classification of proteins encoded in sequenced genomes is critical for making the genome sequences maximally useful for functional and evolutionary studies. The database of Clusters of Orthologous Groups of proteins (COGs) is an attempt on a phylogenetic classification of the proteins encoded in 21 complete genomes of bacteria, archaea and eukaryotes (http://www. ncbi.nlm. nih.gov/COG). The COGs were constructed by applying the criterion of consistency of genome-specific best hits to the results of an exhaustive comparison of all protein sequences from these genomes. The database comprises 2091 COGs that include 56-83% of the gene products from each of the complete bacterial and archaeal genomes and approximately 35% of those from the yeast Saccharomyces cerevisiae genome. The COG database is accompanied by the COGNITOR program that is used to fit new proteins into the COGs and can be applied to functional and phylogenetic annotation of newly sequenced genomes.

Database Management Systems↗

Profiling and functional annotation of mRNA gene expression in pediatric rhabdomyosarcoma and Ewing's sarcoma.

Using Affymetrix oligonucleotide microarrays, we analyzed mRNA gene expression patterns of 12 primary pediatric rhabdomyosarcomas (RMS) and 11 Ewing's sarcomas (EWS), which belong to the small round blue cell tumors (SRBCTs). Diagnostic classification of these cancers is frequently complicated by the highly similar appearance in routine histology, and additional molecular markers could significantly improve tumor classification. A combination of three independent statistical approaches (t-test, SAM, k-nearest neighborhood analysis) resulted in 101 highly significant probe sets that clearly discriminate between EWS and RMS. We identified novel marker transcripts that have not been previously associated with either RMS or EWS yet, including CITED2, glypican 3 (GPC3), and cyclin D1 (CCND1). Expression levels for selected candidate genes were validated by quantitative real-time reverse-transcription PCR. Furthermore, to identify biologically meaningful trends, functional annotations were assigned to 946 genes differentially expressed between EWS and RMS (t-test). Genes involved in protein biosynthesis (n = 28) and complex assembly (n = 9), lipid metabolism (n = 23), energy generation (n = 22), and mRNA processing (n = 11) were expressed significantly higher in EWS. Thus, functional annotation of tumor-specific genes reveals detailed insights into tumor biology and differentiation-specific expression patterns and gives important clues related to the possible cellular origin of these pediatric tumors. Supplementary material for this article is available at the International Journal of Cancer website at http://www.interscience.wiley.com/jpages/0020-7136/suppmat/index.html.

Biomarkers, Tumor↗

Structural domains, protein modules, and sequence similarities enrich our understanding of the Shewanella oneidensis MR-1 proteome.

The protein coding sequences of S. oneidensis MR-1 were analyzed, and new annotations were given to 491 gene products, 306 of which were previously of unknown function. New information was mainly brought in from structural domain predictions for S. oneidensis proteins of the SUPERFAM database (http://supfam.mrc-lmb.cam.ac.uk/SUPERFAMILY/) and newly identified and experimentally verified functions of homologous proteins. Proteins encoded by fused genes were identified and separated into modules, protein units of at least 83 aa with independent functions and distinct evolutionary histories. A reannotation of the fused gene products was done to assign functions to the appropriate module within the protein. Groups of sequence-similar proteins of S. oneidensis were assembled. The fused gene products were represented by their modular entities for the grouping process. The protein groups were analyzed for their size and functions, and they were used to indicate activities that are of importance to the environmental adaptation of this organism. Making use of several approaches not commonly used in annotation, we have been able to enrich our understanding of the functions encoded by the S. oneidensis genome.

Bacterial Proteins↗

Refining protein subcellular localization.

The study of protein subcellular localization is important to elucidate protein function. Even in well-studied organisms such as yeast, experimental methods have not been able to provide a full coverage of localization. The development of bioinformatic predictors of localization can bridge this gap. We have created a Bayesian network predictor called PSLT2 that considers diverse protein characteristics, including the combinatorial presence of InterPro motifs and protein interaction data. We compared the localization predictions of PSLT2 to high-throughput experimental localization datasets. Disagreements between these methods generally involve proteins that transit through or reside in the secretory pathway. We used our multi-compartmental predictions to refine the localization annotations of yeast proteins primarily by distinguishing between soluble lumenal proteins and soluble proteins peripherally associated with organelles. To our knowledge, this is the first tool to provide this functionality. We used these sub-compartmental predictions to characterize cellular processes on an organellar scale. The integration of diverse protein characteristics and protein interaction data in an appropriate setting can lead to high-quality detailed localization annotations for whole proteomes. This type of resource is instrumental in developing models of whole organelles that provide insight into the extent of interaction and communication between organelles and help define organellar functionality.

Amino Acid Motifs↗

An integrated approach to the prediction of domain-domain interactions.

BACKGROUND: The development of high-throughput technologies has produced several large scale protein interaction data sets for multiple species, and significant efforts have been made to analyze the data sets in order to understand protein activities. Considering that the basic units of protein interactions are domain interactions, it is crucial to understand protein interactions at the level of the domains. The availability of many diverse biological data sets provides an opportunity to discover the underlying domain interactions within protein interactions through an integration of these biological data sets. RESULTS: We combine protein interaction data sets from multiple species, molecular sequences, and gene ontology to construct a set of high-confidence domain-domain interactions. First, we propose a new measure, the expected number of interactions for each pair of domains, to score domain interactions based on protein interaction data in one species and show that it has similar performance as the E-value defined by Riley et al. Our new measure is applied to the protein interaction data sets from yeast, worm, fruitfly and humans. Second, information on pairs of domains that coexist in known proteins and on pairs of domains with the same gene ontology function annotations are incorporated to construct a high-confidence set of domain-domain interactions using a Bayesian approach. Finally, we evaluate the set of domain-domain interactions by comparing predicted domain interactions with those defined in iPfam database that were derived based on protein structures. The accuracy of predicted domain interactions are also confirmed by comparing with experimentally obtained domain interactions from H. pylori. As a result, a total of 2,391 high-confidence domain interactions are obtained and these domain interactions are used to unravel detailed protein and domain interactions in several protein complexes. CONCLUSION: Our study shows that integration of multiple biological data sets based on the Bayesian approach provides a reliable framework to predict domain interactions. By integrating multiple data sources, the coverage and accuracy of predicted domain interactions can be significantly increased.

Algorithms↗