Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domains structure, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include format and content enhancements, cross-references to additional databases, new documentation files and improvements to TrEMBL, a computer-annotated supplement to SWISS-PROT. TrEMBL consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDSs) in the EMBL Nucleotide Sequence Database, except the CDSs already included in SWISS-PROT. We also describe the Human Proteomics Initiative (HPI), a major project to annotate all known human sequences according to the quality standards of SWISS-PROT. SWISS-PROT is available at: http://www.expasy.ch/sprot/ and http://www.ebi.ac.uk/swissprot/

Animals↗

Understanding protein trafficking in plant cells through proteomics.

The functions of approximately one-third of the proteins encoded by the Arabidopsis thaliana genome are completely unknown. Moreover, many annotations of the remainder of the genome supply tentative functions, at best. Knowing the ultimate localization of these proteins, as well as the pathways used for getting there, may provide clues as to their functions. The putative localization of most proteins currently relies on in silico-based bioinformatics approaches, which, unfortunately, often result in erroneous predictions. Emerging proteomics techniques coupled with other systems biology approaches now provide researchers with a plethora of methods for elucidating the final location of these proteins on a large scale, as well as the ability to dissect protein-sorting pathways in plants.

Databases, Protein↗

Use of search algorithms to define specificity in Rab GTPase domain function.

The continuing explosion of sequencing data has inspired a corresponding effort in the annotation and classification of protein families. Within a particular protein family, however, individual members may have distinct functions, although they share a common fold and broadly defined physiological role. Rab GTPases are the largest subfamily of the Ras superfamily, yet from early in their discovery, it was apparent that each Rab protein has a unique subcellular localization and regulates a particular stage(s) membrane traffic. To gain insight into the contribution of individual residues to unique protein functions a general strategy is outlined. This method should allow the cell and molecular biologist with no specialist expertise to implement an algorithm that makes use of a combination of experimental and phylogenetic data. The algorithm is applicable to the analysis of any protein domain and here is illustrated with the analysis of residues contributing to the individual functions of a pair of Rab GTPases.

Algorithms↗

The genome sequence of Mannheimia haemolytica A1: insights into virulence, natural competence, and Pasteurellaceae phylogeny.

The draft genome sequence of Mannheimia haemolytica A1, the causative agent of bovine respiratory disease complex (BRDC), is presented. Strain ATCC BAA-410, isolated from the lung of a calf with BRDC, was the DNA source. The annotated genome includes 2,839 coding sequences, 1,966 of which were assigned a function and 436 of which are unique to M. haemolytica. Through genome annotation many features of interest were identified, including bacteriophages and genes related to virulence, natural competence, and transcriptional regulation. In addition to previously described virulence factors, M. haemolytica encodes adhesins, including the filamentous hemagglutinin FhaB and two trimeric autotransporter adhesins. Two dual-function immunoglobulin-protease/adhesins are also present, as is a third immunoglobulin protease. Genes related to iron acquisition and drug resistance were identified and are likely important for survival in the host and virulence. Analysis of the genome indicates that M. haemolytica is naturally competent, as genes for natural competence and DNA uptake signal sequences (USS) are present. Comparison of competence loci and USS in other species in the family Pasteurellaceae indicates that M. haemolytica, Actinobacillus pleuropneumoniae, and Haemophilus ducreyi form a lineage distinct from other Pasteurellaceae. This observation was supported by a phylogenetic analysis using sequences of predicted housekeeping genes.

Actinobacillus pleuropneumoniae↗

A high-throughput approach for subcellular proteome: identification of rat liver proteins using subcellular fractionation coupled with two-dimensional liquid chromatography tandem mass spectrometry and bioinformatic analysis.

Four fractions from rat liver (a crude mitochondria (CM) and cytosol (C) fraction obtained with differential centrifugation, a purified mitochondrial (PM) fraction obtained with nycodenz density gradient centrifugation, and a total liver (TL) fraction) were analyzed with two-dimensional liquid chromatography tandem mass spectrometry analysis. A total of 564 rat proteins were identified and were bioinformatically annotated according to their physicochemical characteristics and functions. While most extreme alkaline ribosomal proteins were identified in the TL fraction, the C fraction mainly included neutral enzymes and the PM fraction enriched alkaline proteins and proteins with electron transfer activity or oxygen binding activity. Such characteristics were more apparent in proteins identified only in the TL, C, or PM fraction. The Swiss-Prot annotation and the bioinformatic prediction results proved that the C and PM fractions had enriched cytoplasmic or mitochondrial proteins, respectively. Combination usage of subcellular fractionation with two-dimensional liquid chromatography tandem mass spectrometry was proved to be a high-throughput, sensitive, and effective analytical approach for subcellular proteomics research. Using such a strategy, we have constructed the largest proteome database to date for rat liver (564 rat proteins) and its cytosol (222 rat proteins) and mitochondrial fractions (227 rat proteins). Moreover, the 352 proteins with Swiss-Prot subcellular location annotation in the 564 identified proteins were used as an actual subcellular proteome dataset to evaluate the widely used bioinformatics tools such as PSORT, TargetP, TMHMM, and GRAVY.

Animals↗

Integration of data for gene annotation using the BioMediator system.

Gene annotation requires integration of data from multiple sources in order to functionally classify genes. We are using BioMediator, a general purpose data-integration solution, to develop a gene annotation system to automate the process of collecting data from disparate genomic databases. Integration of annotation data from multiple sources into a single format will facilitate use of analytic tools for the proper functional classification of genes.

Base Sequence↗

The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1998.

SWISS-PROT (http://www.expasy.ch/) is a curated protein sequence database which strives to provide a high level of annotations (such as the description of the function of a protein, its domains structure, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to two additional databases; a variety of new documentation files and improvements to TrEMBL, a computer annotated supplement to SWISS-PROT. TrEMBL consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except the CDS already included in SWISS-PROT.

Amino Acid Sequence↗

Discovering hidden candidate plastic-degrading enzymes: Combined multi-omics and machine learning strategy.

Plastic pollution poses a major threat to the stability of natural ecosystems as well as human health. Microbial enzymes have long been considered a potential resource for targeted biodegradation but, except for a few successful cases, the discovery of efficient enzymes has proved challenging. Aiming to accelerate the process, we propose an approach combining metagenomics, metatranscriptomics and semi-supervised learning that selects promising plastic-degrading candidate enzymes from the proteome of relevant microorganisms. Tested on a dataset of over 10,000 microbial proteins, ranking models consistently prioritize known plastic-degrading enzymes, achieving an area under the cumulative distribution function curve above 0.96, with leave-one-family-out cross-validation indicating that performance is largely retained across protein families. As a case study, this work focuses on mixed microbial cultures exposed for extended periods to polyethylene, polyethylene terephthalate, and polyurethane substrates. The prevalent species after selective enrichment were functionally characterized, finding Rhodococcus aetherivorans as the most relevant species in two of the five cultures under investigation. Among the top-ranked proteins, several have high structural similarity with known enzymes despite not being identified by sequence similarity search. Moreover, according to metatranscriptomics results, several of these enzymes were found to be expressed at the same level or above that of annotated enzymes, suggesting that they may have functional relevance. Overall, this work highlights the potential of integrating multi-omics with data-driven methods for enzyme discovery and for accelerating the development of biotechnological solutions to plastic pollution.

Biodegradation, Environmental↗

RIKEN Arabidopsis full-length (RAFL) cDNA and its applications for expression profiling under abiotic stress conditions.

Full-length cDNAs are essential for the correct annotation of genomic sequences and for the functional analysis of genes and their products. 155,144 RIKEN Arabidopsis full-length (RAFL) cDNA clones were isolated. The 3'-end expressed sequence tags (ESTs) of all 155,144 RAFL cDNAs were clustered into 14,668 non-redundant cDNA groups, about 60% of predicted genes. The sequence database of the RAFL cDNAs is useful for promoter analysis and the correct annotation of predicted transcription units and gene products. Recently, cDNA microarray analysis has been developed for quantitative analysis of global and simultaneous analysis of expression profiles. RAFL cDNA microarrays were prepared, containing independent full-length cDNA groups for analysing the expression profiles of genes under various stress- and hormone-treatment conditions and in various mutants and transgenic plants. In this review, recent progress on transcriptome analysis using the RAFL cDNA microarray is highlighted.

Arabidopsis↗

The SWISS-PROT protein sequence data bank and its supplement TrEMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotations (such as the description of the function of a protein, structure of its domains, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to two additional databases; a variety of new documentation files and the creation of TrEMBL, a computer annotated supplement to SWISS-PROT. This supplement consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except the CDS already included in SWISS-PROT.

Academies and Institutes↗

Long homopurine*homopyrimidine sequences are characteristic of genes expressed in brain and the pseudoautosomal region.

Homo(purine*pyrimidine) sequences (R*Y tracts) with mirror repeat symmetries form stable triplexes that block replication and transcription and promote genetic rearrangements. A systematic search was conducted to map the location of the longest R*Y tracts in the human genome in order to assess their potential function(s). The 814 R*Y tracts with > or =250 uninterrupted base pairs were preferentially clustered in the pseudoautosomal region of the sex chromosomes and located in the introns of 228 annotated genes whose protein products were associated with functions at the cell membrane. These genes were highly expressed in the brain and particularly in genes associated with susceptibility to mental disorders, such as schizophrenia. The set of 1957 genes harboring the 2886 R*Y tracts with > or =100 uninterrupted base pairs was additionally enriched in proteins associated with phosphorylation, signal transduction, development and morphogenesis. Comparisons of the > or =250 bp R*Y tracts in the mouse and chimpanzee genomes indicated that these sequences have mutated faster than the surrounding regions and are longer in humans than in chimpanzees. These results support a role for long R*Y tracts in promoting recombination and genome diversity during evolution through destabilization of chromosomal DNA, thereby inducing repair and mutation.

Animals↗

Motif-based fold assignment.

Conventional fold recognition techniques rely mainly on the analysis of the entire sequence of a protein. We present an MBA method to improve performance of any conventional sequence-based fold assignment. The method uses sequence motifs, such as those defined in the Prosite database, and the SwissProt annotation of the fold library. When combined with a simple SDP method, the coverage of MBA is comparable to the results obtained with PSI-BLAST. However, the set of the MBA predictions is significantly different from that of PSI-BLAST, leading to a 40% increase of the coverage for the combined MBA/PSI-BLAST method. The MBA approach can be easily adopted to include the results of sequence-independent function prediction methods and alternative motif and annotation databases. The method is available through the web server localized at http://www.doe-mbi.ucla.edu/mba.

Algorithms↗

Caenorhabditis elegans has two genes encoding functional d-aspartate oxidases.

Four cDNA clones that were annotated in the database as encoding d-amino acid oxidase (DAAO) or d-aspartate oxidase (DASPO) were isolated by RT-PCR from Caenorhabditis elegans RNA. The proteins (Y69Ap, C47Ap, F18Ep, and F20Hp) encoded by the cloned cDNAs were expressed in Escherichia coli as recombinant proteins with an N-terminal His-tag. All proteins except F20Hp were recovered in the soluble fractions. The recombinant Y69Ap has functional DAAO activity, as it can deaminate neutral and basic d-amino acids, whereas the recombinants C47Ap and F18Ep have functional DASPO activities, as they can deaminate acidic d-amino acids. Additional experiments using purified recombinant proteins revealed that Y69Ap deaminates d-Arg more efficiently than d-Ala and d-Met, and that C47Ap and F18Ep show distinct kinetic properties against d-Asp, d-Glu, and N-methyl-d-Asp. This is the first time that cDNA cloning of invertebrate DAAO and DASPO genes has been reported. In addition, our study reveals for the first time that C. elegans has at least two genes encoding functional DASPOs and one gene encoding DAAO, although it had previously been thought that organisms only bear one copy each of these genes. The two C. elegans DASPOs differ in their substrate specificities and possibly also in their subcellular localization.

Amino Acid Sequence↗

GDR (Genome Database for Rosaceae): integrated web resources for Rosaceae genomics and genetics research.

BACKGROUND: Peach is being developed as a model organism for Rosaceae, an economically important family that includes fruits and ornamental plants such as apple, pear, strawberry, cherry, almond and rose. The genomics and genetics data of peach can play a significant role in the gene discovery and the genetic understanding of related species. The effective utilization of these peach resources, however, requires the development of an integrated and centralized database with associated analysis tools. DESCRIPTION: The Genome Database for Rosaceae (GDR) is a curated and integrated web-based relational database. GDR contains comprehensive data of the genetically anchored peach physical map, an annotated peach EST database, Rosaceae maps and markers and all publicly available Rosaceae sequences. Annotations of ESTs include contig assembly, putative function, simple sequence repeats, and anchored position to the peach physical map where applicable. Our integrated map viewer provides graphical interface to the genetic, transcriptome and physical mapping information. ESTs, BACs and markers can be queried by various categories and the search result sites are linked to the integrated map viewer or to the WebFPC physical map sites. In addition to browsing and querying the database, users can compare their sequences with the annotated GDR sequences via a dedicated sequence similarity server running either the BLAST or FASTA algorithm. To demonstrate the utility of the integrated and fully annotated database and analysis tools, we describe a case study where we anchored Rosaceae sequences to the peach physical and genetic map by sequence similarity. CONCLUSIONS: The GDR has been initiated to meet the major deficiency in Rosaceae genomics and genetics research, namely a centralized web database and bioinformatics tools for data storage, analysis and exchange. GDR can be accessed at http://www.genome.clemson.edu/gdr/.

Computer Graphics↗

Functional demonstration of reverse transsulfuration in the Mycobacterium tuberculosis complex reveals that methionine is the preferred sulfur source for pathogenic Mycobacteria.

Methionine can be used as the sole sulfur source by the Mycobacterium tuberculosis complex although it is not obvious from examination of the genome annotation how these bacteria utilize methionine. Given that genome annotation is a largely predictive process, key challenges are to validate these predictions and to fill in gaps for known functions for which genes have not been annotated. We have addressed these issues by functional analysis of methionine metabolism. Transport, followed by metabolism of (35)S methionine into the cysteine adduct mycothiol, demonstrated the conversion of exogenous methionine to cysteine. Mutational analysis and cloning of the Rv1079 gene showed it to encode the key enzyme required for this conversion, cystathionine gamma-lyase (CGL). Rv1079, annotated metB, was predicted to encode cystathionine gamma-synthase (CGS), but demonstration of a gamma-elimination reaction with cystathionine as well as the gamma-replacement reaction yielding cystathionine showed it encodes a bifunctional CGL/CGS enzyme. Consistent with this, a Rv1079 mutant could not incorporate sulfur from methionine into cysteine, while a cysA mutant lacking sulfate transport and a methionine auxotroph was hypersensitive to the CGL inhibitor propargylglycine. Thus, reverse transsulfuration alone, without any sulfur recycling reactions, allows M. tuberculosis to use methionine as the sole sulfur source. Intracellular cysteine was undetectable so only the CGL reaction occurs in intact mycobacteria. Cysteine desulfhydrase, an activity we showed to be separable from CGL/CGS, may have a role in removing excess cysteine and could explain the ability of M. tuberculosis to recycle sulfur from cysteine, but not methionine.

Alkynes↗

Protein-protein interactions of the hyperthermophilic archaeon Pyrococcus horikoshii OT3.

BACKGROUND: Although 2,061 proteins of Pyrococcus horikoshii OT3, a hyperthermophilic archaeon, have been predicted from the recently completed genome sequence, the majority of proteins show no similarity to those from other organisms and are thus hypothetical proteins of unknown function. Because most proteins operate as parts of complexes to regulate biological processes, we systematically analyzed protein-protein interactions in Pyrococcus using the mammalian two-hybrid system to determine the function of the hypothetical proteins. RESULTS: We examined 960 soluble proteins from Pyrococcus and selected 107 interactions based on luciferase reporter activity, which was then evaluated using a computational approach to assess the reliability of the interactions. We also analyzed the expression of the assay samples by western blot, and a few interactions by in vitro pull-down assays. We identified 11 hetero-interactions that we considered to be located at the same operon, as observed in Helicobacter pylori. We annotated and classified proteins in the selected interactions according to their orthologous proteins. Many enzyme proteins showed self-interactions, similar to those seen in other organisms. CONCLUSION: We found 13 unannotated proteins that interacted with annotated proteins; this information is useful for predicting the functions of the hypothetical Pyrococcus proteins from the annotations of their interacting partners. Among the heterogeneous interactions, proteins were more likely to interact with proteins within the same ortholog class than with proteins of different classes. The analysis described here can provide global insights into the biological features of the protein-protein interactions in P. horikoshii.

Genes, Archaeal↗

Improved prediction of protein-protein binding sites using a support vector machines approach.

MOTIVATION: Structural genomics projects are beginning to produce protein structures with unknown function, therefore, accurate, automated predictors of protein function are required if all these structures are to be properly annotated in reasonable time. Identifying the interface between two interacting proteins provides important clues to the function of a protein and can reduce the search space required by docking algorithms to predict the structures of complexes. RESULTS: We have combined a support vector machine (SVM) approach with surface patch analysis to predict protein-protein binding sites. Using a leave-one-out cross-validation procedure, we were able to successfully predict the location of the binding site on 76% of our dataset made up of proteins with both transient and obligate interfaces. With heterogeneous cross-validation, where we trained the SVM on transient complexes to predict on obligate complexes (and vice versa), we still achieved comparable success rates to the leave-one-out cross-validation suggesting that sufficient properties are shared between transient and obligate interfaces. AVAILABILITY: A web application based on the method can be found at http://www.bioinformatics.leeds.ac.uk/ppi_pred. The dataset of 180 proteins used in this study is also available via the same web site. CONTACT: westhead@bmb.leeds.ac.uk SUPPLEMENTARY INFORMATION: http://www.bioinformatics.leeds.ac.uk/ppi-pred/supp-material.

Algorithms↗

Cluster analysis of protein array results via similarity of Gene Ontology annotation.

BACKGROUND: With the advent of high-throughput proteomic experiments such as arrays of purified proteins comes the need to analyse sets of proteins as an ensemble, as opposed to the traditional one-protein-at-a-time approach. Although there are several publicly available tools that facilitate the analysis of protein sets, they do not display integrated results in an easily-interpreted image or do not allow the user to specify the proteins to be analysed. RESULTS: We developed a novel computational approach to analyse the annotation of sets of molecules. As proof of principle, we analysed two sets of proteins identified in published protein array screens. The distance between any two proteins was measured as the graph similarity between their Gene Ontology (GO) annotations. These distances were then clustered to highlight subsets of proteins sharing related GO annotation. In the first set of proteins found to bind small molecule inhibitors of rapamycin, we identified three subsets containing four or five proteins each that may help to elucidate how rapamycin affects cell growth whereas the original authors chose only one novel protein from the array results for further study. In a set of phosphoinositide-binding proteins, we identified subsets of proteins associated with different intracellular structures that were not highlighted by the analysis performed in the original publication. CONCLUSION: By determining the distances between annotations, our methodology reveals trends and enrichment of proteins of particular functions within high-throughput datasets at a higher sensitivity than perusal of end-point annotations. In an era of increasingly complex datasets, such tools will help in the formulation of new, testable hypotheses from high-throughput experimental data.

Algorithms↗