Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

Evolution of protein function, from a structural perspective.

The recent growth in structural data, and ensuing analyses, have revealed the structural and functional versatility of protein families. With respect to enzymes, local active-site mutations, variations in surface loops and recruitment of additional domains accommodate the diverse substrate specificities and catalytic activities observed within several superfamilies. Conversely, some functions have more than one structural solution, having evolved independently several times during evolution. Combined with the existence of multi-functional genes, which have arisen by gene recruitment, these phenomena must be considered in the process of genome annotation.

Animals↗

FISH--family identification of sequence homologues using structure anchored hidden Markov models.

The FISH server is highly accurate in identifying the family membership of domains in a query protein sequence, even in the case of very low sequence identities to known homologues. A performance test using SCOP sequences and an E-value cut-off of 0.1 showed that 99.3% of the top hits are to the correct family saHMM. Matches to a query sequence provide the user not only with an annotation of the identified domains and hence a hint to their function, but also with probable 2D and 3D structures, as well as with pairwise and multiple sequence alignments to homologues with low sequence identity. In addition, the FISH server allows users to upload and search their own protein sequence collection or to quarry public protein sequence data bases with individual saHMMs. The FISH server can be accessed at http://babel.ucmp.umu.se/fish/.

Databases, Protein↗

Discussion on the mechanism of Lingguizhugan Decoction in treating hypertension based on network pharmacology and molecular simulation technology.

To explore the mechanism of Lingguizhugan Decoction in treating hypertension based on network pharmacology and molecular simulation. The active ingredients and potential targets were screened by the Systematic Pharmacological Analysis Platform of Traditional Chinese Medicine (TCMSP). Hypertension-related targets were obtained from OMIM and GeneCards databases. Common targets between drug and hypertension were screened in the Venny platform. A protein-protein interaction (PPI) network was constructed in the STRING database using intersection targets. Key targets in PPI network were analyzed by Cytoscape. R language program was used for Gene Ontology (GO) functional annotation and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis. Finally, the binding abilities of the main active ingredients to critical targets were verified by molecular simulation. Naringenin, quercetin, kaempferol, and β-sitosterol in Lingguizhugan Decoction, and potential targets such as STAT3, AKT1, TNF, IL6, JUN, PTGS2, MMP9, CASP3, TP53, and MAPK3, were screened out. KEGG Enrichment analysis revealed that the common targets of Lingguizhugan Decoction and hypertension are mainly involved in the lipid and atherosclerosis signaling pathway, AGE-RAGE signaling pathway in diabetic complications, fluid shear stress and atherosclerosis, and IL17 signaling pathway. The molecular simulation results showed that naringenin-MAPK3, quercetin-MMP9, quercetin-PTGS2, and quercetin-TP53 were the top four in the docking scores. Naringenin-MAPK3 and quercetin-MMP9 were stable, with binding free energies of -27.97 ± 1.41 kcal/mol and -21.15 ± 3.17 kcal/mol, respectively. The possible mechanism of Lingguizhugan Decoction in treating hypertension is characterized of multi-component, multi-target, and multi-pathway.Communicated by Ramaswamy H. Sarma.

Network Pharmacology↗

GeneViTo: visualizing gene-product functional and structural features in genomic datasets.

BACKGROUND: The availability of increasing amounts of sequence data from completely sequenced genomes boosts the development of new computational methods for automated genome annotation and comparative genomics. Therefore, there is a need for tools that facilitate the visualization of raw data and results produced by bioinformatics analysis, providing new means for interactive genome exploration. Visual inspection can be used as a basis to assess the quality of various analysis algorithms and to aid in-depth genomic studies. RESULTS: GeneViTo is a JAVA-based computer application that serves as a workbench for genome-wide analysis through visual interaction. The application deals with various experimental information concerning both DNA and protein sequences (derived from public sequence databases or proprietary data sources) and meta-data obtained by various prediction algorithms, classification schemes or user-defined features. Interaction with a Graphical User Interface (GUI) allows easy extraction of genomic and proteomic data referring to the sequence itself, sequence features, or general structural and functional features. Emphasis is laid on the potential comparison between annotation and prediction data in order to offer a supplement to the provided information, especially in cases of "poor" annotation, or an evaluation of available predictions. Moreover, desired information can be output in high quality JPEG image files for further elaboration and scientific use. A compilation of properly formatted GeneViTo input data for demonstration is available to interested readers for two completely sequenced prokaryotes, Chlamydia trachomatis and Methanococcus jannaschii. CONCLUSIONS: GeneViTo offers an inspectional view of genomic functional elements, concerning data stemming both from database annotation and analysis tools for an overall analysis of existing genomes. The application is compatible with Linux or Windows ME-2000-XP operating systems, provided that the appropriate Java Runtime Environment is already installed in the system.

Bacterial Proton-Translocating ATPases↗

PRODISTIN Web Site: a tool for the functional classification of proteins from interaction networks.

UNLABELLED: The PRODISTIN Web Site is a web service allowing users to functionally classify genes/proteins from any type of interaction network. The resulting computation provides a classification tree in which (1) genes/proteins are clustered according to the identity of their interaction partners and (2) functional classes are delineated in the tree using the Biological Process Gene Ontology annotations. AVAILABILITY: The PRODISTIN Web Site is freely accessible at http://gin.univ-mrs.fr/webdistin

Algorithms↗

Ureaplasma urealyticum: an opportunity for combinatorial genomics.

Examination of genomic or enzymatic activity data alone neither provides a complete picture of metabolic function or potential nor confidently reveals sites amenable to inhibition. Furthermore, in some cases, gene annotation and in aqua assays disagree by describing gene annotation without enzyme activity and enzyme activity without homologous annotation. The newly sequenced genome of Ureaplasma urealyticum (parvum) is another prokaryote example of the class Mollicutes where such confounding differences are observed. The little-considered role of some proteins as multifunctional enzymes - substitutes for 'missing' genes - could both partially explain the apparent anomalies and relate to any inaccurate deductions of inhibitor function. A combinatorial analysis involving available evidence of genomic sequence, transcription, translational phenomena, structure and enzymatic activity gives the best picture of the organism's vital metabolic alternatives.

Bacterial Proteins↗

Annotation of human chromosome 21 for relevance to Down syndrome: gene structure and expression analysis.

Down syndrome is caused by an extra copy of human chromosome 21 and the resultant dosage-related overexpression of genes contained within it. To efficiently direct experiments to determine specific gene-phenotype correlations, it is necessary to identify all genes within 21q and assess their functional associations and expression patterns. Analysis of the complete finished sequence of 21q resulted in annotated 225 genes and gene models, most of which were incomplete and/or had little or no experimental verification. Here we correct or complete the genomic structures of 16 genes, 4 of which were not reported in the annotation of the complete sequence. Our data include the identification of six genes encoding short or ambiguous open reading frames; the identification of three cases in which alternative splicing produces two structurally unrelated protein sequences; and the identification of six genes encoding proteins with functional motifs, two genes with unusually low similarity to their orthologous mouse proteins, and four genes with significant conservation in Drosophila melanogaster. We further demonstrate that an additional nine gene models represent bona fide transcripts and develop expression patterns for these genes plus nine additional novel chromosome 21 genes and four paralogous genes mapping elsewhere in the human genome. These data have implications for generating complete transcript maps of chromosome 21 and for the entire human genome, and for defining expression abnormalities in Down syndrome and mouse models.

Animals↗

Toward consistent assignment of structural domains in proteins.

The assignment of protein domains from three-dimensional structure is critically important in understanding protein evolution and function, yet little quality assurance has been performed. Here, the differences in the assignment of structural domains are evaluated using six common assignment methods. Three human expert methods (AUTHORS (authors' annotation), CATH and SCOP) and three fully automated methods (DALI, DomainParser and PDP) are investigated by analysis of individual methods against the author's assignment as well as analysis based on the consensus among groups of methods (only expert, only automatic, combined). The results demonstrate that caution is recommended in using current domain assignments, and indicates where additional work is needed. Specifically, the major factors responsible for conflicting domain assignments between methods, both experts and automatic, are: (1) the definition of very small domains; (2) splitting secondary structures between domains; (3) the size and number of discontinuous domains; (4) closely packed or convoluted domain-domain interfaces; (5) structures with large and complex architectures; and (6) the level of significance placed upon structural, functional and evolutionary concepts in considering structural domain definitions. A web-based resource that focuses on the results of benchmarking and the analysis of domain assignments is available at

Algorithms↗

An integrated analysis of the genome of the hyperthermophilic archaeon Pyrococcus abyssi.

The hyperthermophilic euryarchaeon Pyrococcus abyssi and the related species Pyrococcus furiosus and Pyrococcus horikoshii, whose genomes have been completely sequenced, are presently used as model organisms in different laboratories to study archaeal DNA replication and gene expression and to develop genetic tools for hyperthermophiles. We have performed an extensive re-annotation of the genome of P. abyssi to obtain an integrated view of its phylogeny, molecular biology and physiology. Many new functions are predicted for both informational and operational proteins. Moreover, several candidate genes have been identified that might encode missing links in key metabolic pathways, some of which have unique biochemical features. The great majority of Pyrococcus proteins are typical archaeal proteins and their phylogenetic pattern agrees with its position near the root of the archaeal tree. However, proteins probably from bacterial origin, including some from mesophilic bacteria, are also present in the P. abyssi genome.

Adaptation, Physiological↗

Phylogenetic web profiler.

SUMMARY: Phylogenetic Web Profiler (PWP) is a web-based service designed to perform phylogenetic profiling of proteins against genomes. The current version offers a selection of 63 completed genomes and available plasmids as annotated in the PEDANT genome database. Unlike currently available applications, this tool offers several choices of ortholog prediction parameters including E-value cutoff, percent length difference tolerance, and annotation similarity. Additional features include tight integration with the PEDANT database and tools to analyze properties of predicted proteins. PWP should prove very useful for the analysis of functional-linkage between proteins.

Amino Acid Sequence↗

MITOP, the mitochondrial proteome database: 2000 update.

MITOP (http://www.mips.biochem.mpg.de/proj/medgen/mitop/) is a comprehensive database for genetic and functional information on both nuclear- and mitochondrial-encoded proteins and their genes. The five species files--Saccharomyces cerevisiae, Mus musculus, Caenorhabditis elegans, Neurospora crassa and Homo sapiens--include annotated data derived from a variety of online resources and the literature. A wide spectrum of search facilities is given in the overlapping sections 'Gene catalogues', 'Protein catalogues', 'Homologies', 'Pathways and metabolism' and 'Human disease catalogue' including extensive references and hyperlinks to other databases. Central features are the results of various homology searches, which should facilitate the investigations into interspecies relationships. Precomputed FASTA searches using all the MITOP yeast protein entries and a list of the best human EST hits with graphical cluster alignments related to the yeast reference sequence are presented. The orthologue tables with cross-listings to all the protein entries for each species in MITOP have been expanded by adding the genomes of Rickettsia prowazeckii and Escherichia coli. To find new mitochondrial proteins the complete yeast genome has been analyzed using the MITOPROT program which identifies mitochondrial targeting sequences. The 'Human disease catalogue' contains tables with a total of 110 human diseases related to mitochondrial protein abnormalities, sorted by clinical criteria and age of onset. MITOP should contribute to the systematic genetic characterization of the mitochondrial proteome in relation to human disease.

Animals↗

Serial analysis of gene expression (SAGE) in the rat limbal and central corneal epithelium.

PURPOSE: To identify genes preferentially expressed in the stem-cell-rich limbal epithelium of the rat cornea. METHODS: The limbal and central corneal epithelial cells of 6-week-old rats were isolated by microdissection. Serial analysis of gene expression (SAGE) libraries were constructed and analyzed, and in situ hybridization, reverse transcription-polymerase chain reaction (RT-PCR) and cDNA cloning were conducted by conventional procedures. RESULTS: The rat limbal and central corneal epithelial SAGE libraries consisted of 41,894 and 40,691 tags, respectively. After annotation, this was reduced to 759 transcripts specific for the limbal library and 844 transcripts specific for the central corneal library; 2292 transcripts overlapped. Transcripts encoding proteins with metabolic functions comprised the major functional category in both libraries. In situ hybridization and/or RT-PCR results of 12 of the most abundant, highly enriched transcripts in the limbal epithelium were in general agreement with the SAGE data and showed that these proteins are also expressed in the conjunctival epithelium. Interesting limbal-enriched transcripts encode WDNM1-like protein (similar to WDNM1/Expi, a putative secreted proteinase and inhibitor of metastasis), mesothelin (a cancer marker), marapsin (a trypsin-like serine protease that may control cell growth and migration), K4 and K15 (both cytokeratins), and membrane-spanning four-domain subfamily A member 8B. WDNM1-like protein was cloned and confirmed as a member of the four-disulfide core family. CONCLUSIONS: The SAGE results extend the database of genes expressed in the rodent cornea and suggest an association between several genes preferentially expressed in the limbal epithelium with cellular proliferation and migration.

Animals↗

Isolation, sequence analysis, and expression studies of florally expressed cDNAs in Arabidopsis.

Molecular genetics has identified dozens of genes that regulate flower development in Arabidopsis. However, the complexity of flower development suggests that many other genes are yet to be uncovered. To identify floral genes that are expressed at low levels in the flower, we have sequenced 1587 cDNA fragments from a subtractive floral cDNA library. A total of 1222 unique genes represented by these ESTs are distributed on all five chromosomes with similar frequencies as all predicted genes in the genome. Among these, 17 genes were shown to be expressed anywhere for the first time because they were not found in previous EST and full-length cDNA datasets. Furthermore, 724 of the genes revealed by this library were not definitively shown to be expressed in the flower by previous floral EST datasets. In addition, 49 transcriptional regulators, 31 protein kinases, 12 zinc-finger proteins and other signaling proteins were found to be present in floral buds. Moreover, the EST sequences likely extended the transcribed regions of 26 previously annotated genes, and may have uncovered several previously unrecognized genes. To obtain additional clues about possible gene function, we hybridized cDNA microarray with probes derived from wild-type Arabidopsis rosette leaves and floral buds. We estimated that over 50% of genes were expressed at levels lower than 1/30 of the highest detectable signal intensity, indicating that many floral genes are expressed at low levels. Furthermore, 97 genes were found to be expressed at a higher level in the flower than the leaf by the Significance Analysis of Microarray (SAM) method with a 1.0% false discovery rate (FDR). Further RT-PCR analyses of selected genes support the microarray results. We suggest that the genes encoding putative regulatory proteins and at least some proteins with currently unknown functions might play important roles during flowering.

Arabidopsis↗

Homology-based gene prediction using neural nets.

We have developed and implemented a method for computational gene identification called GIN (gene identification using neural nets and homology information) that has been particularly designed to avoid false positive predictions. It thus predicts 55% of all genes tested correctly, has a specificity of 99%, but also has an overall accuracy of 92% on a benchmark set of 570 vertebrate genes constructed by Burset and Guigo. The method combines homology searches in protein and expressed sequence tag databases with several neural networks designed to recognize start codons, Poly(A) signals, stop codons, and splice sites. Predicted exons are assembled into genes using a homology-based scoring function. GIN is able to recognize multiple genes within genomic DNA as demonstrated by the identification of a globin gene (gamma-globin-1(G)) that has not been annotated as a coding region in the widely used the test set of Burset and Guigo. Furthermore, GIN identifies more than 107 other protein hits in noncoding regions and classifies them into possible pseudogenes or splice variants.

Animals↗

Phosphoproteome analysis of HeLa cells using stable isotope labeling with amino acids in cell culture (SILAC).

Identification of phosphorylated proteins remains a difficult task despite technological advances in protein purification methods and mass spectrometry. Here, we report identification of tyrosine-phosphorylated proteins by coupling stable isotope labeling with amino acids in cell culture (SILAC) to mass spectrometry. We labeled HeLa cells with stable isotopes of tyrosine, or, a combination of arginine and lysine to identify tyrosine phosphorylated proteins. This allowed identification of 118 proteins, of which only 45 proteins were previously described as tyrosine-phosphorylated proteins. A total of 42 in vivo tyrosine phosphorylation sites were mapped, including 34 novel ones. We validated the phosphorylation status of a subset of novel proteins including cytoskeleton associated protein 1, breast cancer anti-estrogen resistance 3, chromosome 3 open reading frame 6, WW binding protein 2, Nice-4 and RNA binding motif protein 4. Our strategy can be used to identify potential kinase substrates without prior knowledge of the signaling pathways and can also be applied to profiling to specific kinases in cells. Because of its sensitivity and general applicability, our approach will be useful for investigating signaling pathways in a global fashion and for using phosphoproteomics for functional annotation of genomes.

Amino Acid Sequence↗

The histone database: a comprehensive WWW resource for histones and histone fold-containing proteins.

The Histone Database (HDB) is an annotated and searchable collection of all full-length sequences and structures of histone and non-histone proteins containing the histone fold motif. These sequences are both eukaryotic and archaeal in origin. Several new histone fold-containing proteins have been identified, including Spt7p, and a few false positives have been removed from the earlier version of HDB. Database contents include compilations of post-translational modifications for each of the core and linker histones, as well as genomic information in the form of map loci for the human histone gene complement, with the genetic loci linked to Online Mendelian Inheritance in Man (OMIM). Conflicts between similar sequence entries from a number of source databases are also documented. Newly added to the HDB are multiple sequence alignments in which predicted functions of histone fold amino acid residues are annotated. The database is freely accessible through the WWW at http://genome.nhgri.nih.gov/histones/

Amino Acid Sequence↗

Reevaluation of human cytomegalovirus coding potential.

The Bio-Dictionary-based Gene Finder was used to reassess the coding potential of the AD169 laboratory strain of human cytomegalovirus and sequences in the Toledo strain that are missing in the laboratory strain of the virus. The gene-finder algorithm assesses the potential of an ORF to encode a protein based on matches to a database of amino acid patterns derived from a large collection of proteins. The algorithm was used to score all human cytomegalovirus ORFs with the potential to encode polypeptides >/=50 aa in length. As a further test for functionality, the genomes of the chimpanzee, rhesus, and murine cytomegaloviruses were searched for orthologues of the predicted human cytomegalovirus ORFs. The analysis indicates that 37 previously annotated ORFs ought to be discarded, and at least nine previously unrecognized ORFs with relatively strong coding potential should be added. Thus, the human cytomegalovirus genome appears to contain approximately 192 unique ORFs with the potential to encode a protein. Support for several of the predictions of our in silico analysis was obtained by sequencing several domains within a clinical isolate of human cytomegalovirus.

Algorithms↗

High-throughput protein localization in Arabidopsis using Agrobacterium-mediated transient expression of GFP-ORF fusions.

We describe a streamlined and systematic method for cloning green fluorescent protein (GFP)-open reading frame (ORF) fusions and assessing their subcellular localization in Arabidopsis thaliana cells. The sequencing of the Arabidopsis genome has made it feasible to undertake genome-based approaches to determine the function of each protein and define its subcellular localization. This is an essential step towards full functional analysis. The approach described here allows the economical handling of hundreds of expressed plant proteins in a timely fashion. We have integrated recombinational cloning of full-length trimmed ORF clones (available from the SSP consortium) with high-efficiency transient transformation of Arabidopsis cell cultures by a hypervirulent strain of Agrobacterium. To demonstrate its utility, we have used a selection of trimmed ORFs, representing a variety of key cellular processes and have defined the localization patterns of 155 fusion proteins. These patterns have been classified into five main categories, including cytoplasmic, nuclear, nucleolar, organellar and endomembrane compartments. Several genes annotated in GenBank as unknown have been ascribed a protein localization pattern. We also demonstrate the application of flow cytometry to estimate the transformation efficiency and cell cycle phase of the GFP-positive cells. This approach can be extended to functional studies, including the precise cellular localization and the prediction of the role of unknown proteins, the confirmation of bioinformatic predictions and proteomic experiments, such as the determination of protein interactions in vivo, and therefore has numerous applications in the post-genomic analysis of protein function.

Agrobacterium tumefaciens↗