Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Unexpected catalytic site variation in phosphoprotein phosphatase homologues of cofactor-dependent phosphoglycerate mutase.

The cofactor-dependent phosphoglycerate mutase (dPGM) superfamily contains, besides mutases, a variety of phosphatases, both broadly and narrowly substrate-specific. Distant dPGM homologues, conspicuously abundant in microbial genomes, represent a challenge for functional annotation based on sequence comparison alone. Here we carry out sequence analysis and molecular modelling of two families of bacterial dPGM homologues, one the SixA phosphoprotein phosphatases, the other containing various proteins of no known molecular function. The models show how SixA proteins have adapted to phosphoprotein substrate and suggest that the second family may also encode phosphoprotein phosphatases. Unexpected variation in catalytic and substrate-binding residues is observed in the models.

Amino Acid Sequence↗

MannDB - a microbial database of automated protein sequence analyses and evidence integration for protein characterization.

BACKGROUND: MannDB was created to meet a need for rapid, comprehensive automated protein sequence analyses to support selection of proteins suitable as targets for driving the development of reagents for pathogen or protein toxin detection. Because a large number of open-source tools were needed, it was necessary to produce a software system to scale the computations for whole-proteome analysis. Thus, we built a fully automated system for executing software tools and for storage, integration, and display of automated protein sequence analysis and annotation data. DESCRIPTION: MannDB is a relational database that organizes data resulting from fully automated, high-throughput protein-sequence analyses using open-source tools. Types of analyses provided include predictions of cleavage, chemical properties, classification, features, functional assignment, post-translational modifications, motifs, antigenicity, and secondary structure. Proteomes (lists of hypothetical and known proteins) are downloaded and parsed from Genbank and then inserted into MannDB, and annotations from SwissProt are downloaded when identifiers are found in the Genbank entry or when identical sequences are identified. Currently 36 open-source tools are run against MannDB protein sequences either on local systems or by means of batch submission to external servers. In addition, BLAST against protein entries in MvirDB, our database of microbial virulence factors, is performed. A web client browser enables viewing of computational results and downloaded annotations, and a query tool enables structured and free-text search capabilities. When available, links to external databases, including MvirDB, are provided. MannDB contains whole-proteome analyses for at least one representative organism from each category of biological threat organism listed by APHIS, CDC, HHS, NIAID, USDA, USFDA, and WHO. CONCLUSION: MannDB comprises a large number of genomes and comprehensive protein sequence analyses representing organisms listed as high-priority agents on the websites of several governmental organizations concerned with bio-terrorism. MannDB provides the user with a BLAST interface for comparison of native and non-native sequences and a query tool for conveniently selecting proteins of interest. In addition, the user has access to a web-based browser that compiles comprehensive and extensive reports. Access to MannDB is freely available at http://manndb.llnl.gov/.

Algorithms↗

Predicting protein function from sequence and structural data.

When a protein's function cannot be experimentally determined, it can often be inferred from sequence similarity. Should this process fail, analysis of the protein structure can provide functional clues or confirm tentative functional assignments inferred from the sequence. Many structure-based approaches exist (e.g. fold similarity, three-dimensional templates), but as no single method can be expected to be successful in all cases, a more prudent approach involves combining multiple methods. Several automated servers that integrate evidence from multiple sources have been released this year and particular improvements have been seen with methods utilizing the Gene Ontology functional annotation schema.

Binding Sites↗

RIKEN mouse genome encyclopedia.

We have been working to establish the comprehensive mouse full-length cDNA collection and sequence database to cover as many genes as we can, named Riken mouse genome encyclopedia. Recently we are constructing higher-level annotation (Functional ANnoTation Of Mouse cDNA; FANTOM) not only with homology search based annotation but also with expression data profile, mapping information and protein-protein database. More than 1,000,000 clones prepared from 163 tissues were end-sequenced to classify into 159,789 clusters and 60,770 representative clones were fully sequenced. As a conclusion, the 60,770 sequences contained 33,409 unique. The next generation of life science is clearly based on all of the genome information and resources. Based on our cDNA clones we developed the additional system to explore gene function. We developed cDNA microarray system to print all of these cDNA clones, protein-protein interaction screening system, protein-DNA interaction screening system and so on. The integrated database of all the information is very useful not only for analysis of gene transcriptional network and for the connection of gene to phenotype to facilitate positional candidate approach. In this talk, the prospect of the application of these genome resourced should be discussed. More information is available at the web page: http://genome.gsc.riken.go.jp/.

Animals↗

Mapping the Immune cell-specific gene regulatory network in bipolar disorder: A framework from scTWMR to exploratory drug-target annotation.

BACKGROUND: Although the involvement of the immune system in the genetic susceptibility of bipolar disorder (BD) is widely acknowledged, the causal relationship between gene expression in specific immune cell subtypes and BD requires systematic elucidation. METHODS: We implemented an analytical framework integrating single-cell transcriptome-wide Mendelian randomization (scTWMR) with colocalization analysis. This approach utilized cis-expression quantitative trait loci (cis-eQTLs) derived from 14 distinct immune cell types as instrumental variables to interrogate BD genome-wide association study (GWAS) summary statistics (comprising 41,917 cases and 371,549 controls). Subsequent investigations encompassed functional enrichment analysis, protein-protein interaction (PPI) network construction, phenome-wide association study (PheWAS), and performed an exploratory drug-target annotation. RESULTS: Our analysis identified 33 gene-immune cell associations. Colocalization analysis provided robust evidence (PPH4 > 90%) for shared causal variants implicating the MAD1L1, APOM, and NFKBIL1 loci. Significantly enriched biological pathways included cell cycle regulation, circadian rhythm entrainment, and neuroinflammation. The PPI network revealed a core regulatory module centered on histone-encoding and immune-related genes. Exploratory drug-target annotation nominated compounds for further investigation for compounds targeting APOM, TMEM258, and NFKBIL1. CONCLUSION: This study systematically delineates a genetically supported regulatory network of immune cell-specific gene expression in BD, predominantly implicating CD8⁺ effector T cells, plasma cells, and B cells. The findings corroborate established pathological pathways while uncovering novel cell type-specific therapeutic targets, thereby providing a genetic framework for prioritizing candidate targets for future investigation.

Bipolar disorder↗

Globin gene server: a prototype E-mail database server featuring extensive multiple alignments and data compilation for electronic genetic analysis.

The sequence of virtually the entire cluster of beta-like globin genes has been determined from several mammals, and many regulatory regions have been analyzed by mutagenesis, functional assays, and nuclear protein binding studies. This very large amount of sequence and functional data needs to be compiled in a readily accessible and usable manner to optimize data analysis, hypothesis testing, and model building. We report a Globin Gene Server that will provide this service in a constantly updated manner when fully implemented. The Server has two principal functions. The first (currently available) provides an annotated multiple alignment of the DNA sequences throughout the gene cluster from representatives of all species analyzed. The second compiles data on functional and protein binding assays throughout the gene cluster. A prototype of this compilation using the aligned 5' flanking region of beta-globin genes from five species shows examples of (1) well-conserved regions that have demonstrated functions, including cases in which the functional data are in apparent conflict, (2) proposed functional regions that are not well conserved, and (3) conserved regions with no currently assigned function. Such an electronic genetic analysis leads to many readily testable hypotheses that were not immediately apparent without the multiple alignment and compilation. The Server is accessible via E-mail on computer networks, and printed results can be obtained by request to the authors. This prototype will be a helpful guide for developing similar tools for many genomic loci.

Animals↗

Analysis and prediction of functionally important sites in proteins.

The rapidly increasing volume of sequence and structure information available for proteins poses the daunting task of determining their functional importance. Computational methods can prove to be very useful in understanding and characterizing the biochemical and evolutionary information contained in this wealth of data, particularly at functionally important sites. Therefore, we perform a detailed survey of compositional and evolutionary constraints at the molecular and biological function level for a large set of known functionally important sites extracted from a wide range of protein families. We compare the degree of conservation across different functional categories and provide detailed statistical insight to decipher the varying evolutionary constraints at functionally important sites. The compositional and evolutionary information at functionally important sites has been compiled into a library of functional templates. We developed a module that predicts functionally important columns (FIC) of an alignment based on the detection of a significant "template match score" to a library template. Our template match score measures an alignment column's similarity to a library template and combines a term explicitly representing a column's residue composition with various evolutionary conservation scores (information content and position-specific scoring matrix-derived statistics). Our benchmarking studies show good sensitivity/specificity for the prediction of functional sites and high accuracy in attributing correct molecular function type to the predicted sites. This prediction method is based on information derived from homologous sequences and no structural information is required. Therefore, this method could be extremely useful for large-scale functional annotation.

Binding Sites↗

Website review: interPro (the integrated resource of protein domains and functional sites).

The family and motif databases, PROSITE, PRINTS, Pfam and ProDom, have been integrated into a powerful resource for protein secondary annotation. As of June 2000, InterPro had processed 384 572 proteins in SWISS-PROT and TrEMBL. Because the contributing databases have different clustering principles and scoring sensitivities, the combined assignments compliment each other for grouping protein families and delineating domains. The graphic displays of all matches above the scoring thresholds enables judgements to be made on the concordances or differences between the assignments. The website links can be used to analyse novel sequences and for queries across the proteomes of 32 organisms, including the partial human set, by domain and/or protein family. An analysis of selected HtrA/DegQ proteases demonstrates the utility of this website for detailed comparative genomics. Further information on the project can be found at the European Bioinformatics Institute at http://www.ebi.ac.uk/interpro/

Amino Acid Motifs↗

Sequence-based heuristics for faster annotation of non-coding RNA families.

MOTIVATION: Non-coding RNAs (ncRNAs) are functional RNA molecules that do not code for proteins. Covariance Models (CMs) are a useful statistical tool to find new members of an ncRNA gene family in a large genome database, using both sequence and, importantly, RNA secondary structure information. Unfortunately, CM searches are extremely slow. Previously, we created rigorous filters, which provably sacrifice none of a CM's accuracy, while making searches significantly faster for virtually all ncRNA families. However, these rigorous filters make searches slower than heuristics could be. RESULTS: In this paper we introduce profile HMM-based heuristic filters. We show that their accuracy is usually superior to heuristics based on BLAST. Moreover, we compared our heuristics with those used in tRNAscan-SE, whose heuristics incorporate a significant amount of work specific to tRNAs, where our heuristics are generic to any ncRNA. Performance was roughly comparable, so we expect that our heuristics provide a high-quality solution that--unlike family-specific solutions--can scale to hundreds of ncRNA families. AVAILABILITY: The source code is available under GNU Public License at the supplementary web site.

Algorithms↗

Preparation of a set of expression-ready clones of mammalian long cDNAs encoding large proteins by the ORF trap cloning method.

Although we have so far identified and sequenced >2000 human long cDNAs, known as KIAA cDNAs, half of them have yet to be functionally annotated. Expression-ready cDNA clones derived from these genes, where the open reading frame (ORF) of the gene of interest is placed under the control of an appropriate promoter, are critical for functional characterization of these gene products. In this study, we attempted to systematically convert original cDNA clones to expression-ready forms for native and fusion proteins. For this purpose, we developed a new method for ORF cloning based on a homologous recombination in Escherichia coli to avoid laborious manipulations and artificial introduction of mutations in ORF. Using 1589 putative full-length ORFs (from 1002 KIAA genes, 119 human known genes and 468 mouse genes) with an average size of 2.8 kb, we successfully prepared expression plasmids for 1463 native proteins and for 1343 fusion proteins by this method. The resultant expression-ready clones were examined using an in vitro transcription/translation system followed by SDS-polyacrylamide gel electrophoresis and by transient expression of GFP-fusion proteins in human embryonic kidney (HEK) 293 cells. This set of expression-ready clones of long cDNAs encoding large proteins would open a new route to experimentally analyze their functions on a proteomic scale, since unavailability of expression-ready clones for mammalian large proteins has been a major obstacle to the functional analysis of these cDNAs.

Animals↗

The Drosophila gene CG9918 codes for a pyrokinin-1 receptor.

The database from the Drosophila Genome Project contains a gene, CG9918, annotated to code for a G protein-coupled receptor. We cloned the cDNA of this gene and functionally expressed it in Chinese hamster ovary cells. We tested a library of about 25 Drosophila and other insect neuropeptides, and seven insect biogenic amines on the expressed receptor and found that it was activated by low concentrations of the Drosophila neuropeptide, pyrokinin-1 (TGPSASSGLWFGPRLamide; EC50, 5 x 10(-8) M). The receptor was also activated by other Drosophila neuropeptides, terminating with the sequence PRLamide (Hug-gamma, ecdysis-triggering-hormone-1, pyrokinin-2), but in these cases about six to eight times higher concentrations were needed. The receptor was not activated by Drosophila neuropeptides, containing a C-terminal PRIamide sequence (such as ecdysis-triggering-hormone-2), or PRVamide (such as capa-1 and -2), or other neuropeptides and biogenic amines not related to the pyrokinins. This paper is the first conclusive report that CG9918 is a Drosophila pyrokinin-1 receptor gene.

Amino Acid Sequence↗

CandidaDB: a genome database for Candida albicans pathogenomics.

CandidaDB is a database dedicated to the genome of the most prevalent systemic fungal pathogen of humans, Candida albicans. CandidaDB is based on an annotation of the Stanford Genome Technology Center C.albicans genome sequence data by the European Galar Fungail Consortium. CandidaDB Release 2.0 (June 2004) contains information pertaining to Assembly 19 of the genome of C.albicans strain SC5314. The current release contains 6244 annotated entries corresponding to 130 tRNA genes and 5917 protein-coding genes. For these, it provides tentative functional assignments along with numerous pre-run analyses that can assist the researcher in the evaluation of gene function for the purpose of specific or large-scale analysis. CandidaDB is based on GenoList, a generic relational data schema and a World Wide Web interface that has been adapted to the handling of eukaryotic genomes. The interface allows users to browse easily through genome data and retrieve information. CandidaDB also provides more elaborate tools, such as pattern searching, that are tightly connected to the overall browsing system. As the C.albicans genome is diploid and still incompletely assembled, CandidaDB provides tools to browse the genome by individual supercontigs and to examine information about allelic sequences obtained from complementary contigs. CandidaDB is accessible at http://genolist.pasteur.fr/CandidaDB.

Candida albicans↗

SNP@Domain: a web resource of single nucleotide polymorphisms (SNPs) within protein domain structures and sequences.

The single nucleotide polymorphisms (SNPs) in conserved protein regions have been thought to be strong candidates that alter protein functions. Thus, we have developed SNP@Domain, a web resource, to identify SNPs within human protein domains. We annotated SNPs from dbSNP with protein structure-based as well as sequence-based domains: (i) structure-based using SCOP and (ii) sequence-based using Pfam to avoid conflicts from two domain assignment methodologies. Users can investigate SNPs within protein domains with 2D and 3D maps. We expect this visual annotation of SNPs within protein domains will help scientists select and interpret SNPs associated with diseases. A web interface for the SNP@Domain is freely available at http://snpnavigator.net/ and from http://bioportal.net/.

Computer Graphics↗

Analysis of superfamily specific profile-profile recognition accuracy.

BACKGROUND: Annotation of sequences that share little similarity to sequences of known function remains a major obstacle in genome annotation. Some of the best methods of detecting remote relationships between protein sequences are based on matching sequence profiles. We analyse the superfamily specific performance of sequence profile-profile matching. Our benchmark consists of a set of 16 protein superfamilies that are highly diverse at the sequence level. We relate the performance to the number of sequences in the profiles, the profile diversity and the extent of structural conservation in the superfamily. RESULTS: The performance varies greatly between superfamilies with the truncated receiver operating characteristic, ROC10, varying from 0.95 down to 0.01. These large differences persist even when the profiles are trimmed to approximately the same level of diversity. CONCLUSIONS: Although the number of sequences in the profile (profile width) and degree of sequence variation within positions in the profile (profile diversity) contribute to accurate detection there are other superfamily specific factors.

Benchmarking↗

Identification and functional analyses of two cDNAs that encode fatty acid 9-/13-hydroperoxide lyase (CYP74C) in rice.

Fatty acid hydroperoxide lyase (HPL), a member of cytochrome P450 (CYP74), produces aldehydes and oxo-acids involved in plant defensive reactions. In monocots, HPL that cleaves 13-hydroperoxides of fatty acids has been reported, but HPL that cleaves 9-hydroperoxides is still unknown. To find this type of HPL, in silico screening of candidate cDNA clones and subsequent functional analyses of recombinant proteins were performed. We found that AK105964 and AK107161 (Genbank accession numbers), cDNAs previously annotated as allene oxide synthase (AOS) in rice, are distinctively grouped from AOS and 13-HPL. Recombinant proteins of these cDNAs produced in Escherichia. coli cleaved both 9- and 13-hydroperoxide of linoleic and linolenic into aldehydes, while having only a trace level of AOS activity and no divinyl ether synthase activity. Hence we designated AK105964 and AK107161 OsHPL1 and OsHPL2 respectively. They are the first CYP74C family cDNAs to be found in monocots.

Amino Acid Sequence↗

Combining bioinformatics and phylogenetics to identify large sets of single-copy orthologous genes (COSII) for comparative, evolutionary and systematic studies: a test case in the euasterid plant clade.

We report herein the application of a set of algorithms to identify a large number (2869) of single-copy orthologs (COSII), which are shared by most, if not all, euasterid plant species as well as the model species Arabidopsis. Alignments of the orthologous sequences across multiple species enabled the design of "universal PCR primers," which can be used to amplify the corresponding orthologs from a broad range of taxa, including those lacking any sequence databases. Functional annotation revealed that these conserved, single-copy orthologs encode a higher-than-expected frequency of proteins transported and utilized in organelles and a paucity of proteins associated with cell walls, protein kinases, transcription factors, and signal transduction. The enabling power of this new ortholog resource was demonstrated in phylogenetic studies, as well as in comparative mapping across the plant families tomato (family Solanaceae) and coffee (family Rubiaceae). The combined results of these studies provide compelling evidence that (1) the ancestral species that gave rise to the core euasterid families Solanaceae and Rubiaceae had a basic chromosome number of x=11 or 12.2) No whole-genome duplication event (i.e., polyploidization) occurred immediately prior to or after the radiation of either Solanaceae or Rubiaceae as has been recently suggested.

Algorithms↗

[MGAP-A microbe genome annotation platform].

A Microbe Genome Annotation Platform (MGAP) was developed and applied to the cynobacterium PCC7002 genome annotation. Various bioinformatics software tools from sequence analysis to gene identification and function prediction were implemented in MGAP. Protein sequence databases SWISSPROT and PDBseq, protein information resource InterPro and COG were also integrated in the platform. The web interface of MGAP has the functionality to display a circular map of gene distribution and GC contents throughout the genome. Detailed information such as the DNA and protein sequence, the location of genes on chromosomes can be viewed by clicking the corresponding object within the map. MGAP is based on a PC/Linux system affordable for small biological laboratories and has the advantage of using free software tools including MySQL, Apache and Perl.

Cyanobacteria↗