Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Update of NUREBASE: nuclear hormone receptor functional genomics.

Nuclear hormone receptors are an abundant class of ligand-activated transcriptional regulators, found in varying numbers in all animals. Based on our experience of managing the official nomenclature of nuclear receptors, we have developed NUREBASE, a database containing protein and DNA sequences, reviewed protein alignments and phylogenies, taxonomy and annotations for all nuclear receptors. New developments in NUREBASE include explicit declaration of alternative transcripts of each gene, and expression data for human and mouse nuclear receptors. The core of NUREBASE is reviewed, and it is completed by NUREBASE_DAILY, automatically updated every 24 h. All information on accessing and installing NUREBASE may be found at http://www. ens-lyon.fr/LBMC/laudet/nurebase/nurebase.html.

Alternative Splicing↗

Analysis of the mouse transcriptome based on functional annotation of 60,770 full-length cDNAs.

Only a small proportion of the mouse genome is transcribed into mature messenger RNA transcripts. There is an international collaborative effort to identify all full-length mRNA transcripts from the mouse, and to ensure that each is represented in a physical collection of clones. Here we report the manual annotation of 60,770 full-length mouse complementary DNA sequences. These are clustered into 33,409 'transcriptional units', contributing 90.1% of a newly established mouse transcriptome database. Of these transcriptional units, 4,258 are new protein-coding and 11,665 are new non-coding messages, indicating that non-coding RNA is a major component of the transcriptome. 41% of all transcriptional units showed evidence of alternative splicing. In protein-coding transcripts, 79% of splice variations altered the protein product. Whole-transcriptome analyses resulted in the identification of 2,431 sense-antisense pairs. The present work, completely supported by physical clones, provides the most comprehensive survey of a mammalian transcriptome so far, and is a valuable resource for functional genomics.

Alternative Splicing↗

KOBAS server: a web-based platform for automated annotation and pathway identification.

There is an increasing need to automatically annotate a set of genes or proteins (from genome sequencing, DNA microarray analysis or protein 2D gel experiments) using controlled vocabularies and identify the pathways involved, especially the statistically enriched pathways. We have previously demonstrated the KEGG Orthology (KO) as an effective alternative controlled vocabulary and developed a standalone KO-Based Annotation System (KOBAS). Here we report a KOBAS server with a friendly web-based user interface and enhanced functionalities. The server can support input by nucleotide or amino acid sequences or by sequence identifiers in popular databases and can annotate the input with KO terms and KEGG pathways by BLAST sequence similarity or directly ID mapping to genes with known annotations. The server can then identify both frequent and statistically enriched pathways, offering the choices of four statistical tests and the option of multiple testing correction. The server also has a 'User Space' in which frequent users may store and manage their data and results online. We demonstrate the usability of the server by finding statistically enriched pathways in a set of upregulated genes in Alzheimer's Disease (AD) hippocampal cornu ammonis 1 (CA1). KOBAS server can be accessed at http://kobas.cbi.pku.edu.cn.

Alzheimer Disease↗

Mining sequence annotation databanks for association patterns.

MOTIVATION: Millions of protein sequences currently being deposited to sequence databanks will never be annotated manually. Similarity-based annotation generated by automatic software pipelines unavoidably contains spurious assignments due to the imperfection of bioinformatics methods. Examples of such annotation errors include over- and underpredictions caused by the use of fixed recognition thresholds and incorrect annotations caused by transitivity based information transfer to unrelated proteins or transfer of errors already accumulated in databases. One of the most difficult and timely challenges in bioinformatics is the development of intelligent systems aimed at improving the quality of automatically generated annotation. A possible approach to this problem is to detect anomalies in annotation items based on association rule mining. RESULTS: We present the first large-scale analysis of association rules derived from two large protein annotation databases-Swiss-Prot and PEDANT-and reveal novel, previously unknown tendencies of rule strength distributions. Most of the rules are either very strong or very weak, with rules in the medium strength range being relatively infrequent. Based on dynamics of error correction in subsequent Swiss-Prot releases and on our own manual analysis we demonstrate that exceptions from strong rules are, indeed, significantly enriched in annotation errors and can be used to automatically flag them. We identify different strength dependencies of rules derived from different fields in Swiss-Prot. A compositional breakdown of association rules generated from PEDANT in terms of their constituent items indicates that most of the errors that can be corrected are related to gene functional roles. Swiss-Prot errors are usually caused by under-annotation owing to its conservative approach, whereas automatically generated PEDANT annotation suffers from over-annotation. AVAILABILITY: All data generated in this study are available for download and browsing at http://pedant.gsf.de/ARIA/index.htm.

Conserved Sequence↗

CDD: a curated Entrez database of conserved domain alignments.

The Conserved Domain Database (CDD) is now indexed as a separate database within the Entrez system and linked to other Entrez databases such as MEDLINE(R). This allows users to search for domain types by name, for example, or to view the domain architecture of any protein in Entrez's sequence database. CDD can be accessed on the WorldWideWeb at http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=cdd. Users may also employ the CD-Search service to identify conserved domains in new sequences, at http://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi. CD-Search results, and pre-computed links from Entrez's protein database, are calculated using the RPS-BLAST algorithm and Position Specific Score Matrices (PSSMs) derived from CDD alignments. CD-Searches are also run by default for protein-protein queries submitted to BLAST(R) at http://www.ncbi.nlm.nih.gov/BLAST. CDD mirrors the publicly available domain alignment collections SMART and PFAM, and now also contains alignment models curated at NCBI. Structure information is used to identify the core substructure likely to be present in all family members, and to produce sequence alignments consistent with structure conservation. This alignment model allows NCBI curators to annotate 'columns' corresponding to functional sites conserved among family members.

Amino Acid Sequence↗

Improved techniques for the identification of pseudogenes.

MOTIVATION: Pseudogenes are the remnants of genomic sequences of genes which are no longer functional. They are frequent in most eukaryotic genomes, and an important resource for comparative genomics. However, pseudogenes are often mis-annotated as functional genes in sequence databases. Current methods for identifying pseudogenes include methods which rely on the presence of stop codons and frameshifts, as well as methods based on the ratio of non-silent to silent nucleotide substitution rates (dN/dS). A recent survey concluded that 50% of human pseudogenes have no detectable truncation in their pseudo-coding regions, indicating that the former methods lack sensitivity. The latter methods have been used to find sets of genes enriched for pseudogenes, but are not specific enough to accurately separate pseudogenes from expressed genes. RESULTS: We introduce a program called pseudogene inference from loss of constraint (PSILC) which incorporates novel methods for separating pseudogenes from functional genes. The methods calculate the log-odds score that evolution along the final branch of the gene tree to the query gene has been according to the following constraints: A neutral nucleotide model compared to a Pfam domain encoding model (PSILC(nuc/dom)); A protein coding model compared to a Pfam domain encoding model (PSILC(prot/dom)). Using the manual annotation of human chromosome 6, we show that both these methods result in a more accurate classification of pseudogenes than dN/dS when a Pfam domain alignment is available. AVAILABILITY: PSILC is available from http://www.sanger.ac.uk/Software/PSILC

Algorithms↗

What does it mean to identify a protein in proteomics?

The annotation of the human genome indicates the surprisingly low number of approximately 40,000 genes. However, the estimated number of proteins encoded by these genes is two to three orders of magnitude higher. The ability to unambiguously identify the proteins is a prerequisite for their functional investigation. As proteins derived from the same gene can be largely identical, and might differ only in small but functionally relevant details, protein identification tools must not only identify a large number of proteins but also be able to differentiate between close relatives. This information can be generated by mass spectrometry, an approach that identifies proteins by partial analysis of their digestion-derived peptides. Information gleaned from databases fills in the missing sequence information. Because both sequence databases and experimental data are limited, a certain ambiguity often remains concerning which sequence variant(s) and modification(s) are present. As the common denominator of all the isoforms is a gene, in our opinion, it would be more accurate to state that a product of this particular gene rather than a certain protein has been identified by mass spectrometry.

Gene Expression Profiling↗

The fragment transformation method to detect the protein structural motifs.

To identify functional structural motifs from protein structures of unknown function becomes increasingly important in recent years due to the progress of the structural genomics initiatives. Although certain structural patterns such as the Asp-His-Ser catalytic triad are easy to detect because of their conserved residues and stringently constrained geometry, it is usually more challenging to detect a general structural motifs like, for example, the betabetaalpha-metal binding motif, which has a much more variable conformation and sequence. At present, the identification of these motifs usually relies on manual procedures based on different structure and sequence analysis tools. In this study, we develop a structural alignment algorithm combining both structural and sequence information to identify the local structure motifs. We applied our method to the following examples: the betabetaalpha-metal binding motif and the treble clef motif. The betabetaalpha-metal binding motif plays an important role in nonspecific DNA interactions and cleavage in host defense and apoptosis. The treble clef motif is a zinc-binding motif adaptable to diverse functions such as the binding of nucleic acid and hydrolysis of phosphodiester bonds. Our results are encouraging, indicating that we can effectively identify these structural motifs in an automatic fashion. Our method may provide a useful means for automatic functional annotation through detecting structural motifs associated with particular functions.

Amino Acid Motifs↗

Pangenome-wide identification and expression analysis of the chalcone synthase (CHS) gene family in five yellowhorn spp.

Chalcone synthase (CHS) is a pivotal enzyme in flavonoid biosynthesis involved in plant development, defense, and secondary metabolism. Xanthoceras sorbifolium (yellowhorn) is a medicinal and ornamental species with high resistance to environmental stresses, but its CHS gene family remains uncharacterized. We performed a pangenome-wide identification of CHS genes across five yellowhorn genomes (Xzs4, Xwf8, Xjg, Xg11, and Xzg2). Across the five yellowhorn genomes, 27 CHS genes were identified and classified into four core pangenes, present in all five genomes, and two dispensable genes, present only in a subset of genomes. Phylogenetic analysis grouped these genes into three major clades, and chromosomal mapping and duplication analyses identified four tandemly duplicated gene pairs under purifying selection. The analyses of conserved structural features, including protein motifs and exon-intron organization, together with promoter cis-regulatory elements and gene ontology annotation, further indicated the potential involvement of CHS genes in flavonoid biosynthesis and stress-responsive mechanisms. Gene expression profiling identified significant upregulation of Xg11_CHS1 and Xg11_CHS3 under cold and drought stress, with tissue-specific expression patterns. These findings provide valuable insights into the evolution, functional diversification, and stress-responsive roles of the CHS gene family, identifying candidate genes for future studies targeting stress tolerance and flavonoid biosynthesis in yellowhorn.

Acyltransferases↗

BioIE: extracting informative sentences from the biomedical literature.

SUMMARY: BioIE is a rule-based system that extracts informative sentences relating to protein families, their structures, functions and diseases from the biomedical literaturE. Based on manual definition of templates and rules, it aims at precise sentence extraction rather than wide recall. After uploading source text or retrieving abstracts from MEDLINE, users can extract sentences based on predefined or user-defined template categories. BioIE also provides a brief insight into the syntactic and semantic context of the source-text by looking at word, N-gram and MeSH-term distributions. Important Applications of BioIE are in, for example, annotation of microarray data and of protein databases. AVAILABILITY: http://umber.sbs.man.ac.uk/dbbrowser/bioie/

Database Management Systems↗

Whole genomes: the foundation of new biology and medicine.

Our genomic DNA sequence provides a unique glimpse of the provenance and evolution of our species, the migration of peoples, and the causation of disease. Understanding the genome may help resolve previously unanswerable questions, including perhaps which human characteristics are innate or acquired. Such an understanding will make it possible to study how genomic DNA sequence varies among populations and among individuals, including the role of such variation in the pathogenesis of important illnesses and responses to pharmaceuticals. The study of the genome and the associated proteomics of free-living organisms will eventually make it possible to localize and annotate every human gene, as well as the regulatory elements that control the timing, organ-site specificity, extent of gene expression, protein levels, and post-translational modifications. For any given physiological process, we will have a new paradigm for addressing its evolution, development, function, and mechanism.

Animals↗

Identification of the first invertebrate interleukin JAK/STAT receptor, the Drosophila gene domeless.

The JAK/STAT signaling pathway plays important roles in vertebrate development and the regulation of complex cellular processes. Components of the pathway are conserved in Dictyostelium, Caenorhabditis, and Drosophila, yet the complete sequencing and annotation of the D. melanogaster and C. elegans genomes has failed to identify a receptor, raising the possibility that an alternative type of receptor exists for the invertebrate JAK/STAT pathway. Here we show that domeless (dome) codes for a transmembrane protein required for all JAK/STAT functions in the Drosophila embryo. This includes its known requirement for embryonic segmentation and a newly discovered function in trachea specification. The DOME protein has a similar extracellular structure to the vertebrate cytokine class I receptors, although its sequence has greatly diverged. Like many interleukin receptors, DOME has a cytokine binding homology module (CBM) and three extracellular fibronectin-type-III domains (FnIII). Despite its low degree of overall similarity, key amino acids required for signaling in the vertebrate cytokine class I receptors [3] are conserved in the CBM region. DOME is a signal-transducing receptor with most similarities to the IL-6 receptor family, but it also has characteristics found in the IL-3 receptor family. This suggests that the vertebrate families evolved from a single ancestral receptor that also gave rise to dome.

Amino Acid Motifs↗

PINdb: a database of nuclear protein complexes from human and yeast.

SUMMARY: Proteins Interacting in the Nucleus database (PINdb) is a database of protein complexes purified from the nucleus of human and yeast cells. It is compiled from the published literature and existing databases. Currently, PINdb contains mostly protein complexes that may be involved in gene transcription. To facilitate comparative analyses and identification of protein complexes, the compositional information is integrated with standardized gene nomenclature, annotation and protein sequences from public databases. The PINdb web interface provides a number of tools for (1) comparison of protein complexes, (2) search for a protein complex by its published name or by a partial list of its components and (3) browsing specific subsets or a functional classification of the complexes. Availablity: http://pin.mskcc.org

Abstracting and Indexing↗

In vitro and in vivo complementation of the Helicobacter pylori arginase mutant using an intergenic chromosomal site.

BACKGROUND: Gene complementation strategies are important in validating the roles of genes in specific phenotypes. Complementation systems in Helicobacter pylori include shuttle vectors, which transform H. pylori at relatively low frequencies, and chromosomally based approaches. Chromosomal complementation strategies are susceptible to polar effects and disruption of other H. pylori genes, leading to unwanted pleiotropic effects. MATERIALS AND METHODS: A new complementation strategy was developed for H. pylori by utilizing a suicide plasmid vector that contains fragments of an H. pylori intergenic region (hp0203-hp0204), a chloramphenicol acetyltransferase cassette (cat), and a multiple-cloning site. Genes of interest could be cloned into the intergenic plasmid and the genes integrated into H. pylori by homologous recombination into the intergenic chromosomal region without disrupting any annotated H. pylori gene. The complementation system was validated using the gene encoding arginase (rocF). RESULTS: A rocF mutant unable to hydrolyze or consume l-arginine regained these functions by complementation with the wild-type rocF gene. Complemented strains also had restored arginase protein as determined by Western blot analysis. The complementation system could be successfully applied to multiple H. pylori strains. The intergenic region varied in length and sequence across 17 H. pylori strains, but the flanking-3' ends of the hp0203 and hp0204 coding regions were highly conserved. Inserting a cat cassette and wild-type rocF into the intergenic region did not alter the ability of strain SS1 to colonize mice. CONCLUSIONS: This complementation strategy should greatly facilitate genetic experiments in H. pylori.

Animals↗

Nutrition, cancer, and aging: an annotated review. II. Cancer cachexia and aging.

The interactions of cancer and malnutrition are discussed with the focus on aging. To establish whether the elderly are more likely to develop cancer cachexia and its complications, this review encompasses the pathogenesis of malnutrition in cancer; the age-related alterations of appetite, gastrointestinal function, energy expenditure, and protein turnover; the diagnosis of malnutrition; and the effectiveness of nutritional support in the elderly. Although metabolic and physiologic changes induced by cancer and age appear synergistic in causing cachexia, more frequent complications of malnutrition have not been observed in the geriatric cancer patients. This may be due to only a small proportion of the elderly with cancer being enrolled in clinical studies or to a reduced cachexia-inducing ability of tumors in these patients. A limited number of studies indicate nutritional replenishment is obtainable in malnourished elderly by hyperalimentation. As restoration of the lean body mass may be slower in older patients, early institution of nutritional support is recommended in malnourished elderly or elderly at risk for malnutrition during neoplastic treatment.

Aged↗

microRNA target predictions across seven Drosophila species and comparison to mammalian targets.

microRNAs are small noncoding genes that regulate the protein production of genes by binding to partially complementary sites in the mRNAs of targeted genes. Here, using our algorithm PicTar, we exploit cross-species comparisons to predict, on average, 54 targeted genes per microRNA above noise in Drosophila melanogaster. Analysis of the functional annotation of target genes furthermore suggests specific biological functions for many microRNAs. We also predict combinatorial targets for clustered microRNAs and find that some clustered microRNAs are likely to coordinately regulate target genes. Furthermore, we compare microRNA regulation between insects and vertebrates. We find that the widespread extent of gene regulation by microRNAs is comparable between flies and mammals but that certain microRNAs may function in clade-specific modes of gene regulation. One of these microRNAs (miR-210) is predicted to contribute to the regulation of fly oogenesis. We also list specific regulatory relationships that appear to be conserved between flies and mammals. Our findings provide the most extensive microRNA target predictions in Drosophila to date, suggest specific functional roles for most microRNAs, indicate the existence of coordinate gene regulation executed by clustered microRNAs, and shed light on the evolution of microRNA function across large evolutionary distances. All predictions are freely accessible at our searchable Web site http://pictar.bio.nyu.edu.

Journal Article↗

Characterisation of global protein expression by two-dimensional electrophoresis and mass spectrometry: proteomics of Toxoplasma gondii.

The development of tools for the analysis of global gene expression is vital for the optimal exploitation of the data on parasite genomes that are now being generated in abundance. Recent advances in two-dimensional electrophoresis (2-DE), mass spectrometry and bioinformatics have greatly enhanced the possibilities for mapping and characterisation of protein populations. We have employed these developments in a proteomics approach for the analysis of proteins expressed in the tachyzoite stage of Toxoplasma gondii. Over 1000 polypeptides were reproducibly separated by high-resolution 2-DE using the pH ranges 4-7 and 6-11. Further separations using narrow range gels suggest that at least 3000-4000 polypeptides should be resolvable by 2-DE using multiple single pH unit gels. Mass spectrometry was used to characterise a variety of protein spots on the 2-DE gels. Peptide mass fingerprints, acquired by matrix-assisted laser desorption/ionisation-(MALDI) mass spectrometry, enabled unambiguous protein identifications to be made where full gene sequence information was available. However, interpretation of peptide mass fingerprint data using the T. gondii expressed sequence tag (EST) database was less reliable. Peptide fragmentation data, acquired by post-source decay mass spectrometry, proved a more successful strategy for the putative identification of proteins using the T. gondii EST database and protein databases from other organisms. In some instances, several protein spots appeared to be encoded by the same gene, indicating that post-translational modification and/or alternative splicing events may be a common feature of functional gene expression in T. gondii. The data demonstrate that proteomic analyses are now viable for T. gondii and other protozoa for which there are good EST databases, even in the absence of complete genome sequence. Moreover, proteomics is of great value in interpreting and annotating EST databases.

Animals↗

Distinguishing the ORFs from the ELFs: short bacterial genes and the annotation of genomes.

A substantial fraction of hypothetical open reading frames (ORFs) in completely sequenced bacterial genomes are short, suggesting that many are not genes but random stretches of DNA. Although it is not feasible to authenticate the coding capacity of all such regions experimentally, comparisons of ORFs in related genomes can expose those that encode functional proteins.

Bacteria↗