Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Molecular cloning and functional expression of the first two specific insect myosuppressin receptors.

The Drosophila Genome Project database contains the sequences of two genes, CG8985 and CG13803, which are predicted to code for G protein-coupled receptors. We cloned the cDNAs corresponding to these genes and found that their gene structures had not been correctly annotated. We subsequently expressed the coding regions of the two corrected receptor genes in Chinese hamster ovary cells and found that each of them coded for a receptor that could be activated by low concentrations of Drosophila myosuppressin (EC50,4 x 10(-8) M). The insect myosuppressins are decapeptides that generally inhibit insect visceral muscles. Other tested Drosophila neuropeptides did not activate the two receptors. In addition to the two Drosophila myosuppressin receptors, we identified a sequence in the genomic database from the malaria mosquito Anopheles gambiae that also very likely codes for a myosuppressin receptor. To our knowledge, this paper is the first report on the molecular identification of specific insect myosuppressin receptors.

Amino Acid Sequence↗

Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.

We present an analysis of 203 completed genomes in the Gene3D resource (including 17 eukaryotes), which demonstrates that the number of protein families is continually expanding over time and that singleton-sequences appear to be an intrinsic part of the genomes. A significant proportion of the proteomes can be assigned to fewer than 6000 well-characterized domain families with the remaining domain-like regions belonging to a much larger number of small uncharacterized families that are largely species specific. Our comprehensive domain annotation of 203 genomes enables us to provide more accurate estimates of the number of multi-domain proteins found in the three kingdoms of life than previous calculations. We find that 67% of eukaryotic sequences are multi-domain compared with 56% of sequences in prokaryotes. By measuring the domain coverage of genome sequences, we show that the structural genomics initiatives should aim to provide structures for less than a thousand structurally uncharacterized Pfam families to achieve reasonable structural annotation of the genomes. However, in large families, additional structures should be determined as these would reveal more about the evolution of the family and enable a greater understanding of how function evolves.

Algorithms↗

Enhanced genome annotation using structural profiles in the program 3D-PSSM.

A method (three-dimensional position-specific scoring matrix, 3D-PSSM) to recognise remote protein sequence homologues is described. The method combines the power of multiple sequence profiles with knowledge of protein structure to provide enhanced recognition and thus functional assignment of newly sequenced genomes. The method uses structural alignments of homologous proteins of similar three-dimensional structure in the structural classification of proteins (SCOP) database to obtain a structural equivalence of residues. These equivalences are used to extend multiply aligned sequences obtained by standard sequence searches. The resulting large superfamily-based multiple alignment is converted into a PSSM. Combined with secondary structure matching and solvation potentials, 3D-PSSM can recognise structural and functional relationships beyond state-of-the-art sequence methods. In a cross-validated benchmark on 136 homologous relationships unambiguously undetectable by position-specific iterated basic local alignment search tool (PSI-Blast), 3D-PSSM can confidently assign 18 %. The method was applied to the remaining unassigned regions of the Mycoplasma genitalium genome and an additional 13 regions were assigned with 95 % confidence. 3D-PSSM is available to the community as a web server: http://www.bmm.icnet.uk/servers/3dpssm

Algorithms↗

Proteomic hub proteins CDKN2B, TRAPPC2L, WFS1, and ARPP19 drive biochemical recurrence and metastatic progression in prostate cancer: Protein macromolecule action.

The biological characteristics and metastasis mechanism of prostate cancer are complex, involving the important role of many proteins in cell transcriptional regulation. This study focused on the role of the proteomic hub proteins CDKN2B, TRAPPC2L, WFS1 and ARPP19 in the biochemical recurrence and metastasis progression of prostate cancer. Cross-platform transcriptome integration and differential expression analysis were used to evaluate transcriptome characteristics in a prostate cancer cohort. Functional enrichment analysis was performed by gene ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway annotation, and weighted gene co-expression network analysis (WGCNA) was used to investigate cancer progression subtypes. It was found that prostate cancer progression showed significant transcriptome heterogeneity, and low-expression genes dominated. We reveal the important role of epithelial-immune interactions and inflammatory signaling in transcriptional remodeling in prostate cancer. The co-expression network topology analysis showed that the immune-metabolic center module plays a central role in cancer progression. CDKN2B was identified as a key transcriptional determinant in prostate cancer typing, while TRAPPC2L and WFS1 acted as core transcriptional regulators, driving metastatic heterogeneity. ARPP19 and LOC650152 also show important transcriptional driving effects in advanced prostate cancer.

Humans↗

Improving gene annotation of complete viral genomes.

Gene annotation in viruses often relies upon similarity search methods. These methods possess high specificity but some genes may be missed, either those unique to a particular genome or those highly divergent from known homologs. To identify potentially missing viral genes we have analyzed all complete viral genomes currently available in GenBank with a specialized and augmented version of the gene finding program GeneMarkS. In particular, by implementing genome-specific self-training protocols we have better adjusted the GeneMarkS statistical models to sequences of viral genomes. Hundreds of new genes were identified, some in well studied viral genomes. For example, a new gene predicted in the genome of the Epstein-Barr virus was shown to encode a protein similar to alpha-herpesvirus minor tegument protein UL14 with heat shock functions. Convincing evidence of this similarity was obtained after only 12 PSI-BLAST iterations. In another example, several iterations of PSI-BLAST were required to demonstrate that a gene predicted in the genome of Alcelaphine herpesvirus 1 encodes a BALF1-like protein which is thought to be involved in apoptosis regulation and, potentially, carcinogenesis. New predictions were used to refine annotations of viral genomes in the RefSeq collection curated by the National Center for Biotechnology Information. Importantly, even in those cases where no sequence similarities were detected, GeneMarkS significantly reduced the number of primary targets for experimental characterization by identifying the most probable candidate genes. The new genome annotations were stored in VIOLIN, an interactive database which provides access to similarity search tools for up-to-date analysis of predicted viral proteins.

Amino Acid Sequence↗

Identification of immune-relevant genes from atlantic salmon using suppression subtractive hybridization.

In order to probe the interaction between an invading microorganism and its host, we have investigated differential gene expression in Atlantic salmon (Salmo salar) experimentally infected with the pathogen Aeromonas salmonicida, the causative agent of furunculosis. Subtractive cDNA libraries were constructed by suppression subtractive hybridization (SSH) from 3 immune-relevant tissues at 2 time points during the infection process. Both forward- and reverse-subtracted libraries were generated, and approximately 200 clones were sequenced from each library, giving a total of 1778 expressed sequence tags (ESTs), which were annotated according to functional categories and deposited in GenBank (BQ035314-BQ037059). Numerous genes involved in signal transduction, innate immunity, and other processes have been uncovered in the subtractive libraries. These include known acute-phase reactants, along with more novel genes encoding proteins such as tachylectin, hepcidin, precerebellin-like protein, O-methyltransferase, a putative saxitoxin-binding protein, and others. A subset of genes that were represented in the subtracted libraries was further analyzed by virtual Northern, or reverse transcription-polymerase chain reaction (RT-PCR) assays to verify their differential expression as a result of infection.

Aeromonas salmonicida↗

Mitochondrial diseases preferentially involve proteins with prokaryote homologues.

The comparison of each of the 393 nuclear-encoded human mitochondrial proteins annotated in the SwissProt databank with 256,953 proteins from 94 prokaryote species showed that two thirds of the mitochondrial proteome were homologous with prokaryotic proteins, whereas one third was not. Prokaryotic mitochondrial proteins differ markedly from eukaryotic proteins, particularly in regard to their size, localization, function, and mitochondrial-targeting N-terminal sequence. Remarkably, the majority of nuclear genes implicated in respiratory chain mitochondrial diseases were found to be of prokaryotic ancestry. Our study indicates that the investigation of the co-evolution of eukaryotic and prokaryotic mitochondrial proteins should lead to a better understanding of mitochondrial diseases.

Humans↗

Functional screening for proapoptotic genes by reverse transfection cell array technology.

Application of mathematical algorithms to sequenced whole genomes revealed a large number of predicted genes, requiring functional assays for their characterization in a high-throughput manner. Here, we report on the development of a screening assay, which is based on reverse transfection of cellular arrays and subsequent analysis of cell morphology to identify novel proapoptotic genes. Expression plasmids containing full-length cDNAs were cotransfected with the reporter plasmid pEYFP to screen for apoptotic body formation, based on EYFP fluorescence. The assay was validated and applied to 382 human sequence-verified full-length open reading frames, most of them of unknown function. In this initial screening, proapoptotic effects could be demonstrated for 10 of these genes. For 6 of them apoptosis induction could be confirmed both by TUNEL assay and by FACS analysis of cells stained according to Nicoletti: 1 gene was not yet annotated for an apoptotic function (ST6GAL2), while 5 genes were without annotated function (FLJ20551, CXorf12, FAM105A, TMEM66, C19orf4). Our study demonstrates the potential of this method to characterize functionally genes of unknown function in a highly parallel format.

Apoptosis↗

The Protein Information Resource.

The Protein Information Resource (PIR) is an integrated public resource of protein informatics that supports genomic and proteomic research and scientific discovery. PIR maintains the Protein Sequence Database (PSD), an annotated protein database containing over 283 000 sequences covering the entire taxonomic range. Family classification is used for sensitive identification, consistent annotation, and detection of annotation errors. The superfamily curation defines signature domain architecture and categorizes memberships to improve automated classification. To increase the amount of experimental annotation, the PIR has developed a bibliography system for literature searching, mapping, and user submission, and has conducted retrospective attribution of citations for experimental features. PIR also maintains NREF, a non-redundant reference database, and iProClass, an integrated database of protein family, function, and structure information. PIR-NREF provides a timely and comprehensive collection of protein sequences, currently consisting of more than 1 000 000 entries from PIR-PSD, SWISS-PROT, TrEMBL, RefSeq, GenPept, and PDB. The PIR web site (http://pir.georgetown.edu) connects data analysis tools to underlying databases for information retrieval and knowledge discovery, with functionalities for interactive queries, combinations of sequence and text searches, and sorting and visual exploration of search results. The FTP site provides free download for PSD and NREF biweekly releases and auxiliary databases and files.

Amino Acid Sequence↗

ACLAME: a CLAssification of Mobile genetic Elements.

The ACLAME database (http://aclame.ulb.ac.be) is a collection and classification of prokaryotic mobile genetic elements (MGEs) from various sources, comprising all known phage genomes, plasmids and transposons. In addition to providing information on the full genomes and genetic entities, it aims to build a comprehensive classification of the functional modules of MGEs at the protein, gene and higher levels. This first version contains a comprehensive classification of 5069 proteins from 119 DNA bacteriophages into over 400 functional families. This classification was produced automatically using TRIBE-MCL, a graph-theory-based Markov clustering algorithm that uses sequence measures as input, and then manually curated. Manual curation was aided by consulting annotations available in public databases retrieved through additional sequence similarity searches using Psi-Blast and Hidden Markov Models. The database is publicly accessible and open to expert volunteers willing to participate in its curation. Its web interface allows browsing as well as querying the classification. The main objectives are to collect and organize in a rational way the complexity inherent to MGEs, to extend and improve the inadequate annotation currently associated with MGEs and to screen known genomes for the validation and discovery of new MGEs.

Bacteriophages↗

EchoBASE: an integrated post-genomic database for Escherichia coli.

EchoBASE (http://www.ecoli-york.org) is a relational database designed to contain and manipulate information from post-genomic experiments using the model bacterium Escherichia coli K-12. Its aim is to collate information from a wide range of sources to provide clues to the functions of the approximately 1500 gene products that have no confirmed cellular function. The database is built on an enhanced annotation of the updated genome sequence of strain MG1655 and the association of experimental data with the E.coli genes and their products. Experiments that can be held within EchoBASE include proteomics studies, microarray data, protein-protein interaction data, structural data and bioinformatics studies. EchoBASE also contains annotated information on 'orphan' enzyme activities from this microbe to aid characterization of the proteins that catalyse these elusive biochemical reactions.

Databases, Genetic↗

Plant protein annotation in the UniProt Knowledgebase.

The Swiss-Prot, TrEMBL, Protein Information Resource (PIR), and DNA Data Bank of Japan (DDBJ) protein database activities have united to form the Universal Protein Resource (UniProt) Consortium. UniProt presents three database layers: the UniProt Archive, the UniProt Knowledgebase (UniProtKB), and the UniProt Reference Clusters. The UniProtKB consists of two sections: UniProtKB/Swiss-Prot (fully manually curated entries) and UniProtKB/TrEMBL (automated annotation, classification and extensive cross-references). New releases are published fortnightly. A specific Plant Proteome Annotation Program (http://www.expasy.org/sprot/ppap/) was initiated to cope with the increasing amount of data produced by the complete sequencing of plant genomes. Through UniProt, our aim is to provide the scientific community with a single, centralized, authoritative resource for protein sequences and functional information that will allow the plant community to fully explore and utilize the wealth of information available for both plant and non-plant model organisms.

Amino Acid Sequence↗

A probabilistic classifier for olfactory receptor pseudogenes.

BACKGROUND: Olfactory receptors (ORs), the largest mammalian gene superfamily (900-1400 genes), has >50% pseudogenes in humans. While most of these inactive genes are identified via coding frame (nonsense) disruptions, seemingly intact genes may also be inactive due to other deleterious (missense) mutations. An ultimate assessment of the actual size of the functional human OR repertoire thus requires an accurate distinction between genes and pseudogenes. RESULTS: To characterize inactive ORs with intact open reading frame, we have developed a probabilistic Classifier for Olfactory Receptor Pseudogenes (CORP). This algorithm is based on deviations from a functionally crucial consensus, constituting sixty highly conserved positions identified by a comparison of two evolutionarily-constrained OR repertoires (mouse and dog) with a small pseudogene fraction. We used a logistic regression analysis to assign appropriate coefficients to the conserved position and thus achieving maximal separation between active and inactive ORs. Consequently, the algorithms identified only 5% of the mouse functional ORs as pseudogenes, setting an upper limit of 0.05 to the false positive detection. Finally we used this algorithm to classify the 384 purportedly intact human OR genes. Of these, 135 were predicted as likely encoding non-functional proteins, and 38 were segregating between active and inactive forms due to missense polymorphisms. CONCLUSION: We demonstrated that the CORP algorithm is capable to distinguish between functional and non-functional OR genes with high precision even when the encoded protein would differ by a single amino acid. Using the CORP algorithm, we predict that approximately 70% of human OR genes are likely non-functional pseudogenes, a much higher number than hitherto suspected. The method we present may be employed for better annotation of inactive members in other gene families as well. CORP algorithm is available at: http://bioportal.weizmann.ac.il/HORDE/CORP/

Algorithms↗

Genome-wide analysis of the general stress response in Bacillus subtilis.

Bacteria respond to diverse growth-limiting stresses by producing a large set of general stress proteins. In Bacillus subtilis and related Gram-positive pathogens, this response is governed by the sigma(B) transcription factor. To establish the range of cellular functions associated with the general stress response, we compared the transcriptional profiles of wild and mutant strains under conditions that induce sigma(B) activity. Macroarrays representing more than 3900 annotated reading frames of the B. subtilis genome were hybridized to (33)P-labelled cDNA populations derived from (i) wild-type and sigB mutant strains that had been subjected to ethanol stress; and (ii) a strain in which sigma(B) expression was controlled by an inducible promoter. On the basis of their significant sigma(B)-dependent expression in three independent experiments, we identified 127 genes as prime candidates for members of the sigma(B) regulon. Of these genes, 30 were known previously or inferred to be sigma(B) dependent by other means. To assist in the analysis of the 97 new genes, we constructed hidden Markov models (HMM) that identified possible sigma(B) recognition sequences preceding 21 of them. To test the HMM and to provide an independent validation of the hybridization experiments, we mapped the sigma(B)-dependent messages for seven representative genes. For all seven, the 5' end of the message lay near typical sigma(B) recognition sequences, and these had been predicted correctly by the HMM for five of the seven examples. Lastly, all 127 gene products were assigned to functional groups by considering their similarity to known proteins. Notably, products with a direct protective function were in the minority. Instead, the general stress response increased relative message levels for known or predicted regulatory proteins, for transporters controlling solute influx and efflux, including potential drug efflux pumps, and for products implicated in carbon metabolism, envelope function and macromolecular turnover.

Bacillus subtilis↗

Text-based analysis of genes, proteins, aging, and cancer.

The diverse nature of cancer- and aging-related genes presents a challenge for large-scale studies based on molecular sequence and profiling data. An underexplored source of data for modeling and analysis is the textual descriptions and annotations present in curated gene-centered biomedical corpora. Here, 450 genes designated by surveys of the scientific literature as being associated with cancer and aging were analyzed using two complementary approaches. The first, ensemble attribute profile clustering, is a recently formulated, text-based, semi-automated data interpretation strategy that exploits ideas from statistical information retrieval to discover and characterize groups of genes with common structural and functional properties. Groups of genes with shared and unique Gene Ontology terms and protein domains were defined and examined. Human homologs of a group of known Drosphila aging-related genes are candidates for genes that may influence lifespan (hep/MAPK2K7, bsk/MAPK8, puc/LOC285193). These JNK pathway-associated proteins may specify a molecular hub that coordinates and integrates multiple intra- and extracellular processes via space- and time-dependent interactions with proteins in other pathways. The second approach, a qualitative examination of the chromosomal locations of 311 human cancer- and aging-related genes, provides anecdotal evidence for a "phenotype position effect": genes that are proximal in the linear genome often encode proteins involved in the same phenomenon. Comparative genomics was employed to enhance understanding of several genes, including open reading frames, identified as new candidates for genes with roles in aging or cancer. Overall, the results highlight fundamental molecular and mechanistic connections between progenitor/stem cell lineage determination, embryonic morphogenesis, cancer, and aging. Despite diversity in the nature of the molecular and cellular processes associated with these phenomena, they seem related to the architectural hub of tissue polarity and a need to generate and control this property in a timely manner.

Aging↗

The Pfam protein families database.

Pfam is a large collection of protein multiple sequence alignments and profile hidden Markov models. Pfam is available on the World Wide Web in the UK at http://www.sanger.ac.uk/Software/Pfam/, in Sweden at http://www.cgb.ki.se/Pfam/, in France at http://pfam.jouy.inra.fr/ and in the US at http://pfam.wustl.edu/. The latest version (6.6) of Pfam contains 3071 families, which match 69% of proteins in SWISS-PROT 39 and TrEMBL 14. Structural data, where available, have been utilised to ensure that Pfam families correspond with structural domains, and to improve domain-based annotation. Predictions of non-domain regions are now also included. In addition to secondary structure, Pfam multiple sequence alignments now contain active site residue mark-up. New search tools, including taxonomy search and domain query, greatly add to the functionality and usability of the Pfam resource.

Animals↗

Update of NUREBASE: nuclear hormone receptor functional genomics.

Nuclear hormone receptors are an abundant class of ligand-activated transcriptional regulators, found in varying numbers in all animals. Based on our experience of managing the official nomenclature of nuclear receptors, we have developed NUREBASE, a database containing protein and DNA sequences, reviewed protein alignments and phylogenies, taxonomy and annotations for all nuclear receptors. New developments in NUREBASE include explicit declaration of alternative transcripts of each gene, and expression data for human and mouse nuclear receptors. The core of NUREBASE is reviewed, and it is completed by NUREBASE_DAILY, automatically updated every 24 h. All information on accessing and installing NUREBASE may be found at http://www. ens-lyon.fr/LBMC/laudet/nurebase/nurebase.html.

Alternative Splicing↗

Analysis of the mouse transcriptome based on functional annotation of 60,770 full-length cDNAs.

Only a small proportion of the mouse genome is transcribed into mature messenger RNA transcripts. There is an international collaborative effort to identify all full-length mRNA transcripts from the mouse, and to ensure that each is represented in a physical collection of clones. Here we report the manual annotation of 60,770 full-length mouse complementary DNA sequences. These are clustered into 33,409 'transcriptional units', contributing 90.1% of a newly established mouse transcriptome database. Of these transcriptional units, 4,258 are new protein-coding and 11,665 are new non-coding messages, indicating that non-coding RNA is a major component of the transcriptome. 41% of all transcriptional units showed evidence of alternative splicing. In protein-coding transcripts, 79% of splice variations altered the protein product. Whole-transcriptome analyses resulted in the identification of 2,431 sense-antisense pairs. The present work, completely supported by physical clones, provides the most comprehensive survey of a mammalian transcriptome so far, and is a valuable resource for functional genomics.

Alternative Splicing↗