Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Transcriptome analyses of human genes and applications for proteome analyses.

By utilizing recently developed full-length cDNA technologies, large-scale cDNA sequencing was carried out by several cDNA projects. Now full-length cDNA resources cover the major part of the protein-coding human genes. Comprehensive analyses of the collected full-length cDNA data revealed not only the complete sequences of thousands of novel gene transcripts but also novel alternatively spliced isoforms of hitherto identified genes. However, it was not as easy as expected to deduce their encoded amino acid sequences based solely on the full-length cDNA sequences. It was neither always the case that the longest open reading frame corresponded to the real protein coding region nor that the first ATG was the translation initiator codon. Also, proteome-wide mass-spectrometry analysis has shown that there is an unexpectedly large population of small proteins, encoded by so-called upstream open reading frames, within the cell. Since sound manual annotations by experts were still indispensable to address these problems, an international meeting to make transcriptome-wide functional annotations of cDNAs was held, namely the H-invitational. In this meeting, functional annotations were made both manually and computationally for most of the pre-existing full-length cDNAs collected from world-wide cDNA projects. The achieved integrated information for each of the cDNAs was published as a database. It was also shown that the full-length cDNA data were useful for identifying alternative splicing variants, exact transcriptional start sites of the mRNAs and the adjacent promoter regions. Rapidly accumulating genome data as well as versatile use of the transcriptome information will shortly lay a firm foundation for proteome-level understanding of human gene networks.

Alternative Splicing↗

Identification and expression of the cym, cmt, and tod catabolic genes from Pseudomonas putida KL47: expression of the regulatory todST genes as a factor for catabolic adaptation.

Pseudomonas putida KL47 is a natural isolate that assimilates benzene, 1-alkylbenzene (C(1)-C(4)), biphenyl, p-cumate, and p-cymene. The genetic background of strain KL47 underlying the broad range of growth substrates was examined. It was found that the cym and cmt operons are constitutively expressed due to a lack of the cymR gene, and the tod operon is still inducible by toluene and biphenyl. The entire array of gene clusters responsible for the catabolism of toluene and p-cymene/p-cumate has been cloned in a cosmid vector, pLAFR3, and were named pEK6 and pEK27, respectively. The two inserts overlap one another and the nucleotide sequence (42,505 bp) comprising the cym, cmt, and tod operons and its flanking genes in KL47 are almost identical (>99%) to those of P. putida F1. In the cloned DNA fragment, two genes with unknown functions, labeled cymZ and cmtR, were newly identified and show high sequence homology to dienelactone hydrolase and CymR proteins, respectively. The cmtR gene was identified in the place of the cmtI gene of previous annotation. Western blot analysis showed that, in strains F1 and KL47, the todT gene is not expressed during growth on Luria Bertani medium. In minimal basal salt medium, expression of the todT gene is inducible by toluene, but not by biphenyl in strain F1; however, it is constantly expressed in strain KL47, indicating that high levels of expression of the todST genes with one amino acid substitution in TodS might provide strain KL47 with a means of adaptation of the tod catabolic operon to various aromatic hydrocarbons.

Amino Acid Sequence↗

PRED-CLASS: cascading neural networks for generalized protein classification and genome-wide applications.

A cascading system of hierarchical, artificial neural networks (named PRED-CLASS) is presented for the generalized classification of proteins into four distinct classes-transmembrane, fibrous, globular, and mixed-from information solely encoded in their amino acid sequences. The architecture of the individual component networks is kept very simple, reducing the number of free parameters (network synaptic weights) for faster training, improved generalization, and the avoidance of data overfitting. Capturing information from as few as 50 protein sequences spread among the four target classes (6 transmembrane, 10 fibrous, 13 globular, and 17 mixed), PRED-CLASS was able to obtain 371 correct predictions out of a set of 387 proteins (success rate approximately 96%) unambiguously assigned into one of the target classes. The application of PRED-CLASS to several test sets and complete proteomes of several organisms demonstrates that such a method could serve as a valuable tool in the annotation of genomic open reading frames with no functional assignment or as a preliminary step in fold recognition and ab initio structure prediction methods. Detailed results obtained for various data sets and completed genomes, along with a web sever running the PRED-CLASS algorithm, can be accessed over the World Wide Web at http://o2.biol.uoa.gr/PRED-CLASS.

Computational Biology↗

A novel clan of zinc metallopeptidases with possible intramembrane cleavage properties.

Computer-based database searching and protein multiple sequence alignment has identified a novel clan of zinc metallopeptidases, which, by phylogenetic analysis, has been shown to contain six subfamilies. The family is characterized by four common transmembrane segments and three conserved sequence motifs. The combination of topology analysis and motif identification has detected three potential Zn2+ coordinating residues. Only two of the sequences of this novel zinc metallopeptidase clan possess any functional annotation, one of which is able to cleave its substrate within a cytosol/transmembrane segment junction. A number of observations suggest that the remaining members of this novel clan may also cleave their substrates within transmembrane segments.

Animals↗

PartsList: a web-based system for dynamically ranking protein folds based on disparate attributes, including whole-genome expression and interaction information.

As the number of protein folds is quite limited, a mode of analysis that will be increasingly common in the future, especially with the advent of structural genomics, is to survey and re-survey the finite parts list of folds from an expanding number of perspectives. We have developed a new resource, called PartsList, that lets one dynamically perform these comparative fold surveys. It is available on the web at http://bioinfo.mbb.yale.edu/partslist and http://www.partslist.org. The system is based on the existing fold classifications and functions as a form of companion annotation for them, providing 'global views' of many already completed fold surveys. The central idea in the system is that of comparison through ranking; PartsList will rank the approximately 420 folds based on more than 180 attributes. These include: (i) occurrence in a number of completely sequenced genomes (e.g. it will show the most common folds in the worm versus yeast); (ii) occurrence in the structure databank (e.g. most common folds in the PDB); (iii) both absolute and relative gene expression information (e.g. most changing folds in expression over the cell cycle); (iv) protein-protein interactions, based on experimental data in yeast and comprehensive PDB surveys (e.g. most interacting fold); (v) sensitivity to inserted transposons; (vi) the number of functions associated with the fold (e.g. most multi-functional folds); (vii) amino acid composition (e.g. most Cys-rich folds); (viii) protein motions (e.g. most mobile folds); and (ix) the level of similarity based on a comprehensive set of structural alignments (e.g. most structurally variable folds). The integration of whole-genome expression and protein-protein interaction data with structural information is a particularly novel feature of our system. We provide three ways of visualizing the rankings: a profiler emphasizing the progression of high and low ranks across many pre-selected attributes, a dynamic comparer for custom comparisons and a numerical rankings correlator. These allow one to directly compare very different attributes of a fold (e.g. expression level, genome occurrence and maximum motion) in the uniform numerical format of ranks. This uniform framework, in turn, highlights the way that the frequency of many of the attributes falls off with approximate power-law behavior (i.e. according to V(-b), for attribute value V and constant exponent b), with a few folds having large values and most having small values.

Cysteine↗

Collection, mapping, and annotation of over 28,000 cDNA clones from japonica rice.

We collected and completely sequenced 28,469 full-length complementary DNA clones from Oryza sativa L. ssp. japonica cv. Nipponbare. Through homology searches of publicly available sequence data, we assigned tentative protein functions to 21,596 clones (75.86%). Mapping of the cDNA clones to genomic DNA revealed that there are 19,000 to 20,500 transcription units in the rice genome. Protein informatics analysis against the InterPro database revealed the existence of proteins presented in rice but not in Arabidopsis. Sixty-four percent of our cDNAs are homologous to Arabidopsis proteins.

Alternative Splicing↗

Epstein-Barr virus-encoded BILF1 is a constitutively active G protein-coupled receptor.

Both beta- and gammaherpesviruses encode G protein-coupled receptors (GPCRs) with unique pharmacological phenotypes and important biological functions. An example is ORF74, the gamma2-herpesvirus Kaposi's sarcoma-associated herpesvirus (KSHV)-encoded GPCR, which is highly constitutively active and considered the key oncogene in Kaposi's sarcoma pathogenesis. In contrast, the current annotation of the Epstein-Barr virus (EBV) genome does not reveal any GPCR homolog encoded by this human oncogenic gamma1-herpesvirus. However, by employing bioinformatics, we recognized that the previously established EBV open reading frame BILF1 indeed encodes a GPCR. Additionally, BILF1 is a member of a new family of related GPCRs exclusively encoded by gamma1-herpesviruses. Expression of hemagglutinin-tagged BILF1 in the HEK293 epithelial cell line revealed that BILF1 is expressed as an approximately 50-kDa glycosylated protein. Immunocytochemistry and confocal microscopy revealed that BILF1 localizes predominantly to the plasma membrane, similar to the localization of KSHV ORF74. Using chimeric G proteins, we found that human and rhesus EBV-encoded BILF1 are highly potent constitutively active receptors, activating Galphai. Furthermore, BILF1 is able to inhibit forskolin-triggered CREB activation via stimulation of endogenous G proteins in a pertussis toxin-sensitive manner, verifying that BILF1 signals constitutively through Galphai. We suggest that EBV may use BILF1 to regulate Galphai-activated pathways during viral lytic replication, thereby affecting disease progression.

Amino Acid Sequence↗

Protein families and TRIBES in genome sequence space.

Accurate detection of protein families allows assignment of protein function and the analysis of functional diversity in complete genomes. Recently, we presented a novel algorithm called TribeMCL for the detection of protein families that is both accurate and efficient. This method allows family analysis to be carried out on a very large scale. Using TribeMCL, we have generated a resource called TRIBES that contains protein family information, comprising annotations, protein sequence alignments and phylogenetic distributions describing 311 257 proteins from 83 completely sequenced genomes. The analysis of at least 60 934 detected protein families reveals that, with the essential families excluded, paralogy levels are similar between prokaryotes, irrespective of genome size. The number of essential families is estimated to be between 366 and 426. We also show that the currently known space of protein families is scale free and discuss the implications of this distribution. In addition, we show that smaller families are often formed by shorter proteins and discuss the reasons for this intriguing pattern. Finally, we analyse the functional diversity of protein families in entire genome sequences. The TRIBES protein family resource is accessible at http://www.ebi.ac.uk/research/cgg/tribes/.

Algorithms↗

PASS2: an automated database of protein alignments organised as structural superfamilies.

BACKGROUND: The functional selection and three-dimensional structural constraints of proteins in nature often relates to the retention of significant sequence similarity between proteins of similar fold and function despite poor sequence identity. Organization of structure-based sequence alignments for distantly related proteins, provides a map of the conserved and critical regions of the protein universe that is useful for the analysis of folding principles, for the evolutionary unification of protein families and for maximizing the information return from experimental structure determination. The Protein Alignment organised as Structural Superfamily (PASS2) database represents continuously updated, structural alignments for evolutionary related, sequentially distant proteins. DESCRIPTION: An automated and updated version of PASS2 is, in direct correspondence with SCOP 1.63, consisting of sequences having identity below 40% among themselves. Protein domains have been grouped into 628 multi-member superfamilies and 566 single member superfamilies. Structure-based sequence alignments for the superfamilies have been obtained using COMPARER, while initial equivalencies have been derived from a preliminary superposition using LSQMAN or STAMP 4.0. The final sequence alignments have been annotated for structural features using JOY4.0. The database is supplemented with sequence relatives belonging to different genomes, conserved spatially interacting and structural motifs, probabilistic hidden markov models of superfamilies based on the alignments and useful links to other databases. Probabilistic models and sensitive position specific profiles obtained from reliable superfamily alignments aid annotation of remote homologues and are useful tools in structural and functional genomics. PASS2 presents the phylogeny of its members both based on sequence and structural dissimilarities. Clustering of members allows us to understand diversification of the family members. The search engine has been improved for simpler browsing of the database. CONCLUSIONS: The database resolves alignments among the structural domains consisting of evolutionarily diverged set of sequences. Availability of reliable sequence alignments of distantly related proteins despite poor sequence identity and single-member superfamilies permit better sampling of structures in libraries for fold recognition of new sequences and for the understanding of protein structure-function relationships of individual superfamilies. PASS2 is accessible at http://www.ncbs.res.in/~faculty/mini/campass/pass2.html

Amino Acid Sequence↗

Linking publication, gene and protein data.

The computational reconstruction of biological systems, 'systems biology', is necessarily dependent on the existence of well-annotated data sets defining and describing the components of these systems, especially genes and the proteins they encode. Information about these components can be accessed either through structured bioinformatics databases, which store basic chemical and functional information abstracted from (or supplementing) the scientific literature, or through the literature itself, which is richer in content but essentially unstructured.

Amino Acid Sequence↗

Digging for dead genes: an analysis of the characteristics of the pseudogene population in the Caenorhabditis elegans genome.

Pseudogenes are non-functioning copies of genes in genomic DNA, which may either result from reverse transcription from an mRNA transcript (processed pseudogenes) or from gene duplication and subsequent disablement (non-processed pseudogenes). As pseudogenes are apparently 'dead', they usually have a variety of obvious disablements (e.g., insertions, deletions, frameshifts and truncations) relative to their functioning homologs. We have derived an initial estimate of the size, distribution and characteristics of the pseudogene population in the Caenorhabditis elegans genome, performing a survey in 'molecular archaeology'. Corresponding to the 18 576 annotated proteins in the worm (i.e., in Wormpep18), we have found an estimated total of 2168 pseudogenes, about one for every eight genes. Few of these appear to be processed. Details of our pseudogene assignments are available from http://bioinfo.mbb.yale.edu/genome/worm/pseudogene. The population of pseudogenes differs significantly from that of genes in a number of respects: (i) pseudogenes are distributed unevenly across the genome relative to genes, with a disproportionate number on chromosome IV; (ii) the density of pseudogenes is higher on the arms of the chromosomes; (iii) the amino acid composition of pseudogenes is midway between that of genes and (translations of) random intergenic DNA, with enrichment of Phe, Ile, Leu and Lys, and depletion of Asp, Ala, Glu and Gly relative to the worm proteome; and (iv) the most common protein folds and families differ somewhat between genes and pseudogenes-whereas the most common fold found in the worm proteome is the immunoglobulin fold and the most common 'pseudofold' is the C-type lectin. In addition, the size of a gene family bears little overall relationship to the size of its corresponding pseudogene complement, indicating a highly dynamic genome. There are in fact a number of families associated with large populations of pseudogenes. For example, one family of seven-transmembrane receptors (represented by gene B0334.7) has one pseudogene for every four genes, and another uncharacterized family (represented by gene B0403.1) is approximately two-thirds pseudogenic. Furthermore, over a hundred apparent pseudogenic fragments do not have any obvious homologs in the worm.

Amino Acid Sequence↗

A physical map of 30,000 human genes.

A map of 30,181 human gene-based markers was assembled and integrated with the current genetic map by radiation hybrid mapping. The new gene map contains nearly twice as many genes as the previous release, includes most genes that encode proteins of known function, and is twofold to threefold more accurate than the previous version. A redesigned, more informative and functional World Wide Web site (www.ncbi.nlm.nih.gov/genemap) provides the mapping information and associated data and annotations. This resource constitutes an important infrastructure and tool for the study of complex genetic traits, the positional cloning of disease genes, the cross-referencing of mammalian genomes, and validated human transcribed sequences for large-scale studies of gene expression.

Animals↗

Using Kleisli to Bring Out Features in BLASTP Results.

BLASTP gives a good overall indication of what function a protein might have. However, analysis of BLASTP reports to discover various domain features in the protein is still tedious. We address this problem by using the modern data integration system, Kleisli, to bring out annotated features of BLASTP results. We further strengthen our solution by incorporating additional information from SEG, ClustalW, hmmPfam, etc. It is also noteworthy that the codes of our implementation is sufficiently short to be presented in its entirety.

Journal Article↗

Pharmacoproteomics in drug development.

The field of proteomics is taking on increased significance as the relevance of investigating and understanding protein expression in disease and drug development is appreciated. Recent advances in proteomics have been driven by the availability of numerous annotated whole-genome sequences and a broad range of technological and bioinformatic developments that underscore the complexity of the proteome. This review briefly addresses some of the various technologies that comprise Expression Proteomics and Functional Proteomics, citing examples where these emerging approaches have been applied to pharmacology, toxicology, and the development of drugs.

Animals↗

Generation, annotation, and analysis of an extensive Aspergillus niger EST collection.

BACKGROUND: Aspergillus niger, a saprophyte commonly found on decaying vegetation, is widely used and studied for industrial purposes. Despite its place as one of the most important organisms for commercial applications, the lack of available information about its genetic makeup limits research with this filamentous fungus. RESULTS: We present here the analysis of 12,820 expressed sequence tags (ESTs) generated from A. niger cultured under seven different growth conditions. These ESTs identify about 5,108 genes of which 44.5% code for proteins sharing similarity (E < or = 1e(-5)) with GenBank entries of known function, 38% code for proteins that only share similarity with GenBank entries of unknown function and 17.5% encode proteins that do not have a GenBank homolog. Using the Gene Ontology hierarchy, we present a first classification of the A. niger proteins encoded by these genes and compare its protein repertoire with other well-studied fungal species. We have established a searchable web-based database that includes the EST and derived contig sequences and their annotation. Details about this project and access to the annotated A. niger database are available. CONCLUSION: This EST collection and its annotation provide a significant resource for fundamental and applied research with A. niger. The gene set identified in this manuscript will be highly useful in the annotation of the genome sequence of A. niger, the genes described in the manuscript, especially those encoding hydrolytic enzymes will provide a valuable source for researchers interested in enzyme properties and applications.

Aspergillus niger↗

Comparative analysis of protein coding sequences from human, mouse and the domesticated pig.

BACKGROUND: The availability of abundant sequence data from key model organisms has made large scale studies of molecular evolution an exciting possibility. Here we use full length cDNA alignments comprising more than 700,000 nucleotides from human, mouse, pig and the Japanese pufferfish Fugu rubrices in order to investigate 1) the relationships between three major lineages of mammals: rodents, artiodactyls and primates, and 2) the rate of evolution and the occurrence of positive Darwinian selection using codon based models of sequence evolution. RESULTS: We provide evidence that the evolutionary splits among primates, rodents and artiodactyls happened shortly after each other, with most gene trees favouring a topology with rodents as outgroup to primates and artiodactyls. Using an unrooted topology of the three mammalian species we show that since their diversification, the pig and mouse lineages have on average experienced 1.44 and 2.86 times as many synonymous substitutions as humans, respectively, whereas the rates of non-synonymous substitutions are more similar. The analysis shows the highest average dN/dS ratio in the human lineage, followed by the pig and then the mouse lineages. Using codon based models we detect signals of positive Darwinian selection in approximately 5.3%, 4.9% and 6.0% of the genes on the human, pig and mouse lineages respectively. Approximately 16.8% of all the genes studied here are not currently annotated as functional genes in humans. Our analyses indicate that a large fraction of these genes may have lost their function quite recently or may still be functional genes in some or all of the three mammalian species. CONCLUSIONS: We present a comparative analysis of protein coding genes from three major mammalian lineages. Our study demonstrates the usefulness of codon-based likelihood models in detecting selection and it illustrates the value of sequencing organisms at different phylogenetic distances for comparative studies.

Animals↗

Annotating the human proteome: beyond establishing a parts list.

The completion of the human genome has shifted the attention from deciphering the sequence to the identification and characterisation of the functional components, including genes. Improved gene prediction algorithms, together with the existing transcript and protein information, have enabled the identification of most exons in a genome. Availability of the 'parts list' has fostered the development of experimental approaches to systematically interrogate gene function on the genome, transcriptome and proteome level. Studying gene function at the protein level is vital to the understanding of how cells perform their functions as variations in protein isoforms and protein quantity which may underlie a change in phenotype can often not be deduced from sequence or transcript level genomics experiments alone. Recent advancements in proteomics have afforded technologies capable of measuring protein expression, post-translational modifications of these proteins, their subcellular localisation and assembly into complexes and pathways. Although an enormous amount of data already exists on the function of many human proteins, much of it is scattered over multiple resources. Public domain databases are therefore required to manage and collate this information and present it to the user community in both a human and machine readable manner. Of special importance here is the integration of heterogeneous data to facilitate the creation of resources that go beyond a mere parts list.

Humans↗

'Frankenstein genes', or the Mad Magazine version of the human pseudogenome.

Annotation of the human genome is inching forward. Seven human chromosomes have now been fully annotated, covering 17 per cent of the genome, and at least one chromosome has been re-annotated. The enormity of the task forces a dependence on automated tools for detecting and assembling the genes, followed by hand curation to correct errors and polish the gene models. The accuracy of gene prediction algorithms is very good for internal exons from intact genes, but these programs do peculiar and exasperating things to pseudogenes. These programs can actually resurrect pseudogenes from the dead, making them into viable gene models for intact proteins, albeit science-fictional proteins. This process is demonstrated for four human pseudogenes from the cytochrome P450 family and one putatively functional P450 gene, CYP2U1, having a non-consensus intron boundary. These examples are offered as a call-to-arms to improve pseudogene prediction as an art in itself, and not as a by-product of gene annotation. Failure to do so will flood the databases with thousands of false-positive predictions. Indeed, they are already there.

Amino Acid Sequence↗