Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Reevaluating human gene annotation: a second-generation analysis of chromosome 22.

We report a second-generation gene annotation of human chromosome 22. Using expressed sequence databases, comparative sequence analysis, and experimental verification, we have extended genes, fused previously fragmented structures, and identified new genes. The total length in exons of annotation was increased by 74% over our previously published annotation and includes 546 protein-coding genes and 234 pseudogenes. Thirty-two potential protein-coding annotations are partial copies of other genes, and may represent duplications on an evolutionary path to change or loss of function. We also identified 31 non-protein-coding transcripts, including 16 possible antisense RNAs. By extrapolation, we estimate the human genome contains 29,000-36,000 protein-coding genes, 21,300 pseudogenes, and 1500 antisense RNAs. We suggest that our revised annotation criteria provide a paradigm for future annotation of the human genome.

Animals↗

CD-Search: protein domain annotations on the fly.

We describe the Conserved Domain Search service (CD-Search), a web-based tool for the detection of structural and functional domains in protein sequences. CD-Search uses BLAST(R) heuristics to provide a fast, interactive service, and searches a comprehensive collection of domain models. Search results are displayed as domain architecture cartoons and pairwise alignments between the query and domain-model consensus sequences. Search results may be visualized in further detail by embedding the query sequence into multiple alignment displays and by mapping onto three-dimensional molecular graphic displays of known structures within the domain family. CD-Search can be accessed at http://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi.

Amino Acid Sequence↗

Super paramagnetic clustering of protein sequences.

BACKGROUND: Detection of sequence homologues represents a challenging task that is important for the discovery of protein families and the reliable application of automatic annotation methods. The presence of domains in protein families of diverse function, inhomogeneity and different sizes of protein families create considerable difficulties for the application of published clustering methods. RESULTS: Our work analyses the Super Paramagnetic Clustering (SPC) and its extension, global SPC (gSPC) algorithm. These algorithms cluster input data based on a method that is analogous to the treatment of an inhomogeneous ferromagnet in physics. For the SwissProt and SCOP databases we show that the gSPC improves the specificity and sensitivity of clustering over the original SPC and Markov Cluster algorithm (TRIBE-MCL) up to 30%. The three algorithms provided similar results for the MIPS FunCat 1.3 annotation of four bacterial genomes, Bacillus subtilis, Helicobacter pylori, Listeria innocua and Listeria monocytogenes. However, the gSPC covered about 12% more sequences compared to the other methods. The SPC algorithm was programmed in house using C++ and it is available at http://mips.gsf.de/proj/spc. The FunCat annotation is available at http://mips.gsf.de. CONCLUSION: The gSPC calculated to a higher accuracy or covered a larger number of sequences than the TRIBE-MCL algorithm. Thus it is a useful approach for automatic detection of protein families and unsupervised annotation of full genomes.

Algorithms↗

Ab initio prediction of transcription factor targets using structural knowledge.

Current approaches for identification and detection of transcription factor binding sites rely on an extensive set of known target genes. Here we describe a novel structure-based approach applicable to transcription factors with no prior binding data. Our approach combines sequence data and structural information to infer context-specific amino acid-nucleotide recognition preferences. These are used to predict binding sites for novel transcription factors from the same structural family. We demonstrate our approach on the Cys(2)His(2) Zinc Finger protein family, and show that the learned DNA-recognition preferences are compatible with experimental results. We use these preferences to perform a genome-wide scan for direct targets of Drosophila melanogaster Cys(2)His(2) transcription factors. By analyzing the predicted targets along with gene annotation and expression data we infer the function and activity of these proteins.

Journal Article↗

AnoEST: toward A. gambiae functional genomics.

Here, we present an analysis of 215,634 EST and cDNA sequences of a major vector of human malaria Anopheles gambiae structured into the AnoEST database. The expressed sequences are grouped into clusters using genomic sequence as template and associated with inferred functional annotation, including the following: corresponding Ensembl gene prediction, putative orthologous genes in other species, homology to known proteins, protein domains, associated Gene Ontology terms, and corresponding classification into broad GO-slim functional groups. AnoEST is a vital resource for interpretation of expression profiles derived using recently developed A. gambiae cDNA microarrays. Using these cDNA microarrays, we have experimentally confirmed the expression of 7961 clusters during mosquito development. Of these, 3100 are not associated with currently predicted genes. Moreover, we found that clusters with confirmed expression are nonbiased with respect to the current gene annotation or homology to known proteins. Consequently, we expect that many as yet unconfirmed clusters are likely to be actual A. gambiae genes. [AnoEST is publicly available at http://komar.embl.de, and is also accessible as a Distributed Annotation Service (DAS).].

Animals↗

Phylogeny of related functions: the case of polyamine biosynthetic enzymes.

Genome annotation requires explicit identification of gene function. This task frequently uses protein sequence alignments with examples having a known function. Genetic drift, co-evolution of subunits in protein complexes and a variety of other constraints interfere with the relevance of alignments. Using a specific class of proteins, it is shown that a simple data analysis approach can help solve some of the problems posed. The origin of ureohydrolases has been explored by comparing sequence similarity trees, maximizing amino acid alignment conservation. The trees separate agmatinases from arginases but suggest the presence of unknown biases responsible for unexpected positions of some enzymes. Using factorial correspondence analysis, a distance tree between sequences was established, comparing regions with gaps in the alignments. The gap tree gives a consistent picture of functional kinship, perhaps reflecting some aspects of phylogeny, with a clear domain of enzymes encoding two types of ureohydrolases (agmatinases and arginases) and activities related to, but different from ureohydrolases. Several annotated genes appeared to correspond to a wrong assignment if the trees were significant. They were cloned and their products expressed and identified biochemically. This substantiated the validity of the gap tree. Its organization suggests a very ancient origin of ureohydrolases. Some enzymes of eukaryotic origin are spread throughout the arginase part of the trees: they might have been derived from the genes found in the early symbiotic bacteria that became the organelles. They were transferred to the nucleus when symbiotic genes had to escape Muller's ratchet. This work also shows that arginases and agmatinases share the same two manganese-ion-binding sites and exhibit only subtle differences that can be accounted for knowing the three-dimensional structure of arginases. In the absence of explicit biochemical data, extreme caution is needed when annotating genes having similarities to ureohydrolases.

Amino Acid Sequence↗

Silencing the transcriptome's dark matter: mechanisms for suppressing translation of intergenic transcripts.

Large portions of the genomes of higher eukaryotes are transcribed into RNA molecules that are never destined for translation into proteins. Although some of these transcripts have clearly defined biological roles other than protein coding, most arise from genomic regions devoid of functional genes and many are antisense to regions containing annotated genes. A variety of mechanisms exist to prevent adventitious production of proteins from these transcripts, ranging from degradation within the nucleus to translational silencing in the cytosol.

Animals↗

A new family of NAD(P)H-dependent oxidoreductases distinct from conventional Rossmann-fold proteins.

A new family of NAD(P)H-dependent oxidoreductases is now recognized as a protein family distinct from conventional Rossmann-fold proteins. Numerous putative proteins belonging to the family have been annotated as malate dehydrogenase (MDH) or lactate dehydrogenase (LDH) according to the previous classification as type-2 malate/L-lactate dehydrogenases. However, recent biochemical and genetic studies have revealed that the protein family consists of a wide variety of enzymes with unique catalytic activities other than MDH or LDH activity. Based on their sequence homologies and plausible functions, the family proteins can be grouped into eight clades. This classification would be useful for reliable functional annotation of the new family of NAD(P)H-dependent oxidoreductases.

Amino Acid Sequence↗

Hepatitis C databases, principles and utility to researchers.

Part of the effort to develop hepatitis C-specific drugs a nd vaccines is the study of genetic variability of allpublicly available HCV sequences. Three HCV databases are currently available to aid this effort and to provide additional insight into the basic biology, immunology, and evolution of the virus. The Japanese HCV database (http://s2as02.genes.nig.ac.jp) gives access to a genomic mapping of sequences as well as their phylogenetic relationships. The European HCV database (http://euhcvdb.ibcp.fr) offers access to a computer-annotated set of sequences and molecular models of HCV proteins and focuses on protein sequence, structure and function analysis. The HCV database at the Los Alamos National Laboratory in the United States (http://hcv.lanl.gov) provides access to a manually annotated sequence database and a database of immunological epitopes which contains concise descriptions of experimental results. In this paper, we briefly describe each of these databases and their associated websites and tools, and give some examples of their use in furthering HCV research.

Biomedical Research↗

A high-quality chromosome-level genome assembly and annotation of the giant freshwater prawn (Macrobrachium rosenbergii).

The giant freshwater prawn, Macrobrachium rosenbergii, is native to Southeast Asia and is used in aquacultural practices worldwide. It is considered advantageous because of its rapid growth, high nutritional value, and economic benefits. As one of the three major freshwater aquaculture shrimp sources in China, a high-quality genome resource is of great significance for promoting the germplasm improvement of varieties. This study presents a high-quality chromosome-level genome assembly of M. rosenbergii that was generated by combining PacBio, MGI, and Hi-C reads. The assembled genome was 2.96 Gb in size, with a contig N50 of 0.64 Mb and a scaffold N50 of 55.76 Mb, which was positioned on 59 pseudo-chromosomes. The Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis for genome assembly reached 94.37%. In total, 27,111 protein-coding genes were identified, of which 25,470 were functionally annotated. These results provide a foundation for future research into adaptive evolution, genomics, and molecular breeding in M. rosenbergii.

Animals↗

Serine-threonine phosphoregulation by PknB and Stp contributes to quiescence and antibiotic tolerance in Staphylococcus aureus.

Staphylococcus aureus can cause infections that are often chronic and difficult to treat, even when the bacteria are not antibiotic resistant because most antibiotics act only on metabolically active cells. Subpopulations of persister cells are metabolically quiescent, a state associated with delayed growth, reduced protein synthesis, and increased tolerance to antibiotics. Serine-threonine kinases and phosphatases similar to those found in eukaryotes can fine-tune essential bacterial cellular processes, such as metabolism and stress signaling. We found that acid stress-mimicking conditions that S. aureus experiences in host tissues delayed growth, globally altered the serine and threonine phosphoproteome, and increased threonine phosphorylation of the activation loop of the serine-threonine protein kinase B (PknB). The deletion of stp, which encodes the only annotated functional serine-threonine phosphatase in S. aureus, increased the growth delay and phenotypic heterogeneity under different stress challenges, including growth in acidic conditions, the intracellular milieu of human cells, and abscesses in mice. This growth delay was associated with reduced protein translation and intracellular ATP concentrations and increased antibiotic tolerance. Using phosphopeptide enrichment and mass spectrometry-based proteomics, we identified targets of serine-threonine phosphorylation that may regulate bacterial growth and metabolism. Together, our findings highlight the importance of phosphoregulation in mediating bacterial quiescence and antibiotic tolerance and suggest that targeting PknB or Stp might offer a future therapeutic strategy to prevent persister formation during S. aureus infections.

Animals↗

Receptor-defined targeting of a genomically unique melanoma-enriched noncanonical antigen.

Effective T cell-based immunotherapies require functional receptors that can be engineered and redeployed to recognize tumor-restricted antigens. Noncanonical peptides arising from transcription outside annotated protein-coding regions expand the antigenic landscape of cancer; however, systematic strategies to biologically prioritize and functionally validate such targets remain underdeveloped. Here, we integrated de novo transcript analysis, exon-resolved quantification, RNA in situ hybridization, and immunopeptidomics to identify melanoma-associated noncanonical transcripts and advance candidates through receptor-level validation. Among three recurrent melanoma-associated transcripts, EVA003 emerged as a lead target based on its distinct repeat-enriched genomic architecture, consistent tumor-enriched exon-level expression across independent datasets, and a genomically unique immunogenic core sequence. We demonstrate endogenous presentation of EVA003-derived peptides on HLA-A*03:01 and detect specific reactivity in patient-derived tumor-infiltrating lymphocytes. Single-cell transcriptomic profiling identified a dominant peptide-reactive clonotype, enabling isolation of a naturally occurring T cell receptor. Transfer of this receptor into healthy donor T cells conferred antigen-dependent activation and cytotoxicity against both peptide-pulsed targets and melanoma cells expressing EVA003 endogenously. Together, these findings establish a biologically informed strategy for prioritizing noncanonical tumor antigens and demonstrate that genomically unique, tumor-enriched noncanonical peptides can be presented to molecularly defined receptors capable of mediating cancer cell killing. These findings support the integration of prioritized noncanonical antigens into engineered T cell therapeutic strategies.

Humans↗

Large scale analysis of sequences from Neurospora crassa.

After 50 years of analysing Neurospora crassa genes one by one large scale sequence analysis has increased the number of accessible genes tremendously in the last few years. Being the only filamentous fungus for which a comprehensive genomic sequence database is publicly accessible N. crassa serves as the model for this important group of microorganisms. The MIPS N. crassa database currently holds more than 16 Mb of non-redundant data of the chromosomes II and V analysed by the German Neurospora Genome Project. This represents more than one-third of the genome. Open reading frames (ORFs) have been extracted from the sequence and the deduced proteins have been annotated extensively. They are classified according to matches in sequence databases and attributed to functional categories according to their relatives. While 41% of analysed proteins are related to known proteins, 30% are hypothetical proteins with no match to a database entry. The entire genome is expected to comprise some 13000 protein coding genes, more than twice as many as found in yeasts, and reflects the high potential of filamentous fungi to cope with various environmental conditions.

Biotechnology↗

The TIGRFAMs database of protein families.

TIGRFAMs is a collection of manually curated protein families consisting of hidden Markov models (HMMs), multiple sequence alignments, commentary, Gene Ontology (GO) assignments, literature references and pointers to related TIGRFAMs, Pfam and InterPro models. These models are designed to support both automated and manually curated annotation of genomes. TIGRFAMs contains models of full-length proteins and shorter regions at the levels of superfamilies, subfamilies and equivalogs, where equivalogs are sets of homologous proteins conserved with respect to function since their last common ancestor. The scope of each model is set by raising or lowering cutoff scores and choosing members of the seed alignment to group proteins sharing specific function (equivalog) or more general properties. The overall goal is to provide information with maximum utility for the annotation process. TIGRFAMs is thus complementary to Pfam, whose models typically achieve broad coverage across distant homologs but end at the boundaries of conserved structural domains. The database currently contains over 1600 protein families. TIGRFAMs is available for searching or downloading at www.tigr.org/TIGRFAMs.

Animals↗

Combining text mining and sequence analysis to discover protein functional regions.

Recently presented protein sequence classification models can identify relevant regions of the sequence. This observation has many potential applications to detecting functional regions of proteins. However, identifying such sequence regions automatically is difficult in practice, as relatively few types of information have enough annotated sequences to perform this analysis. Our approach addresses this data scarcity problem by combining text and sequence analysis. First, we train a text classifier over the explicit textual annotations available for some of the sequences in the dataset, and use the trained classifier to predict the class for the rest of the unlabeled sequences. We then train a joint sequence text classifier over the text contained in the functional annotations of the sequences, and the actual sequences in this larger, automatically extended dataset. Finally, we project the classifier onto the original sequences to determine the relevant regions of the sequences. We demonstrate the effectiveness of our approach by predicting protein sub-cellular localization and determining localization specific functional regions of these proteins.

Algorithms↗

New friendly tools for users of ESTHER, the database of the alpha/beta-hydrolase fold superfamily of proteins.

The structural alpha/beta-hydrolase fold is characterized by a beta-sheet core of five to eight strands connected by alpha-helices to form a alpha/beta/alpha sandwich. The superfamily members, exemplified by the cholinesterases, diverged from a common ancestor into a number of hydrolytic enzymes displaying a wide range of substrate specificities, along with proteins with no recognized hydrolytic activity. In the enzymes, the catalytic triad residues are presented on loops of which one, the nucleophile elbow, is the most conserved feature of the fold. Of the other proteins, which all lack from one to all of the catalytic residues, some may simply be 'inactive' enzymes while others have been shown to be involved in heterologous surface recognition functions. The ESTHER (for esterases, alpha/beta-hydrolase enzymes and relatives) database (http://bioweb.ensam.inra.fr.esther) gathers and annotates all the published pieces of information (gene and protein sequences; biochemical, pharmacological, and structural data) related to the superfamily, and connects them together to provide the bases for studying structure-function relationships within the superfamily. The most recent developments of the database are presented.

Cholinesterases↗

Annotation, nomenclature and evolution of four novel homeobox genes expressed in the human germ line.

The homeobox genes comprise a large gene superfamily characterised by a conserved DNA motif encoding the homeodomain. Most homeodomain proteins function as transcription factors, and many have important roles in embryonic development and cell differentiation. Here we describe, annotate and name four novel homeobox genes in the human genome: ARGFX, DPRX, TPRX1 and DUXA. Each has generated multiple retrotransposed (processed) pseudogenes; these are reliable indicators of germ-line expression because only in germ-line cells can retrotransposition result in inheritance to the next generation. The retrotransposed sequences were exploited here as a novel means to deduce exon-intron boundaries. All four novel genes show accelerated rates of protein sequence evolution. This fast rate of sequence change may be connected with roles in human reproductive biology. Deducing the evolutionary origins of these genes is not straightforward, but we propose that TPRX1, DPRX and DUXA are highly divergent derivatives of the CRX gene, itself a member of the Otx homeobox gene family.

Evolution, Molecular↗

Arabidopsis thaliana full genome longmer microarrays: a powerful gene discovery tool for agriculture and forestry.

Sequenced plant genomes provide a large reservoir of known genes with potential for use in crop and tree improvement, but assignment of specific functions to annotated genes in sequenced plant genomes remains a challenge. Furthermore, most plant genes belong to families encoding proteins with related but distinct functions. In this commentary, we discuss our development of Arabidopsis spotted whole genome longmer oligonucleotide microarrays, and their use in global transcription profiling. We show that longmer array based transcriptome analysis in Arabidopsis can be used as an efficient and effective gene discovery and functional genomics tool, particularly for functional analyses of members of large gene families. We discuss experiments that focus on gene families involved in phenylpropanoid natural product biosynthesis and fiber differentiation. These analyses have helped to elucidate functions of individual gene family members, and have identified new candidate genes involved in fiber development and differentiation. Results obtained by these studies in Arabidopsis can be used as the basis for gene discovery in commercially important plants, and we have focused our attention on Populus trichocarpa (poplar), a species important in forestry and agroforestry for which complete genome sequence information is available.

Agriculture↗