Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,765 records · Page 98Linked to original sources

The SYSTERS Protein Family Database in 2005.

The SYSTERS project aims to provide a meaningful partitioning of the whole protein sequence space by a fully automatic procedure. A refined two-step algorithm assigns each protein to a family and a superfamily. The sequence data underlying SYSTERS release 4 now comprise several protein sequence databases derived from completely sequenced genomes (ENSEMBL, TAIR, SGD and GeneDB), in addition to the comprehensive Swiss-Prot/TrEMBL databases. The SYSTERS web server (http://systers.molgen.mpg.de) provides access to 158 153 SYSTERS protein families. To augment the automatically derived results, information from external databases like Pfam and Gene Ontology are added to the web server. Furthermore, users can retrieve pre-processed analyses of families like multiple alignments and phylogenetic trees. New query options comprise a batch retrieval tool for functional inference about families based on automatic keyword extraction from sequence annotations. A new access point, PhyloMatrix, allows the retrieval of phylogenetic profiles of SYSTERS families across organisms with completely sequenced genomes.

Algorithms↗

FISH--family identification of sequence homologues using structure anchored hidden Markov models.

The FISH server is highly accurate in identifying the family membership of domains in a query protein sequence, even in the case of very low sequence identities to known homologues. A performance test using SCOP sequences and an E-value cut-off of 0.1 showed that 99.3% of the top hits are to the correct family saHMM. Matches to a query sequence provide the user not only with an annotation of the identified domains and hence a hint to their function, but also with probable 2D and 3D structures, as well as with pairwise and multiple sequence alignments to homologues with low sequence identity. In addition, the FISH server allows users to upload and search their own protein sequence collection or to quarry public protein sequence data bases with individual saHMMs. The FISH server can be accessed at http://babel.ucmp.umu.se/fish/.

Databases, Protein↗

Collaboration system for radiology workstations.

Consultation between radiologists and referring physicians is part of routine medical practice. Nevertheless, a typical picture archiving and communication system contains no provision that will allow this critical interaction to occur on-line. The authors describe an image viewing system designed for real-time interactive consultation over the Internet. The system has two main components: an image viewer and a collaboration server. The image viewer connects to the collaboration server over an Internet-compatible network. Once the image viewer is connected, its display can be synchronized with that of another connected image viewer, so that radiologists can point out image findings and diagnoses in real time to remotely located physicians. The image viewer can retrieve images from any DICOM-compatible archive. In addition to standard image manipulation functions, the image viewer contains a new user interface for image annotation. Developed specifically for medical imaging, this user interface is activated by mouse actions instead of conventional on-screen controls, greatly improving the ease with which annotations can be created. The collaboration system is based on a simple yet flexible programming interface that can be readily generalized to other types of collaborative applications. The system was developed with the Java programming language because of Java's integrated support of Internet-compatible networking capabilities.

Humans↗

Shared genetic architecture of smoking dependence and Crohn's disease: A cross-trait analysis of GWAS summary statistics.

INTRODUCTION: Smoking dependence (SD) and Crohn's disease (CD) are epidemiologically associated, but whether this relationship reflects shared genetic susceptibility remains unclear. METHODS: We conducted a cross-trait genetic analysis of SD and CD using publicly available genome-wide association study (GWAS) summary statistics from European-ancestry populations. Genome-wide genetic correlation was estimated using linkage disequilibrium score regression (LDSC) and high-definition likelihood (HDL). Pleiotropic variants were identified using PLACO and mapped to genomic loci using FUMA. Regional signal sharing was assessed by Bayesian colocalization. Functional analyses included stratified LDSC, Multi-marker Analysis of GenoMic Annotation (MAGMA), GTEx tissue analysis, and Metascape. Expression-linked candidate genes were prioritized using expression quantitative trait locus (eQTL)-based summary-data-based Mendelian randomization (SMR) with heterogeneity in dependent instruments (HEIDI) testing. Genetically informed spatial mapping of cells for complex traits (gsMap) was used for spatial mapping. RESULTS: SD and CD showed positive genetic correlation by LDSC (rg=0.2090, p=0.0008) and HDL (rg=0.3817, p=0.00106). PLACO identified 81 genome-wide significant pleiotropic SNPs, which were mapped by FUMA to three loci at 1p31.3, 5p13.1, and 12q12, represented by rs11209031, rs1395152, and rs17467116, respectively. MAGMA identified 22 FDR-significant genes, four of which remained Bonferroni significant: LRRK2, TNFRSF6B, ZGPAT, and RP4-583P15.15. Cross-trait tissue analysis showed significant enrichment of the shared genetic signal in whole blood and small intestine, while gene-set analysis highlighted inflammatory response (pbon=1.86×10-5) and T-helper 17 cell differentiation (pbon=7.37×10-4). SMR/HEIDI analysis further prioritized RPS6KB1 as a shared expression-linked candidate. Spatial mapping revealed a prominent signal in the embryonic gastrointestinal tract and gene-specific regional patterns involving LRRK2 and SLC2A13 in the adult mouse brain. CONCLUSIONS: SD and CD showed measurable shared genetic susceptibility, with convergent evidence from pleiotropic loci, immune-inflammatory pathway enrichment, tissue-level associations, and spatial transcriptomic mapping.

Crohn's disease↗

Gotrees: predicting go associations from protein domain composition using decision trees.

The Gene Ontology (GO) offers a comprehensive and standardized way to describe a protein's biological role. Proteins are annotated with GO terms based on direct or indirect experimental evidence. Term assignments are also inferred from homology and literature mining. Regardless of the type of evidence used, GO assignments are manually curated or electronic. Unfortunately, manual curation cannot keep pace with the data, available from publications and various large experimental datasets. Automated literature-based annotation methods have been developed in order to speed up the annotation. However, they only apply to proteins that have been experimentally investigated or have close homologs with sufficient and consistent annotation. One of the homology-based electronic methods for GO annotation is provided by the InterPro database. The InterPro2GO/PFAM2GO associates individual protein domains with GO terms and thus can be used to annotate the less studied proteins. However, protein classification via a single functional domain demands stringency to avoid large number of false positives. This work broadens the basic approach. We model proteins via their entire functional domain content and train individual decision tree classifiers for each GO term using known protein assignments. We demonstrate that our approach is sensitive, specific and precise, as well as fairly robust to sparse data. We have found that our method is more sensitive when compared to the InterPro2GO performance and suffers only some precision decrease. In comparison to the InterPro2GO we have improved the sensitivity by 22%, 27% and 50% for Molecular Function, Biological Process and Cellular GO terms respectively.

Algorithms↗

Evolution of distinct EGF domains with specific functions.

EGF domains are extracellular protein modules cross-linked by three intradomain disulfides. Past studies suggest the existence of two types of EGF domain with three-disulfides, human EGF-like (hEGF) domains and complement C1r-like (cEGF) domains, but to date no functional information has been related to the two different types, and they are not differentiated in sequence or structure databases. We have developed new sequence patterns based on the different C-termini to search specifically for the two types of EGF domains in sequence databases. The exhibited sensitivity and specificity of the new pattern-based method represents a significant advancement over the currently available sequence detection techniques. We re-annotated EGF sequences in the latest release of Swiss-Prot looking for functional relationships that might correlate with EGF type. We show that important post-translational modifications of three-disulfide EGFs, including unusual forms of glycosylation and post-translational proteolytic processing, are dependent on EGF subtype. For example, EGF domains that are shed from the cell surface and mediate intercellular signaling are all hEGFs, as are all human EGF receptor family ligands. Additional experimental data suggest that functional specialization has accompanied subtype divergence. Based on our structural analysis of EGF domains with three-disulfide bonds and comparison to laminin and integrin-like EGF domains with an additional inter-domain disulfide, we propose that these hEGF and cEGF domains may have arisen from a four-disulfide ancestor by selective loss of different cysteine residues.

Amino Acid Sequence↗

Strategies for the identification, the assembly and the classification of integrated biological systems in completely sequenced genomes.

The proteins involved in a single biological process may form a stable supra-molecular assembly or be transiently in interaction. Although, the first annotation steps of a complete genome may allow the identification of the different partners, their assembly in a functional system, referred to as an integrated system, is a domain where methodological effort has to be done. Indeed, the knowledge required to assemble partners of such systems should be explicitly included in annotation software. The availability of a complete genome, and therefore of all the proteins encoded by that genome, motivated the development of automated approaches through the coordinated combination of different bio-informatic methods allowing the identification of the different partners, their assembly and the classification of the reconstructed systems in functional categories. In this data flux, the identification of the sequence partners represents the principal bottleneck. Here, we describe and compare the results obtained with different classes of methods (BLASTP2, PSI-BLAST, MAST and META-MEME) applied to the identification in complete genomes of a given family of integrated systems: the ABC transporters. PSI-BLAST appears to significantly outperform motif-based methods, and the results are discussed according to the nature of the proteins and the structure of the sub-families.

Binding Sites↗

Phylogenomic analysis of the GIY-YIG nuclease superfamily.

BACKGROUND: The GIY-YIG domain was initially identified in homing endonucleases and later in other selfish mobile genetic elements (including restriction enzymes and non-LTR retrotransposons) and in enzymes involved in DNA repair and recombination. However, to date no systematic search for novel members of the GIY-YIG superfamily or comparative analysis of these enzymes has been reported. RESULTS: We carried out database searches to identify all members of known GIY-YIG nuclease families. Multiple sequence alignments together with predicted secondary structures of identified families were represented as Hidden Markov Models (HMM) and compared by the HHsearch method to the uncharacterized protein families gathered in the COG, KOG, and PFAM databases. This analysis allowed for extending the GIY-YIG superfamily to include members of COG3680 and a number of proteins not classified in COGs and to predict that these proteins may function as nucleases, potentially involved in DNA recombination and/or repair. Finally, all old and new members of the GIY-YIG superfamily were compared and analyzed to infer the phylogenetic tree. CONCLUSION: An evolutionary classification of the GIY-YIG superfamily is presented for the very first time, along with the structural annotation of all (sub)families. It provides a comprehensive picture of sequence-structure-function relationships in this superfamily of nucleases, which will help to design experiments to study the mechanism of action of known members (especially the uncharacterized ones) and will facilitate the prediction of function for the newly discovered ones.

Archaeal Proteins↗

Complementing genomics with proteomics: the membrane subproteome of Pseudomonas aeruginosa PAO1.

With the completion of many genome projects, a shift is now occurring from the acquisition of gene sequence to understanding the role and context of gene products within the genome. The opportunistic pathogen Pseudomonas aeruginosa is one organism for which a genome sequence is now available, including the annotation of open reading frames (ORFs). However, approximately one third of the ORFs are as yet undefined in function. Proteomics can complement genomics, by characterising gene products and their response to a variety of biological and environmental influences. In this study we have established the first two-dimensional gel electrophoresis reference map of proteins from the membrane fraction of P. aeruginosa strain PA01. A total of 189 proteins have been identified and correlated with 104 genes from the P. aeruginosa genome. Annotated membrane proteins could be grouped into three distinct categories: (i) those with functions previously characterised in P. aeruginosa (38%); (ii) those with significant sequence similarity to proteins with assigned function or hypothetical proteins in other organisms (46%); and (iii) those with unknown function (16%). Transmembrane prediction algorithms showed that each identified protein sequence contained at least one membrane-spanning region. Furthermore, the current methodology used to isolate the membrane fraction was shown to be highly specific since no contaminating cytosolic proteins were characterised. Preliminary analysis showed that at least 15 gel spots may be glycosylated in vivo, including three proteins that have not previously been functionally characterised. The reference map of membrane proteins from this organism is now the basis for determining surface molecules associated with antibiotic resistance and efflux, cell-cell signalling and pathogen-host interactions in a variety of P. aeruginosa strains.

Bacterial Proteins↗

Large-scale testing of bibliome informatics using Pfam protein families.

Literature mining is expected to help not only with automatically sifting through huge biomedical literature and annotation databases, but also with linking bio-chemical entities to appropriate functional hypotheses. However, there has been very limited success in testing literature mining methods due to the lack of large, objectively validated test sets or "gold standards". To improve this situation we created a large-scale test of literature mining methods and resources. We report on a specific implementation of this test: how well can the Pfam protein family classification be replicated from independently mining different literature/annotation resources? We test and compare different keyterm sets as well as different algorithms for issuing protein family predictions. We find that protein families can indeed be automatically predicted from the literature. Using words from PubMed abstracts, of 3663 proteins tested, over 75% were correctly assigned to one of 618 Pfam families. For 90% of proteins the correct Pfam family was among the top 5 ranked families. We found that protein family prediction is far superior with keywords extracted from PubMed abstracts than with GO annotations or MeSH keyterms, suggesting that the text itself (in combination with the vector space model) is superior to GO and MeSH as a literature mining resources, at least for detecting protein family membership. Finally, we show that Shannon's entropy can be exploited to improve prediction by facilitating the integration of the different literature sources tested.

Algorithms↗

The evolutionary analysis of "orphans" from the Drosophila genome identifies rapidly diverging and incorrectly annotated genes.

In genome projects of eukaryotic model organisms, a large number of novel genes of unknown function and evolutionary history ("orphans") are being identified. Since many orphans have no known homologs in distant species, it is unclear whether they are restricted to certain taxa or evolve rapidly, either because of a lack of constraints or positive Darwinian selection. Here we use three criteria for the selection of putatively rapidly evolving genes from a single sequence of Drosophila melanogaster. Thirteen candidate genes were chosen from the Adh region on the second chromosome and 1 from the tip of the X chromosome. We succeeded in obtaining sequence from 6 of these in the closely related species D. simulans and D. yakuba. Only 1 of the 6 genes showed a large number of amino acid replacements and in-frame insertions/deletions. A population survey of this gene suggests that its rapid evolution is due to the fixation of many neutral or nearly neutral mutations. Two other genes showed "normal" levels of divergence between species. Four genes had insertions/deletions that destroy the putative reading frame within exons, suggesting that these exons have been incorrectly annotated. The evolutionary analysis of orphan genes in closely related species is useful for the identification of both rapidly evolving and incorrectly annotated genes.

Animals↗

Cloning and functional analysis of cDNAs with open reading frames for 300 previously undefined genes expressed in CD34+ hematopoietic stem/progenitor cells.

Three hundred cDNAs containing putatively entire open reading frames (ORFs) for previously undefined genes were obtained from CD34+ hematopoietic stem/progenitor cells (HSPCs), based on EST cataloging, clone sequencing, in silico cloning, and rapid amplification of cDNA ends (RACE). The cDNA sizes ranged from 360 to 3496 bp and their ORFs coded for peptides of 58-752 amino acids. Public database search indicated that 225 cDNAs exhibited sequence similarities to genes identified across a variety of species. Homology analysis led to the recognition of 50 basic structural motifs/domains among these cDNAs. Genomic exon-intron organization could be established in 243 genes by integration of cDNA data with genome sequence information. Interestingly, a new gene named as HSPC070 on 3p was found to share a sequence of 105bp in 3' UTR with RAF gene in reversed transcription orientation. Chromosomal localizations were obtained using electronic mapping for 192 genes and with radiation hybrid (RH) for 38 genes. Macroarray technique was applied to screen the gene expression patterns in five hematopoietic cell lines (NB4, HL60, U937, K562, and Jurkat) and a number of genes with differential expression were found. The resource work has provided a wide range of information useful not only for expression genomics and annotation of genomic DNA sequence, but also for further research on the function of genes involved in hematopoietic development and differentiation.

Alternative Splicing↗

From structure to function: YrbI from Haemophilus influenzae (HI1679) is a phosphatase.

The crystal structure of the YrbI protein from Haemophilus influenzae (HI1679) was determined at a 1.67-A resolution. The function of the protein had not been assigned previously, and it is annotated as hypothetical in sequence databases. The protein exhibits the alpha/beta-hydrolase fold (also termed the Rossmann fold) and resembles most closely the fold of the L-2-haloacid dehalogenase (HAD) superfamily. Following this observation, a detailed sequence analysis revealed remote homology to two members of the HAD superfamily, the P-domain of Ca(2+) ATPase and phosphoserine phosphatase. The 19-kDa chains of HI1679 form a tetramer both in solution and in the crystalline form. The four monomers are arranged in a ring such that four beta-hairpin loops, each inserted after the first beta-strand of the core alpha/beta-fold, form an eight-stranded barrel at the center of the assembly. Four active sites are located at the subunit interfaces. Each active site is occupied by a cobalt ion, a metal used for crystallization. The cobalt is octahedrally coordinated to two aspartate side-chains, a backbone oxygen, and three solvent molecules, indicating that the physiological metal may be magnesium. HI1679 hydrolyzes a number of phosphates, including 6-phosphogluconate and phosphotyrosine, suggesting that it functions as a phosphatase in vivo. The physiological substrate is yet to be identified; however the location of the gene on the yrb operon suggests involvement in sugar metabolism.

Amino Acid Sequence↗

The TIGRFAMs database of protein families.

TIGRFAMs is a collection of manually curated protein families consisting of hidden Markov models (HMMs), multiple sequence alignments, commentary, Gene Ontology (GO) assignments, literature references and pointers to related TIGRFAMs, Pfam and InterPro models. These models are designed to support both automated and manually curated annotation of genomes. TIGRFAMs contains models of full-length proteins and shorter regions at the levels of superfamilies, subfamilies and equivalogs, where equivalogs are sets of homologous proteins conserved with respect to function since their last common ancestor. The scope of each model is set by raising or lowering cutoff scores and choosing members of the seed alignment to group proteins sharing specific function (equivalog) or more general properties. The overall goal is to provide information with maximum utility for the annotation process. TIGRFAMs is thus complementary to Pfam, whose models typically achieve broad coverage across distant homologs but end at the boundaries of conserved structural domains. The database currently contains over 1600 protein families. TIGRFAMs is available for searching or downloading at www.tigr.org/TIGRFAMs.

Animals↗

BAGEL: a web-based bacteriocin genome mining tool.

A common problem in the annotation of open reading frames (ORFs) is the identification of genes that are functionally similar but have limited or no sequence homology. This is particularly the case for bacteriocins, a very diverse group of antimicrobial peptides produced by bacteria and usually encoded by small, poorly conserved ORFs. ORFs surrounding bacteriocin genes are often biosynthetic genes. This information can be used to locate putative structural bacteriocin genes. Here, we describe BAGEL, a web server that identifies putative bacteriocin ORFs in a DNA sequence using novel, knowledge-based bacteriocin databases and motif databases. Many bacteriocins are encoded by small genes that are often omitted in the annotation process of bacterial genomes. Thus, we have implemented ORF detection using a number of published ORF prediction tools. In addition, BAGEL takes into account the genomic context, i.e. for each potential bacteriocin-encoding ORF, the sequence of the surrounding region on the genome is analyzed for genes that might encode proteins involved in biosynthesis, transport, regulation and/or immunity. These innovations make BAGEL unique in its ability to detect putative bacteriocin gene clusters in (new) bacterial genomes. BAGEL is freely accessible at: http://bioinformatics.biol.rug.nl/websoftware/bagel.

Amino Acid Motifs↗

Saccharomyces Genome Database (SGD) provides secondary gene annotation using the Gene Ontology (GO).

The Saccharomyces Genome Database (SGD) resources, ranging from genetic and physical maps to genome-wide analysis tools, reflect the scientific progress in identifying genes and their functions over the last decade. As emphasis shifts from identification of the genes to identification of the role of their gene products in the cell, SGD seeks to provide its users with annotations that will allow relationships to be made between gene products, both within Saccharomyces cerevisiae and across species. To this end, SGD is annotating genes to the Gene Ontology (GO), a structured representation of biological knowledge that can be shared across species. The GO consists of three separate ontologies describing molecular function, biological process and cellular component. The goal is to use published information to associate each characterized S.cerevisiae gene product with one or more GO terms from each of the three ontologies. To be useful, this must be done in a manner that allows accurate associations based on experimental evidence, modifications to GO when necessary, and careful documentation of the annotations through evidence codes for given citations. Reaching this goal is an ongoing process at SGD. For information on the current progress of GO annotations at SGD and other participating databases, as well as a description of each of the three ontologies, please visit the GO Consortium page at http://www.geneontology.org. SGD gene associations to GO can be found by visiting our site at http://genome-www.stanford.edu/Saccharomyces/.

Animals↗

Prediction and functional analysis of native disorder in proteins from the three kingdoms of life.

An automatic method for recognizing natively disordered regions from amino acid sequence is described and benchmarked against predictors that were assessed at the latest critical assessment of techniques for protein structure prediction (CASP) experiment. The method attains a Wilcoxon score of 90.0, which represents a statistically significant improvement on the methods evaluated on the same targets at CASP. The classifier, DISOPRED2, was used to estimate the frequency of native disorder in several representative genomes from the three kingdoms of life. Putative, long (>30 residue) disordered segments are found to occur in 2.0% of archaean, 4.2% of eubacterial and 33.0% of eukaryotic proteins. The function of proteins with long predicted regions of disorder was investigated using the gene ontology annotations supplied with the Saccharomyces genome database. The analysis of the yeast proteome suggests that proteins containing disorder are often located in the cell nucleus and are involved in the regulation of transcription and cell signalling. The results also indicate that native disorder is associated with the molecular functions of kinase activity and nucleic acid binding.

Databases, Genetic↗

MeMo: a hybrid SQL/XML approach to metabolomic data management for functional genomics.

BACKGROUND: The genome sequencing projects have shown our limited knowledge regarding gene function, e.g. S. cerevisiae has 5-6,000 genes of which nearly 1,000 have an uncertain function. Their gross influence on the behaviour of the cell can be observed using large-scale metabolomic studies. The metabolomic data produced need to be structured and annotated in a machine-usable form to facilitate the exploration of the hidden links between the genes and their functions. DESCRIPTION: MeMo is a formal model for representing metabolomic data and the associated metadata. Two predominant platforms (SQL and XML) are used to encode the model. MeMo has been implemented as a relational database using a hybrid approach combining the advantages of the two technologies. It represents a practical solution for handling the sheer volume and complexity of the metabolomic data effectively and efficiently. The MeMo model and the associated software are available at http://dbkgroup.org/memo/. CONCLUSION: The maturity of relational database technology is used to support efficient data processing. The scalability and self-descriptiveness of XML are used to simplify the relational schema and facilitate the extensibility of the model necessitated by the creation of new experimental techniques. Special consideration is given to data integration issues as part of the systems biology agenda. MeMo has been physically integrated and cross-linked to related metabolomic and genomic databases. Semantic integration with other relevant databases has been supported through ontological annotation. Compatibility with other data formats is supported by automatic conversion.

Computer Simulation↗