Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

From fold predictions to function predictions: automation of functional site conservation analysis for functional genome predictions.

A database of functional sites for proteins with known structures, SITE, is constructed and used in conjunction with a simple pattern matching program SiteMatch to evaluate possible function conservation in a recently constructed database of fold predictions for Escherichia coli proteins (Rychlewski L et al., 1999, Protein Sci 8:614-624). In this and other prediction databases, fold predictions are based on algorithms that can recognize weak sequence similarities and putatively assign new proteins into already characterized protein families. It is not clear whether such sequence similarities arise from distant homologies or general similarity of physicochemical features along the sequence. Leaving aside the important question of nature of relations within fold superfamilies, it is possible to assess possible function conservation by looking at the pattern of conservation of crucial functional residues. SITE consists of a multilevel function description based on structure annotations and structure analyses. In particular, active site residues, ligand binding residues, and patterns of hydrophobic residues on the protein surface are used to describe different functional features. SiteMatch, a simple pattern matching program, is designed to check the conservation of residues involved in protein activity in alignments generated by any alignment method. Here, this procedure is used to study conservation of functional features in alignments between protein sequences from the E. coli genome and their optimal structural templates. The optimal templates were identified and alignments taken from the database of genomic structural predictions was described in a previous publication (Rychlewski L et al., 1999, Protein Sci 8:614-624). An automated assessment of function conservation is used to analyze the relation between fold and function similarity for a large number of fold predictions. For instance, it is shown that identifying low significance predictions with a high level of functional residue conservations can be used to extend the prediction sensitivity for fold prediction methods. Over 100 new fold/function predictions in this class were obtained in the E. coli genome. At the same time, about 30% of our previous fold predictions are not confirmed as function predictions, further highlighting the problem of function divergence in fold superfamilies.

Algorithms↗

The protein information resource (PIR).

The Protein Information Resource (PIR) produces the largest, most comprehensive, annotated protein sequence database in the public domain, the PIR-International Protein Sequence Database, in collaboration with the Munich Information Center for Protein Sequences (MIPS) and the Japan International Protein Sequence Database (JIPID). The expanded PIR WWW site allows sequence similarity and text searching of the Protein Sequence Database and auxiliary databases. Several new web-based search engines combine searches of sequence similarity and database annotation to facilitate the analysis and functional identification of proteins. New capabilities for searching the PIR sequence databases include annotation-sorted search, domain search, combined global and domain search, and interactive text searches. The PIR-International databases and search tools are accessible on the PIR WWW site at http://pir.georgetown.edu and at the MIPS WWW site at http://www. mips.biochem.mpg.de. The PIR-International Protein Sequence Database and other files are also available by FTP.

Databases, Factual↗

Identifying functional gene sets from hierarchically clustered expression data: map of abiotic stress regulated genes in Arabidopsis thaliana.

We present MultiGO, a web-enabled tool for the identification of biologically relevant gene sets from hierarchically clustered gene expression trees (http://ekhidna.biocenter.helsinki.fi/poxo/multigo). High-throughput gene expression measuring techniques, such as microarrays, are nowadays often used to monitor the expression of thousands of genes. Since these experiments can produce overwhelming amounts of data, computational methods that assist the data analysis and interpretation are essential. MultiGO is a tool that automatically extracts the biological information for multiple clusters and determines their biological relevance, and hence facilitates the interpretation of the data. Since the entire expression tree is analysed, MultiGO is guaranteed to report all clusters that share a common enriched biological function, as defined by Gene Ontology annotations. The tool also identifies a plausible cluster set, which represents the key biological functions affected by the experiment. The performance is demonstrated by analysing drought-, cold- and abscisic acid-related expression data sets from Arabidopsis thaliana. The analysis not only identified known biological functions, but also brought into focus the less established connections to defense-related gene clusters. Thus, in comparison to analyses of manually selected gene lists, the systematic analysis of every cluster can reveal unexpected biological phenomena and produce much more comprehensive biological insights to the experiment of interest.

Abscisic Acid↗

A new measure for functional similarity of gene products based on Gene Ontology.

BACKGROUND: Gene Ontology (GO) is a standard vocabulary of functional terms and allows for coherent annotation of gene products. These annotations provide a basis for new methods that compare gene products regarding their molecular function and biological role. RESULTS: We present a new method for comparing sets of GO terms and for assessing the functional similarity of gene products. The method relies on two semantic similarity measures; simRel and funSim. One measure (simRel) is applied in the comparison of the biological processes found in different groups of organisms. The other measure (funSim) is used to find functionally related gene products within the same or between different genomes. Results indicate that the method, in addition to being in good agreement with established sequence similarity approaches, also provides a means for the identification of functionally related proteins independent of evolutionary relationships. The method is also applied to estimating functional similarity between all proteins in Saccharomyces cerevisiae and to visualizing the molecular function space of yeast in a map of the functional space. A similar approach is used to visualize the functional relationships between protein families. CONCLUSION: The approach enables the comparison of the underlying molecular biology of different taxonomic groups and provides a new comparative genomics tool identifying functionally related gene products independent of homology. The proposed map of the functional space provides a new global view on the functional relationships between gene products or protein families.

Algorithms↗

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa↗

ChromCall: assigning chromatin status to defined genomic regions using epigenomic profiling data.

MOTIVATION: Chromatin regulation is crucial for modulating gene expression and cellular function by altering DNA accessibility. Defining and understanding chromatin regulation across diverse biological conditions, including health and disease, requires quantification of both the presence and enrichment level of diverse DNA-binding factors and chromatin modifications across defined genomic regions. Existing approaches mainly rely on peak-based or genome-wide models, which identify high-signal regions but do not annotate chromatin status at predefined functional genomic regions, such as promoters or enhancers. This lack of region-based annotation limits downstream comparative and integrative analyses across multiple factors and datasets, prompting us to create ChromCall. RESULTS: ChromCall is an R package for region-based chromatin enrichment analysis that provides a robust and extensible foundation for transparent and reproducible epigenomic profiling at predefined genomic regions. We applied ChromCall to ChIP-seq data from glioblastoma (GBM) brain tumours and found that the promoters of genes implicated in treatment resistance are significantly more likely to exhibit a combination of histone marks associated with phenotypic plasticity. This highlights a potential novel mechanism of therapeutic escape in these deadly tumours. AVAILABILITY AND IMPLEMENTATION: The R package is available on https://github.com/GliomaGenomics/ChromCall and the version used in this paper is archived at https://doi.org/10.5281/zenodo.19580967.

Chromatin↗

Http://C. elegans: mining the functional genomic landscape.

Caenorhabditis elegans is a powerful animal model for the study of functional genomics. The completed and well-annotated DNA sequence is available and a systematic study of gene function by RNA-interference-mediated knockdown of every gene is in progress. Full-genome DNA microarrays and DNA chips can be used to determine expression changes at different stages of development and in different mutant backgrounds, and a protein-interaction map based on the yeast two-hybrid approach is in progress. These high-capacity approaches to studying gene function will provide new insights into invertebrate and vertebrate biology.

Animals↗

The Candida Genome Database (CGD), a community resource for Candida albicans gene and protein information.

The Candida Genome Database (CGD) is a new database that contains genomic information about the opportunistic fungal pathogen Candida albicans. CGD is a public resource for the research community that is interested in the molecular biology of this fungus. CGD curators are in the process of combing the scientific literature to collect all C.albicans gene names and aliases; to assign gene ontology terms that describe the molecular function, biological process, and subcellular localization of each gene product; to annotate mutant phenotypes; and to summarize the function and biological context of each gene product in free-text description lines. CGD also provides community resources, including a reservation system for gene names and a colleague registry through which Candida researchers can share contact information and research interests. CGD is publicly funded (by NIH grant R01 DE15873-01 from the NIDCR) and is freely available at http://www.candidagenome.org/.

Candida albicans↗

Analysis of superfamily specific profile-profile recognition accuracy.

BACKGROUND: Annotation of sequences that share little similarity to sequences of known function remains a major obstacle in genome annotation. Some of the best methods of detecting remote relationships between protein sequences are based on matching sequence profiles. We analyse the superfamily specific performance of sequence profile-profile matching. Our benchmark consists of a set of 16 protein superfamilies that are highly diverse at the sequence level. We relate the performance to the number of sequences in the profiles, the profile diversity and the extent of structural conservation in the superfamily. RESULTS: The performance varies greatly between superfamilies with the truncated receiver operating characteristic, ROC10, varying from 0.95 down to 0.01. These large differences persist even when the profiles are trimmed to approximately the same level of diversity. CONCLUSIONS: Although the number of sequences in the profile (profile width) and degree of sequence variation within positions in the profile (profile diversity) contribute to accurate detection there are other superfamily specific factors.

Benchmarking↗

FuGE: Functional Genomics Experiment Object Model.

This is an interim report on the Functional Genomics Experiment (FuGE) Object Model. FuGE is a framework for creating data standards for high-throughput biological experiments, developed by a consortium of researchers from academia and industry. FuGE supports rich annotation of samples, protocols, instruments, and software, as well as providing extension points for technology specific details. It has been adopted by microarray and proteomics standards bodies as a basis for forthcoming standards. It is hoped that standards developers for other omics techniques will join this collaborative effort; widespread adoption will allow uniform annotation of common parts of functional genomics workflows, reduce standard development and learning times through the sharing of consistent practice, and ease the construction of software for accessing and integrating functional genomics data.

Computer Simulation↗

Genome-scale models of microbial cells: evaluating the consequences of constraints.

Microbial cells operate under governing constraints that limit their range of possible functions. With the availability of annotated genome sequences, it has become possible to reconstruct genome-scale biochemical reaction networks for microorganisms. The imposition of governing constraints on a reconstructed biochemical network leads to the definition of achievable cellular functions. In recent years, a substantial and growing toolbox of computational analysis methods has been developed to study the characteristics and capabilities of microorganisms using a constraint-based reconstruction and analysis (COBRA) approach. This approach provides a biochemically and genetically consistent framework for the generation of hypotheses and the testing of functions of microbial cells.

Bacterial Physiological Phenomena↗

A database and tools for 3-D protein structure comparison and alignment using the Combinatorial Extension (CE) algorithm.

The database reported here is derived using the Combinatorial Extension (CE) algorithm which compares pairs of protein polypeptide chains and provides a list of structurally similar proteins along with their structure alignments. Using CE, structure-structure alignments can provide insights into biological function. When a protein of known function is shown to be structurally similar to a protein of unknown function, a relationship might be inferred; a relationship not necessarily detectable from sequence comparison alone. Establishing structure-structure relationships in this way is of great importance as we enter an era of structural genomics where there is a likelihood of an increasing number of structures with unknown functions being determined. Thus the CE database is an example of a useful tool in the annotation of protein structures of unknown function. Comparisons can be performed on the complete PDB or on a structurally representative subset of proteins. The source protein(s) can be from the PDB (updated monthly) or uploaded by the user. CE provides sequence alignments resulting from structural alignments and Cartesian coordinates for the aligned structures, which may be analyzed using the supplied Compare3D Java applet, or downloaded for further local analysis. Searches can be run from the CE web site, http://cl.sdsc.edu/ce.html, or the database and software downloaded from the site for local use.

Algorithms↗

[Envisioning the inner body in Edo-Japan: Inshoku yojo kagami and Boji yojo kagami].

There are two ukiyoe, Japanese woodblock prints, presumed to have been produced around 1850 by the artist Utagawa Kunisada (1786-1864), or possibly, an understudy at his shop. One of the two ukiyoe, titled Inshoku yojo kagami (Rules of Dietary Life) shows a man drinking sake, holding a goblet in his hand. The other, titled Boji yojo kagami (Rules of Sexual Life) shows a woman, apparently a courtesan, holding a tobacco pipe to her mouth. These prints give a good picture of the images of the inside of human body, which were widely accepted among the common people after the end of the seventeenth century in the Edo period, because ukiyoe was a popular art produced by the common people in the Edo period, and the market for ukiyoe prints was primarily the general populace of the cities. The contrivance of the two Rules of Life prints lies in their fusion of two formats. One is the format of see-through body displaying the internal organs. The other is that of explaining the functions of the various internal organs in the form of familiar scenes from the living space of cities and households. Miniature sketches of people at work can be seen in them, performing the tasks believed to be that of each organ. By observing the work being carried out by the people, one could understand the organ's function. The purpose of the two annotated prints is explained in the notes as twofold. One was to educate viewers about the functions of the five viscera and six entrails, i.e., the principal inner organs in the traditional East Asian conception of the body. The other was to admonish them against excessive eating, drinking and sexual intercourse.

Anatomy↗

Proteome annotations and identifications of the human pulmonary fibroblast.

We hereby report on a three year project initiative undertaken by our research team encompassing large-scale protein expression profiling and annotations of human primary lung fibroblast cells. An overview is given of proteomic studies of the fibroblast target cell involved in several diseases such as asthma, idiopatic pulmonary disease, and COPD. It has been the objective within our research team to map and identify the protein expressions occurring in both activated-, as well as resting cell states. The JGGL database www.2DDB.org has been built around these data, allowing advanced hypothesis building using the interactive query bioinformatic tools developed. Gene ontology has been applied to these annotations, classifying and correlating protein expressions to function. The localization as well as the biological processes involved for the annotations are being presented including an annotation-, and sequence-identification strategy, resulting in close to 2000 protein identities. Both gel based, high resolution 2D-gels, and liquid-phase separation (three-dimensional HPLC), as well as the combination of gel- and LC-based approaches (1D-gels and nano-capillary LC, reversed-phase) were utilized. Protein sequencing and structure identities were acquired by a combination of MALDI-, and electrospray-mass spectrometry techniques. Phenotypical and morphological characterizations were also made for this human disease target cell in both stimulated- and resting-cell states. The use of functional assays that demonstrate the key regulating role of growth factors and cytokine stimuli such as PDGF, TGF-beta, and EGF and the effect of ECM molecules such as Biglycan, are also presented and discussed.

Amino Acid Sequence↗

Holter recordings with continuous marker annotations: a new tool in pacemaker diagnostics.

UNLABELLED: Pacemakers provide marker annotations to facilitate the interpretation of pacemaker electrocardiograms (ECGs) and can be used in cases of suspected pacemaker malfunction or to understand pacemaker behavior. Due to the need for a programmer, only short-term evaluations are possible. We evaluated a prototype Telemetry Data Logger (TDL) designed to continuously transfer markers from the pacemaker to a conventional Holter recorder. A miniaturized telemetry receiving coil was attached to patient's skin above the pacemaker, which was programmed to transmit markers continuously. The TDL, which receives and converts markers into eight positive and eight negative deflections, ranging from -2.5 to +2.5 mV in amplitude, was connected to one channel of a conventional Holter recorder (Tracker 2). We performed 20 Holters in 13 patients who had implanted VDDR or DDDR devices from the same manufacturer and evaluated three versions of software. Marker transmission was possible in all patients, producing Holter ECGs with complete marker annotations. Artifacts occurred < 4% of the time. A 50-ms rectangular pulse was optimal for marker interpretation. The device, which was easy to use and well accepted by the patients, assisted in the diagnosis of inappropriate pacemaker programming, even when the surface ECG seemed to show regular pacemaker function. In the presence of low quality surface ECGs, marker annotations allowed the assessment of pacemaker function. The capability to annotate the onset of special algorithms, like tachycardia termination algorithms or mode switching, facilitates interpretation of pacemaker behavior, enabling a reliable assessment of the appropriateness of such algorithms. CONCLUSION: The TDL effectively enables pacemaker markers to be inscribed onto a conventional Holter recording, facilitating the interpretation of pacemaker ECGs and the diagnosis of inappropriate pacemaker programming even when not discernible from the surface ECG alone.

Algorithms↗

The use of edge-betweenness clustering to investigate biological function in protein interaction networks.

BACKGROUND: This paper describes an automated method for finding clusters of interconnected proteins in protein interaction networks and retrieving protein annotations associated with these clusters. RESULTS: Protein interaction graphs were separated into subgraphs of interconnected proteins, using the JUNG implementation of Girvan and Newman's Edge-Betweenness algorithm. Functions were sought for these subgraphs by detecting significant correlations with the distribution of Gene Ontology terms which had been used to annotate the proteins within each cluster. The method was implemented using freely available software (JUNG and the R statistical package). Protein clusters with significant correlations to functional annotations could be identified and included groups of proteins know to cooperate in cell metabolism. The method appears to be resilient against the presence of false positive interactions. CONCLUSION: This method provides a useful tool for rapid screening of small to medium size protein interaction datasets.

Algorithms↗

Mitoproteome: human heart mitochondrial protein sequence database.

The human mitochondrial proteome database has been developed by deriving data from a combination of public repositories and experimental and computational prediction methods. The experimental data is derived from highly purified mitochondria from human heart tissue, whereas predictions have been performed by MITOPRED, a genome-scale method for the prediction of nucleus-encoded mitochondrial proteins. Mitochondrial protein sequences from different sources have been clustered to generate a nonredundant dataset. Annotations related to the protein function, structure, disease association, pathways, and so on are collected from a number of public databases using commonly used UNIX and Perl scripts. This chapter provides a detailed description of various data sources and methods used to download, curate, parse, and generate meaningful annotations from primary as well as derived databases.

Computational Biology↗

metaSHARK: software for automated metabolic network prediction from DNA sequence and its application to the genomes of Plasmodium falciparum and Eimeria tenella.

The metabolic SearcH And Reconstruction Kit (metaSHARK) is a new fully automated software package for the detection of enzyme-encoding genes within unannotated genome data and their visualization in the context of the surrounding metabolic network. The gene detection package (SHARKhunt) runs on a Linux system and requires only a set of raw DNA sequences (genomic, expressed sequence tag and/or genome survey sequence) as input. Its output may be uploaded to our web-based visualization tool (SHARKview) for exploring and comparing data from different organisms. We first demonstrate the utility of the software by comparing its results for the raw Plasmodium falciparum genome with the manual annotations available at the PlasmoDB and PlasmoCyc websites. We then apply SHARKhunt to the unannotated genome sequences of the coccidian parasite Eimeria tenella and observe that, at an E-value cut-off of 10(-20), our software makes 142 additional assertions of enzymatic function compared with a recent annotation package working with translated open reading frame sequences. The ability of the software to cope with low levels of sequence coverage is investigated by analyzing assemblies of the E.tenella genome at estimated coverages from 0.5x to 7.5x. Lastly, as an example of how metaSHARK can be used to evaluate the genomic evidence for specific metabolic pathways, we present a study of coenzyme A biosynthesis in P.falciparum and E.tenella.

Animals↗