Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Functional annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Structure-based functional annotation: yeast ymr099c codes for a D-hexose-6-phosphate mutarotase.

Despite the generation of a large amount of sequence information over the last decade, more than 40% of well characterized enzymatic functions still lack associated protein sequences. Assigning protein sequences to documented biochemical functions is an interesting challenge. We illustrate here that structural genomics may be a reasonable approach in addressing these questions. We present the crystal structure of the Saccharomyces cerevisiae YMR099cp, a protein of unknown function. YMR099cp adopts the same fold as galactose mutarotase and shares the same catalytic machinery necessary for the interconversion of the alpha and beta anomers of galactose. The structure revealed the presence in the active site of a sulfate ion attached by an arginine clamp made by the side chain from two strictly conserved arginine residues. This sulfate is ideally positioned to mimic the phosphate group of hexose 6-phosphate. We have subsequently successfully demonstrated that YMR099cp is a hexose-6-phosphate mutarotase with broad substrate specificity. We solved high resolution structures of some substrate enzyme complexes, further confirming our functional hypothesis. The metabolic role of a hexose-6-phosphate mutarotase is discussed. This work illustrates that structural information has been crucial to assign YMR099cp to the orphan EC activity: hexose-phosphate mutarotase.

Amino Acid Sequence↗

Functional annotation and analysis of Korean patented biological sequences using bioinformatics.

A recent report of the Korean Intellectual Property Office (KIPO) showed that the number of biological sequence-based patents is rapidly increasing in Korea. We present biological features of Korean patented sequences though bioinformatic analysis. The analysis is divided into two steps. The first is an annotation step in which the patented sequences were annotated with the Reference Sequence (RefSeq) database. The second is an association step in which the patented sequences were linked to genes, diseases, pathway, and biological functions. We used Entrez Gene, Online Mendelian Inheritance in Man (OMIM), Kyoto Encyclopedia of Genes and Genomes (KEGG), and Gene Ontology (GO) databases. Through the association analysis, we found that nearly 2.6% of human genes were associated with Korean patenting, compared to 20% of human genes in the U.S. patent. The association between the biological functions and the patented sequences indicated that genes whose products act as hormones on defense responses in the extra-cellular environments were the most highly targeted for patenting. The analysis data are available at http://www.patome.net.

Base Sequence↗

GeneTools--application for functional annotation and statistical hypothesis testing.

BACKGROUND: Modern biology has shifted from "one gene" approaches to methods for genomic-scale analysis like microarray technology, which allow simultaneous measurement of thousands of genes. This has created a need for tools facilitating interpretation of biological data in "batch" mode. However, such tools often leave the investigator with large volumes of apparently unorganized information. To meet this interpretation challenge, gene-set, or cluster testing has become a popular analytical tool. Many gene-set testing methods and software packages are now available, most of which use a variety of statistical tests to assess the genes in a set for biological information. However, the field is still evolving, and there is a great need for "integrated" solutions. RESULTS: GeneTools is a web-service providing access to a database that brings together information from a broad range of resources. The annotation data are updated weekly, guaranteeing that users get data most recently available. Data submitted by the user are stored in the database, where it can easily be updated, shared between users and exported in various formats. GeneTools provides three different tools: i) NMC Annotation Tool, which offers annotations from several databases like UniGene, Entrez Gene, SwissProt and GeneOntology, in both single- and batch search mode. ii) GO Annotator Tool, where users can add new gene ontology (GO) annotations to genes of interest. These user defined GO annotations can be used in further analysis or exported for public distribution. iii) eGOn, a tool for visualization and statistical hypothesis testing of GO category representation. As the first GO tool, eGOn supports hypothesis testing for three different situations (master-target situation, mutually exclusive target-target situation and intersecting target-target situation). An important additional function is an evidence-code filter that allows users, to select the GO annotations for the analysis. CONCLUSION: GeneTools is the first "all in one" annotation tool, providing users with a rapid extraction of highly relevant gene annotation data for e.g. thousands of genes or clones at once. It allows a user to define and archive new GO annotations and it supports hypothesis testing related to GO category representations. GeneTools is freely available through www.genetools.no

Algorithms↗

Functional annotation of putative aminoglycoside antibiotic modifying proteins in Mycobacterium tuberculosis H37Rv.

The growing availability of sequences of bacterial genomes has revealed a number of open reading frames predicted by sequence alignment to encode antibiotic resistance proteins. The presence of these putative resistance genes within bacterial genomes raises important questions regarding potential reservoirs of resistance elements and their evolution. Here we examine four gene products encoding predicted aminoglycoside-aminocyclitol antibiotic modifying enzymes, two phosphotransferases and two acetyltransferases, derived from analysis of the genome sequence of Mycobacterium tuberculosis strain H37Rv with the goal of assigning biochemical function by purification of each protein and characterization of their ability to modify aminoglycoside antibiotics. Only one of these enzymes, the previously characterized aminoglycoside acetyltransferase AAC(2')-Ic, displayed compelling aminoglycoside modifying activity. While the putative phosphotransferase encoded by the Rv3225c gene did display low levels of aminoglycoside kinase activity, the predicted kinase encoded by the Rv3817 gene lacked any such activity. A potential aminoglycoside 6'-acetyltransferase, encoded by the Rv1347c gene, did not show antibiotic acylation activity but did demonstrate selective thioesterase activity with numerous acyl-CoAs. This activity, together with the genomic environment of the Rv1347c gene in a likely polyketide synthesis cluster, suggests a role for this protein in secondary metabolism and not in antibiotic modification. It was thus shown that only one of four putative aminoglycosides modifying enzymes derived from the whole genome sequencing of M. tuberculosis H37Rv showed sufficient predicted enzyme activity to be annotated as an aminoglycoside resistance element. This study demonstrates the necessity of biochemical annotation methods as a follow up to in silico sequence alignment-based methods of assigning gene product function.

Acetyltransferases↗

Protein surface analysis for function annotation in high-throughput structural genomics pipeline.

Structural genomics (SG) initiatives are expanding the universe of protein fold space by rapidly determining structures of proteins that were intentionally selected on the basis of low sequence similarity to proteins of known structure. Often these proteins have no associated biochemical or cellular functions. The SG success has resulted in an accelerated deposition of novel structures. In some cases the structural bioinformatics analysis applied to these novel structures has provided specific functional assignment. However, this approach has also uncovered limitations in the functional analysis of uncharacterized proteins using traditional sequence and backbone structure methodologies. A novel method, named pvSOAR (pocket and void Surface of Amino Acid Residues), of comparing the protein surfaces of geometrically defined pockets and voids was developed. pvSOAR was able to detect previously unrecognized and novel functional relationships between surface features of proteins. In this study, pvSOAR is applied to several structural genomics proteins. We examined the surfaces of YecM, BioH, and RpiB from Escherichia coli as well as the CBS domains from inosine-5'-monosphate dehydrogenase from Streptococcus pyogenes, conserved hypothetical protein Ta549 from Thermoplasm acidophilum, and CBS domain protein mt1622 from Methanobacterium thermoautotrophicum with the goal to infer information about their biochemical function.

Adenine Nucleotides↗

Functional annotation of mouse mutations in embryonic stem cells by use of expression profiling.

Expression profiling offers a potential high-throughput phenotype screen for mutant mouse embryonic stem (ES) cells. We have assessed the ability of expression arrays to distinguish among heterozygous mutant ES cell lines and to accurately reflect the normal function of the mutated genes. Two ES cell lines hemizygous for overlapping regions of mouse Chromosome (Chr) 5 differed substantially from the wildtype parental line and from each other. Expression differences included frequent downregulation of hemizygous genes and downstream effects on genes mapping to other chromosomes. Some genes were affected similarly in each deletion line, consistent with the overlap of the deletions. To determine whether such downstream effects reveal pathways impacted by a mutation, we examined ES cell lines heterozygous for mutations in either of two well-characterized genes. A heterozygous mutation in the gene encoding the cell cycle regulator, cyclin D kinase 4 ( Cdk4), affected expression of many genes involved in cell growth and proliferation. A heterozygous mutation in the ATP binding cassette transporter family A, member 1 ( Abca1) gene, altered genes associated with lipid homeostasis, the cytoskeleton, and vesicle trafficking. Heterozygous Abca1 mutation had similar effects in liver, indicating that ES cell expression profile reflects changes in fundamental processes relevant to mutant gene function in multiple cell types.

ATP Binding Cassette Transporter 1↗

The Nuclear Protein Database (NPD): sub-nuclear localisation and functional annotation of the nuclear proteome.

The Nuclear Protein Database (NPD) is a curated database that contains information on more than 1300 vertebrate proteins that are thought, or are known, to localise to the cell nucleus. Each entry is annotated with information on predicted protein size and isoelectric point, as well as any repeats, motifs or domains within the protein sequence. In addition, information on the sub-nuclear localisation of each protein is provided and the biological and molecular functions are described using Gene Ontology (GO) terms. The database is searchable by keyword, protein name, sub-nuclear compartment and protein domain/motif. Links to other databases are provided (e.g. Entrez, SWISS-PROT, OMIM, PubMed, PubMed Central). Thus, NPD provides a gateway through which the nuclear proteome may be explored. The database can be accessed at http://npd.hgu.mrc.ac.uk and is updated monthly.

Amino Acid Sequence↗

Functional annotation of two orphan G-protein-coupled receptors, Drostar1 and -2, from Drosophila melanogaster and their ligands by reverse pharmacology.

By combining a Drosophila genome data base search and reverse transcriptase-PCR-based cDNA isolation, two G-protein-coupled receptors were cloned, which are the closest known invertebrate homologs of the mammalian opioid/somatostatin receptors. However, when functionally expressed in Xenopus oocytes by injection of Drosophila orphan receptor RNAs together with a coexpressed potassium channel, neither receptor was activated by known mammalian agonists. By applying a reverse pharmacological approach, the physiological ligands were isolated from peptide extracts from adult flies and larvae. Edman sequencing and mass spectrometry of the purified ligands revealed two decapentapeptides, which differ only by an N-terminal pyroglutamate/glutamine. The peptides align to a hormone precursor sequence of the Drosophila genome data base and are almost identical to allatostatin C from Manduca sexta. Both receptors were activated by the synthetic peptides irrespective of the N-terminal modification. Site-directed mutagenesis of a residue in transmembrane region 3 and the loop between transmembrane regions 6 and 7 affect ligand binding, as previously described for somatostatin receptors. The two receptor genes each containing three exons and transcribed in opposite directions are separated by 80 kb with no other genes predicted between. Localization of receptor transcripts identifies a role of the new transmitter system in visual information processing as well as endocrine regulation.

Amino Acid Sequence↗

Cross genome phylogenetic analysis of human and Drosophila G protein-coupled receptors: application to functional annotation of orphan receptors.

BACKGROUND: The cell-membrane G-protein coupled receptors (GPCRs) are one of the largest known superfamilies and are the main focus of intense pharmaceutical research due to their key role in cell physiology and disease. A large number of putative GPCRs are 'orphans' with no identified natural ligands. The first step in understanding the function of orphan GPCRs is to identify their ligands. Phylogenetic clustering methods were used to elucidate the chemical nature of receptor ligands, which led to the identification of natural ligands for many orphan receptors. We have clustered human and Drosophila receptors with known ligands and orphans through cross genome phylogenetic analysis and hypothesized higher relationship of co-clustered members that would ease ligand identification, as related receptors share ligands with similar structure or class. RESULTS: Cross-genome phylogenetic analyses were performed to identify eight major groups of GPCRs dividing them into 32 clusters of 371 human and 113 Drosophila proteins (excluding olfactory, taste and gustatory receptors) and reveal unexpected levels of evolutionary conservation across human and Drosophila GPCRs. We also observe that members of human chemokine receptors, involved in immune response, and most of nucleotide-lipid receptors (except opsins) do not have counterparts in Drosophila. Similarly, a group of Drosophila GPCRs (methuselah receptors), associated in aging, is not present in humans. CONCLUSION: Our analysis suggests ligand class association to 52 unknown Drosophila receptors and 95 unknown human GPCRs. A higher level of phylogenetic organization was revealed in which clusters with common domain architecture or cellular localization or ligand structure or chemistry or a shared function are evident across human and Drosophila genomes. Such analyses will prove valuable for identifying the natural ligands of Drosophila and human orphan receptors that can lead to a better understanding of physiological and pathological roles of these receptors.

Amino Acid Sequence↗

Dividing the large glycoside hydrolase family 13 into subfamilies: towards improved functional annotations of alpha-amylase-related proteins.

Family GH13, also known as the alpha-amylase family, is the largest sequence-based family of glycoside hydrolases and groups together a number of different enzyme activities and substrate specificities acting on alpha-glycosidic bonds. This polyspecificity results in the fact that the simple membership of this family cannot be used for the prediction of gene function based on sequence alone. In order to establish robust groups that show an improved correlation between sequence and enzymatic specificity, we have performed a large-scale analysis of 1691 family GH13 sequences by combining clustering, similarity search and phylogenetic methods. About 80% of the sequences could be reliably classified into 35 subfamilies. Most subfamilies appear monofunctional (i.e. contain enzymes with the same substrate and the same product). The close examination of the other, apparently polyspecific, subfamilies revealed that they actually group together enzymes with strongly related (or even sometimes virtually identical) activities. Overall our subfamily assignment allows to set the limits for genomic function prediction on this large family of biologically and industrially important enzymes.

Amino Acid Sequence↗

Functional annotation of proteins identified in human brain during the HUPO Brain Proteome Project pilot study.

The HUPO Brain Proteome Project is an initiative coordinating proteomics studies to characterise human and mouse brain proteomes. Proteins identified in human brain samples during the project's pilot phase were put into biological context through integration with various annotation sources followed by a bioinformatics analysis. The data set was related to the genome sequence via the genes encoding identified proteins including an assessment of splice variant identification as well as an analysis of tissue specificity of the respective transcripts. Proteins were furthermore categorised according to subcellular localisation, molecular function and biological process, grouped into protein families and mapped to biological pathways they are known to act in. Involvement in pathological conditions was examined based on association with entries in the online version of Mendelian Inheritance in Man and an interaction network was derived from curated protein-proteininteraction data. Overall a non-redundant set of 1804 proteins was identified in human brain samples. In the majority of cases splice variants could be unambiguously identified by unique peptides, including matches to several hypothetical transcripts of known as well as predicted genes.

Alternative Splicing↗

Whole-Genome Sequence Dataset of Rhodococcus qingshengii IEGM 267-Terpenoid Biotransformer Toward Genetic Functional Annotation.

Background/Objectives: Microbial biotransformation of monoterpenoids is a promising approach for obtaining bioactive compounds. Rhodococcus species are attractive biocatalysts due to their metabolic versatility and ability to transform hydrophobic substrates. In this study, we investigated the catalytic potential of Rhodococcus qingshengii IEGM 267 toward carveol isomers and explored genomic features that may underlie this activity. Methods: The strain was cultivated in mineral medium supplemented with (-)-trans-carveol. Biotransformation products were analyzed by TLC and GC-MS. The draft genome was sequenced, assembled, taxonomically assigned, and annotated using standard bioinformatics tools. Results: Rhodococcus qingshengii IEGM 267 efficiently converted (-)-trans-carveol to carvone. Genome analysis confirmed the taxonomic assignment of the strain and revealed a large repertoire of oxidoreductases, including monooxygenases, hydroxylases, and dehydrogenases. Seven genes encoding cytochrome P450-dependent oxygenases were identified as candidate enzymes potentially involved in carveol oxidation. Conclusions: R. qingshengii IEGM 267 is an efficient and stereoselective biocatalyst for (-)-trans-carveol oxidation. The results of bioinformatics analysis suggest an alternative enzymatic basis for this transformation and provide a foundation for future functional characterization.

Rhodococcus↗

Sequencing and functional annotation of the Bacillus subtilis genes in the 200 kb rrnB-dnaB region.

The 200 kb region of the Bacillus subtilis chromosome spanning from 255 to 275 degrees on the genetic map was sequenced. The strategy applied, based on use of yeast artificial chromosomes and multiplex Long Accurate PCR, proved to be very efficient for sequencing a large bacterial chromosome area. A total of 193 genes of this part of the chromosome was classified by level of knowledge and biological category of their functions. Five levels of gene function understanding are defined. These are: (i) experimental evidence is available of gene product or biological function; (ii) strong homology exists for the putative gene product with proteins from other organisms; (iii) some indication of the function can be derived from homologies with known proteins; (iv) the gene product can be clustered with hypothetical proteins; (v) no indication on the gene function exists. The percentage of detected genes in each category was: 20, 28, 20, 15 and 17, respectively. In the sequenced region, a high percentage of genes are implicated in transport and metabolic linking of glycolysis and the citric acid cycle. A functional connection of several genes from this region and the genes close to 140 degrees in the chromosome was also observed.

Bacillus subtilis↗

Functional annotation and network reconstruction through cross-platform integration of microarray data.

The rapid accumulation of microarray data translates into a need for methods to effectively integrate data generated with different platforms. Here we introduce an approach, 2(nd)-order expression analysis, that addresses this challenge by first extracting expression patterns as meta-information from each data set (1(st)-order expression analysis) and then analyzing them across multiple data sets. Using yeast as a model system, we demonstrate two distinct advantages of our approach: we can identify genes of the same function yet without coexpression patterns and we can elucidate the cooperativities between transcription factors for regulatory network reconstruction by overcoming a key obstacle, namely the quantification of activities of transcription factors. Experiments reported in the literature and performed in our lab support a significant number of our predictions.

Algorithms↗

Array2BIO: from microarray expression data to functional annotation of co-regulated genes.

BACKGROUND: There are several isolated tools for partial analysis of microarray expression data. To provide an integrative, easy-to-use and automated toolkit for the analysis of Affymetrix microarray expression data we have developed Array2BIO, an application that couples several analytical methods into a single web based utility. RESULTS: Array2BIO converts raw intensities into probe expression values, automatically maps those to genes, and subsequently identifies groups of co-expressed genes using two complementary approaches: (1) comparative analysis of signal versus control and (2) clustering analysis of gene expression across different conditions. The identified genes are assigned to functional categories based on Gene Ontology classification and KEGG protein interaction pathways. Array2BIO reliably handles low-expressor genes and provides a set of statistical methods for quantifying expression levels, including Benjamini-Hochberg and Bonferroni multiple testing corrections. An automated interface with the ECR Browser provides evolutionary conservation analysis for the identified gene loci while the interconnection with Crème allows prediction of gene regulatory elements that underlie observed expression patterns. CONCLUSION: We have developed Array2BIO - a web based tool for rapid comprehensive analysis of Affymetrix microarray expression data, which also allows users to link expression data to Dcode.org comparative genomics tools and integrates a system for translating co-expression data into mechanisms of gene co-regulation. Array2BIO is publicly available at http://array2bio.dcode.org.

Algorithms↗

Functional annotation of IFN-alpha-stimulated gene expression profiles from sensitive and resistant renal cell carcinoma cell lines.

The antiproliferative, antiviral, and immunomodulatory properties of interferons (IFNs) have led to its therapeutic implementation. IFNs effects are mediated by a complex network of signal transducers, culminating in IFN-stimulated gene (ISG) induction. This complexity leads to diverse clinical responses to IFN, from no response to complete regression of disease. Elucidation of ISG induction patterns is, therefore, essential to understand and maximize its therapeutic potential. To correlate ISG expression profiles with IFN responsiveness, two renal cell carcinoma (RCC) cell lines differing in antiviral and apoptotic response to IFN were treated with IFN-alpha for different times, and expression profiles were analyzed using a customized microarray containing 850 unique putative ISGs. Genes with similar kinetics of induction in both cell lines were clustered and analyzed for gene function. Seven sets of coordinately regulated genes were identified by k-means cluster analysis, and significant functional similarities were identified for five of the seven sets. Strikingly, expression of genes associated with transcription temporally preceded expression of those involved in signal transduction. Enhanced antiviral sensitivity to IFN was coincident with sustained expression of ISGs involved in transcriptional regulation. However, no difference in Stat1 activation was observed between the cell lines. Analysis of ISG expression patterns suggests that subtle differences in transcription profiles contribute to differences in IFN responsiveness.

Antineoplastic Agents↗

Analysis and functional annotation of expressed sequence tags from the fall armyworm Spodoptera frugiperda.

BACKGROUND: Little is known about the genome sequences of lepidopteran insects, although this group of insects has been studied extensively in the fields of endocrinology, development, immunity, and pathogen-host interactions. In addition, cell lines derived from Spodoptera frugiperda and other lepidopteran insects are routinely used for baculovirus foreign gene expression. This study reports the results of an expressed sequence tag (EST) sequencing project in cells from the lepidopteran insect S. frugiperda, the fall armyworm. RESULTS: We have constructed an EST database using two cDNA libraries from the S. frugiperda-derived cell line, SF-21. The database consists of 2,367 ESTs which were assembled into 244 contigs and 951 singlets for a total of 1,195 unique sequences. CONCLUSION: S. frugiperda is an agriculturally important pest insect and genomic information will be instrumental for establishing initial transcriptional profiling and gene function studies, and for obtaining information about genes manipulated during infections by insect pathogens such as baculoviruses.

Animals↗

Functional Annotation of the Major Histocompatibility Complex Locus.

The human major histocompatibility complex (MHC) locus has the greatest density of disease-associations in the human genome, including links to over 100 polygenic disorders. Its complex haplotype structure, rich gene density, and high degree of linkage disequilibrium combine to make deciphering the gene regulatory logic of the MHC locus extremely challenging. Employing complementary high-throughput CRISPR interference (CRISPRi) and activation (CRISPRa) epigenetic screens coupled with single-cell transcriptome profiling across three distinct human cell types, we identified hundreds of new connections between cis -regulatory elements (CREs) and their target genes in this locus. These CRE-gene links are largely cell type-specific and act as enhancers. Additionally, some CREs have complex features, including harboring both active and repressive histone marks, lacking chromatin accessibility, targeting multiple genes, or acting as silencers. Computational methods fail to predict a majority of these CRE-gene connections. These findings emphasize the potential for functional perturbation experiments to dissect complex loci and reveal shared and cell type-specific regulatory mechanisms relevant to genomics of complex diseases. Collectively, this study provides a unique resource for understanding the complex regulatory landscape within the MHC locus and supports the need for creating new models that encompass CRE-gene interactions, cell type-specific gene expression, and disease genetics in the noncoding genome.

Journal Article↗