Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Probabilistic annotation of protein sequences based on functional classifications.

BACKGROUND: One of the most evident achievements of bioinformatics is the development of methods that transfer biological knowledge from characterised proteins to uncharacterised sequences. This mode of protein function assignment is mostly based on the detection of sequence similarity and the premise that functional properties are conserved during evolution. Most automatic approaches developed to date rely on the identification of clusters of homologous proteins and the mapping of new proteins onto these clusters, which are expected to share functional characteristics. RESULTS: Here, we inverse the logic of this process, by considering the mapping of sequences directly to a functional classification instead of mapping functions to a sequence clustering. In this mode, the starting point is a database of labelled proteins according to a functional classification scheme, and the subsequent use of sequence similarity allows defining the membership of new proteins to these functional classes. In this framework, we define the Correspondence Indicators as measures of relationship between sequence and function and further formulate two Bayesian approaches to estimate the probability for a sequence of unknown function to belong to a functional class. This approach allows the parametrisation of different sequence search strategies and provides a direct measure of annotation error rates. We validate this approach with a database of enzymes labelled by their corresponding four-digit EC numbers and analyse specific cases. CONCLUSION: The performance of this method is significantly higher than the simple strategy consisting in transferring the annotation from the highest scoring BLAST match and is expected to find applications in automated functional annotation pipelines.

Algorithms↗

MutDB: annotating human variation with functionally relevant data.

SUMMARY: We have developed a resource, MutDB (http://mutdb.org/), to aid in determining which single nucleotide polymorphisms (SNPs) are likely to alter the function of their associated protein product. MutDB contains protein structure annotations and comparative genomic annotations for 8000 disease-associated mutations and SNPs found in the UCSC Annotated Genome and the human RefSeq gene set. MutDB provides interactive mutation maps at the gene and protein levels, and allows for ranking of their predicted functional consequences based on conservation in multiple sequence alignments. AVAILABILITY: http://mutdb.org/ SUPPLEMENTARY INFORMATION: http://mutdb.org/about/about.html

Database Management Systems↗

Linking enzyme sequence to function using Conserved Property Difference Locator to identify and annotate positions likely to control specific functionality.

BACKGROUND: Families of homologous enzymes evolved from common progenitors. The availability of multiple sequences representing each activity presents an opportunity for extracting information specifying the functionality of individual homologs. We present a straightforward method for the identification of residues likely to determine class specific functionality in which multiple sequence alignments are converted to an annotated graphical form by the Conserved Property Difference Locator (CPDL) program. RESULTS: Three test cases, each comprised of two groups of functionally-distinct homologs, are presented. Of the test cases, one is a membrane and two are soluble enzyme families. The desaturase/hydroxylase data was used to design and test the CPDL algorithm because a comparative sequence approach had been successfully applied to manipulate the specificity of these enzymes. The other two cases, ATP/GTP cyclases, and MurD/MurE synthases were chosen because they are well characterized structurally and biochemically. For the desaturase/hydroxylase enzymes, the ATP/GTP cyclases and the MurD/MurE synthases, groups of 8 (of approximately 400), 4 (of approximately 150) and 10 (of >400) residues, respectively, of interest were identified that contain empirically defined specificity determining positions. CONCLUSION: CPDL consistently identifies positions near enzyme active sites that include those predicted from structural and/or biochemical studies to be important for specificity and/or function. This suggests that CPDL will have broad utility for the identification of potential class determining residues based on multiple sequence analysis of groups of homologous proteins. Because the method is sequence, rather than structure, based it is equally well suited for designing structure-function experiments to investigate membrane and soluble proteins.

Algorithms↗

Mycoplasma genes: a case for reflective annotation.

Although function can be assigned to genome sequence by homology at a macroscopic level, this can be misleading in the absence of data on enzyme activities. Together, such data can reveal whether open reading frames are expressed, identify multienzyme function and point to 'orphan' function. Because of their small size and small genomes, the genome sequences of some Mycoplasma spp. are very amenable to detailed analyses.

Genes, Bacterial↗

Large language models improve annotation of prokaryotic viral proteins.

Viral genomes are poorly annotated in metagenomic samples, representing an obstacle to understanding viral diversity and function. Current annotation approaches rely on alignment-based sequence homology methods, which are limited by the paucity of characterized viral proteins and divergence among viral sequences. Here we show that protein language models can capture prokaryotic viral protein function, enabling new portions of viral sequence space to be assigned biologically meaningful labels. When applied to global ocean virome data, our classifier expanded the annotated fraction of viral protein families by 29%. Among previously unannotated sequences, we highlight the identification of an integrase defining a mobile element in marine picocyanobacteria and a capsid protein that anchors globally widespread viral elements. Furthermore, improved high-level functional annotation provides a means to characterize similarities in genomic organization among diverse viral sequences. Protein language models thus enhance remote homology detection of viral proteins, serving as a useful complement to existing approaches.

Viral Proteins↗

Systematic identification of functional orthologs based on protein network comparison.

Annotating protein function across species is an important task that is often complicated by the presence of large paralogous gene families. Here, we report a novel strategy for identifying functionally related proteins that supplements sequence-based comparisons with information on conserved protein-protein interactions. First, the protein interaction networks of two species are aligned by assigning proteins to sequence homology clusters using the Inparanoid algorithm. Next, probabilistic inference is performed on the aligned networks to identify pairs of proteins, one from each species, that are likely to retain the same function based on conservation of their interacting partners. Applying this method to Drosophila melanogaster and Saccharomyces cerevisiae, we analyze 121 cases for which functional orthology assignment is ambiguous when sequence similarity is used alone. In 61 of these cases, the network supports a different protein pair than that favored by sequence comparisons. These results suggest that network analysis can be used to provide a key source of information for refining sequence-based homology searches.

Animals↗

Automated methods of predicting the function of biological sequences using GO and BLAST.

BACKGROUND: With the exponential increase in genomic sequence data there is a need to develop automated approaches to deducing the biological functions of novel sequences with high accuracy. Our aim is to demonstrate how accuracy benchmarking can be used in a decision-making process evaluating competing designs of biological function predictors. We utilise the Gene Ontology, GO, a directed acyclic graph of functional terms, to annotate sequences with functional information describing their biological context. Initially we examine the effect on accuracy scores of increasing the allowed distance between predicted and a test set of curator assigned terms. Next we evaluate several annotator methods using accuracy benchmarking. Given an unannotated sequence we use the Basic Local Alignment Search Tool, BLAST, to find similar sequences that have already been assigned GO terms by curators. A number of methods were developed that utilise terms associated with the best five matching sequences. These methods were compared against a benchmark method of simply using terms associated with the best BLAST-matched sequence (best BLAST approach). RESULTS: The precision and recall of estimates increases rapidly as the amount of distance permitted between a predicted term and a correct term assignment increases. Accuracy benchmarking allows a comparison of annotation methods. A covering graph approach performs poorly, except where the term assignment rate is high. A term distance concordance approach has a similar accuracy to the best BLAST approach, demonstrating lower precision but higher recall. However, a discriminant function method has higher precision and recall than the best BLAST approach and other methods shown here. CONCLUSION: Allowing term predictions to be counted correct if closely related to a correct term decreases the reliability of the accuracy score. As such we recommend using accuracy measures that require exact matching of predicted terms with curator assigned terms. Furthermore, we conclude that competing designs of BLAST-based GO term annotators can be effectively compared using an accuracy benchmarking approach. The most accurate annotation method was developed using data mining techniques. As such we recommend that designers of term annotators utilise accuracy benchmarking and data mining to ensure newly developed annotators are of high quality.

Benchmarking↗

Biological function made crystal clear - annotation of hypothetical proteins via structural genomics.

Many of the gene products of completely sequenced organisms are 'hypothetical' - they cannot be related to any previously characterized proteins - and so are of completely unknown function. Structural studies provide one means of obtaining functional information in these cases. A 'structural genomics' project has been initiated aimed at determining the structures of 50 hypothetical proteins from Haemophilus influenzae to gain an understanding of their function. Each stage of the project - target selection, protein production, crystallization, structure determination, and structure analysis - makes use of recent advances to streamline procedures. Early results from this and similar projects are encouraging in that some level of functional understanding can be deduced from experimentally solved structures.

Bacterial Proteins↗

Predicting function: from genes to genomes and back.

Predicting function from sequence using computational tools is a highly complicated procedure that is generally done for each gene individually. This review focuses on the added value that is provided by completely sequenced genomes in function prediction. Various levels of sequence annotation and function prediction are discussed, ranging from genomic sequence to that of complex cellular processes. Protein function is currently best described in the context of molecular interactions. In the near future it will be possible to predict protein function in the context of higher order processes such as the regulation of gene expression, metabolic pathways and signalling cascades. The analysis of such higher levels of function description uses, besides the information from completely sequenced genomes, also the additional information from proteomics and expression data. The final goal will be to elucidate the mapping between genotype and phenotype.

Bacterial Proteins↗

Evolutionary rate variation in eukaryotic lineage specific human intronless proteins.

The present study examines 783 human-mouse orthologous gene pairs for their pattern of sequence evolution, contrasting mammalia, eukaryota, coelomata, and bilateria specific human intronless genes. Such comparisons may be of use in understanding the general evolution of human genome. Evolutionary rate analyses indicate that mammalia specific human intronless genes are evolving faster as compared to other intronless genes specific to eukaryotic lineage, indicating towards their rapid evolution. The observations indicates that the genes conserved in eukaryota, coelomata, and bilateria, that is, proteins that arose earlier in evolution as compared to mammalia specific genes evolve slowly and are subjected to negative selection. The cause underlying rate variations was also explored. Although mutational bias might slightly fasten the nonsynonymous rates in mammalia specific genes, it is unlikely to be major cause of rate difference between the various categories. Furthermore, rate of divergence of mammalia specific intronless genes has been related to functional classification using the protein family annotation. Protein function was found in some cases to have larger impact on the rate of evolution of genes. Also, the codon usage pattern of mammalia specific intronless genes do not seem to differ much from those of other intronless genes conserved solely in eukaryotic lineage.

Animals↗

EUCLID: automatic classification of proteins in functional classes by their database annotations.

UNLABELLED: A tool is described for the automatic classification of sequences in functional classes using their database annotations. The Euclid system is based on a simple learning procedure from examples provided by human experts. AVAILABILITY: Euclid is freely available for academics at http://www.gredos.cnb.uam.es/EUCLID, with the corresponding dictionaries for the generation of three, eight and 14 functional classes. CONTACT: E-mail: valencia@cnb.uam.es SUPPLEMENTARY INFORMATION: The results of the EUCLID classification of different genomes are available at http://www.sander.ebi.ac. uk/genequiz/. A detailed description of the different applications mentioned in the text is available at http://www.gredos.cnb.uam. es/EUCLID/Full_Paper

Computational Biology↗

Microarray profiling of human white adipose tissue after exogenous leptin injection.

BACKGROUND: Leptin is a secreted adipocyte hormone that plays a key role in the regulation of body weight homeostasis. The leptin effect on human white adipose tissue (WAT) is still debated. OBJECTIVE: The aim of this study was to assess whether the administration of polyethylene glycol-leptin (PEG-OB) in a single supraphysiological dose has transcriptional effects on genes of WAT and to identify its target genes and functional pathways in WAT. MATERIALS AND METHODS: Blood samples and WAT biopsies were obtained from 10 healthy nonobese men before treatment and 72 h after the PEG-OB injection, leading to an approximate 809-fold increase in circulating leptin. The WAT gene expression profile before and after the PEG-OB injection was compared using pangenomic microarrays. Functional gene annotations based on the gene ontology of the PEG-OB regulated genes were performed using both an 'in house' automated procedure and GenMAPP (Gene Microarray Pathway Profiler), designed for viewing and analyzing gene expression data in the context of biological pathways. RESULTS: Statistical analysis of microarray data revealed that PEG-OB had a major down-regulated effect on WAT gene expression, as we obtained 1,822 and 100 down- and up-regulated genes, respectively. Microarray data were validated using reverse transcription quantitative PCR. Functional gene annotations of PEG-OB regulated genes revealed that the functional class related to immunity and inflammation was among the most mobilized PEG-OB pathway in WAT. These genes are mainly expressed in the cell of the stroma vascular fraction in comparison with adipocytes. CONCLUSION: Our observations support the hypothesis that leptin could act on WAT, particularly on genes related to inflammation and immunity, which may suggest a novel leptin target pathway in human WAT.

Adipocytes↗

Automated structure-based prediction of functional sites in proteins: applications to assessing the validity of inheriting protein function from homology in genome annotation and to protein docking.

A major problem in genome annotation is whether it is valid to transfer the function from a characterised protein to a homologue of unknown activity. Here, we show that one can employ a strategy that uses a structure-based prediction of protein functional sites to assess the reliability of functional inheritance. We have automated and benchmarked a method based on the evolutionary trace approach. Using a multiple sequence alignment, we identified invariant polar residues, which were then mapped onto the protein structure. Spatial clusters of these invariant residues formed the predicted functional site. For 68 of 86 proteins examined, the method yielded information about the observed functional site. This algorithm for functional site prediction was then used to assess the validity of transferring the function between homologues. This procedure was tested on 18 pairs of homologous proteins with unrelated function and 70 pairs of proteins with related function, and was shown to be 94 % accurate. This automated method could be linked to schemes for genome annotation. Finally, we examined the use of functional site prediction in protein-protein and protein-DNA docking. The use of predicted functional sites was shown to filter putative docked complexes with a discrimination similar to that obtained by manually including biological information about active sites or DNA-binding residues.

Algorithms↗

The comparative metabolism of the mollicutes (Mycoplasmas): the utility for taxonomic classification and the relationship of putative gene annotation and phylogeny to enzymatic function in the smallest free-living cells.

Mollicutes or mycoplasmas are a class of wall-less bacteria descended from low G + C% Gram-positive bacteria. Some are exceedingly small, about 0.2 micron in diameter, and are examples of the smallest free-living cells known. Their genomes are equally small; the smallest in Mycoplasma genitalium is sequenced and is 0.58 mb with 475 ORFs, compared with 4.639 mb and 4288 ORFs for Escherichia coli. Because of their size and apparently limited metabolic potential, Mollicutes are models for describing the minimal metabolism necessary to sustain independent life. Mollicutes have no cytochromes or the TCA cycle except for malate dehydrogenase activity. Some uniquely require cholesterol for growth, some require urea and some are anaerobic. They fix CO2 in anaplerotic or replenishing reactions. Some require pyrophosphate not ATP as an energy source for reactions, including the rate-limiting step of glycolysis: 6-phosphofructokinase. They scavenge for nucleic acid precursors and apparently do not synthesize pyrimidines or purines de novo. Some genera uniquely lack dUTPase activity and some species also lack uracil-DNA glycosylase. The absence of the latter two reactions that limit the incorporation of uracil or remove it from DNA may be related to the marked mutability of the Mollicutes and their tachytelic or rapid evolution. Approximately 150 cytoplasmic activities have been identified in these organisms, 225 to 250 are presumed to be present. About 100 of the core reactions are graphically linked in a metabolic map, including glycolysis, pentose phosphate pathway, arginine dihydrolase pathway, transamination, and purine, pyrimidine, and lipid metabolism. Reaction sequences or loci of particular importance are also described: phosphofructokinases, NADH oxidase, thioredoxin complex, deoxyribose-5-phosphate aldolase, and lactate, malate, and glutamate dehydrogenases. Enzymatic activities of the Mollicutes are grouped according to metabolic similarities that are taxonomically discriminating. The arrangements attempt to follow phylogenetic relationships. The relationships of putative gene assignments and enzymatic function in My. genitalium, My. pneumoniae, and My. capricolum subsp. capricolum are specially analyzed. The data are arranged in four tables. One associates gene annotations with congruent reports of the enzymatic activity in these same Mollicutes, and hence confirms the annotations. Another associates putative annotations with reports of the enzyme activity but from different Mollicutes. A third identifies the discrepancies represented by those enzymatic activities found in Mollicutes with sequenced genomes but without any similarly annotated ORF. This suggests that the gene sequence is significantly different from those already deposited in the databanks and putatively annotated with the same function. Another comparison lists those enzymatic activities that are both undetected in Mollicutes and not associated with any ORF. Evidence is presented supporting the theory that there are relatively small gene sequences that code for functional centers of multiple enzymatic activity. This property is seemingly advantageous for an organism with a small genome and perhaps under some coding restraint. The data suggest that a concept of "remnant" or "useless genes" or "useless enzymes" should be considered when examining the relationship of gene annotation and enzymatic function. It also suggests that genes in addition to representing what cells are doing or what they may do, may also identify what they once might have done and may never do again.

Adenosine Triphosphate↗

Quantitative analysis of bristle number in Drosophila mutants identifies genes involved in neural development.

BACKGROUND: The identification of the function of all genes that contribute to specific biological processes and complex traits is one of the major challenges in the postgenomic era. One approach is to employ forward genetic screens in genetically tractable model organisms. In Drosophila melanogaster, P element-mediated insertional mutagenesis is a versatile tool for the dissection of molecular pathways, and there is an ongoing effort to tag every gene with a P element insertion. However, the vast majority of P element insertion lines are viable and fertile as homozygotes and do not exhibit obvious phenotypic defects, perhaps because of the tendency for P elements to insert 5' of transcription units. Quantitative genetic analysis of subtle effects of P element mutations that have been induced in an isogenic background may be a highly efficient method for functional genome annotation. RESULTS: Here, we have tested the efficacy of this strategy by assessing the extent to which screening for quantitative effects of P elements on sensory bristle number can identify genes affecting neural development. We find that such quantitative screens uncover an unusually large number of genes that are known to function in neural development, as well as genes with yet uncharacterized effects on neural development, and novel loci. CONCLUSIONS: Our findings establish the use of quantitative trait analysis for functional genome annotation through forward genetics. Similar analyses of quantitative effects of P element insertions will facilitate our understanding of the genes affecting many other complex traits in Drosophila.

Animals↗

PDBSite: a database of the 3D structure of protein functional sites.

The PDBSite database provides comprehensive structural and functional information on various protein sites (post-translational modification, catalytic active, organic and inorganic ligand binding, protein-protein, protein-DNA and protein-RNA interactions) in the Protein Data Bank (PDB). The PDBSite is available online at http://wwwmgs.bionet.nsc.ru/mgs/gnw/pdbsite/. It consists of functional sites extracted from PDB using the SITE records and of an additional set containing the protein interaction sites inferred from the contact residues in heterocomplexes. The PDBSite was set up by automated processing of the PDB. The PDBSite database can be queried through the functional description and the structural characteristics of the site and its environment. The PDBSite is integrated with the PDBSiteScan tool allowing structural comparisons of a protein against the functional sites. The PDBSite enables the recognition of functional sites in protein tertiary structures, providing annotation of function through structure. The PDBSite is updated after each new PDB release.

Binding Sites↗

Mining gene expression data based on template theory.

MOTIVATION: It is understood that clustering genes are useful for exploring scientific knowledge from DNA microarray gene expression data. The explored knowledge can be finally used for annotating biological function for novel genes. Representing the explored knowledge in an efficient manner is then closely related to the classification accuracy. However, this issue has not yet been paid the attention it deserves. RESULT: A novel method based on template theory in cognitive psychology and pattern recognition is developed in this study for representing knowledge extracted from cluster analysis effectively. The basic principle is to represent knowledge according to the relationship between genes and a found cluster structure. Based on this novel knowledge representation method, a pattern recognition algorithm (the decision tree algorithm C4.5) is then used to construct a classifier for annotating biological functions of novel genes. The experiments on five published datasets show that this method has improved the classification performance compared with the conventional method. The statistical tests indicate that this improvement is significant. AVAILABILITY: The software package can be obtained upon request from the author.

Algorithms↗

Comprehensive profiling of antibiotic resistance genes and functional clusters of orthologous groups annotation of gut microbiota in Indonesian Kedu chickens.

Antibiotic resistance is a growing global health concern, with poultry systems acting as important reservoirs of antibiotic resistance genes (ARGs). However, resistome and functional profiles of indigenous chickens raised under traditional systems remain underexplored. This study aimed to characterize the antibiotic resistome, virulence factor genes, and metabolic potential of gut microbiota in Indonesian Kedu chickens using a shotgun metagenomic approach. Digesta samples from five gastrointestinal segments of 21 healthy adult chickens were analyzed through high-throughput sequencing. ARGs were identified using the Comprehensive Antibiotic Resistance Database (CARD) and Antibiotic Resistance Genes Databases (ARDB), while virulence factors and functional genes were annotated using Virulence Factor Database (VFDB), Clusters of Orthologous Groups (COG), and Carbohydrate-Active EnZymes (CAZy) databases. Results revealed a diverse resistome dominated by multidrug resistance and efflux pump mechanisms, with prominent genes associated with fluoroquinolone, tetracycline, β-lactam, and glycopeptide resistance. The detection of clinically relevant ARGs suggests that genetic determinants associated with antimicrobial resistance are present in the gut microbiota of traditionally raised Kedu chickens, although metagenomic data alone cannot determine whether these genes are actively expressed or confer phenotypic resistance. Virulence factor analysis showed functions related to adherence, immune evasion, iron acquisition, quorum sensing, and efflux activity, reflecting strong microbial adaptability. Functional profiling demonstrated enrichment in translation, carbohydrate and amino acid metabolism, genome maintenance, and cell envelope biogenesis. Additionally, CAZyme analysis indicated a high capacity for complex polysaccharide degradation, supporting efficient utilization of fiber-rich traditional diets. In conclusion, this study provides a comprehensive metagenomic overview of antibiotic resistance and functional potential in Kedu chicken gut microbiota, emphasizing the importance of incorporating indigenous poultry into antimicrobial resistance surveillance within a One Health framework.

Antibiotic resistance genes↗