Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

Variation resources at UC Santa Cruz.

The variation resources within the University of California Santa Cruz Genome Browser include polymorphism data drawn from public collections and analyses of these data, along with their display in the context of other genomic annotations. Primary data from dbSNP is included for many organisms, with added information including genomic alleles and orthologous alleles for closely related organisms. Display filtering and coloring is available by variant type, functional class or other annotations. Annotation of potential errors is highlighted and a genomic alignment of the variant's flanking sequence is displayed. HapMap allele frequencies and linkage disequilibrium (LD) are available for each HapMap population, along with non-human primate alleles. The browsing and analysis tools, downloadable data files and links to documentation and other information can be found at http://genome.ucsc.edu/.

Alleles↗

ECgene: genome annotation for alternative splicing.

ECgene provides annotation for gene structure, function and expression, taking alternative splicing events into consideration. The gene-modeling algorithm combines the genome-based expressed sequence tag (EST) clustering and graph-theoretic transcript assembly procedures. The website provides several viewers and applications that have many unique features useful for the analysis of the transcript structure and gene expression. The summary viewer shows the gene summary and the essence of other annotation programs. The genome browser and the transcript viewer are available for comparing the gene structure of splice variants. Changes in the functional domains by alternative splicing can be seen at a glance in the transcript viewer. We also provide two unique ways of analyzing gene expression. The SAGE tags deduced from the assembled transcripts are used to delineate quantitative expression patterns from SAGE libraries available publically. Furthermore, the cDNA libraries of EST sequences in each cluster are used to infer qualitative expression patterns. It should be noted that the ECgene website provides annotation for the whole transcriptome, not just the alternatively spliced genes. Currently, ECgene supports the human, mouse and rat genomes. The ECgene suite of tools and programs is available at http://genome.ewha.ac.kr/ECgene/.

Algorithms↗

Comparative profiling of microbial community structure, enzyme potential, metabolic features, and volatile composition in craft and Jiafan Huangjiu processes.

Craft Huangjiu and Jiafan Huangjiu represent two distinct industrial Huangjiu product outcomes with contrasting volatile profiles. This study compared craft Huangjiu (L70) and Jiafan Huangjiu (L79) to characterize their physicochemical, microbial, gene-level functional, metabolic, and volatile features. Because L70 involved mid-fermentation addition of finished Huangjiu, this comparison was not intended to isolate the sole effect of fermentation interruption versus continued fermentation. L79 showed more extensive carbon and nitrogen utilization, with lower residual substrates and higher ethanol and acetic acid contents than L70, whereas L70 retained a less complete fermentation state. At the volatile level, GC-MS and volatile metabolomics consistently showed an ester-enriched profile in L79 and a more alcohol-dominant profile in L70. FlavorDB-based putative annotation and threshold-based OAV analysis further indicated distinct database-assigned descriptor distributions and potential odor-active compounds, with more OAV > 1 ester-related compounds in L79. Metagenomic analysis showed that L70 was dominated by Lactobacillus acetotolerans, whereas L79 contained higher relative abundances of Saccharomyces cerevisiae, Aspergillus oryzae, Aspergillus flavus, and Fructilactobacillus fructivorans. Metagenomic functional annotation showed higher representation of hydrolysis-related CAZy genes and ester-related enzyme annotations in L79. KEGG-based pathway mapping further indicated greater gene-level potential for ethanol-, acetate-, and acetyl-CoA-related metabolism in L79. Accordingly, the L70 profile should be interpreted as the integrated final-product outcome of process intervention, exogenous input, and subsequent fermentation. The findings provide a comparative basis for future flavor regulation and process optimization in Huangjiu and other fermented alcoholic beverages.

Volatile Organic Compounds↗

The functional genomic distribution of protein divergence in two animal phyla: coevolution, genomic conflict, and constraint.

We compare the functional spectrum of protein evolution in two separate animal lineages with respect to two hypotheses: (1) rates of divergence are distributed similarly among functional classes within both lineages, indicating that selective pressure on the proteome is largely independent of organismic-level biological requirements; and (2) rates of divergence are distributed differently among functional classes within each lineage, indicating species-specific selective regimes impact genome-wide substitutional patterns. Integrating comparative genome sequence with data from tissue-specific expressed-sequence-tag (EST) libraries and detailed database annotations, we find a functional genomic signature of rapid evolution and selective constraint shared between mammalian and nematode lineages despite their extensive morphological and ecological differences and distant common ancestry. In both phyla, we find evidence of accelerated evolution among components of molecular systems involved in coevolutionary change. In mammals, lineage-specific fast evolving genes include those involved in reproduction, immunity, and possibly, maternal-fetal conflict. Likelihood ratio tests provide evidence for positive selection in these rapidly evolving functional categories in mammals. In contrast, slowly evolving genes, in terms of amino acid or insertion/deletion (indel) change, in both phyla are involved in core molecular processes such as transcription, translation, and protein transport. Thus, strong purifying selection appears to act on the same core cellular processes in both mammalian and nematode lineages, whereas positive and/or relaxed selection acts on different biological processes in each lineage.

Amino Acid Substitution↗

Towards a reliable objective function for multiple sequence alignments.

Multiple sequence alignment is a fundamental tool in a number of different domains in modern molecular biology, including functional and evolutionary studies of a protein family. Multiple alignments also play an essential role in the new integrated systems for genome annotation and analysis. Thus, the development of new multiple alignment scores and statistics is essential, in the spirit of the work dedicated to the evaluation of pairwise sequence alignments for database searching techniques. We present here norMD, a new objective scoring function for multiple sequence alignments. NorMD combines the advantages of the column-scoring techniques with the sensitivity of methods incorporating residue similarity scores. In addition, norMD incorporates ab initio sequence information, such as the number, length and similarity of the sequences to be aligned. The sensitivity and reliability of the norMD objective function is demonstrated using structural alignments in the SCOP and BAliBASE databases. The norMD scores are then applied to the multiple alignments of the complete sequences (MACS) detected by BlastP with E-value<10, for a set of 734 hypothetical proteins encoded by the Vibrio cholerae genome. Unrelated or badly aligned sequences were automatically removed from the MACS, leaving a high-quality multiple alignment which could be reliably exploited in a subsequent functional and/or structural annotation process. After removal of unreliable sequences, 176 (24 %) of the alignments contained at least one sequence with a functional annotation. 103 of these new matches were supported by significant hits to the Interpro domain and motif database.

Amino Acid Motifs↗

Automated ortholog inference from phylogenetic trees and calculation of orthology reliability.

MOTIVATION: Orthologous proteins in different species are likely to have similar biochemical function and biological role. When annotating a newly sequenced genome by sequence homology, the most precise and reliable functional information can thus be derived from orthologs in other species. A standard method of finding orthologs is to compare the sequence tree with the species tree. However, since the topology of phylogenetic tree is not always reliable one might get incorrect assignments. RESULTS: Here we present a novel method that resolves this problem by analyzing a set of bootstrap trees instead of the optimal tree. The frequency of orthology assignments in the bootstrap trees can be interpreted as a support value for the possible orthology of the sequences. Our method is efficient enough to analyze data in the scale of whole genomes. It is implemented in Java and calculates orthology support levels for all pairwise combinations of homologous sequences of two species. The method was tested on simulated datasets and on real data of homologous proteins.

Algorithms↗

Genomes for medicine.

We have the human genome sequence. It is freely available, accurate and nearly complete. But is the genome ready for medicine? The new resource is already changing genetic research strategies to find information of medical value. Now we need high-quality annotation of all the functionally important sequences and the variations within them that contribute to health and disease. To achieve this, we need more genome sequences, systematic experimental analyses, and extensive information on human phenotypes. Flexible and user-friendly access to well-annotated genomes will create an environment for innovation, and the potential for unlimited use of sequencing in biomedical research and practice.

Genetic Variation↗

Evaluation of human-readable annotation in biomolecular sequence databases with biological rule libraries.

MOTIVATION: Computer-based selection of entries from sequence databases with respect to a related functional description, e.g. with respect to a common cellular localization or contributing to the same phenotypic function, is a difficult task. Automatic semantic analysis of annotations is not only hampered by incomplete functional assignments. A major problem is that annotations are written in a rich, non-formalized language and are meant for reading by a human expert. This person can extract from the text considerably more information than is immediately apparent due to his extended biological background knowledge and logical reasoning. APPROACH: A technique of automated annotation evaluation based on a combination of lexical analysis and the usage of biological rule libraries has been developed. The proposed algorithm generates new functional descriptors from the annotation of a given entry using the semantic units of the annotation as prepositions for implications executed in accordance with the rule library. RESULTS: The prototype of a software system, the Meta_A(nnotator) program, is described and the results of its application to sequence attribute assignment and sequence selection problems, such as cellular localization and sequence domain annotation of SWISS-PROT entries, are presented. The current software version assigns useful subcellular localization qualifiers to approximately 88% of all SWISS-PROT entries. As shown by demonstrative examples, the combination of sequence and annotation analysis is a powerful approach for the detection of mutual annotation/sequence inconsistencies. AVAILABILITY: Results for the cellular localization assignment can be viewed at the URL http://www.bork. embl-heidelberg.de/CELL_LOC/CELL_LOC.html.

Algorithms↗

Query3d: a new method for high-throughput analysis of functional residues in protein structures.

BACKGROUND: The identification of local similarities between two protein structures can provide clues of a common function. Many different methods exist for searching for similar subsets of residues in proteins of known structure. However, the lack of functional and structural information on single residues, together with the low level of integration of this information in comparison methods, is a limitation that prevents these methods from being fully exploited in high-throughput analyses. RESULTS: Here we describe Query3d, a program that is both a structural DBMS (Database Management System) and a local comparison method. The method conserves a copy of all the residues of the Protein Data Bank annotated with a variety of functional and structural information. New annotations can be easily added from a variety of methods and known databases. The algorithm makes it possible to create complex queries based on the residues' function and then to compare only subsets of the selected residues. Functional information is also essential to speed up the comparison and the analysis of the results. CONCLUSION: With Query3d, users can easily obtain statistics on how many and which residues share certain properties in all proteins of known structure. At the same time, the method also finds their structural neighbours in the whole PDB. Programs and data can be accessed through the PdbFun web interface.

Algorithms↗

Functional metaproteomics for enzyme discovery.

Discovery of microbial biocatalysts traditionally relied on activity screening of isolated bacterial strains. However, since most microorganisms cannot be cultivated in the lab, such an approach leaves the majority of the microbial enzyme diversity untapped. Metagenomic approaches, in which the DNA from a microbial community is directly isolated and then used either for the creation of an expression library or for sequencing and metagenome annotation have alleviated this shortcoming to an extent, but have their own limitations: the generation of large expression libraries is time-consuming and their screening is costly, while metagenome annotation can infer biocatalytic function only from prior knowledge. We have thus developed a functional metaproteomic approach, which combines the immediacy of traditional activity screening with the comprehensiveness of a meta-omics approach. Briefly, the whole metaproteome of an environmental sample is separated on a 2-D gel, biocatalytically active proteins are visualized in-gel through zymography, and those candidate biocatalysts are then identified through mass spectrometry, searching against a metagenome-derived database obtained from the very same environmental sample. Here we explain the process in detail, with a focus on esterases, and give guidelines on how to develop a functional metaproteomic workflow for enzyme discovery.

Proteomics↗

Dendritic mRNAs encode diversified functionalities in hippocampal pyramidal neurons.

BACKGROUND: Targeted transport of messenger RNA and local protein synthesis near the synapse are important for synaptic plasticity. In order to gain an overview of the composition of the dendritic mRNA pool, we dissected out stratum radiatum (dendritic lamina) from rat hippocampal CA1 region and compared its mRNA content with that of stratum pyramidale (cell body layer) using a set of cDNA microarrays. RNAs that have over-representation in the dendritic fraction were annotated and sorted into function groups. RESULTS: We have identified 154 dendritic mRNA candidates, which can be arranged into the categories of receptors and channels, signaling molecules, cytoskeleton and adhesion molecules, and factors that are involved in membrane trafficking, in protein synthesis, in posttranslational protein modification, and in protein degradation. Previously known dendritic mRNAs such as MAP2, calmodulin, and G protein gamma subunit were identified from our screening, as were mRNAs that encode proteins known to be important for synaptic plasticity and memory, such as spinophilin, Pumilio, eEF1A, and MHC class I molecules. Furthermore, mRNAs coding for ribosomal proteins were also found in dendrites. CONCLUSION: Our results suggest that neurons transport a variety of mRNAs to dendrites, not only those directly involved in modulating synaptic plasticity, but also others that play more common roles in cellular metabolism.

Animals↗

High-quality protein knowledge resource: SWISS-PROT and TrEMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domain structure, post-translational modifications, variants, etc.), a minimal level of redundancy and a high level of integration with other databases. Together with its automatically annotated supplement TrEMBL, it provides a comprehensive and high-quality view of the current state of knowledge about proteins. Ongoing developments include the further improvement of functional and automatic annotation in the databases including evidence attribution with particular emphasis on the human, archaeal and bacterial proteomes and the provision of additional resources such as the International Protein Index (IPI) and XML format of SWISS-PROT and TrEMBL to the user community.

Amino Acid Sequence↗

Functional replacement of the FabA and FabB proteins of Escherichia coli fatty acid synthesis by Enterococcus faecalis FabZ and FabF homologues.

The anaerobic unsaturated fatty acid synthetic pathway of Escherichia coli requires two specialized proteins, FabA and FabB. However, the fabA and fabB genes are found only in the Gram-negative alpha- and gamma-proteobacteria, and thus other anaerobic bacteria must synthesize these acids using different enzymes. We report that the Gram-positive bacterium Enterococcus faecalis encodes a protein, annotated as FabZ1, that functionally replaces the E. coli FabA protein, although the sequence of this protein aligns much more closely with E. coli FabZ, a protein that plays no specific role in unsaturated fatty acid synthesis. Therefore E. faecalis FabZ1 is a bifunctional dehydratase/isomerase, an enzyme activity heretofore confined to a group of Gram-negative bacteria. The FabZ2 protein is unable to replace the function of E. coli FabZ, although FabZ2, a second E. faecalis FabZ homologue, has this ability. Moreover, an E. faecalis FabF homologue (FabF1) was found to replace the function of E. coli FabB, whereas a second FabF homologue was inactive. From these data it is clear that bacterial fatty acid biosynthetic pathways cannot be deduced solely by sequence comparisons.

3-Oxoacyl-(Acyl-Carrier-Protein) Synthase↗

Comparative analysis of eukaryotic-type protein phosphatases in two streptomycete genomes.

Inspection of the genomes of Streptomyces coelicolor A3(2) and Streptomyces avermitilis reveals that each contains 55 putative eukaryotic-type protein phosphatases (PPs), the largest number ever identified from any single prokaryotic organism. Unlike most other prokaryotic genomes that have only one or two superfamilies of eukaryotic-type PPs, the streptomycete genomes possess the eukaryotic-type PPs that belong to four superfamilies: 2 phosphoprotein phosphatases and 2 low-molecular-mass protein tyrosine phosphatases in each species, 49 Mg(2+)- or Mn(2+)-dependent protein phosphatases (PPMs) and 2 conventional protein tyrosine phosphatases (CPTPs) in S. coelicolor A3(2), and 48 PPMs and 3 CPTPs in S. avermitilis. Sixty-four percent of the PPs found in S. coelicolor A3(2) have orthologues in S. avermitilis, indicating that they originated from a common ancestor and might be involved in the regulation of more conserved metabolic activities. The genes of eukaryotic-type PP unique to each surveyed streptomycete genome are mainly located in two arms of the linear chromosomes and their evolution might be involved in gene acquisition or duplication to adapt to the extremely variable soil environments where these organisms live. In addition, 56 % of the PPs from S. coelicolor A3(2) and 65 % of the PPs from S. avermitilis possess at least one additional domain having a putative biological function. These include the domains involved in the detection of redox potential, the binding of cyclic nucleotides, mRNA, DNA and ATP, and the catalysis of phosphorylation reactions. Because they contained multiple functional domains, most of them were assigned functions other than PPs in previous annotations. Although few studies have been conducted on the physiological functions of the PPs in streptomycetes, the existence of large numbers of putative PPs in these two streptomycete genomes strongly suggests that eukaryotic-type PPs play important regulatory roles in primary or secondary metabolic pathways. The identification and analysis of such a large number of putative eukaryotic-type PPs from S. coelicolor A3(2) and S. avermitilis constitute a basis for further exploration of the signal transduction pathways mediated by these phosphatases in industrially important strains of streptomycetes.

Computational Biology↗

Synergistic computational and experimental proteomics approaches for more accurate detection of active serine hydrolases in yeast.

An analysis of the structurally and catalytically diverse serine hydrolase protein family in the Saccharomyces cerevisiae proteome was undertaken using two independent but complementary, large-scale approaches. The first approach is based on computational analysis of serine hydrolase active site structures; the second utilizes the chemical reactivity of the serine hydrolase active site in complex mixtures. These proteomics approaches share the ability to fractionate the complex proteome into functional subsets. Each method identified a significant number of sequences, but 15 proteins were identified by both methods. Eight of these were unannotated in the Saccharomyces Genome Database at the time of this study and are thus novel serine hydrolase identifications. Three of the previously uncharacterized proteins are members of a eukaryotic serine hydrolase family, designated as Fsh (family of serine hydrolase), identified here for the first time. OVCA2, a potential human tumor suppressor, and DYR-SCHPO, a dihydrofolate reductase from Schizosaccharomyces pombe, are members of this family. Comparing the combined results to results of other proteomic methods showed that only four of the 15 proteins were identified in a recent large-scale, "shotgun" proteomic analysis and eight were identified using a related, but similar, approach (neither identifies function). Only 10 of the 15 were annotated using alternate motif-based computational tools. The results demonstrate the precision derived from combining complementary, function-based approaches to extract biological information from complex proteomes. The chemical proteomics technology indicates that a functional protein is being expressed in the cell, while the computational proteomics technology adds details about the specific type of function and residue that is likely being labeled. The combination of synergistic methods facilitates analysis, enriches true positive results, and increases confidence in novel identifications. This work also highlights the risks inherent in annotation transfer and the use of scoring functions for determination of correct annotations.

Amino Acid Sequence↗

QuasiMotiFinder: protein annotation by searching for evolutionarily conserved motif-like patterns.

Sequence signature databases such as PROSITE, which include amino acid segments that are indicative of a protein's function, are useful for protein annotation. Lamentably, the annotation is not always accurate. A signature may be falsely detected in a protein that does not carry out the associated function (false positive prediction, FP) or may be overlooked in a protein that does carry out the function (false negative prediction, FN). A new approach has emerged in which a signature is replaced with a sequence profile, calculated based on multiple sequence alignment (MSA) of homologous proteins that share the same function. This approach, which is superior to the simple pattern search, essentially searches with the sequence of the query protein against an MSA library. We suggest here an alternative approach, implemented in the QuasiMotiFinder web server (http://quasimotifinder.tau.ac.il/), which is based on a search with an MSA of homologous query proteins against the original PROSITE signatures. The explicit use of the average evolutionary conservation of the signature in the query proteins significantly reduces the rate of FP prediction compared with the simple pattern search. QuasiMotiFinder also has a reduced rate of FN prediction compared with simple pattern searches, since the traditional search for precise signatures has been replaced by a permissive search for signature-like patterns that are physicochemically similar to known signatures. Overall, QuasiMotiFinder and the profile search are comparable to each other in terms of performance. They are also complementary to each other in that signatures that are falsely detected in (or overlooked by) one may be correctly detected by the other.

Amino Acid Motifs↗

Biological function of unannotated transcription during the early development of Drosophila melanogaster.

Many animal and plant genomes are transcribed much more extensively than current annotations predict. However, the biological function of these unannotated transcribed regions is largely unknown. Approximately 7% and 23% of the detected transcribed nucleotides during D. melanogaster embryogenesis map to unannotated intergenic and intronic regions, respectively. Based on computational analysis of coordinated transcription, we conservatively estimate that 29% of all unannotated transcribed sequences function as missed or alternative exons of well-characterized protein-coding genes. We estimate that 15.6% of intergenic transcribed regions function as missed or alternative transcription start sites (TSS) used by 11.4% of the expressed protein-coding genes. Identification of P element mutations within or near newly identified 5' exons provides a strategy for mapping previously uncharacterized mutations to their respective genes. Collectively, these data indicate that at least 85% of the fly genome is transcribed and processed into mature transcripts representing at least 30% of the fly genome.

Amino Acid Sequence↗

The GermOnline cross-species systems browser provides comprehensive information on genes and gene products relevant for sexual reproduction.

We report a novel release of the GermOnline knowledgebase covering genes relevant for the cell cycle, gametogenesis and fertility. GermOnline was extended into a cross-species systems browser including information on DNA sequence annotation, gene expression and the function of gene products. The database covers eight model organisms and Homo sapiens, for which complete genome annotation data are available. The database is now built around a sophisticated genome browser (Ensembl), our own microarray information management and annotation system (MIMAS) used to extensively describe experimental data obtained with high-density oligonucleotide microarrays (GeneChips) and a comprehensive system for online editing of database entries (MediaWiki). The RNA data include results from classical microarrays as well as tiling arrays that yield information on RNA expression levels, transcript start sites and lengths as well as exon composition. Members of the research community are solicited to help GermOnline curators keep database entries on genes and gene products complete and accurate. The database is accessible at http://www.germonline.org/.

Animals↗