Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

Flower proteome: changes in protein spectrum during the advanced stages of rose petal development.

Flowering is a unique and highly programmed process, but hardly anything is known about the developmentally regulated proteome changes in petals. Here, we employed proteomic technologies to study petal development in rose (Rosa hybrida). Using two-dimensional polyacrylamide gel electrophoresis, we generated stage-specific (closed bud, mature flower and flower at anthesis) petal protein maps with ca. 1,000 unique protein spots. Expression analyses of all resolved protein spots revealed that almost 30% of them were stage-specific, with ca. 90 protein spots for each stage. Most of the proteins exhibited differential expression during petal development, whereas only ca. 6% were constitutively expressed. Eighty-two of the resolved proteins were identified by mass spectrometry and annotated. Classification of the annotated proteins into functional groups revealed energy, cell rescue, unknown function (including novel sequences) and metabolism to be the largest classes, together comprising ca. 90% of all identified proteins. Interestingly, a large number of stress-related proteins were identified in developing petals. Analyses of the expression patterns of annotated proteins and their corresponding RNAs confirmed the importance of proteome characterization.

Flowers↗

TRED: a Transcriptional Regulatory Element Database and a platform for in silico gene regulation studies.

In order to understand gene regulation, accurate and comprehensive knowledge of transcriptional regulatory elements is essential. Here, we report our efforts in building a mammalian Transcriptional Regulatory Element Database (TRED) with associated data analysis functions. It collects cis- and trans-regulatory elements and is dedicated to easy data access and analysis for both single-gene-based and genome-scale studies. Distinguishing features of TRED include: (i) relatively complete genome-wide promoter annotation for human, mouse and rat; (ii) availability of gene transcriptional regulation information including transcription factor binding sites and experimental evidence; (iii) data accuracy is ensured by hand curation; (iv) efficient user interface for easy and flexible data retrieval; and (v) implementation of on-the-fly sequence analysis tools. TRED can provide good training datasets for further genome-wide cis-regulatory element prediction and annotation, assist detailed functional studies and facilitate the decipher of gene regulatory networks (http://rulai.cshl.edu/TRED).

Animals↗

Whole-genome annotation by using evidence integration in functional-linkage networks.

The advent of high-throughput biology has catalyzed a remarkable improvement in our ability to identify new genes. A large fraction of newly discovered genes have an unknown functional role, particularly when they are specific to a particular lineage or organism. These genes, currently labeled "hypothetical," might support important biological cell functions and could potentially serve as targets for medical, diagnostic, or pharmacogenomic studies. An important challenge to the scientific community is to associate these newly predicted genes with a biological function that can be validated by experimental screens. In the absence of sequence or structural homology to known genes, we must rely on advanced biotechnological methods, such as DNA chips and protein-protein interaction screens as well as computational techniques to assign putative functions to these genes. In this article, we propose an effective methodology for combining biological evidence obtained in several high-throughput experimental screens and integrating this evidence in a way that provides consistent functional assignments to hypothetical genes. We use the visualization method of propagation diagrams to illustrate the flow of functional evidence that supports the functional assignments produced by the algorithm. Our results contain a number of predictions and furnish strong evidence that integration of functional information is indeed a promising direction for improving the accuracy and robustness of functional genomics.

Chromosome Mapping↗

Discovery, validation, and genetic dissection of transcription factor binding sites by comparative and functional genomics.

Completing the annotation of a genome sequence requires identifying the regulatory sequences that control gene expression. To identify these sequences, we developed an algorithm that searches for short, conserved sequence motifs in the genomes of related species. The method is effective in finding motifs de novo and for refining known regulatory motifs in Saccharomyces cerevisiae. We tested one novel motif prediction of the algorithm and found it to be the binding site of Stp2; it is significantly different from the previously predicted Stp2 binding site. We show that Stp2 physically interacts with this sequence motif, and that stp2 mutations affect the expression of genes associated with the motif. We demonstrate that the Stp2 binding site also interacts genetically with Stp1, a regulator of amino acid permease genes and, with Sfp1, a key regulator of cell growth. These results illuminate an important transcriptional circuit that regulates cell growth through external nutrient uptake.

Arylsulfotransferase↗

Analysis and visualization of functional relationships between RNA expression and clinical annotation using PathlinX.

We have analyzed a publicly available dataset consisting of gene-expression measurements from 105 lung carcinomas joined with clinical parameters describing the age, smoking history, and survival statistics for the patients that the tumors originated in. Our aim was to demonstrate how the unsupervised analysis technique embodied in PathlinX allows researchers to quickly gain an intuition for the most significant relationships between heterogeneous data elements. A variety of metrics were evaluated empirically by their ability to distinguish biological signal in the data from random noise; this was accomplished by random permutation of the data rows followed by comprehensive pair-wise comparison of all experimental elements. Thresholds of significance were established based on the metric scores for the permuted data. Sub-threshold associations were then removed. The remaining associations were then grouped by a transitive closure process to generate undirected graphs of associations called PathlinX networks. We discuss the various features of each generated PathlinX network and demonstrate the ability of the technique to highlight biological features in large heterogeneous datasets.

Algorithms↗

A sentence sliding window approach to extract protein annotations from biomedical articles.

BACKGROUND: Within the emerging field of text mining and statistical natural language processing (NLP) applied to biomedical articles, a broad variety of techniques have been developed during the past years. Nevertheless, there is still a great ned of comparative assessment of the performance of the proposed methods and the development of common evaluation criteria. This issue was addressed by the Critical Assessment of Text Mining Methods in Molecular Biology (BioCreative) contest. The aim of this contest was to assess the performance of text mining systems applied to biomedical texts including tools which recognize named entities such as genes and proteins, and tools which automatically extract protein annotations. RESULTS: The "sentence sliding window" approach proposed here was found to efficiently extract text fragments from full text articles containing annotations on proteins, providing the highest number of correctly predicted annotations. Moreover, the number of correct extractions of individual entities (i.e. proteins and GO terms) involved in the relationships used for the annotations was significantly higher than the correct extractions of the complete annotations (protein-function relations). CONCLUSION: We explored the use of averaging sentence sliding windows for information extraction, especially in a context where conventional training data is unavailable. The combination of our approach with more refined statistical estimators and machine learning techniques might be a way to improve annotation extraction for future biomedical text mining applications.

Biomedical Research↗

Microarray analysis of radiation response genes in primary human fibroblasts.

PURPOSE: To identify radiation-induced early transcriptional responses in primary human fibroblasts and understand cellular pathways leading to damage correction. METHODS AND MATERIALS: Primary human fibroblast cell lines were irradiated with 2 Gy gamma-radiation and RNA isolated 2 h later. Radiation-induced transcriptional alterations were investigated with microarrays covering the entire human genome. Time- and dose dependent radiation responses were studied by quantitative real-time polymerase chain reaction (RT-PCR). RESULTS: About 200 genes responded to ionizing radiation on the transcriptional level in primary human fibroblasts. The expression profile depended on individual genetic backgrounds. Thirty genes (28 up- and 2 down-regulated) responded to radiation in identical manner in all investigated cells. Twenty of these consensus radiation response genes were functionally categorized: most of them belong to the DNA damage response (GADD45A, BTG2, PCNA, IER5), regulation of cell cycle and cell proliferation (CDKN1A, PPM1D, SERTAD1, PLK2, PLK3, CYR61), programmed cell death (BBC3, TP53INP1) and signaling (SH2D2A, SLIC1, GDF15, THSD1) pathways. Four genes (SEL10, FDXR, CYP26B1, OR11A1) were annotated to other functional groups. Many of the consensus radiation response genes are regulated by, or regulate p53. Time- and dose-dependent expression profiles of selected consensus genes (CDKN1A, GADD45A, IER5, PLK3, CYR61) were investigated by quantitative RT-PCR. Transcriptional alterations depended on the applied dose, and on the time after irradiation. CONCLUSIONS: The data presented here could help in the better understanding of early radiation responses and the development of biomarkers to identify radiation susceptible individuals.

Aged↗

A top-level ontology of functions and its application in the Open Biomedical Ontologies.

MOTIVATION: A clear understanding of functions in biology is a key component in accurate modelling of molecular, cellular and organismal biology. Using the existing biomedical ontologies it has been impossible to capture the complexity of the community's knowledge about biological functions. RESULTS: We present here a top-level ontological framework for representing knowledge about biological functions. This framework lends greater accuracy, power and expressiveness to biomedical ontologies by providing a means to capture existing functional knowledge in a more formal manner. An initial major application of the ontology of functions is the provision of a principled way in which to curate functional knowledge and annotations in biomedical ontologies. Further potential applications include the facilitation of ontology interoperability and automated reasoning. A major advantage of the proposed implementation is that it is an extension to existing biomedical ontologies, and can be applied without substantial changes to these domain ontologies. AVAILABILITY: The Ontology of Functions (OF) can be downloaded in OWL format from http://onto.eva.mpg.de/. Additionally, a UML profile and supplementary information and guides for using the OF can be accessed from the same website.

Biomedical Engineering↗

Phylogeny of related functions: the case of polyamine biosynthetic enzymes.

Genome annotation requires explicit identification of gene function. This task frequently uses protein sequence alignments with examples having a known function. Genetic drift, co-evolution of subunits in protein complexes and a variety of other constraints interfere with the relevance of alignments. Using a specific class of proteins, it is shown that a simple data analysis approach can help solve some of the problems posed. The origin of ureohydrolases has been explored by comparing sequence similarity trees, maximizing amino acid alignment conservation. The trees separate agmatinases from arginases but suggest the presence of unknown biases responsible for unexpected positions of some enzymes. Using factorial correspondence analysis, a distance tree between sequences was established, comparing regions with gaps in the alignments. The gap tree gives a consistent picture of functional kinship, perhaps reflecting some aspects of phylogeny, with a clear domain of enzymes encoding two types of ureohydrolases (agmatinases and arginases) and activities related to, but different from ureohydrolases. Several annotated genes appeared to correspond to a wrong assignment if the trees were significant. They were cloned and their products expressed and identified biochemically. This substantiated the validity of the gap tree. Its organization suggests a very ancient origin of ureohydrolases. Some enzymes of eukaryotic origin are spread throughout the arginase part of the trees: they might have been derived from the genes found in the early symbiotic bacteria that became the organelles. They were transferred to the nucleus when symbiotic genes had to escape Muller's ratchet. This work also shows that arginases and agmatinases share the same two manganese-ion-binding sites and exhibit only subtle differences that can be accounted for knowing the three-dimensional structure of arginases. In the absence of explicit biochemical data, extreme caution is needed when annotating genes having similarities to ureohydrolases.

Amino Acid Sequence↗

Rice Annotation Database (RAD): a contig-oriented database for map-based rice genomics.

A contig-oriented database for annotation of the rice genome has been constructed to facilitate map-based rice genomics. The Rice Annotation Database has the following functional features: (i) extensive effort of manual annotations of P1-derived artificial chromosome/bacterial artificial chromosome clones can be merged at chromosome and contig-level; (ii) concise visualization of the annotation information such as the predicted genes, results of various prediction programs (RiceHMM, Genscan, Genscan+, Fgenesh, GeneMark, etc.), homology to expressed sequence tag, full-length cDNA and protein; (iii) user-friendly clone / gene query system; (iv) download functions for nucleotide, amino acid and coding sequences; (v) analysis of various features of the genome (GC-content, average value, etc.); and (vi) genome-wide homology search (BLAST) of contig- and chromosome-level genome sequence to allow comparative analysis with the genome sequence of other organisms. As of October 2004, the database contains a total of 215 Mb sequence with relevant annotation results including 30 000 manually curated genes. The database can provide the latest information on manual annotation as well as a comprehensive structural analysis of various features of the rice genome. The database can be accessed at http://rad.dna.affrc.go.jp/.

Chromosomes, Plant↗

Protein coding potential of retroviruses and other transposable elements in vertebrate genomes.

We suggest an annotation strategy for genes encoded by retroviruses and transposable elements (RETRA genes) based on a set of marker protein domains. Usually RETRA genes are masked in vertebrate genomes prior to the application of automated gene prediction pipelines under the assumption that they provide no selective advantage to the host. Yet, we show that about 1000 genes in four vertebrate gene sets analyzed contain at least one RETRA gene marker domain. Using the conservation of genomic neighborhood (synteny), we were able to discriminate between RETRA genes with putative functionality in the vertebrates and those that probably function only in the context of mobile elements. We identified 35 such genes in human, along with their corresponding mouse and rat orthologs; which included almost all known human genes with similarity to mobile elements. The results also imply that the vast majority of the remaining RETRA genes in current gene sets are unlikely to encode vertebrate functions. To automatically annotate RETRA genes in other vertebrate genomes, we provide as a tool a set of marker protein domains and a manually refined list of domesticated or ancestral RETRA genes for rescuing genes with vertebrate functions.

Animals↗

ParPEST: a pipeline for EST data analysis based on parallel computing.

BACKGROUND: Expressed Sequence Tags (ESTs) are short and error-prone DNA sequences generated from the 5' and 3' ends of randomly selected cDNA clones. They provide an important resource for comparative and functional genomic studies and, moreover, represent a reliable information for the annotation of genomic sequences. Because of the advances in biotechnologies, ESTs are daily determined in the form of large datasets. Therefore, suitable and efficient bioinformatic approaches are necessary to organize data related information content for further investigations. RESULTS: We implemented ParPEST (Parallel Processing of ESTs), a pipeline based on parallel computing for EST analysis. The results are organized in a suitable data warehouse to provide a starting point to mine expressed sequence datasets. The collected information is useful for investigations on data quality and on data information content, enriched also by a preliminary functional annotation. CONCLUSION: The pipeline presented here has been developed to perform an exhaustive and reliable analysis on EST data and to provide a curated set of information based on a relational database. Moreover, it is designed to reduce execution time of the specific steps required for a complete analysis using distributed processes and parallelized software. It is conceived to run on low requiring hardware components, to fulfill increasing demand, typical of the data used, and scalability at affordable costs.

Algorithms↗

Phytophthora functional genomics database (PFGD): functional genomics of phytophthora-plant interactions.

The Phytophthora Functional Genomics Database (PFGD; http://www.pfgd.org), developed by the National Center for Genome Resources in collaboration with The Ohio State University-Ohio Agricultural Research and Development Center (OSU-OARDC), is a publicly accessible information resource for Phytophthora-plant interaction research. PFGD contains transcript, genomic, gene expression and functional assay data for Phytophthora infestans, which causes late blight of potato, and Phytophthora sojae, which affects soybeans. Automated analyses are performed on all sequence data, including consensus sequences derived from clustered and assembled expressed sequence tags. The PFGD search filter interface allows intuitive navigation of transcript and genomic data organized by library and derived queries using modifiers, annotation keywords or sequence names. BLAST services are provided for libraries built from the transcript and genomic sequences. Transcript data visualization tools include Quality Screening, Multiple Sequence Alignment and Features and Annotations viewers. A genomic browser that supports comparative analysis via novel dynamic functional annotation comparisons is also provided. PFGD is integrated with the Solanaceae Genomics Database (SolGD; http://www.solgd.org) to help provide insight into the mechanisms of infection and resistance, specifically as they relate to the genus Phytophthora pathogens and their plant hosts.

Algal Proteins↗

Dietary effects of arachidonate-rich fungal oil and fish oil on murine hepatic and hippocampal gene expression.

BACKGROUND: The functions, actions, and regulation of tissue metabolism affected by the consumption of long chain polyunsaturated fatty acids (LC-PUFA) from fish oil and other sources remain poorly understood; particularly how LC-PUFAs affect transcription of genes involved in regulating metabolism. In the present work, mice were fed diets containing fish oil rich in eicosapentaenoic acid and docosahexaenoic acid, fungal oil rich in arachidonic acid, or the combination of both. Liver and hippocampus tissue were then analyzed through a combined gene expression- and lipid- profiling strategy in order to annotate the molecular functions and targets of dietary LC-PUFA. RESULTS: Using microarray technology, 329 and 356 dietary regulated transcripts were identified in the liver and hippocampus, respectively. All genes selected as differentially expressed were grouped by expression patterns through a combined k-means/hierarchical clustering approach, and annotated using gene ontology classifications. In the liver, groups of genes were linked to the transcription factors PPARalpha, HNFalpha, and SREBP-1; transcription factors known to control lipid metabolism. The pattern of differentially regulated genes, further supported with quantitative lipid profiling, suggested that the experimental diets increased hepatic beta-oxidation and gluconeogenesis while decreasing fatty acid synthesis. Lastly, novel hippocampal gene changes were identified. CONCLUSIONS: Examining the broad transcriptional effects of LC-PUFAs confirmed previously identified PUFA-mediated gene expression changes and identified novel gene targets. Gene expression profiling displayed a complex and diverse gene pattern underlying the biological response to dietary LC-PUFAs. The results of the studied dietary changes highlighted broad-spectrum effects on the major eukaryotic lipid metabolism transcription factors. Further focused studies, stemming from such transcriptomic data, will need to dissect the transcription factor signaling pathways to fully explain how fish oils and arachidonic acid achieve their specific effects on health.

Journal Article↗

Large-scale benchmarking of prokaryotic annotation tools across thousands of species.

BACKGROUND: Genome annotation is an important step in deriving functional meaning from prokaryotic sequencing data, yet systematic evaluations guiding tool selection are lacking. We present the first large-scale investigation of four prominent open-source annotation tools (Prokka, Bakta, EggNOG-mapper, and PGAP) across 156,033 diverse genomes. This includes Escherichia coli strains for baseline performance, thousands of archaea and bacteria genomes, as well as frameshifted and metagenome-assembled genomes. RESULTS: Bakta excels in annotating high-quality bacterial genomes, while PGAP was better for archaeal genomes and challenging bacterial assemblies, including metagenome-assembled, fragmented, or contaminated samples. For Gene Ontology annotation, PGAP consistently provides broader term coverage, whereas EggNOG-mapper offers more terms per feature. CONCLUSIONS: Our findings highlight tool-specific strengths crucial for selecting optimal solutions based on genome quality, taxonomy, and origin (e.g. MAGs). This study provides an evidence-based guide for users and informs future tool development.

Molecular Sequence Annotation↗

T1DBase, a community web-based resource for type 1 diabetes research.

T1DBase (http://T1DBase.org) is a public website and database that supports the type 1 diabetes (T1D) research community. The site is currently focused on the molecular genetics and biology of T1D susceptibility and pathogenesis. It includes the following datasets: annotated genome sequence for human, rat and mouse; information on genetically identified T1D susceptibility regions in human, rat and mouse, and genetic linkage and association studies pertaining to T1D; descriptions of NOD mouse congenic strains; the Beta Cell Gene Expression Bank, which reports expression levels of genes in beta cells under various conditions, and annotations of gene function in beta cells; data on gene expression in a variety of tissues and organs; and biological pathways from KEGG and BioCarta. Tools on the site include the GBrowse genome browser, site-wide context dependent search, Connect-the-Dots for connecting gene and other identifiers from multiple data sources, Cytoscape for visualizing and analyzing biological networks, and the GESTALT workbench for genome annotation. All data are open access and all software is open source.

Animals↗

UTRdb and UTRsite: a collection of sequences and regulatory motifs of the untranslated regions of eukaryotic mRNAs.

The 5' and 3' untranslated regions of eukaryotic mRNAs play crucial roles in the post-transcriptional regulation of gene expression through the modulation of nucleo-cytoplasmic mRNA transport, translation efficiency, subcellular localization and message stability. UTRdb is a curated database of 5' and 3' untranslated sequences of eukaryotic mRNAs, derived from several sources of primary data. Experimentally validated functional motifs are annotated (and also collated as the UTRsite database) and cross-links to genomic and protein data are provided. The integration of UTRdb with genomic and protein data has allowed the implementation of a powerful retrieval resource for the selection and extraction of UTR subsets based on their genomic coordinates and/or features of the protein encoded by the relevant mRNA (e.g. GO term, PFAM domain, etc.). All internet resources implemented for retrieval and functional analysis of 5' and 3' untranslated regions of eukaryotic mRNAs are accessible at http://www.ba.itb.cnr.it/UTR/.

3' Untranslated Regions↗

Assessing systems properties of yeast mitochondria through an interaction map of the organelle.

Mitochondria carry out specialized functions; compartmentalized, yet integrated into the metabolic and signaling processes of the cell. Although many mitochondrial proteins have been identified, understanding their functional interrelationships has been a challenge. Here we construct a comprehensive network of the mitochondrial system. We integrated genome-wide datasets to generate an accurate and inclusive mitochondrial parts list. Together with benchmarked measures of protein interactions, a network of mitochondria was constructed in their cellular context, including extra-mitochondrial proteins. This network also integrates data from different organisms to expand the known mitochondrial biology beyond the information in the existing databases. Our network brings together annotated and predicted functions into a single framework. This enabled, for the entire system, a survey of mutant phenotypes, gene regulation, evolution, and disease susceptibility. Furthermore, we experimentally validated the localization of several candidate proteins and derived novel functional contexts for hundreds of uncharacterized proteins. Our network thus advances the understanding of the mitochondrial system in yeast and identifies properties of genes underlying human mitochondrial disorders.

Disease Susceptibility↗