Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

Genome-scale functional profiling of the mammalian AP-1 signaling pathway.

Large-scale functional genomics approaches are fundamental to the characterization of mammalian transcriptomes annotated by genome sequencing projects. Although current high-throughput strategies systematically survey either transcriptional or biochemical networks, analogous genome-scale investigations that analyze gene function in mammalian cells have yet to be fully realized. Through transient overexpression analysis, we describe the parallel interrogation of approximately 20,000 sequence annotated genes in cancer-related signaling pathways. For experimental validation of these genome data, we apply an integrative strategy to characterize previously unreported effectors of activator protein-1 (AP-1) mediated growth and mitogenic response pathways. These studies identify the ADP-ribosylation factor GTPase-activating protein Centaurin alpha1 and a Tudor domain-containing hypothetical protein as putative AP-1 regulatory oncogenes. These results provide insight into the composition of the AP-1 signaling machinery and validate this approach as a tractable platform for genome-wide functional analysis.

Animals↗

Mapping Gene Ontology to proteins based on protein-protein interaction data.

MOTIVATION: Gene Ontology (GO) consortium provides structural description of protein function that is used as a common language for gene annotation in many organisms. Large-scale techniques have generated many valuable protein-protein interaction datasets that are useful for the study of protein function. Combining both GO and protein-protein interaction data allows the prediction of function for unknown proteins. RESULT: We apply a Markov random field method to the prediction of yeast protein function based on multiple protein-protein interaction datasets. We assign function to unknown proteins with a probability representing the confidence of this prediction. The functions are based on three general categories of cellular component, molecular function and biological process defined in GO. The yeast proteins are defined in the Saccharomyces Genome Database (SGD). The protein-protein interaction datasets are obtained from the Munich Information Center for Protein Sequences (MIPS), including physical interactions and genetic interactions. The efficiency of our prediction is measured by applying the leave-one-out validation procedure to a functional path matching scheme, which compares the prediction with the GO description of a protein's function from the abstract level to the detailed level along the GO structure. For biological process, the leave-one-out validation procedure shows 52% precision and recall of our method, much better than that of the simple guilty-by-association methods.

Chromosome Mapping↗

SynDB: a Synapse protein DataBase based on synapse ontology.

A synapse is the junction across which a nerve impulse passes from an axon terminal to a neuron, muscle cell or gland cell. The functions and building molecules of the synapse are essential to almost all neurobiological processes. To describe synaptic structures and functions, we have developed Synapse Ontology (SynO), a hierarchical representation that includes 177 terms with hundreds of synonyms and branches up to eight levels deep. associated 125 additional protein keywords and 109 InterPro domains with these SynO terms. Using a combination of automated keyword searches, domain searches and manual curation, we collected 14,000 non-redundant synapse-related proteins, including 3000 in human. We extensively annotated the proteins with information about sequence, structure, function, expression, pathways, interactions and disease associations and with hyperlinks to external databases. The data are stored and presented in the Synapse protein DataBase (SynDB, http://syndb.cbi.pku.edu.cn). SynDB can be interactively browsed by SynO, Gene Ontology (GO), domain families, species, chromosomal locations or Tribe-MCL clusters. It can also be searched by text (including Boolean operators) or by sequence similarity. SynDB is the most comprehensive database to date for synaptic proteins.

Animals↗

Expression profiling of the Leishmania life cycle: cDNA arrays identify developmentally regulated genes present but not annotated in the genome.

As genomic sequencing of Leishmania nears completion, functional analyses that provide a global genetic perspective on biological processes are important. Despite polycistronic transcription, RNA transcript abundance can be measured using microarrays. To provide a resource to evaluate cDNA arrays, we undertook 5' expressed sequence tag analysis of 2183 full-length randomly selected cDNAs from Leishmania major promastigote (days 3, 7, 10 of culture in vitro), and lesion-derived amastigote libraries. PCR-amplified inserts from 1830 of these cDNA representing 1001 unique genes were spotted onto microarrays, and compared internally with PCR-amplified open reading frames (ORFs) from 904 genes representing 842 unique genes annotated in the L. major genome. Microarrays were screened with RNA from procyclic, metacyclic and amastigote populations of L. major. Redundant clones on the array gave highly reproducible results, providing confidence in identification of stage-specific gene expression. Four hundred and thirty unique (i.e. non-redundant) stage-specific genes were identified. A higher percentage of stage-specific gene expression was observed in amastigotes ( approximately 35%) compared to metacyclics ( approximately 12%) for both cDNAs and ORFs, but cDNAs provided a richer source of regulated genes than currently annotated ORFs from the Leishmania genome. In mapping cDNAs onto the Leishmania genome, we noted that approximately 42% aligned to regions not recognised as genes using current predictive annotation tools. These genes are highly represented in our stage-specific genes, and therefore represent important drug targets and vaccine candidates. Careful annotation of cDNAs onto the Leishmania genome will be important before producing the next generation of oligonucleotide arrays based on annotated genes of the genomic sequencing project.

Animals↗

Differential genome analysis applied to the species-specific features of Helicobacter pylori.

We introduce a simple and rapid strategy to identify genes that are responsible for species-specific phenotypes. The genome of a species that has a specific phenotype is compared with at least one, closely related, species that lacks this phenotype. Homologous genes that are shared among the species compared are identified and discarded from the list of candidates for species-specific genes. The process is automated and rapidly yields a small subset of the genome that likely contains genes responsible for the species-specific features. Functions are assigned to the genes, and dubious annotations are filtered out. Information is extracted not only from the presence of genes, but also from their absence with respect to known phenotypes. We have applied the technique to identify a set of species-specific genes in Helicobacter pylori by comparing it with its closest relatives for which complete genome sequences are available, Haemophilus influenzae and Escherichia coli. Of the genes of this set for which functional features can be obtained, a large fraction (63%, 123 proteins) is (potentially) involved in H. pylori's interaction with its host. We hypothesize that a family of outer membrane proteins is critical for the ability of H. pylori to colonize host cells in highly acidic environments.

Amino Acid Sequence↗

Protein tyrosine and serine-threonine phosphatases in the sea urchin, Strongylocentrotus purpuratus: identification and potential functions.

Protein phosphatases, in coordination with protein kinases, play crucial roles in regulation of signaling pathways. To identify protein tyrosine phosphatases (PTPs) and serine-threonine (ser-thr) phosphatases in the Strongylocentrotus purpuratus genome, 179 annotated sequences were studied (122 PTPs, 57 ser-thr phosphatases). Sequence analysis identified 91 phosphatases (33 conventional PTPs, 31 dual specificity phosphatases, 1 Class III Cysteine-based PTP, 1 Asp-based PTP, and 25 ser-thr phosphatases). Using catalytic sites, levels of conservation and constraint in amino acid sequence were examined. Nine of 25 receptor PTPs (RPTPs) corresponded to human, nematode, or fly homologues. Domain structure revealed that sea urchin-specific RPTPs including two, PTPRLec and PTPRscav, may act in immune defense. Embryonic transcription of each phosphatase was recorded from a high-density oligonucleotide tiling microarray experiment. Most RPTPs are expressed at very low levels, whereas nonreceptor PTPs (NRPTPs) are generally expressed at moderate levels. High expression was detected in MAP kinase phosphatases (MKPs) and numerous ser-thr phosphatases. For several expressed NRPTPs, MKPs, and ser-thr phosphatases, morpholino antisense-mediated knockdowns were performed and phenotypes obtained. Finally, to assess roles of annotated phosphatases in endomesoderm formation, a literature review of phosphatase functions in model organisms was superimposed on sea urchin developmental pathways to predict areas of functional activity.

Animals↗

ESTHER, the database of the alpha/beta-hydrolase fold superfamily of proteins.

The alpha/beta-hydrolase fold is characterized by a beta-sheet core of five to eight strands connected by alpha-helices to form a alpha/beta/alpha sandwich. In most of the family members the beta-strands are parallels, but some show an inversion in the order of the first strands, resulting in antiparallel orientation. The members of the superfamily diverged from a common ancestor into a number of hydrolytic enzymes with a wide range of substrate specificities, together with other proteins with no recognized catalytic activity. In the enzymes the catalytic triad residues are presented on loops, of which one, the nucleophile elbow, is the most conserved feature of the fold. Of the other proteins, which all lack from one to all of the catalytic residues, some may simply be 'inactive' enzymes while others are known to be involved in surface recognition functions. The ESTHER database (http://bioweb.ensam.inra.fr/esther) gathers and annotates all the published information related to gene and protein sequences of this superfamily, as well as biochemical, pharmacological and structural data, and connects them so as to provide the bases for studying structure-function relationships within the family. The most recent developments of the database, which include a section on human diseases related to members of the family, are described.

Animals↗

From fold to function predictions: an apoptosis regulator protein BID.

With the rapidly increasing pace of genome sequencing projects and the resulting flood of predicted amino acid sequences of uncharacterized proteins, protein sequence analysis, and in particular, protein structure prediction is quickly gaining in importance. Prediction algorithms can be used for preliminary annotation of newly sequenced proteins and, at least in some cases, provide insights into their function and specific mode of action. Such annotations for several microbial genomes were performed by several groups and placed in public domain for evaluation. An example presented in this work comes from a related project of structural and functional predictions for proteins involved in the process of controlled cell death (apoptosis). The BID protein belongs to an important class of regulators of apoptosis identified by short sequence motifs. Here, several fold prediction methods are used to build a series of three-dimensional models. Structure analysis of the models with reference to the biological data available allows selection of the most appropriate model. It is found that the most likely structural model of BID is built on the structure of Bcl-X(L). The model is discussed in terms of experimental data on specific proteolytic cleavage of BID and its effect on BID interactions with other proteins and membranes.

Algorithms↗

FISH analysis of Drosophila melanogaster heterochromatin using BACs and P elements.

The heterochromatin of chromosomes 2 and 3 of Drosophila melanogaster contains about 30 essential genes defined by genetic analysis. In the last decade only a few of these genes have been molecularly characterized and found to correspond to protein-coding genes involved in important cellular functions. Moreover, several predicted genes have been identified by annotation of genomic sequence that are associated with polytene chromosome divisions 40, 41 and 80 but their locations on the cytogenetic map of the heterochromatin are still uncertain. To expand our current knowledge of the genetic functions located in heterochromatin, we have performed fluorescence in situ hybridization (FISH) mapping to mitotic chromosomes of nine bacterial artificial chromosomes (BACs) carrying several predicted genes and of 13 P element insertions assigned to the proximal regions of 2R and 3L. We found that 22 predicted genes map to the h46 region of 2R and eight map to the h47 regions of 3L. This amounts to at least 30 predicted genes located in these heterochromatic regions, whereas previous studies detected only seven vital genes. Finally, another 58 genes localize either in the euchromatin-heterochromatin transition regions or in the proximal euchromatin of 2R and 3L.

Animals↗

Analysis of duplication and possible sub-functionalization of wing gene network components in pea aphids.

A fundamental focus of evolutionary-developmental biology is uncovering the genetic mechanisms responsible for the gain and loss of characters. One approach to this question is to investigate changes in the coordinated expression of a group of genes important for the development of a character of interest (a gene regulatory network). Here we consider the possibility that modifications to the wing gene regulatory network (wGRN), as defined by work primarily done in Drosophila melanogaster, were involved in the evolution of wing dimorphisms of the pea aphid (Acyrthosiphon pisum). We hypothesize that this may have occurred via changes in expression levels or duplication followed by sub-functionalization of wGRN components. To test this, we annotated members of the wGRN in the pea aphid genome and assessed their expression levels in first and third nymphal instars of winged and wingless morphs of males and asexual females. We find that only two of the 32 assessed genes exhibit morph-biased expression. We also find that three wing genes (apterous (ap), warts (wts), and decapentaplegic (dpp)) have undergone gene duplication. In each case, the resulting paralogs show signs of functional divergence, exhibiting either sex-, morph-, or stage-specific expression. Two gene duplicates, wts2 and dpp3, are of particular interest with respect to wing dimorphism, as they exhibit a wingless male-specific isoform and wingless male-biased expression, respectively. These results supplement our understanding of trends in developmental gene network evolution, such as side-stepping pleiotropic constraint via duplication and sub-functionalization, underlying the emergence of novel phenotypes.

dimorphism↗

Annotation transfer between genomes: protein-protein interologs and protein-DNA regulogs.

Proteins function mainly through interactions, especially with DNA and other proteins. While some large-scale interaction networks are now available for a number of model organisms, their experimental generation remains difficult. Consequently, interolog mapping--the transfer of interaction annotation from one organism to another using comparative genomics--is of significant value. Here we quantitatively assess the degree to which interologs can be reliably transferred between species as a function of the sequence similarity of the corresponding interacting proteins. Using interaction information from Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster, and Helicobacter pylori, we find that protein-protein interactions can be transferred when a pair of proteins has a joint sequence identity >80% or a joint E-value <10(-70). (These "joint" quantities are the geometric means of the identities or E-values for the two pairs of interacting proteins.) We generalize our interolog analysis to protein-DNA binding, finding such interactions are conserved at specific thresholds between 30% and 60% sequence identity depending on the protein family. Furthermore, we introduce the concept of a "regulog"--a conserved regulatory relationship between proteins across different species. We map interologs and regulogs from yeast to a number of genomes with limited experimental annotation (e.g., Arabidopsis thaliana) and make these available through an online database at http://interolog.gersteinlab.org. Specifically, we are able to transfer approximately 90,000 potential protein-protein interactions to the worm. We test a number of these in two-hybrid experiments and are able to verify 45 overlaps, which we show to be statistically significant.

Amino Acid Sequence↗

Characteristics of the Lotus japonicus gene repertoire deduced from large-scale expressed sequence tag (EST) analysis.

To perform a comprehensive analysis of genes expressed in a model legume, Lotus japonicus, a total of 74472 3'-end expressed sequence tags (EST) were generated from cDNA libraries produced from six different organs. Clustering of sequences was performed with an identity criterion of 95% for 50 bases, and a total of 20457 non-redundant sequences, 8503 contigs and 11954 singletons were generated. EST sequence coverage was analyzed by using the annotated L. japonicus genomic sequence and 1093 of the 1889 predicted protein-encoding genes (57.9%) were hit by the EST sequence(s). Gene content was compared to several plant species. Among the 8503 contigs, 471 were identified as sequences conserved only in leguminous species and these included several disease resistance-related genes. This suggested that in legumes, these genes may have evolved specifically to resist pathogen attack. The rate of gene sequence divergence was assessed by comparing similarity level and functional category based on the Gene Ontology (GO) annotation of Arabidopsis genes. This revealed that genes encoding ribosomal proteins, as well as those related to translation, photosynthesis, and cellular structure were more abundantly represented in the highly conserved class, and that genes encoding transcription factors and receptor protein kinases were abundantly represented in the less conserved class. To make the sequence information and the cDNA clones available to the research community, a Web database with useful services was created at http://www.kazusa.or.jp/en/plant/lotus/EST/.

DNA, Complementary↗

Annotation of a 95-kb Populus deltoides genomic sequence reveals a disease resistance gene cluster and novel class I and class II transposable elements.

Poplar has become a model system for functional genomics in woody plants. Here, we report the sequencing and annotation of the first large contiguous stretch of genomic sequence (95 kb) of poplar, corresponding to a bacterial artificial chromosome clone mapped 0.6 centiMorgan from the Melampsora larici-populina resistance locus. The annotation revealed 15 putative genetic objects, of which five were classified as hypothetical genes that were similar only with expressed sequence tags from poplar. Ten putative objects showed similarity with known genes, of which one was similar to a kinase. Three other objects corresponded to the toll/interleukin-1 receptor/nucleotide-binding site/leucine-rich repeat class of plant disease resistance genes, of which two were predicted to encode an amino terminal nuclear localization signal. Four objects were homologous to the Ty1/ copia family of class I transposable elements, one of which was designated Retropop and interrupted one of the disease resistance genes. Two other objects constituted a novel Spm-like class II transposable element, which we designated Magali.

Amino Acid Sequence↗

GOLEM: an interactive graph-based gene-ontology navigation and analysis tool.

BACKGROUND: The Gene Ontology has become an extremely useful tool for the analysis of genomic data and structuring of biological knowledge. Several excellent software tools for navigating the gene ontology have been developed. However, no existing system provides an interactively expandable graph-based view of the gene ontology hierarchy. Furthermore, most existing tools are web-based or require an Internet connection, will not load local annotations files, and provide either analysis or visualization functionality, but not both. RESULTS: To address the above limitations, we have developed GOLEM (Gene Ontology Local Exploration Map), a visualization and analysis tool for focused exploration of the gene ontology graph. GOLEM allows the user to dynamically expand and focus the local graph structure of the gene ontology hierarchy in the neighborhood of any chosen term. It also supports rapid analysis of an input list of genes to find enriched gene ontology terms. The GOLEM application permits the user either to utilize local gene ontology and annotations files in the absence of an Internet connection, or to access the most recent ontology and annotation information from the gene ontology webpage. GOLEM supports global and organism-specific searches by gene ontology term name, gene ontology id and gene name. CONCLUSION: GOLEM is a useful software tool for biologists interested in visualizing the local directed acyclic graph structure of the gene ontology hierarchy and searching for gene ontology terms enriched in genes of interest. It is freely available both as an application and as an applet at http://function.princeton.edu/GOLEM.

Computer Graphics↗

Processing sequence annotation data using the Lua programming language.

The data processing language in a graphical software tool that manages sequence annotation data from genome databases should provide flexible functions for the tasks in molecular biology research. Among currently available languages we adopted the Lua programming language. It fulfills our requirements to perform computational tasks for sequence map layouts, i.e. the handling of data containers, symbolic reference to data, and a simple programming syntax. Upon importing a foreign file, the original data are first decomposed in the Lua language while maintaining the original data schema. The converted data are parsed by the Lua interpreter and the contents are stored in our data warehouse. Then, portions of annotations are selected and arranged into our catalog format to be depicted on the sequence map. Our sequence visualization program was successfully implemented, embedding the Lua language for processing of annotation data and layout script. The program is available at http://staff.aist.go.jp/yutaka.ueno/guppy/.

Computational Biology↗

Improving genome annotations using phylogenetic profile anomaly detection.

MOTIVATION: A promising strategy for refining genome annotations is to detect features that conflict with known functional or evolutionary relationships between groups of genes. Previous work in this area has been focused on investigating the absence of 'housekeeping' genes or components of well-studied pathways. We have sought to develop a method for improving new annotations that can automatically synthesize and use the information available in a database of other annotated genomes. RESULTS: We show that a probabilistic model of phylogenetic profiles, trained from a database of curated genome annotations, can be used to reliably detect errors in new annotations. We use our method to identify 22 genes that were missed in previously published annotations of prokaryotic genomes. AVAILABILITY: The method was evaluated using MATLAB and open source software referenced in this work. Scripts and datasets are available from the authors upon request. CONTACT: tarjei@broad.mit.edu.

Algorithms↗

Genomic approaches to the genetics of alcoholism.

When studying complex diseases such as alcoholism that develop as a result of numerous genetic and environmental factors, researchers can use the sequence data that have become available both for the human and for animal genomes. For these analyses, investigators are being aided by efforts to identify and characterize functionally relevant DNA sequences in the entire genomic DNA sequence--a process called annotation. Various bioinformatics and annotation tools can help in this enterprise. These include four primary approaches: (1) precomputed, annotated public Web sites that provide a plethora of information; (2) in-house analyses from which users can choose the appropriate analyses for their purposes; (3) Web-based annotation systems that analyze a user's DNA sequence; and (4) private resources that provide access to annotated genomic sequences at cost. In addition to careful study of the DNA sequence for clues about function, expression studies of mRNA levels using gene chips provide information about the activity levels of thousands of genes that may vary in different tissues, different animals and people, or under different environmental conditions.

Alcoholism↗

miRGen: a database for the study of animal microRNA genomic organization and function.

miRGen is an integrated database of (i) positional relationships between animal miRNAs and genomic annotation sets and (ii) animal miRNA targets according to combinations of widely used target prediction programs. A major goal of the database is the study of the relationship between miRNA genomic organization and miRNA function. This is made possible by three integrated and user friendly interfaces. The Genomics interface allows the user to explore where whole-genome collections of miRNAs are located with respect to UCSC genome browser annotation sets such as Known Genes, Refseq Genes, Genscan predicted genes, CpG islands and pseudogenes. These miRNAs are connected through the Targets interface to their experimentally supported target genes from TarBase, as well as computationally predicted target genes from optimized intersections and unions of several widely used mammalian target prediction programs. Finally, the Clusters interface provides predicted miRNA clusters at any given inter-miRNA distance and provides specific functional information on the targets of miRNAs within each cluster. All of these unique features of miRGen are designed to facilitate investigations into miRNA genomic organization, co-transcription and targeting. miRGen can be freely accessed at http://www.diana.pcbi.upenn.edu/miRGen.

Animals↗