Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

RIO: analyzing proteomes by automated phylogenomics using resampled inference of orthologs.

BACKGROUND: When analyzing protein sequences using sequence similarity searches, orthologous sequences (that diverged by speciation) are more reliable predictors of a new protein's function than paralogous sequences (that diverged by gene duplication). The utility of phylogenetic information in high-throughput genome annotation ("phylogenomics") is widely recognized, but existing approaches are either manual or not explicitly based on phylogenetic trees. RESULTS: Here we present RIO (Resampled Inference of Orthologs), a procedure for automated phylogenomics using explicit phylogenetic inference. RIO analyses are performed over bootstrap resampled phylogenetic trees to estimate the reliability of orthology assignments. We also introduce supplementary concepts that are helpful for functional inference. RIO has been implemented as Perl pipeline connecting several C and Java programs. It is available at http://www.genetics.wustl.edu/eddy/forester/. A web server is at http://www.rio.wustl.edu/. RIO was tested on the Arabidopsis thaliana and Caenorhabditis elegans proteomes. CONCLUSION: The RIO procedure is particularly useful for the automated detection of first representatives of novel protein subfamilies. We also describe how some orthologies can be misleading for functional inference.

Animals↗

Systematic analysis of human kinase genes: a large number of genes and alternative splicing events result in functional and structural diversity.

BACKGROUND: Protein kinases are a well defined family of proteins, characterized by the presence of a common kinase catalytic domain and playing a significant role in many important cellular processes, such as proliferation, maintenance of cell shape, apoptosis. In many members of the family, additional non-kinase domains contribute further specialization, resulting in subcellular localization, protein binding and regulation of activity, among others. About 500 genes encode members of the kinase family in the human genome, and although many of them represent well known genes, a larger number of genes code for proteins of more recent identification, or for unknown proteins identified as kinase only after computational studies. RESULTS: A systematic in silico study performed on the human genome, led to the identification of 5 genes, on chromosome 1, 11, 13, 15 and 16 respectively, and 1 pseudogene on chromosome X; some of these genes are reported as kinases from NCBI but are absent in other databases, such as KinBase. Comparative analysis of 483 gene regions and subsequent computational analysis, aimed at identifying unannotated exons, indicates that a large number of kinase may code for alternately spliced forms or be incorrectly annotated. An InterProScan automated analysis was performed to study domain distribution and combination in the various families. At the same time, other structural features were also added to the annotation process, including the putative presence of transmembrane alpha helices, and the cystein propensity to participate into a disulfide bridge. CONCLUSION: The predicted human kinome was extended by identifying both additional genes and potential splice variants, resulting in a varied panorama where functionality may be searched at the gene and protein level. Structural analysis of kinase proteins domains as defined in multiple sources together with transmembrane alpha helices and signal peptide prediction provides hints to function assignment. The results of the human kinome analysis are collected in the KinWeb database, available for browsing and searching over the internet, where all results from the comparative analysis and the gene structure annotation are made available, alongside the domain information. Kinases may be searched by domain combinations and the relative genes may be viewed in a graphic browser at various level of magnification up to gene organization on the full chromosome set.

Algorithms↗

PUMA2--grid-based high-throughput analysis of genomes and metabolic pathways.

The PUMA2 system (available at http://compbio.mcs.anl.gov/puma2) is an interactive, integrated bioinformatics environment for high-throughput genetic sequence analysis and metabolic reconstructions from sequence data. PUMA2 provides a framework for comparative and evolutionary analysis of genomic data and metabolic networks in the context of taxonomic and phenotypic information. Grid infrastructure is used to perform computationally intensive tasks. PUMA2 currently contains precomputed analysis of 213 prokaryotic, 22 eukaryotic, 650 mitochondrial and 1493 viral genomes and automated metabolic reconstructions for >200 organisms. Genomic data is annotated with information integrated from >20 sequence, structural and metabolic databases and ontologies. PUMA2 supports both automated and interactive expert-driven annotation of genomes, using a variety of publicly available bioinformatics tools. It also contains a suite of unique PUMA2 tools for automated assignment of gene function, evolutionary analysis of protein families and comparative analysis of metabolic pathways. PUMA2 allows users to submit batch sequence data for automated functional analysis and construction of metabolic models. The results of these analyses are made available to the users in the PUMA2 environment for further interactive sequence analysis and annotation.

Computational Biology↗

Detecting functional modules in the yeast protein-protein interaction network.

MOTIVATION: Identification of functional modules in protein interaction networks is a first step in understanding the organization and dynamics of cell functions. To ensure that the identified modules are biologically meaningful, network-partitioning algorithms should take into account not only topological features but also functional relationships, and identified modules should be rigorously validated. RESULTS: In this study we first integrate proteomics and microarray datasets and represent the yeast protein-protein interaction network as a weighted graph. We then extend a betweenness-based partition algorithm, and use it to identify 266 functional modules in the yeast proteome network. For validation we show that the functional modules are indeed densely connected subgraphs. In addition, genes in the same functional module confer a similar phenotype. Furthermore, known protein complexes are largely contained in the functional modules in their entirety. We also analyze an example of a functional module and show that functional modules can be useful for gene annotation. CONTACT: yuan.33@osu.edu SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

GeneTrees: a phylogenomics resource for prokaryotes.

The GeneTrees phylogenomics system pursues comparative genomic analyses from the perspective of gene phylogenies for individual genes. The GeneTrees project has the goal of providing detailed evolutionary models for all protein-coding gene components of the fully sequenced genomes. Currently, a database of alignments and trees for all protein sequences for 325 fully sequenced and annotated prokaryote genomes is available. The prokaryote database contains 890,000 protein sequences organized into over 100,000 alignments, each described by a phylogenetic tree. An original homology group discovery tool assembles sets of related proteins from all versus all pairwise alignments. Multiple alignments for each homology group are stored and subjected to phylogenetic tree inference. A graphical web interface provides visual exploration of the GeneTrees database. Homology groups can be queried by sequence identifiers or annotation terms. Genomes can be browsed visually on a gene map of each chromosome or plasmid. Phylogenetic trees with support values are displayed in conjunction with the associated sequence alignment. A variety of classes of information can be selected to label the tree tips to aid in visual evaluation of annotation and gene function. This web interface is available at http://genetrees.vbi.vt.edu.

Bacterial Proteins↗

TAIR: a resource for integrated Arabidopsis data.

The Arabidopsis Information Resource (TAIR; http://arabidopsis.org) provides an integrated view of genomic data for Arabidopsis thaliana. The information is obtained from a battery of sources, including the Arabidopsis user community, the literature, and the major genome centers. Currently TAIR provides information about genes, markers, polymorphisms, maps, sequences, clones, DNA and seed stocks, gene families and proteins. In addition, users can find Arabidopsis publications and information about Arabidopsis researchers. Our emphasis is now on incorporating functional annotations of genes and gene products, genome-wide expression, and biochemical pathway data. Among the tools developed at TAIR, the most notable is the Sequence Viewer, which displays gene annotation, clones, transcripts, markers and polymorphisms on the Arabidopsis genome, and allows zooming in to the nucleotide level. A tool recently released is AraCyc, which is designed for visualization of biochemical pathways. We are also developing tools to extract information from the literature in a systematic way, and building controlled vocabularies to describe biological concepts in collaboration with other database groups. A significant new feature is the integration of the ABRC database functions and stock ordering system, which allows users to place orders for seed and DNA stocks directly from the TAIR site.

Arabidopsis↗

Prediction of the coding sequences of unidentified human genes. XXI. The complete sequences of 60 new cDNA clones from brain which code for large proteins.

As an extension of a sequencing project of human cDNA clones which encode large proteins of unidentified genes, we herein present the entire sequences of 60 cDNA clones for the genes named KIAA1879-KIAA1938. The cDNA clones were isolated from size-fractionated cDNA libraries derived from human fetal brain, adult whole brain and amygdala, and their protein-coding sequences were predicted. Thirty-seven cDNA clones entirely sequenced in this study were selected as cDNAs which have coding potentiality by in vitro transcription/translation experiments, and the remaining 23 cDNA clones were chosen by computer-assisted analysis of terminal sequences of cDNAs. The average sizes of the inserts and corresponding open reading frames of cDNA clones analyzed here were 4.5 kb and 2.2 kb (733 amino acid residues), respectively. Sequence analyses against the public databases enabled us to annotate the functions of the predicted products of the 25 genes; 84% of these predicted gene products (21 gene products) were classified into proteins related to cell signaling/communication, nucleic acid management, and cell structure/motility. In addition to the sequence information about these 60 genes, their expression profiles were also studied in some human tissues including brain regions by reverse transcription-coupled polymerase chain reaction, products of which were quantified by enzyme-linked immunosorbent assay.

Adult↗

Rice Annotation Database (RAD): a contig-oriented database for map-based rice genomics.

A contig-oriented database for annotation of the rice genome has been constructed to facilitate map-based rice genomics. The Rice Annotation Database has the following functional features: (i) extensive effort of manual annotations of P1-derived artificial chromosome/bacterial artificial chromosome clones can be merged at chromosome and contig-level; (ii) concise visualization of the annotation information such as the predicted genes, results of various prediction programs (RiceHMM, Genscan, Genscan+, Fgenesh, GeneMark, etc.), homology to expressed sequence tag, full-length cDNA and protein; (iii) user-friendly clone / gene query system; (iv) download functions for nucleotide, amino acid and coding sequences; (v) analysis of various features of the genome (GC-content, average value, etc.); and (vi) genome-wide homology search (BLAST) of contig- and chromosome-level genome sequence to allow comparative analysis with the genome sequence of other organisms. As of October 2004, the database contains a total of 215 Mb sequence with relevant annotation results including 30 000 manually curated genes. The database can provide the latest information on manual annotation as well as a comprehensive structural analysis of various features of the rice genome. The database can be accessed at http://rad.dna.affrc.go.jp/.

Chromosomes, Plant↗

Inferring direct regulatory targets from expression and genome location analyses: a comparison of transcription factor deletion and overexpression.

BACKGROUND: Effects on gene expression due to environmental or genetic changes can be easily measured using microarrays. However, indirect effects on expression can be substantial. The indirect effects of a perturbation need to be distinguished from the direct effects if we are to understand the structure and behavior of regulatory networks. RESULTS: The most direct way to perturb a transcriptional network is to alter transcription factor activity. Here, for the first time, we compare expression changes and genomic binding in a simple regulon under conditions of both low and high transcription factor activity. Specifically, we assessed the effects on expression and binding due to deletion of the yeast LEU3 transcription factor gene and effects due to elevation of Leu3 activity. Leu3 activity was elevated through overexpression and the introduction of a mutation that renders the protein constitutively active. Genes that are bound and/or regulated by Leu3 under one or both conditions were characterized in terms of their functional annotations and their predicted potential to be bound by Leu3. We also assessed the evolutionary conservation of the predicted binding potential using a novel alignment-independent method. Both perturbations yield genes that are likely to be direct targets of Leu3, including most of the classically defined targets. Additional direct targets are identified by each of the methods. However, experimental and computational criteria suggest that most genes whose expression is affected by the Leu3 genotype are unlikely to be regulated by binding of the protein. CONCLUSION: Most genes that are differentially expressed by Leu3 are not direct targets despite the exceptional simplicity of the regulon, and the unusually direct nature of the perturbations investigated. These conclusions are reached through computational analyses that support and extend chromatin immunoprecipitation data on the identities of direct targets. These results have implications for the interpretation of expression experiments, especially in cases for which chromatin immunoprecipitation data are unavailable, incomplete, or ambiguous.

2-Isopropylmalate Synthase↗

Gene discovery and gene expression in the rice blast fungus, Magnaporthe grisea: analysis of expressed sequence tags.

Over 28,000 expressed sequence tags (ESTs) were produced from cDNA libraries representing a variety of growth conditions and cell types. Several Magnaporthe grisea strains were used to produce the libraries, including a nonpathogenic strain bearing a mutation in the PMK1 mitogen-activated protein kinase. Approximately 23,000 of the ESTs could be clustered into 3,050 contigs, leaving 5,127 singleton sequences. The estimate of 8,177 unique sequences indicates that over half of the genes of the fungus are represented in the ESTs. Analysis of EST frequency reveals growth and cell type-specific patterns of gene expression. This analysis establishes criteria for identification of fungal genes involved in pathogenesis. A large fraction of the genes represented by ESTs have no known function or described homologs. Manual annotation of the most abundant cDNAs with no known homologs allowed us to identify a family of metallothionein proteins present in M. grisea, Neurospora crassa, and Fusarium graminearum. In addition, multiply represented ESTs permitted the identification of alternatively spliced mRNA species. Alternative splicing was rare, and in most cases, the alternate mRNA forms were unspliced, although alternative 5' splice sites were also observed.

Expressed Sequence Tags↗

Crystallization and preliminary X-ray diffraction studies on the bicupin YwfC from Bacillus subtilis.

A central tenet of evolutionary biology is that proteins with diverse biochemical functions evolved from a single ancestral protein. A variation on this theme is that the functional repertoire of proteins in a living organism is enhanced by the evolution of single-chain multidomain polypeptides by gene-fusion or gene-duplication events. Proteins with a double-stranded beta-helix (cupin) scaffold perform a diverse range of functions. Bicupins are proteins with two cupin domains. There are four bicupins in Bacillus subtilis, encoded by the genes yvrK, yoaN, yxaG and ywfC. The extensive phylogenetic information on these four proteins makes them a good model system to study the evolution of function. The proteins YvrK and YoaN are oxalate decarboxylases, whereas YxaG is a quercetin dioxygenase. In an effort to aid the functional annotation of YwfC as well as to obtain a complete structure-function data set of bicupins, it was proposed to determine the crystal structure of YwfC. The bicupin YwfC was crystallized in two crystal forms. Preliminary crystallographic studies were performed on the diamond-shaped crystals, which belonged to the tetragonal space group P422. These crystals were grown using the microbatch method at 298 K. Native X-ray diffraction data from these crystals were collected to 2.2 A resolution on a home source. These crystals have unit-cell parameters a = b = 68.7, c = 211.5 A. Assuming the presence of two molecules per asymmetric unit, the V(M) value was 2.3 A3 Da(-1) and the solvent content was approximately 45%. Although the crystals appeared less frequently than the tetragonal form, YwfC also crystallizes in the monoclinic space group P2(1), with unit-cell parameters a = 46.7, b = 106.3, c = 48.7 A, beta = 92.7 degrees.

Amino Acid Sequence↗

Structure modeling of all identified G protein-coupled receptors in the human genome.

G protein-coupled receptors (GPCRs), encoded by about 5% of human genes, comprise the largest family of integral membrane proteins and act as cell surface receptors responsible for the transduction of endogenous signal into a cellular response. Although tertiary structural information is crucial for function annotation and drug design, there are few experimentally determined GPCR structures. To address this issue, we employ the recently developed threading assembly refinement (TASSER) method to generate structure predictions for all 907 putative GPCRs in the human genome. Unlike traditional homology modeling approaches, TASSER modeling does not require solved homologous template structures; moreover, it often refines the structures closer to native. These features are essential for the comprehensive modeling of all human GPCRs when close homologous templates are absent. Based on a benchmarked confidence score, approximately 820 predicted models should have the correct folds. The majority of GPCR models share the characteristic seven-transmembrane helix topology, but 45 ORFs are predicted to have different structures. This is due to GPCR fragments that are predominantly from extracellular or intracellular domains as well as database annotation errors. Our preliminary validation includes the automated modeling of bovine rhodopsin, the only solved GPCR in the Protein Data Bank. With homologous templates excluded, the final model built by TASSER has a global C(alpha) root-mean-squared deviation from native of 4.6 angstroms, with a root-mean-squared deviation in the transmembrane helix region of 2.1 angstroms. Models of several representative GPCRs are compared with mutagenesis and affinity labeling data, and consistent agreement is demonstrated. Structure clustering of the predicted models shows that GPCRs with similar structures tend to belong to a similar functional class even when their sequences are diverse. These results demonstrate the usefulness and robustness of the in silico models for GPCR functional analysis. All predicted GPCR models are freely available for noncommercial users on our Web site (http://www.bioinformatics.buffalo.edu/GPCR).

Algorithms↗

Analysis of two large functionally uncharacterized regions in the Methanopyrus kandleri AV19 genome.

BACKGROUND: For most sequenced prokaryotic genomes, about a third of the protein coding genes annotated are "orphan proteins", that is, they lack homology to known proteins. These hypothetical genes are typically short and randomly scattered throughout the genome. This trend is seen for most of the bacterial and archaeal genomes published to date. RESULTS: In contrast we have found that a large fraction of the genes coding for such orphan proteins in the Methanopyrus kandleri AV19 genome occur within two large regions. These genes have no known homologs except from other M. kandleri genes. However, analysis of their lengths, codon usage, and Ribosomal Binding Site (RBS) sequences shows that they are most likely true protein coding genes and not random open reading frames. CONCLUSIONS: Although these regions can be considered as candidates for massive lateral gene transfer, our bioinformatics analysis suggests that this is not the case. We predict many of the organism specific proteins to be transmembrane and belong to protein families that are non-randomly distributed between the regions. Consistent with this, we suggest that the two regions are most likely unrelated, and that they may be integrated plasmids.

Amino Acids↗

Microarray analysis of orthologous genes: conservation of the translational machinery across species at the sequence and expression level.

BACKGROUND: Genome projects have provided a vast amount of sequence information. Sequence comparison between species helps to establish functional catalogues within organisms and to study how they are maintained and modified across phylogenetic groups during evolution. Microarray studies allow us to determine groups of genes with similar temporal regulation and perhaps also common regulatory upstream regions for binding of transcription factors. The integration of sequence and expression data is expected to refine our current annotations and provide some insight into the evolution of gene regulation across organisms. RESULTS: We have investigated how well the protein subcellular localization and functional categories established from clustering of orthologous genes agree with gene-expression data in Saccharomyces cerevisiae. An increase in the resolution of biologically meaningful classes is observed upon the combination of experiments under different conditions. The functional categories deduced by sequence comparison approaches are, in general, preserved at the level of expression and can sometimes interact into larger co-regulated networks, such as the protein translation process. Differences and similarities in the expression between cytoplasmic-mitochondrial and interspecies translation machineries complement evolutionary information from sequence similarity. CONCLUSIONS: Combination of several microarray experiments is a powerful tool for the identification of upstream regulatory motifs of yeast genes involved in protein synthesis. Comparison of these yeast co-regulated genes against the archaeal and bacterial operons indicates that the components of the protein translation process are conserved across organisms at the expression level with minor specific adaptations.

Archaea↗

High genetic diversity in the chemoreceptor superfamily of Caenorhabditis elegans.

We investigated genetic polymorphism in the Caenorhabditis elegans srh and str chemoreceptor gene families, each of which consists of approximately 300 genes encoding seven-pass G-protein-coupled receptors. Almost one-third of the genes in each family are annotated as pseudogenes because of apparent functional defects in N2, the sequenced wild-type strain of C. elegans. More than half of these "pseudogenes" have only one apparent defect, usually a stop codon or deletion. We sequenced the defective region for 31 such genes in 22 wild isolates of C. elegans. For 10 of the 31 genes, we found an apparently functional allele in one or more wild isolates, suggesting that these are not pseudogenes but instead functional genes with a defective allele in N2. We suggest the term "flatliner" to describe genes whose functional vs. pseudogene status is unclear. Investigations of flatliner gene positions, d(N)/d(S) ratios, and phylogenetic trees indicate that they are not readily distinguished from functional genes in N2. We also report striking heterogeneity in the frequency of other polymorphisms among these genes. Finally, the large majority of polymorphism was found in just two strains from geographically isolated islands, Hawaii and Madeira. This suggests that our sampling of wild diversity in C. elegans is narrow and that identification of additional strains from similarly isolated regions will greatly expand the diversity available for study.

Alleles↗

Functional coverage of the human genome by existing structures, structural genomics targets, and homology models.

The bias in protein structure and function space resulting from experimental limitations and targeting of particular functional classes of proteins by structural biologists has long been recognized, but never continuously quantified. Using the Enzyme Commission and the Gene Ontology classifications as a reference frame, and integrating structure data from the Protein Data Bank (PDB), target sequences from the structural genomics projects, structure homology derived from the SUPERFAMILY database, and genome annotations from Ensembl and NCBI, we provide a quantified view, both at the domain and whole-protein levels, of the current and projected coverage of protein structure and function space relative to the human genome. Protein structures currently provide at least one domain that covers 37% of the functional classes identified in the genome; whole structure coverage exists for 25% of the genome. If all the structural genomics targets were solved (twice the current number of structures in the PDB), it is estimated that structures of one domain would cover 69% of the functional classes identified and complete structure coverage would be 44%. Homology models from existing experimental structures extend the 37% coverage to 56% of the genome as single domains and 25% to 31% for complete structures. Coverage from homology models is not evenly distributed by protein family, reflecting differing degrees of sequence and structure divergence within families. While these data provide coverage, conversely, they also systematically highlight functional classes of proteins for which structures should be determined. Current key functional families without structure representation are highlighted here; updated information on the "most wanted list" that should be solved is available on a weekly basis from http://function.rcsb.org:8080/pdb/function_distribution/index.html.

Databases, Protein↗

NMPdb: Database of Nuclear Matrix Proteins.

The nuclear matrix (NM) is a structure resulting from the aggregation of proteins and RNA in the nucleus of eukaryotic cells; it is the 'sticky bit' that remains after aggressive DNAse digestion and salt extraction protocols. Owing to the important role of the NM in DNA replication, DNA transcription and RNA splicing, the expression pattern of NM proteins has become an important early indicator for numerous cancers/tumors. Recent descriptions of the NM structure distinguish between a network-like 'internal nuclear matrix' (INM) and a 'nuclear shell' that connects the INM to the inner and outer nuclear membranes. A cautious NM preparation protocol reveals a coat of proteins on top of the INM; these proteins are usually referred to as the 'nuclear matrix-associated proteins'. Here, we describe a new database (NMPdb at http://www.rostlab.org/db/NMPdb/) that currently contains details of 398 NM proteins. We collected these data through a semi-automated analysis of over 3000 scientific articles in PubMed. We could match these 398 proteins to 302 protein sequences in UniProt or GenBank. Our NMPdb repository annotates these links along with the following annotations: organism, cell type, PubMed identifier, sequence-based predictions of structural and functional features and for some entries the explicit sequence segment that is responsible for localization (nuclear matrix targeting signal).

Amino Acid Sequence↗

Protein tyrosine and serine-threonine phosphatases in the sea urchin, Strongylocentrotus purpuratus: identification and potential functions.

Protein phosphatases, in coordination with protein kinases, play crucial roles in regulation of signaling pathways. To identify protein tyrosine phosphatases (PTPs) and serine-threonine (ser-thr) phosphatases in the Strongylocentrotus purpuratus genome, 179 annotated sequences were studied (122 PTPs, 57 ser-thr phosphatases). Sequence analysis identified 91 phosphatases (33 conventional PTPs, 31 dual specificity phosphatases, 1 Class III Cysteine-based PTP, 1 Asp-based PTP, and 25 ser-thr phosphatases). Using catalytic sites, levels of conservation and constraint in amino acid sequence were examined. Nine of 25 receptor PTPs (RPTPs) corresponded to human, nematode, or fly homologues. Domain structure revealed that sea urchin-specific RPTPs including two, PTPRLec and PTPRscav, may act in immune defense. Embryonic transcription of each phosphatase was recorded from a high-density oligonucleotide tiling microarray experiment. Most RPTPs are expressed at very low levels, whereas nonreceptor PTPs (NRPTPs) are generally expressed at moderate levels. High expression was detected in MAP kinase phosphatases (MKPs) and numerous ser-thr phosphatases. For several expressed NRPTPs, MKPs, and ser-thr phosphatases, morpholino antisense-mediated knockdowns were performed and phenotypes obtained. Finally, to assess roles of annotated phosphatases in endomesoderm formation, a literature review of phosphatase functions in model organisms was superimposed on sea urchin developmental pathways to predict areas of functional activity.

Animals↗