Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

Transcriptional coordination of the metabolic network in Arabidopsis.

Patterns of coexpression can reveal networks of functionally related genes and provide deeper understanding of processes requiring multiple gene products. We performed an analysis of coexpression networks for 1,330 genes from the AraCyc database of metabolic pathways in Arabidopsis (Arabidopsis thaliana). We found that genes associated with the same metabolic pathway are, on average, more highly coexpressed than genes from different pathways. Positively coexpressed genes within the same pathway tend to cluster close together in the pathway structure, while negatively correlated genes typically occupy more distant positions. The distribution of coexpression links per gene is highly skewed, with a small but significant number of genes having numerous coexpression partners but most having fewer than 10. Genes with multiple connections (hubs) tend to be single-copy genes, while genes with multiple paralogs are coexpressed with fewer genes, on average, than single-copy genes, suggesting that the network expands through gene duplication, followed by weakening of coexpression links involving duplicate nodes. Using a network-analysis algorithm based on coexpression with multiple pathway members (pathway-level coexpression), we identified and prioritized novel candidate pathway members, regulators, and cross pathway transcriptional control points for over 140 metabolic pathways. To facilitate exploration and analysis of the results, we provide a Web site (http://www.transvar.org/at_coexpress/analysis/web) listing analyzed pathways with links to regression and pathway-level coexpression results. These methods and results will aid in the prioritization of candidates for genetic analysis of metabolism in plants and contribute to the improvement of functional annotation of the Arabidopsis genome.

Arabidopsis↗

Phylogenetic inference in protein superfamilies: analysis of SH2 domains.

This work focuses on the inference of evolutionary relationships in protein superfamilies, and the uses of these relationships to identify key positions in the structure, to infer attributes on the basis of evolutionary distance, and to identify potential errors in sequence annotations. Relative entropy, a distance metric from information theory, is used in combination with Dirichlet mixture priors to estimate a phylogenetic tree for a set of proteins. This method infers key structural or functional positions in the molecule, and guides the tree topology to preserve these important positions within subtrees. Minimum-description-length principles are used to determine a cut of the tree into subtrees, to identify the subfamilies in the data. This method is demonstrated on SH2-domain containing proteins, resulting in a new subfamily assignment for Src2-drome and a suggested evolutionary relationship between Nck_human and Drk_drome, Sem5_caeel, Grb2_human and Grb2_chick.

Amino Acid Sequence↗

Correspondence of function and phylogeny of ABC proteins based on an automated analysis of 20 model protein data sets.

Using our BLAST-based procedure RiPE (Retrieval-induced Phylogeny Environment), which automates the evolutionary analysis of a protein family, we assembled a set of 1138 ABC protein components [adenosine triphosphate (ATP)-binding cassette and transmembrane domain] from the protein data sets of 20 model organisms and subjected them to phylogenetic and functional analysis. For maximum speed, we based the alignment directly on a homology search with a profile of all known human ABC proteins and used neighbor-joining tree estimation. All but 11 sequences from Homo sapiens, Arabidopsis thaliana, Drosophila melanogaster, and Saccharomyces cerevisiae were placed into the correct subtree/subfamily, reproducing published classifications of the individual organisms. By following a simple "function transfer rule", our comparative phylogenetic analysis successfully predicted the known function of human ABC proteins in 19 of 22 cases. Three functional predictions did not correspond, and 10 were novel. Predictions based on BLAST alone were inferior in five cases and superior in two. Bacterial sequences were placed close to the root of most subtrees. This placement coincides with domain architecture, suggesting an early diversification of the ABC family before the kingdoms split apart. Our approach can, in principle, be used to annotate any protein family of any organism included in the study.

ATP-Binding Cassette Transporters↗

FREP: a database of functional repeats in mouse cDNAs.

The FREP database (http://facts.gsc.riken.go.jp/FREP/) contains 31 396 RepeatMasker-identified non-redundant variant repeat sequences derived from 16,527 mouse cDNAs with protein-coding potential. The repeats were computationally associated with potential effects on transcriptional variation, translation, protein function or involvement in disease to identify Functional REPeats (FREPs). FREPs are defined by the (i) occurrence of exon-exon boundaries in repeats, (ii) presence of polyadenylation sites in 3'UTR-located repeats, (iii) effect on translation, (iv) position in the protein- coding region or protein domains or (v) conditional association with disease MeSH terms. Currently the database contains 9261 (29.5%) inferred FREPs derived from 6861 (41.5%) mouse cDNAs. Integrated evidence of the functional assignments and dynamically generated sequence similarity search results support the exploration and annotation of functional, ancestral or taxon-specific repeats. Keyword and pre-selected feature searches (e.g. coding sequence-repeat or splice site-repeat relations) support intuitive database querying as well as the retrieval of repeat sequences. Integrated sequence search and alignment tools allow the analysis of known or identification of new functional repeat candidates. FREP is a unique resource for illuminating the role of transposons and repetitive sequences in shaping the coding part of the mouse transcriptome and for selecting the appropriate experimental model to study diseases with suspected repeat etiology contributions.

Animals↗

Cataloging transcription factor and major signaling molecule genes for functional genomic studies in Ciona intestinalis.

The ascidian Ciona intestinalis provides an excellent experimental system for functional genomic studies because (1) its genome has been sequenced, (2) the transcription factor genes and genes for major signal transduction molecules have been extensively screened and annotated on a genome-wide scale using the molecular phylogenetical method, and (3) their embryonic expression profiles have been almost completely determined. However, the entire genetic structure, including the 5' and 3' untranslated regions and the protein-coding regions, of most gene models used in these prior studies is not always supported by cDNA evidence, and thus, these gene models are potentially imprecise. To facilitate functional genomic studies based on precise gene structures, our present study determined 406 cDNA sequences for 357 transcription factor genes and 112 cDNA sequences for 107 signal transduction molecule genes, greatly improving the previous gene models and revealing transcript variants for 44 genes. Considering these data alongside those of previously characterized genes deposited in the DNA Data Bank of Japan/European Molecular Biology Laboratory/GENBANK databases, 95.6% of the catalogued transcription factor genes (373/390) and 98.3% of the catalogued signal transduction molecule genes (117/119) have now been verified by cDNA sequences. Thus, the present study greatly improves the resources available for functional genomic studies in C. intestinalis.

Animals↗

SOX4 expression in bladder carcinoma: clinical aspects and in vitro functional characterization.

The human transcription factor SOX4 was 5-fold up-regulated in bladder tumors compared with normal tissue based on whole-genome expression profiling of 166 clinical bladder tumor samples and 27 normal urothelium samples. Using a SOX4-specific antibody, we found that the cancer cells expressed the SOX4 protein and, thus, did an evaluation of SOX4 protein expression in 2,360 bladder tumors using a tissue microarray with clinical annotation. We found a correlation (P < 0.05) between strong SOX4 expression and increased patient survival. When overexpressed in the bladder cell line HU609, SOX4 strongly impaired cell viability and promoted apoptosis. To characterize downstream target genes and SOX4-induced pathways, we used a time-course global expression study of the overexpressed SOX4. Analysis of the microarray data showed 130 novel SOX4-related genes, some involved in signal transduction (MAP2K5), angiogenesis (NRP2), and cell cycle arrest (PIK3R3) and others with unknown functions (CGI-62). Among the genes regulated by SOX4, 25 contained at least one SOX4-binding motif in the promoter sequence, suggesting a direct binding of SOX4. The gene set identified in vitro was analyzed in the clinical bladder material and a small subset of the genes showed a high correlation to SOX4 expression. The present data suggest a role of SOX4 in the bladder cancer disease.

Apoptosis↗

MMDB: Entrez's 3D-structure database.

Three-dimensional structures are now known within many protein families and it is quite likely, in searching a sequence database, that one will encounter a homolog with known structure. The goal of Entrez's 3D-structure database is to make this information, and the functional annotation it can provide, easily accessible to molecular biologists. To this end Entrez's search engine provides three powerful features. (i) Sequence and structure neighbors; one may select all sequences similar to one of interest, for example, and link to any known 3D structures. (ii) Links between databases; one may search by term matching in MEDLINE, for example, and link to 3D structures reported in these articles. (iii) Sequence and structure visualization; identifying a homolog with known structure, one may view molecular-graphic and alignment displays, to infer approximate 3D structure. In this article we focus on two features of Entrez's Molecular Modeling Database (MMDB) not described previously: links from individual biopolymer chains within 3D structures to a systematic taxonomy of organisms represented in molecular databases, and links from individual chains (and compact 3D domains within them) to structure neighbors, other chains (and 3D domains) with similar 3D structure. MMDB may be accessed at http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Structure.

Animals↗

Proteome Analyst: custom predictions with explanations in a web-based tool for high-throughput proteome annotations.

Proteome Analyst (PA) (http://www.cs.ualberta.ca/~bioinfo/PA/) is a publicly available, high-throughput, web-based system for predicting various properties of each protein in an entire proteome. Using machine-learned classifiers, PA can predict, for example, the GeneQuiz general function and Gene Ontology (GO) molecular function of a protein. In addition, PA is currently the most accurate and most comprehensive system for predicting subcellular localization, the location within a cell where a protein performs its main function. Two other capabilities of PA are notable. First, PA can create a custom classifier to predict a new property, without requiring any programming, based on labeled training data (i.e. a set of examples, each with the correct classification label) provided by a user. PA has been used to create custom classifiers for potassium-ion channel proteins and other general function ontologies. Second, PA provides a sophisticated explanation feature that shows why one prediction is chosen over another. The PA system produces a Naïve Bayes classifier, which is amenable to a graphical and interactive approach to explanations for its predictions; transparent predictions increase the user's confidence in, and understanding of, PA.

Internet↗

Polymorphism, shared functions and convergent evolution of genes with sequences coding for polyalanine domains.

Mutations causing expansions of polyalanine domains are responsible for nine hereditary diseases. Other GC-rich sequences coding for some polyalanine domains were found to be polymorphic in human. These observations prompted us to identify all sequences in the human genome coding for polyalanine stretches longer than four alanines and establish their degree of polymorphism. We identified 494 annotated human proteins containing 604 polyalanine domains. Thirty-two percent (31/98) of tested sequences coding for more than seven alanines were polymorphic. The length of the polyalanine-coding sequence and its GCG or GCC repeat content are the major predictors of polymorphism. GCG codons are over-represented in human polyalanine coding sequences. Our data suggest that GCG and GCC codons play a key role in polyalanine-coding sequence appearance and polymorphism. The grouping by shared function of polyalanine-containing proteins in Homo sapiens, Drosophila melanogaster and Caenorhabditis elegans shows that the majority are involved in transcriptional regulation. Phylogenetic analyses of HOX, GATA and EVX protein families demonstrate that polyalanine domains arose independently in different members of these families, suggesting that convergent molecular evolution may have played a role. Finally polyalanine domains in vertebrates are conserved between mammals and are rarer and shorter in Gallus gallus and Danio rerio. Together our results show that the polymorphic nature of sequences coding for polyalanine domains makes them prime candidates for mutations in hereditary diseases and suggests that they have appeared in many different protein families through convergent evolution.

Amino Acid Sequence↗

Revisiting the prediction of protein function at CASP6.

The ability to predict the function of a protein, given its sequence and/or 3D structure, is an essential requirement for exploiting the wealth of data made available by genomics and structural genomics projects and is therefore raising increasing interest in the computational biology community. To foster developments in the area as well as to establish the state of the art of present methods, a function prediction category was tentatively introduced in the 6th edition of the Critical Assessment of Techniques for Protein Structure Prediction (CASP) worldwide experiment. The assessment of the performance of the methods was made difficult by at least two factors: (a) the experimentally determined function of the targets was not available at the time of assessment; (b) the experiment is run blindly, preventing verification of whether the convergence of different predictions towards the same functional annotation was due to the similarity of the methods or to a genuine signal detectable by different methodologies. In this work, we collected information about the methods used by the various predictors and revisited the results of the experiment by verifying how often and in which cases a convergent prediction was obtained by methods based on different rationale. We propose a method for classifying the type and redundancy of the methods. We also analyzed the cases in which a function for the target protein has become available. Our results show that predictions derived from a consensus of different methods can reach an accuracy as high as 80%. It follows that some of the predictions submitted to CASP6, once reanalyzed taking into account the type of converging methods, can provide very useful information to researchers interested in the function of the target proteins.

Caspase 6↗

Neurobiology and the Drosophila genome.

The sequencing and annotation of the euchromatin of the Drosophila melanogaster genome provides an important foundation that allows neurobiologists to work back from the complete gene set of neuronal proteins to an eventual understanding of how they function to produce cognition and behavior. Here we provide a brief survey of some of the key insights that have emerged from analyzing the complete gene set in Drosophila. Not surprisingly, both the Caenorhabditis elegans and Drosophila genomes contain a conserved repertoire of neuronal signaling proteins that are also present in mammals. This includes a large number of neuronal cell adhesion receptors, synapse-organizing proteins, ion channels and neurotransmitter receptors, and synaptic vesicle-trafficking proteins. In addition, there are a significant number of fly homologs of human neurological disease loci, suggesting that Drosophila is likely to be an important disease model for human neuropathology in the near future. The experimental analysis of the Drosophila neuronal gene set will provide important insights into how the nervous system functions at the cellular level, allowing the field to integrate this information into the framework of ultimately understanding how neuronal ensembles mediate cognition and behavior.

Animals↗

Integrating biological databases.

Recent years have seen an explosion in the amount of available biological data. More and more genomes are being sequenced and annotated, and protein and gene interaction data are accumulating. Biological databases have been invaluable for managing these data and for making them accessible. Depending on the data that they contain, the databases fulfil different functions. But, although they are architecturally similar, so far their integration has proved problematic.

Animals↗

Protein interaction analysis of SCF ubiquitin E3 ligase subunits from Arabidopsis.

Ubiquitin E3 ligases are a diverse family of protein complexes that mediate the ubiquitination and subsequent proteolytic turnover of proteins in a highly specific manner. Among the several classes of ubiquitin E3 ligases, the Skp1-Cullin-F-box (SCF) class is generally comprised of three 'core' subunits: Skp1 and Cullin, plus at least one F-box protein (FBP) subunit that imparts specificity for the ubiquitination of selected target proteins. Recent genetic and biochemical evidence in Arabidopsis thaliana suggests that post-translational turnover of proteins mediated by SCF complexes is important for the regulation of diverse developmental and environmental response pathways. In this report, we extend upon a previous annotation of the Arabidopsis Skp1-like (ASK) and FBP gene families to include the Cullin family of proteins. Analysis of the protein interaction profiles involving the products of all three gene families suggests a functional distinction between ASK proteins in that selected members of the protein family interact generally while others interact more specifically with members of the F-box protein family. Analysis of the interaction of Cullins with FBPs indicates that CUL1 and CUL2, but not CUL3A, persist as components of selected SCF complexes, suggesting some degree of functional specialization for these proteins. Yeast two-hybrid analyses also revealed binary protein interactions between selected members of the FBP family in Arabidopsis. These and related results are discussed in terms of their implications for subunit composition, stoichiometry and functional diversity of SCF complexes in Arabidopsis.

Amino Acid Sequence↗

FireDB--a database of functionally important residues from proteins of known structure.

The FireDB database is a databank for functional information relating to proteins with known structures. It contains the most comprehensive and detailed repository of known functionally important residues, bringing together both ligand binding and catalytic residues in one site. The platform integrates biologically relevant data filtered from the close atomic contacts in Protein Data Bank crystal structures and reliably annotated catalytic residues from the Catalytic Site Atlas. The interface allows users to make queries by protein, ligand or keyword. Relevant biologically important residues are displayed in a simple and easy to read manner that allows users to assess binding site similarity across homologous proteins. Binding site residue variations can also be viewed with molecular visualization tools. The database is available at http://firedb.bioinfo.cnio.es.

Amino Acids↗

The fasciclin-like arabinogalactan proteins of Arabidopsis. A multigene family of putative cell adhesion molecules.

Fasciclin-like arabinogalactan proteins (FLAs) are a subclass of arabinogalactan proteins (AGPs) that have, in addition to predicted AGP-like glycosylated regions, putative cell adhesion domains known as fasciclin domains. In other eukaryotes (e.g. fruitfly [Drosophila melanogaster] and humans [Homo sapiens]), fasciclin domain-containing proteins are involved in cell adhesion. There are at least 21 FLAs in the annotated Arabidopsis genome. Despite the deduced proteins having low overall similarity, sequence analysis of the fasciclin domains in Arabidopsis FLAs identified two highly conserved regions that define this motif, suggesting that the cell adhesion function is conserved. We show that FLAs precipitate with beta-glucosyl Yariv reagent, indicating that they share structural characteristics with AGPs. Fourteen of the FLA family members are predicted to be C-terminally substituted with a glycosylphosphatidylinositol anchor, a cleavable form of membrane anchor for proteins, indicating different FLAs may have different developmental roles. Publicly available microarray and expressed sequence tag data were used to select FLAs for further expression analysis. RNA gel blots for a number of FLAs indicate that they are likely to be important during plant development and in response to abiotic stress. FLAs 1,2, and 8 show a rapid decrease in mRNA abundance in response to the phytohormone abscisic acid. Also, the accumulation of FLA1 and FLA2 transcripts differs during callus and shoot development, indicating that the proteins may be significant in the process of competence acquisition and induction of shoot development.

Abscisic Acid↗

From information management to protein annotation: preparing protein structures for drug discovery.

In contrast to academic pursuits of structural genomics, Structural GenomiX (SGX) solves protein structures at high throughput for the main purpose of enhancing drug-discovery projects, either internally or in partnership with pharmaceutical/biotechnology companies. This involves a radical redesign of the pipeline of methods that turn a gene sequence into a three-dimensional protein structure. The various processes all report electronically to a Laboratory Information Management System (LIMS) to make sure all the parameters of the experiment are recorded in an accessible and 'mineable' form, helping guarantee reproducibility of results. Quality control at several key points keeps the process from branching out on a wrong hypothesis. Protein annotation, in a broad sense, takes care of the interpretation of a protein crystal structure or the crystal structure of one or several protein-ligand complexes. This interpretation both gathers all necessary biological information (protein function, mechanism, specific features within a protein family etc.) and hands over this information in a form accessible to medicinal chemistry teams designing specific small-molecule agonists or antagonists.

Crystallization↗

Transcriptome analyses of human genes and applications for proteome analyses.

By utilizing recently developed full-length cDNA technologies, large-scale cDNA sequencing was carried out by several cDNA projects. Now full-length cDNA resources cover the major part of the protein-coding human genes. Comprehensive analyses of the collected full-length cDNA data revealed not only the complete sequences of thousands of novel gene transcripts but also novel alternatively spliced isoforms of hitherto identified genes. However, it was not as easy as expected to deduce their encoded amino acid sequences based solely on the full-length cDNA sequences. It was neither always the case that the longest open reading frame corresponded to the real protein coding region nor that the first ATG was the translation initiator codon. Also, proteome-wide mass-spectrometry analysis has shown that there is an unexpectedly large population of small proteins, encoded by so-called upstream open reading frames, within the cell. Since sound manual annotations by experts were still indispensable to address these problems, an international meeting to make transcriptome-wide functional annotations of cDNAs was held, namely the H-invitational. In this meeting, functional annotations were made both manually and computationally for most of the pre-existing full-length cDNAs collected from world-wide cDNA projects. The achieved integrated information for each of the cDNAs was published as a database. It was also shown that the full-length cDNA data were useful for identifying alternative splicing variants, exact transcriptional start sites of the mRNAs and the adjacent promoter regions. Rapidly accumulating genome data as well as versatile use of the transcriptome information will shortly lay a firm foundation for proteome-level understanding of human gene networks.

Alternative Splicing↗

Identification and expression of the cym, cmt, and tod catabolic genes from Pseudomonas putida KL47: expression of the regulatory todST genes as a factor for catabolic adaptation.

Pseudomonas putida KL47 is a natural isolate that assimilates benzene, 1-alkylbenzene (C(1)-C(4)), biphenyl, p-cumate, and p-cymene. The genetic background of strain KL47 underlying the broad range of growth substrates was examined. It was found that the cym and cmt operons are constitutively expressed due to a lack of the cymR gene, and the tod operon is still inducible by toluene and biphenyl. The entire array of gene clusters responsible for the catabolism of toluene and p-cymene/p-cumate has been cloned in a cosmid vector, pLAFR3, and were named pEK6 and pEK27, respectively. The two inserts overlap one another and the nucleotide sequence (42,505 bp) comprising the cym, cmt, and tod operons and its flanking genes in KL47 are almost identical (>99%) to those of P. putida F1. In the cloned DNA fragment, two genes with unknown functions, labeled cymZ and cmtR, were newly identified and show high sequence homology to dienelactone hydrolase and CymR proteins, respectively. The cmtR gene was identified in the place of the cmtI gene of previous annotation. Western blot analysis showed that, in strains F1 and KL47, the todT gene is not expressed during growth on Luria Bertani medium. In minimal basal salt medium, expression of the todT gene is inducible by toluene, but not by biphenyl in strain F1; however, it is constantly expressed in strain KL47, indicating that high levels of expression of the todST genes with one amino acid substitution in TodS might provide strain KL47 with a means of adaptation of the tod catabolic operon to various aromatic hydrocarbons.

Amino Acid Sequence↗