Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

XenDB: full length cDNA prediction and cross species mapping in Xenopus laevis.

BACKGROUND: Research using the model system Xenopus laevis has provided critical insights into the mechanisms of early vertebrate development and cell biology. Large scale sequencing efforts have provided an increasingly important resource for researchers. To provide full advantage of the available sequence, we have analyzed 350,468 Xenopus laevis Expressed Sequence Tags (ESTs) both to identify full length protein encoding sequences and to develop a unique database system to support comparative approaches between X. laevis and other model systems. DESCRIPTION: Using a suffix array based clustering approach, we have identified 25,971 clusters and 40,877 singleton sequences. Generation of a consensus sequence for each cluster resulted in 31,353 tentative contig and 4,801 singleton sequences. Using both BLASTX and FASTY comparison to five model organisms and the NR protein database, more than 15,000 sequences are predicted to encode full length proteins and these have been matched to publicly available IMAGE clones when available. Each sequence has been compared to the KOG database and approximately 67% of the sequences have been assigned a putative functional category. Based on sequence homology to mouse and human, putative GO annotations have been determined. CONCLUSION: The results of the analysis have been stored in a publicly available database XenDB http://bibiserv.techfak.uni-bielefeld.de/xendb/. A unique capability of the database is the ability to batch upload cross species queries to identify potential Xenopus homologues and their associated full length clones. Examples are provided including mapping of microarray results and application of 'in silico' analysis. The ability to quickly translate the results of various species into 'Xenopus-centric' information should greatly enhance comparative embryological approaches.

Animals↗

Prolinks: a database of protein functional linkages derived from coevolution.

The advent of whole-genome sequencing has led to methods that infer protein function and linkages. We have combined four such algorithms (phylogenetic profile, Rosetta Stone, gene neighbor and gene cluster) in a single database--Prolinks--that spans 83 organisms and includes 10 million high-confidence links. The Proteome Navigator tool allows users to browse predicted linkage networks interactively, providing accompanying annotation from public databases. The Prolinks database and the Proteome Navigator tool are available for use online at http://dip.doe-mbi.ucla.edu/pronav.

ATP Synthetase Complexes↗

A structure-based method for identifying DNA-binding proteins and their sites of DNA-interaction.

A classification model of a DNA-binding protein chain was created based on identification of alpha helices within the chain likely to bind to DNA. Using the model, all chains in the Protein Data Bank were classified. For many of the chains classified with high confidence, previous documentation for DNA-binding was found, yet no sequence homology to the structures used to train the model was detected. The result indicates that the chain model can be used to supplement sequence based methods for annotating the function of DNA-binding. Four new candidates for DNA-binding were found, including two structures solved through structural genomics efforts. For each of the candidate structures, possible sites of DNA-binding are indicated by listing the residue ranges of alpha helices likely to interact with DNA.

Binding Sites↗

Origin of the tetraspanin uroplakins and their co-evolution with associated proteins: implications for uroplakin structure and function.

Genome level information coupled with phylogenetic analysis of specific genes and gene families allow for a better understanding of the structure and function of their protein products. In this study, we examine the mammalian uroplakins (UPs) Ia and Ib, members of the tetraspanin superfamily, that interact with uroplakins UPII and UPIIIa/IIIb, respectively, using a phylogenetic approach of these genes from whole genome sequences. These proteins interact to form urothelial plaques that play a central role in the permeability barrier function of the apical urothelial surface of the urinary bladder. Since these plaques are found exclusively in mammalian urothelium, it is enigmatic that UP-like genomic sequences were recently found in lower vertebrates without a typical urothelium. We have cloned full-length UP-related cDNAs from frog (Xenopus laevis), chicken (Gallus gallus), and zebrafish (Danio rerio), and combined these data with sequence information from their orthologs in all the available fully sequenced and annotated animal genomes. Phylogenetic analyses of all the available uroplakin sequences, and an understanding of their distribution in several animal taxa, suggest that: (i) the UPIa/UPIb and UPII/UPIII genes evolved by gene duplication in the common ancestor of vertebrates; (ii) uroplakins can be lost in different combinations in vertebrate lineages; and (iii) there is a strong co-evolutionary relationship between UPIa and UPIb and their partners UPII and UPIIIa/IIIb, respectively. The co-evolution of the tetraspanin UPs and their associated proteins may fine-tune the structure and function of uroplakin complexes enabling them to perform diverse species- and tissue-specific functions. The structure and function of uroplakins, which are also expressed in Xenopus kidney, oocytes and fat body, are much more versatile than hitherto appreciated.

Amino Acid Sequence↗

Homology-based annotation yields 1,042 new candidate genes in the Drosophila melanogaster genome.

The approach to annotating a genome critically affects the number and accuracy of genes identified in the genome sequence. Genome annotation based on stringent gene identification is prone to underestimate the complement of genes encoded in a genome. In contrast, over-prediction of putative genes followed by exhaustive computational sequence, motif and structural homology search will find rarely expressed, possibly unique, new genes at the risk of including non-functional genes. We developed a two-stage approach that combines the merits of stringent genome annotation with the benefits of over-prediction. First we identify plausible genes regardless of matches with EST, cDNA or protein sequences from the organism (stage 1). In the second stage, proteins predicted from the plausible genes are compared at the protein level with EST, cDNA and protein sequences, and protein structures from other organisms (stage 2). Remote but biologically meaningful protein sequence or structure homologies provide supporting evidence for genuine genes. The method, applied to the Drosophila melanogaster genome, validated 1,042 novel candidate genes after filtering 19,410 plausible genes, of which 12,124 matched the original 13,601 annotated genes. This annotation strategy is applicable to genomes of all organisms, including human.

Animals↗

A large scale analysis of cDNA in Arabidopsis thaliana: generation of 12,028 non-redundant expressed sequence tags from normalized and size-selected cDNA libraries.

For comprehensive analysis of genes expressed in the model dicotyledonous plant, Arabidopsis thaliana, expressed sequence tags (ESTs) were accumulated. Normalized and size-selected cDNA libraries were constructed from aboveground organs, flower buds, roots, green siliques and liquid-cultured seedlings, respectively, and a total of 14,026 5'-end ESTs and 39,207 3'-end ESTs were obtained. The 3'-end ESTs could be clustered into 12,028 non-redundant groups. Similarity search of the non-redundant ESTs against the public non-redundant protein database indicated that 4816 groups show similarity to genes of known function, 1864 to hypothetical genes, and the remaining 5348 are novel sequences. Gene coverage by the non-redundant ESTs was analyzed using the annotated genomic sequences of approximately 10 Mb on chromosomes 3 and 5. A total of 923 regions were hit by at least one EST, among which only 499 regions were hit by the ESTs deposited in the public database. The result indicates that the EST source generated in this project complements the EST data in the public database and facilitates new gene discovery.

Arabidopsis↗

Disruption of the neuronal PAS3 gene in a family affected with schizophrenia.

Schizophrenia and its subtypes are part of a complex brain disorder with multiple postulated aetiologies. There is evidence that this common disease is genetically heterogeneous, with many loci involved. In this report, we describe a mother and daughter affected with schizophrenia, who are carriers of a t(9;14)(q34;q13) chromosome. By mapping on flow sorted aberrant chromosomes isolated from lymphoblast cell lines, both subjects were found to have a translocation breakpoint junction between the markers D14S730 and D14S70, a 683 kb interval on chromosome 14q13. This interval was found to contain the neuronal PAS3 gene (NPAS3), by annotating the genomic sequence for ESTs and performing RACE and cDNA library screenings. The NPAS3 gene was characterised with respect to the genomic structure, human expression profile, and protein cellular localisation to gain insight into gene function. The translocation breakpoint junction lies within the third intron of NPAS3, resulting in the disruption of the coding potential. The fact that the bHLH and PAS domains are disrupted from the remaining parts of the encoded protein suggests that the DNA binding and dimerisation functions of this protein are destroyed. The daughter (proband), who is more severely affected, has an additional microdeletion in the second intron of NPAS3. On chromosome 9q34, the translocation breakpoint junction was defined between D9S752 and D9S972 and no genes were found to be disrupted. We propose that haploinsufficiency of NPAS3 contributes to the cause of mental illness in this family.

ATP-Binding Cassette Transporters↗

The Key Trichoderma-Induced Gene Encoding a DUF568 Domain-Containing Protein Mediates Defense Responses in Wheat.

Genes encoding DUF568 domain-containing proteins participate in plant stress adaptation. To elucidate the functional role of DUF568 domain-containing genes in Trichoderma-induced wheat defense responses against wheat Fusarium crown rot, we performed a genome-wide identification and characterization of the TaDUF568 gene family in hexaploid wheat (Triticum aestivum L.). In this study, a total of 33 TaDUF568 family genes were systematically identified and characterized at the genome-wide level, exhibiting uneven chromosomal distribution and diverse physicochemical properties. Phylogenetic, structural, and collinearity analyses revealed conserved family characteristics among monocot species. Segmental duplication was verified as the primary driver of gene family expansion. Expression profiling revealed divergent tissue-specific expression patterns among TaDUF568 family members, among which TaDUF568.18 was strongly induced by Trichoderma M2. Subcellular localization assays confirmed that TaDUF568.18 is a plasma membrane-localized protein. Functional validation via stable transgenes demonstrated that overexpression of TaDUF568.18 restricted lesion expansion, improved agronomic traits, and enhanced disease resistance. This study is the first to characterize the wheat DUF568 family and confirm that TaDUF568.18 (annotated as TaAIR12) acts as a positive regulator of Trichoderma-mediated wheat defense, providing a valuable gene resource for wheat disease-resistance breeding.

DUF568↗

Motif-based fold assignment.

Conventional fold recognition techniques rely mainly on the analysis of the entire sequence of a protein. We present an MBA method to improve performance of any conventional sequence-based fold assignment. The method uses sequence motifs, such as those defined in the Prosite database, and the SwissProt annotation of the fold library. When combined with a simple SDP method, the coverage of MBA is comparable to the results obtained with PSI-BLAST. However, the set of the MBA predictions is significantly different from that of PSI-BLAST, leading to a 40% increase of the coverage for the combined MBA/PSI-BLAST method. The MBA approach can be easily adopted to include the results of sequence-independent function prediction methods and alternative motif and annotation databases. The method is available through the web server localized at http://www.doe-mbi.ucla.edu/mba.

Algorithms↗

WebFEATURE: An interactive web tool for identifying and visualizing functional sites on macromolecular structures.

WebFEATURE (http://feature.stanford.edu/webfeature/) is a web-accessible structural analysis tool that allows users to scan query structures for functional sites in both proteins and nucleic acids. WebFEATURE is the public interface to the scanning algorithm of the FEATURE package, a supervised learning algorithm for creating and identifying 3D, physicochemical motifs in molecular structures. Given an input structure or Protein Data Bank identifier (PDB ID), and a statistical model of a functional site, WebFEATURE will return rank-scored 'hits' in 3D space that identify regions in the structure where similar distributions of physicochemical properties occur relative to the site model. Users can visualize and interactively manipulate scored hits and the query structure in web browsers that support the Chime plug-in. Alternatively, results can be downloaded and visualized through other freely available molecular modeling tools, like RasMol, PyMOL and Chimera. A major application of WebFEATURE is in rapid annotation of function to structures in the context of structural genomics.

Algorithms↗

Gene expression profiling and analysis of signaling pathways involved in priming and differentiation of human neural stem cells.

Human neural stem cells have the ability to differentiate into all three major cell types in the CNS including neurons, astrocytes and oligodendrocytes. The multipotency of human neural stem cells shed a light on the possibility of using stem cells as a therapeutic tool for various neurological disorders including neurodegenerative diseases and neurotrauma that involve a loss of functional neurons. We have discovered previously a priming procedure to direct primarily cultured human neural stem cells to differentiate into almost pure neurons when grafted into adult CNS. However, the molecular mechanism underlying this phenomenon is still unknown. To unravel transcriptional changes of human neural stem cells upon priming, cDNA microarray was used to study temporal changes in human neural stem cell gene expression profile during priming and differentiation. As a result, transcriptional levels of 520 annotated genes were detected changed in at least at two time points during the priming process. In addition, transcription levels of more than 3000 hypothetical protein encoding genes and EST genes were modulated during the priming and differentiation processes of human neural stem cells. We further analyzed the named genes and grouped them into 14 functional categories. Of particular interest, key cell signal transduction pathways, including the G-protein-mediated signaling pathways (heterotrimeric and small monomeric GTPase pathways), the Wnt signaling pathway and the TGF-beta pathway, are modulated by the neural stem cell priming, suggesting important roles of these key signaling pathways in priming and differentiation of human neural stem cells.

Bone Morphogenetic Proteins↗

Identification of novel membrane proteins by searching for patterns in hydropathy profiles.

A technique has been developed to search a proteome database for new members of a functional class of membrane protein. It takes advantage of the highly conserved secondary structure of functionally related membrane proteins. Such proteins typically have the same number of transmembrane domains located at similar relative positions in their polypeptide sequence. This gives rise to a characteristic pattern of peaks in their hydropathy profiles. To conduct a search, each member of a polypeptide database is converted to a hydropathy profile, peaks are automatically detected, and the pattern of peaks is compared with a template. A template was designed for the acetylcholine (ACh) and glycine receptors of the cys-loop receptor superfamily. The key feature was a closely spaced triplet of hydropathy peaks bracketed by deep valleys. When applied to the human proteome the search procedure retrieved 153 profiles with a receptor-like triplet of peaks. The approach was highly selective with 70% of the retrieved profiles annotated as known or putative receptors. These included ACh, glycine, gamma-amino butyric acid and serotonin receptors, which are all related by sequence. However, ionotropic glutamate receptors, which have almost no sequence homology with ACh receptors, were also retrieved. Thus, the strategy can find members of a functional class that cannot be identified by sequence alignment. To demonstrate that the strategy can easily be extended to other membrane protein families, a template was developed for the neurotransmitter/Na+ symporter family, and similar results were obtained. This approach should prove a useful adjunct to sequence-based retrieval tools when searching for novel membrane proteins.

Databases, Protein↗

Clustering protein sequences with a novel metric transformed from sequence similarity scores and sequence alignments with neural networks.

BACKGROUND: The sequencing of the human genome has enabled us to access a comprehensive list of genes (both experimental and predicted) for further analysis. While a majority of the approximately 30,000 known and predicted human coding genes are characterized and have been assigned at least one function, there remains a fair number of genes (about 12,000) for which no annotation has been made. The recent sequencing of other genomes has provided us with a huge amount of auxiliary sequence data which could help in the characterization of the human genes. Clustering these sequences into families is one of the first steps to perform comparative studies across several genomes. RESULTS: Here we report a novel clustering algorithm (CLUGEN) that has been used to cluster sequences of experimentally verified and predicted proteins from all sequenced genomes using a novel distance metric which is a neural network score between a pair of protein sequences. This distance metric is based on the pairwise sequence similarity score and the similarity between their domain structures. The distance metric is the probability that a pair of protein sequences are of the same Interpro family/domain, which facilitates the modelling of transitive homology closure to detect remote homologues. The hierarchical average clustering method is applied with the new distance metric. CONCLUSION: Benchmarking studies of our algorithm versus those reported in the literature shows that our algorithm provides clustering results with lower false positive and false negative rates. The clustering algorithm is applied to cluster several eukaryotic genomes and several dozens of prokaryotic genomes.

Algorithms↗

Computational comparison of two mouse draft genomes and the human golden path.

BACKGROUND: The availability of both mouse and human draft genomes has marked the beginning of a new era of comparative mammalian genomics. The two available mouse genome assemblies, from the public mouse genome sequencing consortium and Celera Genomics, were obtained using different clone libraries and different assembly methods. RESULTS: We present here a critical comparison of the two latest mouse genome assemblies. The utility of the combined genomes is further demonstrated by comparing them with the human 'golden path' and through a subsequent analysis of a resulting conserved sequence element (CSE) database, which allows us to identify over 6,000 potential novel genes and to derive independent estimates of the number of human protein-coding genes. CONCLUSION: The Celera and public mouse assemblies differ in about 10% of the mouse genome. Each assembly has advantages over the other: Celera has higher accuracy in base-pairs and overall higher coverage of the genome; the public assembly, however, has higher sequence quality in some newly finished bacterial artificial chromosome clone (BAC) regions and the data are freely accessible. Perhaps most important, by combining both assemblies, we can get a better annotation of the human genome; in particular, we can obtain the most complete set of CSEs, one third of which are related to known genes and some others are related to other functional genomic regions. More than half the CSEs are of unknown function. From the CSEs, we estimate the total number of human protein-coding genes to be about 40,000. This searchable publicly available online CSEdb will expedite new discoveries through comparative genomics.

Animals↗

Structure and function of cytokinin oxidase/dehydrogenase genes of maize, rice, Arabidopsis and other species.

Cytokinin oxidases/dehydrogenases (CKX) catalyze the irreversible degradation of the cytokinins isopentenyladenine, zeatin, and their ribosides in a single enzymatic step by oxidative side chain cleavage. To date the sequences of 17 fully annotated CKX genes are known, including two prokaryotic genes. The CKX gene families of Arabidopsis thaliana and rice comprise seven and at least ten members, respectively. The main features of CKX genes and proteins are summarized in this review. Individual proteins differ in their catalytic properties, their subcellular localization and their expression domains. The evolutionary development of cytokinin-catabolizing gene families and the individual properties of their members indicate an important role for the fine-tuned control of catabolism to assure proper regulation of cytokinin functions. The use of CKX genes as a tool in studies of cytokinin biology and biotechnological applications is discussed.

Amino Acid Sequence↗

Thermophile-specific proteins: the gene product of aq_1292 from Aquifex aeolicus is an NTPase.

BACKGROUND: To identify thermophile-specific proteins, we performed phylogenetic patterns searches of 66 completely sequenced microbial genomes. This analysis revealed a cluster of orthologous groups (COG1618) which contains a protein from every thermophile and no sequence from 52 out of 53 mesophilic genomes. Thus, COG1618 proteins belong to the group of thermophile-specific proteins (THEPs) and therefore we here designate COG1618 proteins as THEP1s. Since no THEP1 had been analyzed biochemically thus far, we characterized the gene product of aq_1292 which is THEP1 from the hyperthermophilic bacterium Aquifex aeolicus (aaTHEP1). RESULTS: aaTHEP1 was cloned in E. coli, expressed and purified to homogeneity. At a temperature optimum between 70 and 80 degrees C, aaTHEP1 shows enzymatic activity in hydrolyzing ATP to ADP + Pi with kcat = 5 x 10(-3) s(-1) and Km = 5.5 x 10(-6) M. In addition, the enzyme exhibits GTPase activity (kcat = 9 x 10(-3) s(-1) and Km= 45 x 10(-6) M). aaTHEP1 is inhibited competitively by CTP, UTP, dATP, dGTP, dCTP, and dTTP. As shown by gel filtration, aaTHEP1 in its purified state appears as a monomer. The enzyme is resistant to limited proteolysis suggesting that it consists of a single domain. Although THEP1s are annotated as "predicted nucleotide kinases" we could not confirm such an activity experimentally. CONCLUSION: Since aaTHEP1 is the first member of COG1618 that is characterized biochemically and functional information about one member of a COG may be transferred to the entire COG, we conclude that COG1618 proteins are a family of thermophilic NTPases.

Adenosine Triphosphate↗

aCHEdb: the database system for ESTHER, the alpha/beta fold family of proteins and the Cholinesterase gene server.

Acetylcholinesterase belongs to a family of proteins, the alpha/beta hydrolase fold family, whose constituents evolutionarily diverged from a common ancestor and share a similar structure of a central beta sheet surrounded by alpha helices. These proteins fulfil a wide range of physiological functions (hydrolases, adhesion molecules, hormone precursors) [Krejci,E., Duval,N., Chatonnet,A., Vincens,P. and Massoulié,J. (1991) Proc. Natl. Acad. Sci. USA , 88, 6647-6651]. ESTHER (for esterases, alpha/beta hydrolase enzymes and relatives) is a database aimed at collecting in one information system, sequence data together with biological annotations and experimental biochemical results related to the structure-function analysis of the enzymes of the family. The major upgrade of the database comes from the use of a new database management system: aCHEdb which uses the ACeDB program designed by Richard Durbin and Jean Thierry-Mieg. It can be found at http://www.ensam.inra.fr/cholinesterase

Animals↗

Chemosensory proteins in the honey bee: Insights from the annotated genome, comparative analyses and expressional profiling.

Small chemosensory proteins (CSPs) belong to a conserved, but poorly understood protein family that has been implicated in transporting chemical stimuli within insect sensilla. However, their expression patterns suggest that these molecules are also critical for other functions including early development. Here we used both bioinformatics and experimental approaches to characterize the CSP gene family in a social insect, the Western honey bee Apis mellifera, and then compared its members to CSPs in other arthropods. The number of CSPs in the honey bee genome (six) is similar to that found in the sequenced dipteran species (four-seven), but is much lower than the number of CSPs in the moth or in the beetle (around 20 each). These differences seem to be the result of lineage specific expansions. Our analysis of CSPs in a number of arthropods reveals a conserved gene family found in both Mandibulates and Chelicerates. Expressional profiling in diverse tissues and throughout development reveals broader than expected patterns of expression with none of the CSPs restricted to the antennae and one found only in the queen ovaries and in embryos. We conclude that CSPs are multifunctional context-dependent proteins involved in diverse cellular processes ranging from embryonic development to chemosensory signal transduction. Some CSPs may function in cuticle synthesis, consistent with their evolutionary origins in the arthropods.

Amino Acid Sequence↗