Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

High-throughput protein analysis integrating bioinformatics and experimental assays.

The wealth of transcript information that has been made publicly available in recent years requires the development of high-throughput functional genomics and proteomics approaches for its analysis. Such approaches need suitable data integration procedures and a high level of automation in order to gain maximum benefit from the results generated. We have designed an automatic pipeline to analyse annotated open reading frames (ORFs) stemming from full-length cDNAs produced mainly by the German cDNA Consortium. The ORFs are cloned into expression vectors for use in large-scale assays such as the determination of subcellular protein localization or kinase reaction specificity. Additionally, all identified ORFs undergo exhaustive bioinformatic analysis such as similarity searches, protein domain architecture determination and prediction of physicochemical characteristics and secondary structure, using a wide variety of bioinformatic methods in combination with the most up-to-date public databases (e.g. PRINTS, BLOCKS, INTERPRO, PROSITE SWISSPROT). Data from experimental results and from the bioinformatic analysis are integrated and stored in a relational database (MS SQL-Server), which makes it possible for researchers to find answers to biological questions easily, thereby speeding up the selection of targets for further analysis. The designed pipeline constitutes a new automatic approach to obtaining and administrating relevant biological data from high-throughput investigations of cDNAs in order to systematically identify and characterize novel genes, as well as to comprehensively describe the function of the encoded proteins.

Automation↗

Expression, purification and crystallization of the ammonium transporter Amt-1 from Archaeoglobus fulgidus.

Ammonium transporters (Amts) are a class of membrane-integral transport proteins found in organisms from all kingdoms of life. Their key function is the transport of nitrogen in its reduced bioavailable form, ammonia, across cellular membranes, a crucial step in nitrogen assimilation for biosynthetic purposes. The genome of the hyperthermophilic archaeon Archaeoglobus fulgidus has been annotated with three individual genes for ammonium transporters, amt1-3, the roles of which are as yet unknown. The amt1 gene product has been produced by heterologous overexpression in Escherichia coli and the resulting protein has been purified to electrophoretic homogeneity. Crystals of Amt-1 have been obtained by sitting-drop vapour diffusion and diffraction data have been collected.

Archaeal Proteins↗

Quantitative evolutionary genomics: differential gene expression and male reproductive success in Drosophila melanogaster.

We combined traditional quantitative genetics and oligonucleotide microarrays to examine within-population genetic variation in a trait closely related to fitness. The trait, male reproductive success under competitive conditions (MCRS), is of central importance to both life-history and sexual-selection theory. We identified 27 candidate genes whose expression levels were associated with within-population variation in MCRS. "High" MCRS was associated with low expression of a cytochrome P450 that causes pesticide resistance, suggesting a fitness cost to resistance. Two groups of metabolic proteins (glutathione transferases and phosphatases) were significantly over-represented, and a large portion of the candidates are genes involved in oxidative stress resistance, energy acquisition or energy storage. Genes expressed in accessory glands and testes were not over-represented among differentially expressed genes, but testis-expressed genes were significantly more likely to be upregulated in high MCRS genotypes. Finally, nine candidate genes that we identified had no previous functional annotation, and this experiment suggests that they play a role in male reproductive success.

Animals↗

Functional discrimination of gene expression patterns in terms of the gene ontology.

The ever-growing amount of experimental data in molecular biology and genetics requires its automated analysis, by employing sophisticated knowledge discovery tools. We use an Inductive Logic Programming (ILP) learner to induce functional discrimination rules between genes studied using microarrays and found to be differentially expressed in three recently discovered subtypes of adenocarcinoma of the lung. The discrimination rules involve functional annotations from the Proteome HumanPSD database in terms of the Gene Ontology, whose hierarchical structure is essential for this task. While most of the lower levels of gene expression data (pre)processing have been automated, our work can be seen as a step toward automating the higher level functional analysis of the data. We view our application not just as a prototypical example of applying more sophisticated machine learning techniques to the functional analysis of genes, but also as an incentive for developing increasingly more sophisticated functional annotations and ontologies, that can be automatically processed by such learning algorithms.

Adenocarcinoma↗

Information decay in molecular docking screens against holo, apo, and modeled conformations of enzymes.

Molecular docking uses the three-dimensional structure of a receptor to screen a small molecule database for potential ligands. The dependence of docking screens on the conformation of the binding site remains an open question. To evaluate the information loss that occurs as the active site conformation becomes less defined, a small molecule database was docked against the holo (ligand bound), apo, and homology modeled structures of 10 different enzyme binding sites. The holo and apo representations were crystallographic structures taken from the Protein Data Bank (PDB), and the homology-modeled structures were taken from the publicly available resource ModBase. The database docked was the MDL Drug Data Report (MDDR), a functionally annotated database of 95000 small molecules that contained at least 35 ligands for each of the 10 systems. In all sites, at least 99% of the molecules in the MDDR were treated as nonbinding decoys. For each system, the holo, apo, and modeled structures were used to screen the MDDR, and the ability of each structure to enrich the known ligands for that system over random selection was evaluated. The best overall enrichment was produced by the holo structure in seven systems, the apo structure in two systems, and the modeled structure in one system. These results suggest that the performance of the docking calculation is affected by the particular representation of the receptor used in the screen, and that the holo structure is the one most likely to yield the best discrimination between known ligands and decoy molecules, but important exceptions to this rule also emerge from this study. Although each of the holo, apo, and modeled conformations led to enrichment of known ligands in all systems, the enrichment did not always rise to a level judged to be sufficient to justify the effort of a docking screen. Using a 20-fold enrichment of known ligands over random selection as a rough guideline for what might be enough to justify a docking screen, the holo conformation of the enzyme met this criterion in eight of 10 sites, whereas the apo conformation met this criterion in only two sites and the modeled conformation in three.

Animals↗

NovelFam3000--uncharacterized human protein domains conserved across model organisms.

BACKGROUND: Despite significant efforts from the research community, an extensive portion of the proteins encoded by human genes lack an assigned cellular function. Most metazoan proteins are composed of structural and/or functional domains, of which many appear in multiple proteins. Once a domain is characterized in one protein, the presence of a similar sequence in an uncharacterized protein serves as a basis for inference of function. Thus knowledge of a domain's function, or the protein within which it arises, can facilitate the analysis of an entire set of proteins. DESCRIPTION: From the Pfam domain database, we extracted uncharacterized protein domains represented in proteins from humans, worms, and flies. A data centre was created to facilitate the analysis of the uncharacterized domain-containing proteins. The centre both provides researchers with links to dispersed internet resources containing gene-specific experimental data and enables them to post relevant experimental results or comments. For each human gene in the system, a characterization score is posted, allowing users to track the progress of characterization over time or to identify for study uncharacterized domains in well-characterized genes. As a test of the system, a subset of 39 domains was selected for analysis and the experimental results posted to the NovelFam3000 system. For 25 human protein members of these 39 domain families, detailed sub-cellular localizations were determined. Specific observations are presented based on the analysis of the integrated information provided through the online NovelFam3000 system. CONCLUSION: Consistent experimental results between multiple members of a domain family allow for inferences of the domain's functional role. We unite bioinformatics resources and experimental data in order to accelerate the functional characterization of scarcely annotated domain families.

Animals↗

Cloning and initial characterization of the Arabidopsis thaliana endoplasmic reticulum oxidoreductins.

The oxidation and isomerization of disulfide bonds is necessary for the growth of all organisms. In yeast, the oxidative folding of secretory pathway proteins is catalyzed by protein disulfide isomerase (PDI), which requires Ero1p (endoplasmic reticulum oxidoreductin) for its own oxidation. In Homo sapiens, two homologues of Ero1p, Ero1-Lalpha and Ero1-Lbeta, have been cloned. Both Ero1-Lalpha and Ero1-Lbeta interact via disulfide bonds with PDI and support the oxidation of immunoglobulin light chains. However, the function of Ero proteins in plants has not yet been analyzed. In this article, we report the cloning of the two Ero1p homologues present in Arabidopsis thaliana, demonstrating that one of the cDNAs has a shorter terminal exon than predicted and differs from the annotated sequence found in the genome database. Sequence analysis of the Arabidopsis endoplasmic reticulum oxidoreductins (AEROs) reveals that both AERO1 and AERO2 are more closely related to each other than to either of the human Eros. Both in vitro translated AERO proteins are targeted to the endoplasmic reticulum and glycosylated. The ability to use a genetically tractable multicellular organism in combination with biochemical approaches should further our understanding of redox networks and Ero function in both plants and animals.

Arabidopsis↗

Prioritizing regions of candidate genes for efficient mutation screening.

The availability of the complete sequence of the human genome has dramatically facilitated the search for disease-causing sequence variations. In fact, the rate-limiting step has shifted from the discovery and characterization of candidate genes to the actual screening of human populations and the subsequent interpretation of observed variations. In this study we tested the hypothesis that some segments of candidate genes are more likely than others to contain disease-causing variations and that these segments can be predicted bioinformatically. A bioinformatic technique, prioritization of annotated regions (PAR), was developed to predict the likelihood that a specific coding region of a gene will harbor a disease-causing mutation based on conserved protein functional domains and protein secondary structures. This method was evaluated by using it to analyze 710 genes that collectively harbor 4,498 previously identified mutations. Nearly 50% of the genes were recognized as disease-associated after screening only 9% of the complete coding sequence. The PAR technique identified 90% of the genes as containing at least one mutation, with less than 40% of the screening resources that traditional approaches would require. These results suggest that prioritization strategies such as PAR can accelerate disease-gene identification through more efficient use of screening resources.

Computational Biology↗

Transposable elements as a significant source of transcription regulating signals.

Transposable elements (TEs) are major components of eukaryotic genomes, contributing about 50% to the size of mammalian genomes. TEs serve as recombination hot spots and may acquire specific cellular functions, such as controlling protein translation and gene transcription. The latter is the subject of the analysis presented. We scanned TE sequences located in promoter regions of all annotated genes in the human genome for their content in potential transcription regulating signals. All investigated signals are likely to be over-represented in at least one TE class, which shows that TEs have an important potential to contribute to pre-transcriptional gene regulation, especially by moving transcriptional signals within the genome and thus potentially leading to new gene expression patterns. We also found that some TE classes are more likely than others to carry transcription regulating signals, which can explain why they have different retention rates in regions neighboring genes.

Base Sequence↗

mPPP1R16B is a novel mouse protein phosphatase 1 targeting subunit whose mRNA is located in cell bodies and dendrites of neurons in four distinct regions of the brain.

We cloned a cDNA encoding a novel mouse protein whose human homolog has been annotated in GenBank as a regulatory subunit of protein phosphatase 1, PPP1R16B. Both the primary protein sequence and the domain structure are highly conserved between PPP1R16B and proteins of unknown function from other species, such as Caenorhabditis elegans and Drosphila melanogaster. Besides a protein phosphatase 1 interaction motif, mouse PPP1R16B (mPPP1R16B) and the related proteins contain ankyrin repeats that may constitute binding sites for other proteins and C-terminal prenylation signals that are likely to target the proteins to the plasma membrane. In the adult mouse, Ppp1r16b mRNA is expressed in most tissues examined, with highest expression levels in kidney and brain. In the brain, Ppp1r16b message is particularly enriched in the olfactory bulb, striatum, dentate gyrus, and cerebellum. During postnatal cerebellar development, Ppp1r16b mRNA expression levels increase gradually and are maximal around postnatal day 30. In situ hybridization revealed that Ppp1r16b message is found in both the cell bodies and the dendrites in Purkinje cells of the cerebellum and granule neurons of the dentate gyrus.

Amino Acid Sequence↗

Identification of a novel protein with guanylyl cyclase activity in Arabidopsis thaliana.

Guanylyl cyclases (GCs) catalyze the formation of the second messenger guanosine 3',5'-cyclic monophosphate (cGMP) from guanosine 5'-triphosphate (GTP). While many cGMP-mediated processes in plants have been reported, no plant molecule with GC activity has been identified. When the Arabidopsis thaliana genome is queried with GC sequences from cyanobacteria, lower and higher eukaryotes no unassigned proteins with significant similarity are found. However, a motif search of the A. thaliana genome based on conserved and functionally assigned amino acids in the catalytic center of annotated GCs returns one candidate that also contains the adjacent glycine-rich domain typical for GCs. In this molecule, termed AtGC1, the catalytic domain is in the N-terminal part. AtGC1 contains the arginine or lysine that participates in hydrogen bonding with guanine and the cysteine that confers substrate specificity for GTP. When AtGC1 is expressed in Escherichia coli, cell extracts yield >2.5 times more cGMP than control extracts and this increase is not nitric oxide dependent. Furthermore, purified recombinant AtGC1 has Mg(2+)-dependent GC activity in vitro and >3 times less adenylyl cyclase activity when assayed with ATP as substrate in the absence of GTP. Catalytic activity in vitro proves that AtGC1 can function either as a monomer or homo-oligomer. AtGC1 is thus not only the first functional plant GC but also, due to its unusual domain organization, a member of a new class of GCs.

Amino Acid Sequence↗

Transmembrane proteins in the Protein Data Bank: identification and classification.

MOTIVATION: Integral membrane proteins play important roles in living cells. Although these proteins are estimated to constitute 25% of proteins at a genomic scale, the Protein Data Bank (PDB) contains only a few hundred membrane proteins due to the difficulties with experimental techniques. The presence of transmembrane proteins in the structure data bank, however, is quite invisible, as the annotation of these entries is rather poor. Even if a protein is identified as a transmembrane one, the possible location of the lipid bilayer is not indicated in the PDB because these proteins are crystallized without their natural lipid bilayer, and currently no method is publicly available to detect the possible membrane plane using the atomic coordinates of membrane proteins. RESULTS: Here, we present a new geometrical approach to distinguish between transmembrane and globular proteins using structural information only and to locate the most likely position of the lipid bilayer. An automated algorithm (TMDET) is given to determine the membrane planes relative to the position of atomic coordinates, together with a discrimination function which is able to separate transmembrane and globular proteins even in cases of low resolution or incomplete structures such as fragments or parts of large multi chain complexes. This method can be used for the proper annotation of protein structures containing transmembrane segments and paves the way to an up-to-date database containing the structure of all known transmembrane proteins and fragments (PDB_TM) which can be automatically updated. The algorithm is equally important for the purpose of constructing databases purely of globular proteins.

Algorithms↗

PatSearch: A program for the detection of patterns and structural motifs in nucleotide sequences.

Regulation of gene expression at transcriptional and post-transcriptional level involves the interaction between short DNA or RNA tracts and the corresponding trans-acting protein factors. Detection of such cis-acting elements in genome-wide screenings may significantly contribute to genome annotation and comparative analysis as well as to target functional characterization experiments. We present here PatSearch, a flexible and fast pattern matcher able to search for specific combinations of oligonucleotide consensus sequences, secondary structure elements and position-weight matrices. It can also allow for mismatches/mispairings below a user fixed threshold. We report three different applications of the program in the search of complex patterns such as those of the iron responsive element hairpin-loop structure, the p53 responsive element and a promoter module containing CAAT-, TATA- and cap-boxes. PatSearch is available on the web at http://bighost.area.ba.cnr.it/BIG/PatSearch/.

Base Sequence↗

Functional annotation of a full-length Arabidopsis cDNA collection.

Full-length complementary DNAs (cDNAs) are essential for the correct annotation of genomic sequences and for the functional analysis of genes and their products. We isolated 155,144 RIKEN Arabidopsis full-length (RAFL) cDNA clones. The 3'-end expressed sequence tags (ESTs) of 155,144 RAFL cDNAs were clustered into 14,668 nonredundant cDNA groups, about 60% of predicted genes. We also obtained 5' ESTs from 14,034 nonredundant cDNA groups and constructed a promoter database. The sequence database of the RAFL cDNAs is useful for promoter analysis and correct annotation of predicted transcription units and gene products. Furthermore, the full-length cDNAs are useful resources for analyses of the expression profiles, functions, and structures of plant proteins.

Arabidopsis↗

Pscroph, a parasitic plant EST database enriched for parasite associated transcripts.

BACKGROUND: Parasitic plants in the Orobanchaceae develop invasive root haustoria upon contact with host roots or root factors. The development of haustoria can be visually monitored and is rapid, highly synchronous, and strongly dependent on host factor exposure; therefore it provides a tractable system for studying chemical communications between roots of different plants. DESCRIPTION: Triphysaria is a facultative parasitic plant that initiates haustorium development within minutes after contact with host plant roots, root exudates, or purified haustorium-inducing phenolics. In order to identify genes associated with host root identification and early haustorium development, we sequenced suppression subtractive libraries (SSH) enriched for transcripts regulated in Triphysaria roots within five hours of exposure to Arabidopsis roots or the purified haustorium-inducing factor 2,6 dimethoxybenzoquinone. The sequences of over nine thousand ESTs from three SSH libraries and their subsequent assemblies are available at the Pscroph database http://pscroph.ucdavis.edu. The web site also provides BLAST functions and allows keyword searches of functional annotations. CONCLUSION: Libraries prepared from Triphysaria roots treated with host roots or haustorium inducing factors were enriched for transcripts predicted to function in stress responses, electron transport or protein metabolism. In addition to parasitic plant investigations, the Pscroph database provides a useful resource for investigations in rhizosphere interactions, chemical signaling between organisms, and plant development and evolution.

Arabidopsis↗

Comparison of the oxidative phosphorylation (OXPHOS) nuclear genes in the genomes of Drosophila melanogaster, Drosophila pseudoobscura and Anopheles gambiae.

BACKGROUND: In eukaryotic cells, oxidative phosphorylation (OXPHOS) uses the products of both nuclear and mitochondrial genes to generate cellular ATP. Interspecies comparative analysis of these genes, which appear to be under strong functional constraints, may shed light on the evolutionary mechanisms that act on a set of genes correlated by function and subcellular localization of their products. RESULTS: We have identified and annotated the Drosophila melanogaster, D. pseudoobscura and Anopheles gambiae orthologs of 78 nuclear genes encoding mitochondrial proteins involved in oxidative phosphorylation by a comparative analysis of their genomic sequences and organization. We have also identified 47 genes in these three dipteran species each of which shares significant sequence homology with one of the above-mentioned OXPHOS orthologs, and which are likely to have originated by duplication during evolution. Gene structure and intron length are essentially conserved in the three species, although gain or loss of introns is common in A. gambiae. In most tissues of D. melanogaster and A. gambiae the expression level of the duplicate gene is much lower than that of the original gene, and in D. melanogaster at least, its expression is almost always strongly testis-biased, in contrast to the soma-biased expression of the parent gene. CONCLUSIONS: Quickly achieving an expression pattern different from the parent genes may be required for new OXPHOS gene duplicates to be maintained in the genome. This may be a general evolutionary mechanism for originating phenotypic changes that could lead to species differentiation.

Animals↗

Evolution of protein function, from a structural perspective.

The recent growth in structural data, and ensuing analyses, have revealed the structural and functional versatility of protein families. With respect to enzymes, local active-site mutations, variations in surface loops and recruitment of additional domains accommodate the diverse substrate specificities and catalytic activities observed within several superfamilies. Conversely, some functions have more than one structural solution, having evolved independently several times during evolution. Combined with the existence of multi-functional genes, which have arisen by gene recruitment, these phenomena must be considered in the process of genome annotation.

Animals↗