Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Computational detection of genomic cis-regulatory modules applied to body patterning in the early Drosophila embryo.

BACKGROUND: Regulation of gene transcription is crucial for the function and development of all organisms. While gene prediction programs that identify protein coding sequence are used with remarkable success in the annotation of genomes, the development of computational methods to analyze noncoding regions and to delineate transcriptional control elements is still in its infancy. RESULTS: Here we present novel algorithms to detect cis-regulatory modules through genome wide scans for clusters of transcription factor binding sites using three levels of prior information. When binding sites for the factors are known, our statistical segmentation algorithm, Ahab, yields about 150 putative gap gene regulated modules, with no adjustable parameters other than a window size. If one or more related modules are known, but no binding sites, repeated motifs can be found by a customized Gibbs sampler and input to Ahab, to predict genes with similar regulation. Finally using only the genome, we developed a third algorithm, Argos, that counts and scores clusters of overrepresented motifs in a window of sequence. Argos recovers many of the known modules, upstream of the segmentation genes, with no training data. CONCLUSIONS: We have demonstrated, in the case of body patterning in the Drosophila embryo, that our algorithms allow the genome-wide identification of regulatory modules. We believe that Ahab overcomes many problems of recent approaches and we estimated the false positive rate to be about 50%. Argos is the first successful attempt to predict regulatory modules using only the genome without training data. Complete results and module predictions across the Drosophila genome are available at http://uqbar.rockefeller.edu/~siggia/.

Algorithms↗

Proteomic approaches to studying drug targets and resistance in Plasmodium.

Ever increasing drug resistance by Plasmodium falciparum, the most virulent of human malaria parasites, is creating new challenges in malaria chemotherapy. The entire genome sequences of P. falciparum and the rodent malaria parasite, P. yoelii yoelii are now available. Extensive genome sequence data from other Plasmodium species including another important human malaria parasite, P. vivax are also available. Powerful research techniques coupled to genomic resources are needed to help identify new drug and vaccine targets against malaria. Applied to Plasmodium, proteomics combines high-resolution protein or peptide separation with mass spectrometry and computer software to rapidly identify large numbers of proteins expressed from various stages of parasite development. Proteomic methods can be applied to study sub-cellular localization, cell function, organelle composition, changes in protein expression patterns in response to drug exposure, drug-protein binding and validation of data from genomic annotation and transcript expression studies. Recent high-throughput proteomic approaches have provided a wealth of protein expression data on P. falciparum, while smaller-scale studies examining specific drug-related hypotheses are also appearing. Of particular interest is the study of mechanisms of action and resistance of drugs such as the quinolines, whose targets currently may not be predictable from genomic data. Coupling the Plasmodium sequence data with bioinformatics, proteomics and RNA transcript expression profiling opens unprecedented opportunities for exploring new malaria control strategies. This review will focus on pharmacological research in malaria and other intracellular parasites using proteomic techniques, emphasizing resources and strategies available for Plasmodium.

Animals↗

Biological profiling of gene groups utilizing Gene Ontology.

Increasingly used high throughput experimental techniques, like DNA or protein microarrays give as a result groups of interesting, e.g. differentially regulated genes which require further biological interpretation. With the systematic functional annotation provided by the Gene Ontology the information required to automate the interpretation task is now accessible. However, the determination of statistical significance of a biological process within these groups is still an open question. In answering this question, multiple testing issues must be taken into account to avoid misleading results. Here we present a statistical framework that tests whether functions, processes or locations described in the Gene Ontology are significantly enriched within a group of interesting genes when compared to a reference group. First we define an exact analytical expression for the expected number of false positives that allows us to calculate adjusted p-values to control the false discovery rate. Next, we demonstrate and discuss the capabilities of our approach using publicly available microarray data on cell-cycle regulated genes. Further, we analyze the robustness of our framework with respect to the exact gene group composition and compare the performance with earlier approaches. The software package GOSSIP implements our method and is made freely available at http://gossip.gene-groups.net/.

Binding Sites↗

Sec and Tat Mediated Secretion Safeguards Mycobacterium tuberculosis Membrane Homeostasis.

Protein secretion is essential for the growth and virulence of Mycobacterium tuberculosis, yet the organization and function of its secretion pathways remain poorly understood. We reviewed the existing literature, combined it with systematic queries, and finalized annotations based on experimental data and computational predictions to compile a curated list of 92 secretory components and 198 reactions involved in Sec, twin-arginine translocation (Tat), and ESX pathways. Using CRISPRi, targeted depletion of SecA1 or TatAC impaired both in vitro growth and ex vivo survival. Label-free quantitative secretome analysis revealed decreased export of substrates dependent on SecA1 and TatAC, with enrichment of cytosolic proteins in culture filtrates, indicating increased membrane dysbiosis. Membrane proteomics showed elevated levels of proteins engaged in intermediary and lipid metabolism, while proteins associated with the cell wall and cell processes decreased, suggesting weakened membrane integrity. Loss of SecA1 or TatAC increased membrane permeability, with the effect being more pronounced in the case of TatAC, and caused structural abnormalities seen under electron microscopy. Overall, our integrated multi-omics and functional genetics studies demonstrate that the SecA1 and Tat pathways are essential for maintaining membrane homeostasis in Mycobacterium tuberculosis. These results suggest that essential secretory proteins may be promising targets for therapeutic intervention.

Mycobacterium tuberculosis↗

An integrated, functionally annotated gene map of the DXS8026-ELK1 interval on human Xp11.3-Xp11.23: potential hotspot for neurogenetic disorders.

Human chromosome Xp11.3-Xp11.23 encompasses the map location for a growing number of diseases with a genetic basis or genetic component. These include several eye disorders, syndromic and nonsyndromic forms of X-linked mental retardation (XLMR), X-linked neuromuscular diseases and susceptibility loci for schizophrenia, type 1 diabetes, and Graves' disease. We have constructed an approximately 2.7-Mb high-resolution physical map extending from DXS8026 to ELK1, corresponding to a genetic distance of approximately 5.5 cM. A combination of chromosome walking and sequence-tagged site (STS)-content mapping resulted in an integrated framework and transcript map, precisely positioning 10 polymorphic microsatellites (one of which is novel), 16 ESTs, and 12 known genes (RP2, PCTK1, UHX1, UBE1, RBM10, ZNF157, SYN1, ARAF1, TIMP1, PFC, ELK1, UXT). The composite map is currently anchored with 89 STSs to give an average resolution of approximately 1 STS every 30 kb. By a combination of EST database searches and in silico detection of UniGene clusters within genomic sequence generated from this template map, we have mapped several novel genes within this interval: a Na+/H+ exchanger (SLC9A7), at least two zincfinger transcription factors (KIAA0215 and Hs.68318), carbohydrate sulfotransferase-7 (CHST7), regucalcin (RGN), inactivation-escape-1 (INE1), the human ortholog of mouse neuronal protein 15.6, and four putative novel genes. Further genomic analysis enabled annotation of the sequence interval with 20 predicted pseudogenes and 21 UniGene clusters of unknown function. The combined PAC/BAC transcript map and YAC scaffold presented here clarifies previously conflicting data for markers and genes within the Xp11.3-Xp11.23 interval and provides a powerful integrated resource for functional characterization of this clonally unstable, yet gene-rich and clinically significant region of proximal Xp.

Chromosome Mapping↗

Predicting functions from protein sequences--where are the bottlenecks?

The exponential growth of sequence data does not necessarily lead to an increase in knowledge about the functions of genes and their products. Prediction of function using comparative sequence analysis is extremely powerful but, if not performed appropriately, may also lead to the creation and propagation of assignment errors. While current homology detection methods can cope with the data flow, the identification, verification and annotation of functional features need to be drastically improved.

Amino Acid Sequence↗

Protein model representation and construction.

Crystallographic studies play a major role in current efforts towards protein structure determination. However, despite recent advances in computational tools for molecular modeling and graphics, the task of constructing a protein model from crystallographic data remains complex and time-consuming, requiring extensive expert intervention. This paper describes an approach to automating the process of model construction, where a model is represented as an annotated trace (or partial trace) of the three-dimensional backbone of the structure. Potential models are generated using an evolutionary algorithm, which incorporates multiple fitness functions tailored to different structural levels in the protein. Preliminary experimental results, which demonstrate the viability of the approach, are reported.

Algorithms↗

Aphid biology: expressed genes from alate Toxoptera citricida, the brown citrus aphid.

The brown citrus aphid, Toxoptera citricida (Kirkaldy), is considered the primary vector of citrus tristeza virus, a severe pathogen which causes losses to citrus industries worldwide. The alate (winged) form of this aphid can readily fly long distances with the wind, thus spreading citrus tristeza virus in citrus growing regions. To better understand the biology of the brown citrus aphid and the emergence of genes expressed during wing development, we undertook a large-scale 5' end sequencing project of cDNA clones from alate aphids. Similar large-scale expressed sequence tag (EST) sequencing projects from other insects have provided a vehicle for answering biological questions relating to development and physiology. Although there is a growing database in GenBank of ESTs from insects, most are from Drosophila melanogaster and Anopheles gambiae, with relatively few specifically derived from aphids. However, important morphogenetic processes are exclusively associated with piercing-sucking insect development and sap feeding insect metabolism. In this paper, we describe the first public data set of ESTs from the brown citrus aphid, T. citricida. The cDNA library was derived from alate adults due to their significance in spreading viruses (e.g., citrus tristeza virus). Over 5180 cDNA clones were sequenced, resulting in 4263 high-quality ESTs. Contig alignment of these ESTs resulted in 2124 total assembled sequences, including both contiguous sequences and singlets. Approximately 33% of the ESTs currently have no significant match in either the non-redundant protein or nucleic acid databases. Sequences returning matches with an E-value of < or = -10 using BLASTX, BLASTN, or TBLASTX were annotated based on their putative molecular function and biological process using the Gene Ontology classification system. These data will aid research efforts in the identification of important genes within insects, specifically aphids and other sap feeding insects within the Order Hemiptera.

Animals↗

Genome-based bioinformatic selection of chromosomal Bacillus anthracis putative vaccine candidates coupled with proteomic identification of surface-associated antigens.

Bacillus anthracis (Ames strain) chromosome-derived open reading frames (ORFs), predicted to code for surface exposed or virulence related proteins, were selected as B. anthracis-specific vaccine candidates by a multistep computational screen of the entire draft chromosome sequence (February 2001 version, 460 contigs, The Institute for Genomic Research, Rockville, Md.). The selection procedure combined preliminary annotation (sequence similarity searches and domain assignments), prediction of cellular localization, taxonomical and functional screen and additional filtering criteria (size, number of paralogs). The reductive strategy, combined with manual curation, resulted in selection of 240 candidate ORFs encoding proteins with putative known function, as well as 280 proteins of unknown function. Proteomic analysis of two-dimensional gels of a B. anthracis membrane fraction, verified the expression of some gene products. Matrix-assisted laser desorption ionization-time-of-flight mass spectrometry analyses allowed identification of 38 spots cross-reacting with sera from B. anthracis immunized animals. These spots were found to represent eight in vivo immunogens, comprising of EA1, Sap, and 6 proteins whose expression and immunogenicity was not reported before. Five of these 8 immunogens were preselected by the bioinformatic analysis (EA1, Sap, 2 novel SLH proteins and peroxiredoxin/AhpC), as vaccine candidates. This study demonstrates that a combination of the bioinformatic and proteomic strategies may be useful in promoting the development of next generation anthrax vaccine.

Adhesins, Bacterial↗

Crystal structure of THEP1 from the hyperthermophile Aquifex aeolicus: a variation of the RecA fold.

BACKGROUND: aaTHEP1, the gene product of aq_1292 from Aquifex aeolicus, shows sequence homology to proteins from most thermophiles, hyperthermophiles, and higher organisms such as man, mouse, and fly. In contrast, there are almost no homologous proteins in mesophilic unicellular microorganisms. aaTHEP1 is a thermophilic enzyme exhibiting both ATPase and GTPase activity in vitro. Although annotated as a nucleotide kinase, such an activity could not be confirmed for aaTHEP1 experimentally and the in vivo function of aaTHEP1 is still unknown. RESULTS: Here we report the crystal structure of selenomethionine substituted nucleotide-free aaTHEP1 at 1.4 A resolution using a multiple anomalous dispersion phasing protocol. The protein is composed of a single domain that belongs to the family of 3-layer (alpha/beta/alpha)-structures consisting of nine central strands flanked by six helices. The closest structural homologue as determined by DALI is the RecA family. In contrast to the latter proteins, aaTHEP1 possesses an extension of the beta-sheet consisting of four additional beta-strands. CONCLUSION: We conclude that the structure of aaTHEP1 represents a variation of the RecA fold. Although the catalytic function of aaTHEP1 remains unclear, structural details indicate that it does not belong to the group of GTPases, kinases or adenosyltransferases. A mainly positive electrostatic surface indicates that aaTHEP1 might be a DNA/RNA modifying enzyme. The resolved structure of aaTHEP1 can serve as paradigm for the complete THEP1 family.

Adenosine Triphosphatases↗

Synonymous codon usage and gene function are strongly related in Oryza sativa.

The relationship between codon usage and gene function was investigated while considering a dataset of 2106 nuclear genes of Oryza sativa. The results of standard chi(2) test and F-statistic showed that for every 59 synonymous codons, a strongly significant association with gene functional categories existed in rice, indicating that codon usage was generally coordinated with gene function whether it was at the level of individual amino acids or at the level of nucleotides. However, it could not be directly said that the use of every codons differed significantly between any two functional categories. Notably, there existed large difference both in selection for biased codons or selection intensity among functional categories. Therefore, we identified at least two classes of genes: one group of genes, mainly belonging to the "METABOLISM" category, was tended to use G- and/or C-ending codons while the other was more biased to choose codons ending with A and/or U. The latter group contained genes of various functions, especially those genes classified into the "Nuclear Structure" category. These observations will be more important for molecular genetic engineering and genome functional annotation.

Chromosome Mapping↗

MIPS: a database for protein sequences, homology data and yeast genome information.

The MIPS group (Martinsried Institute for Protein Sequences) at the Max-Planck-Institute for Biochemistry, Martinsried near Munich, Germany, collects, processes and distributes protein sequence data within the framework of the tripartite association of the PIR-International Protein Sequence Database (,). MIPS contributes nearly 50% of the data input to the PIR-International Protein Sequence Database. The database is distributed on CD-ROM together with PATCHX, an exhaustive supplement of unique, unverified protein sequences from external sources compiled by MIPS. Through its WWW server (http://www.mips.biochem.mpg.de/ ) MIPS permits internet access to sequence databases, homology data and to yeast genome information. (i) Sequence similarity results from the FASTA program () are stored in the FASTA database for all proteins from PIR-International and PATCHX. The database is dynamically maintained and permits instant access to FASTA results. (ii) Starting with FASTA database queries, proteins have been classified into families and superfamilies (PROT-FAM). (iii) The HPT (hashed position tree) data structure () developed at MIPS is a new approach for rapid sequence and pattern searching. (iv) MIPS provides access to the sequence and annotation of the complete yeast genome (), the functional classification of yeast genes (FunCat) and its graphical display, the 'Genome Browser' (). A CD-ROM based on the JAVA programming language providing dynamic interactive access to the yeast genome and the related protein sequences has been compiled and is available on request.

Academies and Institutes↗

InterPro--an integrated documentation resource for protein families, domains and functional sites.

MOTIVATION: InterPro is a new integrated documentation resource for protein families, domains and functional sites, developed initially as a means of rationalising the complementary efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. RESULTS: Merged annotations from PRINTS, PROSITE and Pfam form the InterPro core. Each combined InterPro entry includes functional descriptions and literature references, and links are made back to the relevant parent database(s), allowing users to see at a glance whether a particular family or domain has associated patterns, profiles, fingerprints, etc. Merged and individual entries (i.e. those that have no counterpart in the companion resources) are assigned unique accession numbers. Release 1.2 of InterPro (June 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification (PTMs) encoded by 6581 different regular expressions, profiles, fingerprints and Hidden Markov Models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1000000 hits from 264333 different proteins out of 384572 in SWISS-PROT and TrEMBL).

Computational Biology↗

Yeast genomic expression studies using DNA microarrays.

The exploration and characterization of yeast genomic expression programs is providing a wealth of information about yeast biology, as well as other organisms. The intriguing biology of yeast species invites characterization of genomic expression patterns to illuminate the details of cellular physiology. In addition to its value as an interesting organism, yeast maintains its role as an excellent model in which to characterize genomic expression programs. Microarray studies are quickly spreading to plant, animal, and microbial organisms that remain in the early stages of characterization. The extensive knowledge of yeast biology, as well as the relative ease with which yeast studies can be performed and controlled, facilitates interpretation of the genomic expression data. Importantly, existing information about yeast biology, including functional annotations for each gene, is captured and efficiently presented in databases such as the Saccharomyces Genome Database (SGD), the Munich Information Center Yeast Genome Database (MIPS), the Yeast and Pombe Protein Databases (YPD and PPD, respectively), and others. A number of databases also allow the exploration of published genomic expression studies, including the "Expression Connection" at SGD and the Microarray Global Viewer (yMGV) organized by Marc et al. Consulting these databases to retrieve known details about gene function and regulation vastly facilitates interpretation of the genomic expression data, allowing biological hypotheses to be formulated and tested. These hypotheses can be applied to other organisms that may execute genomic expression programs similar to those seen in yeast. Furthermore, as more genomic expression studies in multiple organisms emerge, large-scale data comparisons can be conducted, within and across organisms. Incorporating the results of yeast studies into such comparisons is certain to increase our understanding about the function, regulation, and evolution of genomic expression programs.

Carbocyanines↗

Targeted overexpression of the Escherichia coli MinC protein in higher plants results in abnormal chloroplasts.

Higher plant chloroplast division involves some of the same types of proteins that are required in prokaryotic cell division. These include two of the three Min proteins, MinD and MinE, encoded by the min operon in bacteria. Noticeably absent from annotated sequences from higher plants is a MinC homologue. A higher plant functional MinC homologue that would interfere with FtsZ polymerization, has yet to be identified. We sought to determine whether expression of the bacterial MinC in higher plants could affect chloroplast division. The Escherichia coli minC (EcMinC) gene was isolated and inserted behind the Arabidopsis thaliana RbcS transit peptide sequence for chloroplast targeting. This TP-EcMinC gene driven by the CaMV 35S(2) constitutive promoter was then transformed into tobacco (Nicotiana tabacum L.). Abnormally large chloroplasts were observed in the transgenic plants suggesting that overexpression of the E. coli MinC perturbed higher plant chloroplast division.

Chloroplasts↗

Predicted roles for hypothetical proteins in the low-temperature expressed proteome of the Antarctic archaeon Methanococcoides burtonii.

Using liquid chromatography-mass spectrometry, 528 proteins were identified that are expressed during growth at 4 degrees C in the cold adapted archaeon, Methanococcoides burtonii. Of those, 135 were annotated previously as unique or conserved hypothetical proteins. We have performed a comprehensive, integrated analysis of the latter proteins using threading, InterProScan, predicted subcellular localization and visualization of conserved gene context across multiple prokaryotic genomes. Functional information was obtained for 55 proteins, providing new insight into the physiology of M. burtonii. Many of the proteins were predicted to be involved in DNA/RNA binding or modification and cell signaling, suggesting a complex, uncharacterized regulatory network controlling cellular processes during growth at low-temperature. Novel enzymatic functions were predicted for several proteins, including a putative candidate gene for the posttranslational modification of the key methanogenesis enzyme coenzyme M methyl reductase. A bacterial-like CRISPR locus was identified as a strong candidate for archaeal-bacterial lateral gene transfer. Gene context analysis proved a valuable augmentation to the other predictive methods in several cases, by revealing conserved gene associations and annotations in other microbial genomes. Our results underscore the importance of addressing the "hypothetical protein problem" for a complete understanding of cell physiology.

Adaptation, Physiological↗

ProTeus: identifying signatures in protein termini.

ProTeus (PROtein TErminUS) is a web-based tool for the identification of short linear signatures in protein termini. It is based on a position-based search method for revealing short signatures in termini of all proteins. The initial step in ProTeus development was to collect all signature groups (SIGs) based on their relative positions at the termini. The initial set of SIGs went through a sequential process of inspection and removal of SIGs, which did not meet the attributed statistical thresholds. The SIGs that were found significant represent protein sets with minimal or no overall sequence similarity besides the similarity found at the termini. These SIGs were archived and are presented at ProTeus. The SIGs are sorted by their strong correspondence to functional annotation from external databases such as GO. ProTeus provides rich search and visualization tools for evaluating the quality of different SIGs. A search option allows the identification of terminal signatures in new sequences. ProTeus (ver 1.2) is available at http://www.proteus.cs.huji.ac.il.

Databases, Protein↗

Small genes/gene-products in Escherichia coli K-12.

Forty-two protein spots of observed M(r) 6-15 kDa were resolved by two-dimensional gel electrophoresis, stained by Coomassie blue and subjected to Edman microsequencing. All of the proteins could be related back to their encoding open reading frames, thereby vindicating the bioinformatic tools currently utilised in their identification. However, only 14/42 gene-products were expressed as annotated. Translation was confirmed for 14 open reading frames with no attributed function (EcoGene Y-entries), while N-terminal sequence allowed the start codon to be accurately annotated for the genes yigF, yccU, yqiC, ynfD, and yeeX. The methionine start codon was cleaved in 11 gene-products (AtpE, Hns, RpoZ, RplL, CspC, YccJ, YggX, YjgF, HimA, InfA, RpsQ) and a further five showed loss of a signal peptide (PspE, HdeB, HdeA, YnfD, YkfE). Internal (Tig, AtpA, TufA) and N-terminal fragmentation (CspD, RpsF, AtcU) of much larger proteins was also detected, which may have resulted from physiological or translational processes. M(r) and pI isoforms were detected respectively for PtsH and GatB, each being phosphoproteins, as well as RplY which manifested differences with respect to predicted M(r) and pI. In addition, YjgF was shown to belong to a small gene family of unknown function with ancient conserved regions across procaryotes and eucaryotes. YgiN was revealed to have a paralogue and orthologues in Bacillus subtilis, Synechocystis sp., Mycobacterium tuberculosis, Neisseria gonorrhoea, and Rhodococcus erythropolis. Orthologues are also reported for YihD, YccU and YeeX. Of the 14 Y-genes, only YkfE possessed no detectable orthologues. These results highlight the need to complement genomic analysis with detailed proteomics in order to gain a better understanding of cellular molecular biology, while the confirmation of the open reading frame start codon using Edman degradation protein microsequencing has yet to be superseded by recent advances in mass spectrometry.

Amino Acid Sequence↗