Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Counting the zinc-proteins encoded in the human genome.

Metalloproteins are proteins capable of binding one or more metal ions, which may be required for their biological function, or for regulation of their activities or for structural purposes. Genome sequencing projects have provided a huge number of protein primary sequences, but, even though several different elaborate analyses and annotations have been enabled by a rich and ever-increasing portfolio of bioinformatic tools, metal-binding properties remain difficult to predict as well as to investigate experimentally. Consequently, the present knowledge about metalloproteins is only partial. The present bioinformatic research proposes a strategy to answer the question of how many and which proteins encoded in the human genome may require zinc for their physiological function. This is achieved by a combination of approaches, which include: (i) searching in the proteome for the zinc-binding patterns that, on their turn, are obtained from all available X-ray data; (ii) using libraries of metal-binding protein domains based on multiple sequence alignments of known metalloproteins obtained from the Pfam database; and (iii) mining the annotations of human gene sequences, which are based on any type of information available. It is found that 1684 proteins in the human proteome are independently identified by all three approaches as zinc-proteins, 746 are identified by two, and 777 are identified by only one method. By assuming that all proteins identified by at least two approaches are truly zinc-binding and inspecting the proteins identified by a single method, it can be proposed that ca. 2800 human proteins are potentially zinc-binding in vivo, corresponding to 10% of the human proteome, with an uncertainty of 400 sequences. Available functional information suggests that the large majority of human zinc-binding proteins are involved in the regulation of gene expression. The most abundant class of zinc-binding proteins in humans is that of zinc-fingers, with Cys4 and Cys2His2 being the most common types of coordination environment.

Computational Biology↗

The Paris-Sud yeast structural genomics pilot-project: from structure to function.

We present here the outlines and results from our yeast structural genomics (YSG) pilot-project. A lab-scale platform for the systematic production and structure determination is presented. In order to validate this approach, 250 non-membrane proteins of unknown structure were targeted. Strategies and final statistics are evaluated. We finally discuss the opportunity of structural genomics programs to contribute to functional biochemical annotation.

Genomics↗

Pinpointing genomic regions conferring herbicide tolerance in cassava via genome-wide association mapping.

Cassava (Manihot esculenta Crantz) is a tropical crop of major socioeconomic importance, whose productivity can be limited by sensitivity to herbicides used for weed management. This study aimed to perform a genome-wide association study (GWAS) in 194 cassava genotypes to identify genomic regions associated with tolerance to the herbicides mesotrione, S-metolachlor, and chloransulam-methyl. The evaluations performed at 3, 6, 9, 15, and 30 days after application (DAA) were used to characterize the temporal progression of phytotoxicity. Based on this analysis, the phenotype obtained at 9 days after application (PhytoX9DAA) was selected for genome-wide association analyses because it represented the period of greatest symptom expression and the highest discrimination among genotypes. GWAS analyses were performed using de-regressed BLUPs and the MLM, MLMM, and BLINK models, incorporating kinship (K) and population structure (Q) matrices. Significant markers were detected across multiple chromosomes, and the corresponding genomic windows contained candidate genes with functional annotations related to herbicide response. The predominant functional categories included membrane transport, channel activity, signal peptide processing, protein phosphorylation, cellular signaling, and metabolic regulation. Key candidate genes included Manes.02G151900 and Manes.02G152700 (chromosome 2), associated with transmembrane transport and signal peptide processing; Manes.09G060900 (chromosome 9), associated with protein kinase activity, ATP binding, and protein phosphorylation; and Manes.15G083800 and Manes.15G084000 (chromosome 15), associated with S-adenosylmethionine-dependent methyltransferase activity, membrane-related functions, and protein phosphorylation. These genes participate in biochemical pathways involved in cellular signaling, membrane transport, and metabolic regulation that may contribute to herbicide tolerance. Overall, the results demonstrate that herbicide tolerance in cassava is a quantitative and polygenic trait governed by numerous small-effect loci. The integration of cellular signaling, metabolic regulation, and membrane transport supports the physiological resilience of the species under chemical exposure, providing valuable insights for breeding strategies and marker-assisted selection.

Genome-Wide Association Study↗

Complete nucleotide sequence and genome analysis of bacteriophage BFK20--a lytic phage of the industrial producer Brevibacterium flavum.

The entire double-stranded DNA genome of bacteriophage BFK20, a lytic phage of the Brevibacterium flavum CCM 251--industrial producer of L-lysine--was sequenced and analyzed. It consists of 42,968 base pairs with an overall molar G + C content of 56.2%. Fifty-five potential open reading frames were identified and annotated using various bioinformatics tools. Clusters of functionally related putative genes were defined (structural, lytic, replication and regulatory). To verify the annotation of structural proteins, they were resolved by 2D gel electrophoresis and were submitted to N-terminal amino acid sequencing. Structural proteins identified included the portal and major and minor tail proteins. Based on the overall genome sequence comparison, similarities with other known bacteriophage genomes include primarily bacteriophages from Mycobacterium spp. and some regions of Corynebacterium spp. genomes--possible prophages. Our results support the theory that phage genomes are mosaics with respect to each other.

Bacteriophages↗

Multifunctional proteins: examples of gene sharing.

Adding to the difficulty of interpreting the human genome sequence and annotating protein sequence databases is the observation that a single protein can 'moonlight' or perform multiple, apparently unrelated, functions. This review summarizes examples of moonlighting proteins in cellular activities and biochemical pathways important in cancer and other diseases. The proteins include a variety of combinations of functions and mechanisms to switch between functions. Moonlighting proteins can be beneficial to the organism, such as by coordinating cellular activities. However, moonlighting proteins can potentially make more difficult the determination of the molecular mechanisms of disease and the process of rational drug design.

Binding Sites↗

YPD-A database for the proteins of Saccharomyces cerevisiae.

YPD is a database for the proteins of the budding yeast, Saccharomyces cerevisiae. YPD has two formats: (i) a spreadsheet which tabulates many of the physical and functional properties of yeast proteins, and (ii) the YPD Protein Reports which are formatted pages containing the protein properties, annotations gathered from the literature, and references with titles. YPD is available through the World-Wide Web, through an Email server, and by anonymous FTP. New releases of the YPD spreadsheet are produced every two to four months, and the on-line information is updated daily.

Computer Communication Networks↗

Transcriptome analysis of the barley-Fusarium graminearum interaction.

Fusarium head blight (FHB) of barley (Hordeum vulgare L.) is caused by Fusarium graminearum. FHB causes yield losses and reduction in grain quality primarily due to the accumulation of trichothecene mycotoxins such as deoxynivalenol (DON). To develop an understanding of the barley-F. graminearum interaction, we examined the relationship among the infection process, DON concentration, and host transcript accumulation for 22,439 genes in spikes from the susceptible cv. Morex from 0 to 144 h after F. graminearum and water control inoculation. We detected 467 differentially accumulating barley gene transcripts in the F. graminearum-treated plants compared with the water control-treated plants. Functional annotation of the transcripts revealed a variety of infection-induced host genes encoding defense response proteins, oxidative burst-associated enzymes, and phenylpropanoid pathway enzymes. Of particular interest was the induction of transcripts encoding potential trichothecene catabolic enzymes and transporters, and the induction of the tryptophan biosynthetic and catabolic pathway enzymes. Our results define three stages of E graminearum infection. An early stage, between 0 and 48 h after inoculation (hai), exhibited limited fungal development, low DON accumulation, and little change in the transcript accumulation status. An intermediate stage, between 48 and 96 hai, showed increased fungal development and active infection, higher DON accumulation, and increased transcript accumulation. A majority of the host gene transcripts were detected by 72 hai, suggesting that this is an important timepoint for the barley-F. graminearum interaction. A late stage also identified between 96 and 144 hai, exhibiting development of hyphal mats, high DON accumulation, and a reduction in the number of transcripts observed. Our study provides a baseline and hypothesis-generating dataset in barley during F. graminearum infection and in other grasses during pathogen infection.

Fusarium↗

Protein sequence annotation in the genome era: the annotation concept of SWISS-PROT+TREMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation, a minimal level of redundancy and high level of integration with other databases. Ongoing genome sequencing projects have dramatically increased the number of protein sequences to be incorporated into SWISS-PROT. Since we do not want to dilute the quality standards of SWISS-PROT by incorporating sequences without proper sequence analysis and annotation, we cannot speed up the incorporation of new incoming data indefinitely. However, as we also want to make the sequences available as fast as possible, we introduced TREMBL (TRanslation of EMBL nucleotide sequence database), a supplement to SWISS-PROT. TREMBL consists of computer-annotated entries in SWISS-PROT format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except for CDS already included in SWISS-PROT. While TREMBL is already of immense value, its computer-generated annotation does not match the quality of SWISS-PROTs. The main difference is in the protein functional information attached to sequences. With this in mind, we are dedicating substantial effort to develop and apply computer methods to enhance the functional information attached to TREMBL entries.

Amino Acid Sequence↗

Role of functional genes for seed vigor related traits through genome-wide association mapping in finger millet (Eleusine coracana L. Gaertn.).

Finger millet (Eleusine coracana (L.) Gaertn.) is a calcium-rich, nutritious and resilient crop that thrives even in harsh environmental conditions. In such ecologies, seed longevity and seedling vigor are crucial for sustainable crop production amid climate change. The current study explores the genetics of accelerated aging on seed longevity traits across 221 diverse accessions of finger millet through genome-wide association approach (GWAS). A significant variation was identified in germination percentage, germination rate indices, mean germination time, seedling vigor indices and dry weight upon aging treatment. GWAS model from 11,832 high-quality SNPs identified through Genotyping-by-Sequencing (GBS) approach produced 491 marker-trait associations (MTAs) for 27 traits, of which 54 were FDR-corrected. A pleiotropic SNP, FM_SNP_9478 identified on chromosome 7B was associated with the traits viz., germination after aging, germination index after aging and their relative measures. Functional annotation revealed DET1 and expansin-A2 influenced seed coat integrity, critical for germination and aging resilience. Probable protein phosphatase 2C3 and piezo-type ion channels contributed to mechanical sensing and stress adaptation in seeds. Beta-amylase and acetyl-CoA carboxylase 2 were identified for seed metabolism and stress response. These insights lay the framework for targeted breeding efforts to improve seed quality and resilience under diverse production conditions.

Eleusine↗

PDB-UF: database of predicted enzymatic functions for unannotated protein structures from structural genomics.

BACKGROUND: The number of protein structures from structural genomics centers dramatically increases in the Protein Data Bank (PDB). Many of these structures are functionally unannotated because they have no sequence similarity to proteins of known function. However, it is possible to successfully infer function using only structural similarity. RESULTS: Here we present the PDB-UF database, a web-accessible collection of predictions of enzymatic properties using structure-function relationship. The assignments were conducted for three-dimensional protein structures of unknown function that come from structural genomics initiatives. We show that 4 hypothetical proteins (with PDB accession codes: 1VH0, 1NS5, 1O6D, and 1TO0), for which standard BLAST tools such as PSI-BLAST or RPS-BLAST failed to assign any function, are probably methyltransferase enzymes. CONCLUSION: We suggest that the structure-based prediction of an EC number should be conducted having the different similarity score cutoff for different protein folds. Moreover, performing the annotation using two different algorithms can reduce the rate of false positive assignments. We believe, that the presented web-based repository will help to decrease the number of protein structures that have functions marked as "unknown" in the PDB file. AVAILABILITY: http://paradox.harvard.edu/PDB-UF and http://bioinfo.pl/PDB-UF.

Chromosome Mapping↗

Universal protein families and the functional content of the last universal common ancestor.

The phylogenetic distribution of Methanococcus jannaschii proteins can provide, for the first time, an estimate of the genome content of the last common ancestor of the three domains of life. Relying on annotation and comparison with reference to the species distribution of sequence similarities results in 324 proteins forming the universal family set. This set is very well characterized and relatively small and nonredundant, containing 301 biochemical functions, of which 246 are unique. This universal function set contains mostly genes coding for energy metabolism or information processing. It appears that the Last Universal Common Ancestor was an organism with metabolic networks and genetic machinery similar to those of extant unicellular organisms.

Archaeal Proteins↗

Systematic analysis of snake neurotoxins' functional classification using a data warehousing approach.

MOTIVATION: Sequence annotations, functional and structural data on snake venom neurotoxins (svNTXs) are scattered across multiple databases and literature sources. Sequence annotations and structural data are available in the public molecular databases, while functional data are almost exclusively available in the published articles. There is a need for a specialized svNTXs database that contains NTX entries, which are organized, well annotated and classified in a systematic manner. RESULTS: We have systematically analyzed svNTXs and classified them using structure-function groups based on their structural, functional and phylogenetic properties. Using conserved motifs in each phylogenetic group, we built an intelligent module for the prediction of structural and functional properties of unknown NTXs. We also developed an annotation tool to aid the functional prediction of newly identified NTXs as an additional resource for the venom research community. AVAILABILITY: We created a searchable online database of NTX proteins sequences (http://research.i2r.a-star.edu.sg/Templar/DB/snake_neurotoxin). This database can also be found under Swiss-Prot Toxin Annotation Project website (http://www.expasy.org/sprot/).

Animals↗

Efficient gene-driven germ-line point mutagenesis of C57BL/6J mice.

BACKGROUND: Analysis of an allelic series of point mutations in a gene, generated by N-ethyl-N-nitrosourea (ENU) mutagenesis, is a valuable method for discovering the full scope of its biological function. Here we present an efficient gene-driven approach for identifying ENU-induced point mutations in any gene in C57BL/6J mice. The advantage of such an approach is that it allows one to select any gene of interest in the mouse genome and to go directly from DNA sequence to mutant mice. RESULTS: We produced the Cryopreserved Mutant Mouse Bank (CMMB), which is an archive of DNA, cDNA, tissues, and sperm from 4,000 G1 male offspring of ENU-treated C57BL/6J males mated to untreated C57BL/6J females. Each mouse in the CMMB carries a large number of random heterozygous point mutations throughout the genome. High-throughput Temperature Gradient Capillary Electrophoresis (TGCE) was employed to perform a 32-Mbp sequence-driven screen for mutations in 38 PCR amplicons from 11 genes in DNA and/or cDNA from the CMMB mice. DNA sequence analysis of heteroduplex-forming amplicons identified by TGCE revealed 22 mutations in 10 genes for an overall mutation frequency of 1 in 1.45 Mbp. All 22 mutations are single base pair substitutions, and nine of them (41%) result in nonconservative amino acid substitutions. Intracytoplasmic sperm injection (ICSI) of cryopreserved spermatozoa into B6D2F1 or C57BL/6J ova was used to recover mutant mice for nine of the mutations to date. CONCLUSIONS: The inbred C57BL/6J CMMB, together with TGCE mutation screening and ICSI for the recovery of mutant mice, represents a valuable gene-driven approach for the functional annotation of the mammalian genome and for the generation of mouse models of human genetic diseases. The ability of ENU to induce mutations that cause various types of changes in proteins will provide additional insights into the functions of mammalian proteins that may not be detectable by knockout mutations.

Animals↗

Creation and disruption of protein features by alternative splicing -- a novel mechanism to modulate function.

BACKGROUND: Alternative splicing often occurs in the coding sequence and alters protein structure and function. It is mainly carried out in two ways: by skipping exons that encode a certain protein feature and by introducing a frameshift that changes the downstream protein sequence. These mechanisms are widespread and well investigated. RESULTS: Here, we propose an additional mechanism of alternative splicing to modulate protein function. This mechanism creates a protein feature by putting together two non-consecutive exons or destroys a feature by inserting an exon in its body. In contrast to other mechanisms, the individual parts of the feature are present in both splice variants but the feature is only functional in the splice form where both parts are merged. We provide evidence for this mechanism by performing a genome-wide search with four protein features: transmembrane helices, phosphorylation and glycosylation sites, and Pfam domains. CONCLUSION: We describe a novel type of event that creates or removes a protein feature by alternative splicing. Current data suggest that these events are rare. Besides the four features investigated here, this mechanism is conceivable for many other protein features, especially for small linear protein motifs. It is important for the characterization of functional differences of two splice forms and should be considered in genome-wide annotation efforts. Furthermore, it offers a novel strategy for ab initio prediction of alternative splice events.

Alternative Splicing↗

Molecular cloning and functional expression of a Drosophila receptor for the neuropeptides capa-1 and -2.

The Drosophila Genome Project website contains an annotated gene (CG14575) for a G protein-coupled receptor. We cloned this receptor and found that the cloned cDNA did not correspond to the annotated gene; it partly contained different exons and additional exons located at the 5(')-end of the annotated gene. We expressed the coding part of the cloned cDNA in Chinese hamster ovary cells and found that the receptor was activated by two neuropeptides, capa-1 and -2, encoded by the Drosophila capability gene. Database searches led to the identification of a similar receptor in the genome from the malaria mosquito Anopheles gambiae (58% amino acid residue identities; 76% conserved residues; and 5 introns at identical positions within the two insect genes). Because capa-1 and -2 and related insect neuropeptides stimulate fluid secretion in insect Malpighian (renal) tubules, the identification of this first insect capa receptor will advance our knowledge on insect renal function.

Amino Acid Sequence↗

Unexpected catalytic site variation in phosphoprotein phosphatase homologues of cofactor-dependent phosphoglycerate mutase.

The cofactor-dependent phosphoglycerate mutase (dPGM) superfamily contains, besides mutases, a variety of phosphatases, both broadly and narrowly substrate-specific. Distant dPGM homologues, conspicuously abundant in microbial genomes, represent a challenge for functional annotation based on sequence comparison alone. Here we carry out sequence analysis and molecular modelling of two families of bacterial dPGM homologues, one the SixA phosphoprotein phosphatases, the other containing various proteins of no known molecular function. The models show how SixA proteins have adapted to phosphoprotein substrate and suggest that the second family may also encode phosphoprotein phosphatases. Unexpected variation in catalytic and substrate-binding residues is observed in the models.

Amino Acid Sequence↗

MannDB - a microbial database of automated protein sequence analyses and evidence integration for protein characterization.

BACKGROUND: MannDB was created to meet a need for rapid, comprehensive automated protein sequence analyses to support selection of proteins suitable as targets for driving the development of reagents for pathogen or protein toxin detection. Because a large number of open-source tools were needed, it was necessary to produce a software system to scale the computations for whole-proteome analysis. Thus, we built a fully automated system for executing software tools and for storage, integration, and display of automated protein sequence analysis and annotation data. DESCRIPTION: MannDB is a relational database that organizes data resulting from fully automated, high-throughput protein-sequence analyses using open-source tools. Types of analyses provided include predictions of cleavage, chemical properties, classification, features, functional assignment, post-translational modifications, motifs, antigenicity, and secondary structure. Proteomes (lists of hypothetical and known proteins) are downloaded and parsed from Genbank and then inserted into MannDB, and annotations from SwissProt are downloaded when identifiers are found in the Genbank entry or when identical sequences are identified. Currently 36 open-source tools are run against MannDB protein sequences either on local systems or by means of batch submission to external servers. In addition, BLAST against protein entries in MvirDB, our database of microbial virulence factors, is performed. A web client browser enables viewing of computational results and downloaded annotations, and a query tool enables structured and free-text search capabilities. When available, links to external databases, including MvirDB, are provided. MannDB contains whole-proteome analyses for at least one representative organism from each category of biological threat organism listed by APHIS, CDC, HHS, NIAID, USDA, USFDA, and WHO. CONCLUSION: MannDB comprises a large number of genomes and comprehensive protein sequence analyses representing organisms listed as high-priority agents on the websites of several governmental organizations concerned with bio-terrorism. MannDB provides the user with a BLAST interface for comparison of native and non-native sequences and a query tool for conveniently selecting proteins of interest. In addition, the user has access to a web-based browser that compiles comprehensive and extensive reports. Access to MannDB is freely available at http://manndb.llnl.gov/.

Algorithms↗

Predicting protein function from sequence and structural data.

When a protein's function cannot be experimentally determined, it can often be inferred from sequence similarity. Should this process fail, analysis of the protein structure can provide functional clues or confirm tentative functional assignments inferred from the sequence. Many structure-based approaches exist (e.g. fold similarity, three-dimensional templates), but as no single method can be expected to be successful in all cases, a more prudent approach involves combining multiple methods. Several automated servers that integrate evidence from multiple sources have been released this year and particular improvements have been seen with methods utilizing the Gene Ontology functional annotation schema.

Binding Sites↗