Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

Prediction of orthologous relationship by functionally important sites.

Making accurate functional predictions plays an important role in the era of proteomics. Reliable functional information can be extracted from orthologs in other species when annotating an unknown gene. Here a site-based approach called PORFIS is proposed to predict orthologous relationship. When applied to the bacterial transcription factor PurR/LacI family and the protein kinase AGC family, our method was able to identify, with few false positives, the important sites that agree with those verified by biological experiments. We also tested it on the alpha-proteasome family, the glycoprotein hormone family and the growth hormone family to demonstrate its ability to predict orthologous relationship. Compared with other prediction methods based on phylogenetic analysis or hidden Markov models, PORFIS not only has competitive prediction accuracy, but also provides valuable biological information of functionally important sites associated with orthologs which can be further studied in biological experiments.

Humans↗

Gene expression profiling of muscle tissue in Brahman steers during nutritional restriction.

Expression profiling using microarrays allows for the detailed characterization of the gene networks that regulate an animal's response to environmental stresses. During nutritional restriction, processes such as protein turnover, connective tissue remodeling, and muscle atrophy take place in the skeletal muscle of the animal. These processes and their regulation are of interest in the context of managing livestock for optimal production efficiency and product quality. Here we expand on recent research applying complementary DNA (cDNA) microarray technology to the study of the effect of nutritional restriction on bovine skeletal muscle. Using a custom cDNA microarray of 9,274 probes from cattle muscle and s.c. fat libraries, we examined the differential gene expression profile of the LM from 10 Brahman steers under three different dietary treatments. The statistical approach was based on mixed-model ANOVA and model-based clustering of the BLUP solutions for the gene x diet interaction effect. From the results, we defined a transcript profile of 156 differentially expressed array elements between the weight loss and weight gain diet substrates. After sequence and annotation analyses, the 57 upregulated elements represented 29 unique genes, and the 99 downregulated elements represented 28 unique genes. Most of these co-regulated genes cluster into groups with distinct biological function related to protein turnover and cytoskeletal metabolism and contribute to our mechanistic understanding of the processes associated with remodeling of muscle tissue in response to nutritional stress.

Analysis of Variance↗

Protein classification artificial neural system.

A neural network classification method is developed as an alternative approach to the large database search/organization problem. The system, termed Protein Classification Artificial Neural System (ProCANS), has been implemented on a Cray supercomputer for rapid superfamily classification of unknown proteins based on the information content of the neural interconnections. The system employs an n-gram hashing function that is similar to the k-tuple method for sequence encoding. A collection of modular back-propagation networks is used to store the large amount of sequence patterns. The system has been trained and tested with the first 2,148 of the 8,309 entries of the annotated Protein Identification Resource protein sequence database (release 29). The entries included the electron transfer proteins and the six enzyme groups (oxidoreductases, transferases, hydrolases, lyases, isomerases, and ligases), with a total of 620 superfamilies. After a total training time of seven Cray central processing unit (CPU) hours, the system has reached a predictive accuracy of 90%. The classification is fast (i.e., 0.1 Cray CPU second per sequence), as it only involves a forward-feeding through the networks. The classification time on a full-scale system embedded with all known superfamilies is estimated to be within 1 CPU second. Although the training time will grow linearly with the number of entries, the classification time is expected to remain low even if there is a 10-100-fold increase of sequence entries. The neural database, which consists of a set of weight matrices of the networks, together with the ProCANS software, can be ported to other computers and made available to the genome community. The rapid and accurate superfamily classification would be valuable to the organization of protein sequence databases and to the gene recognition in large sequencing projects.

Computers, Mainframe↗

Approaches to defining the ancestral eukaryotic protein complexome.

In this paper, we integrate and summarize the currently available information on the ancestral eukaryotic protein complexome, which is defined as the set of protein complexes that extant eukaryotes inherited from their last common ancestor. From the literature, we compiled lists of complexes with three or more distinct protein components from well-studied eukaryotic model organisms. Combinatorial complexes of membrane-associated signalling proteins and specific transcription factors were disregarded. A stringent but sensitive novel orthology detection algorithm, complemented with manual sequence similarity searches and with published data on whole genome or segmental and tandem gene duplications, enabled us to map the vast majority of these complexes to a virtual primitive eukaryote termed Eukaryotic Virtual Ancestor (EVA). EVA is intended to resemble the last common eukaryotic ancestor and to emulate the biological common denominator of the major extent eukaryotic lineages at the molecular level. The dataset was then used for the functional and domain annotation of the ancestral eukaryotic complexome. Furthermore, we illustrate its usefulness for inferring complexes of poorly studied eukaryotes and for the recognition of highly divergent orthologs. We also discuss the evolution of the circa 1,400 complex-associated ancestral proteins. As about 90% of these proteins have been conserved in all thirteen studied free-living eukaryotes, the evolutionary reduction and loss of complexes seems minimal. Moreover, the available data suggest that, in general, the acquisition of stable complexes of novel design occurs too slowly to be a major contributor to evolutionary innovation. Finally, given the stability of the ancestral eukarotic complexome we propose its use in the formulation of the mathematical systems that aim to simulate biological processes. Our data suggest that these simplified formulations can apply to most free-living model eukaryotes.

Animals↗

inGeno--an integrated genome and ortholog viewer for improved genome to genome comparisons.

BACKGROUND: Systematic genome comparisons are an important tool to reveal gene functions, pathogenic features, metabolic pathways and genome evolution in the era of post-genomics. Furthermore, such comparisons provide important clues for vaccines and drug development. Existing genome comparison software often lacks accurate information on orthologs, the function of similar genes identified and genome-wide reports and lists on specific functions. All these features and further analyses are provided here in the context of a modular software tool "inGeno" written in Java with Biojava subroutines. RESULTS: InGeno provides a user-friendly interactive visualization platform for sequence comparisons (comprehensive reciprocal protein--protein comparisons) between complete genome sequences and all associated annotations and features. The comparison data can be acquired from several different sequence analysis programs in flexible formats. Automatic dot-plot analysis includes output reduction, filtering, ortholog testing and linear regression, followed by smart clustering (local collinear blocks; LCBs) to reveal similar genome regions. Further, the system provides genome alignment and visualization editor, collinear relationships and strain-specific islands. Specific annotations and functions are parsed, recognized, clustered, logically concatenated and visualized and summarized in reports. CONCLUSION: As shown in this study, inGeno can be applied to study and compare in particular prokaryotic genomes against each other (gram positive and negative as well as close and more distantly related species) and has been proven to be sensitive and accurate. This modular software is user-friendly and easily accommodates new routines to meet specific user-defined requirements.

Base Sequence↗

Text-mining and information-retrieval services for molecular biology.

Text-mining in molecular biology -- defined as the automatic extraction of information about genes, proteins and their functional relationships from text documents -- has emerged as a hybrid discipline on the edges of the fields of information science, bioinformatics and computational linguistics. A range of text-mining applications have been developed recently that will improve access to knowledge for biologists and database annotators.

Computational Biology↗

The mouse genome database (MGD): new features facilitating a model system.

The mouse genome database (MGD, http://www.informatics.jax.org/), the international community database for mouse, provides access to extensive integrated data on the genetics, genomics and biology of the laboratory mouse. The mouse is an excellent and unique animal surrogate for studying normal development and disease processes in humans. Thus, MGD's primary goals are to facilitate the use of mouse models for studying human disease and enable the development of translational research hypotheses based on comparative genotype, phenotype and functional analyses. Core MGD data content includes gene characterization and functions, phenotype and disease model descriptions, DNA and protein sequence data, polymorphisms, gene mapping data and genome coordinates, and comparative gene data focused on mammals. Data are integrated from diverse sources, ranging from major resource centers to individual investigator laboratories and the scientific literature, using a combination of automated processes and expert human curation. MGD collaborates with the bioinformatics community on the development of data and semantic standards, and it incorporates key ontologies into the MGD annotation system, including the Gene Ontology (GO), the Mammalian Phenotype Ontology, and the Anatomical Dictionary for Mouse Development and the Adult Anatomy. MGD is the authoritative source for mouse nomenclature for genes, alleles, and mouse strains, and for GO annotations to mouse genes. MGD provides a unique platform for data mining and hypothesis generation where one can express complex queries simultaneously addressing phenotypic effects, biochemical function and process, sub-cellular location, expression, sequence, polymorphism and mapping data. Both web-based querying and computational access to data are provided. Recent improvements in MGD described here include the incorporation of single nucleotide polymorphism data and search tools, the addition of PIR gene superfamily classifications, phenotype data for NIH-acquired knockout mice, images for mouse phenotypic genotypes, new functional graph displays of GO annotations, and new orthology displays including sequence information and graphic displays.

Animals↗

Functional analysis and annotation of the virulence plasmid pMUM001 from Mycobacterium ulcerans.

The presence of a 174 kb plasmid called pMUM001 in Mycobacterium ulcerans, the first example of a mycobacterial plasmid encoding a virulence determinant, was recently reported. Over half of pMUM001 is devoted to six genes, three of which encode giant polyketide synthases (PKS) that produce mycolactone, an unusual cytotoxic lipid produced by M. ulcerans. In this present study the remaining 75 non-PKS-associated protein-coding sequences (CDS) are analysed and it is shown that pMUM001 is a low-copy-number element with a functional ori that supports replication in Mycobacterium marinum but not in the fast-growing mycobacteria Mycobacterium smegmatis and Mycobacterium fortuitum. Sequence analyses revealed a highly mosaic plasmid gene structure that is reminiscent of other large plasmids. Insertion sequences (IS) and fragments of IS, some previously unreported, are interspersed among functional gene clusters, such as those genes involved in plasmid replication, the synthesis of mycolactone, and a potential phosphorelay signal transduction system. Among the IS present on pMUM001 were multiple copies of the high-copy-number M. ulcerans elements IS2404 and IS2606. No plasmid transfer systems were identified, suggesting that trans-acting factors are required for mobilization. The results presented here provide important insights into this unusual virulence plasmid from an emerging but neglected human pathogen.

Bacterial Proteins↗

Structure-based functional annotation: yeast ymr099c codes for a D-hexose-6-phosphate mutarotase.

Despite the generation of a large amount of sequence information over the last decade, more than 40% of well characterized enzymatic functions still lack associated protein sequences. Assigning protein sequences to documented biochemical functions is an interesting challenge. We illustrate here that structural genomics may be a reasonable approach in addressing these questions. We present the crystal structure of the Saccharomyces cerevisiae YMR099cp, a protein of unknown function. YMR099cp adopts the same fold as galactose mutarotase and shares the same catalytic machinery necessary for the interconversion of the alpha and beta anomers of galactose. The structure revealed the presence in the active site of a sulfate ion attached by an arginine clamp made by the side chain from two strictly conserved arginine residues. This sulfate is ideally positioned to mimic the phosphate group of hexose 6-phosphate. We have subsequently successfully demonstrated that YMR099cp is a hexose-6-phosphate mutarotase with broad substrate specificity. We solved high resolution structures of some substrate enzyme complexes, further confirming our functional hypothesis. The metabolic role of a hexose-6-phosphate mutarotase is discussed. This work illustrates that structural information has been crucial to assign YMR099cp to the orphan EC activity: hexose-phosphate mutarotase.

Amino Acid Sequence↗

The SYSTERS Protein Family Database in 2005.

The SYSTERS project aims to provide a meaningful partitioning of the whole protein sequence space by a fully automatic procedure. A refined two-step algorithm assigns each protein to a family and a superfamily. The sequence data underlying SYSTERS release 4 now comprise several protein sequence databases derived from completely sequenced genomes (ENSEMBL, TAIR, SGD and GeneDB), in addition to the comprehensive Swiss-Prot/TrEMBL databases. The SYSTERS web server (http://systers.molgen.mpg.de) provides access to 158 153 SYSTERS protein families. To augment the automatically derived results, information from external databases like Pfam and Gene Ontology are added to the web server. Furthermore, users can retrieve pre-processed analyses of families like multiple alignments and phylogenetic trees. New query options comprise a batch retrieval tool for functional inference about families based on automatic keyword extraction from sequence annotations. A new access point, PhyloMatrix, allows the retrieval of phylogenetic profiles of SYSTERS families across organisms with completely sequenced genomes.

Algorithms↗

Annotation of glycoproteins in the SWISS-PROT database.

SWISS-PROT is a protein sequence database, which aims to be nonredundant, fully annotated and highly cross-referenced. Most eukaryotic gene products undergo co- and/or post-translational modifications, and these need to be included in the database in order to describe the mature protein. SWISS-PROT includes information on many types of different protein modifications. As glycosylation is the most common type of post-translational protein modification, we are currently placing an emphasis on annotation of protein glycosylation in SWISS-PROT. Information on the position of the sugar within the polypeptide chain, the reducing terminal linkage as well as additional information on biological function of the sugar is included in the database. In this paper we describe how we account for the different types of protein glycosylation, namely N-linked glycosylation, O-linked glycosylation, proteoglycans, C-linked glycosylation and the attachment of glycosyl-phosphatidylinosital anchors to proteins.

Amino Acid Sequence↗

Mapping the surface properties of macromolecules.

Methods are presented for the rapid computation of schematic projections of the surfaces of macromolecules, similar to the "roadmaps" used to illustrate the surfaces of viruses (Rossmann, M.G. & Palmenberg, A.C., 1988, Virology 164, 373-382). Several types of projections are described, extending the application of "roadmaps" to the external surfaces of all macromolecules and their interior binding pockets and pores. The surface projections, showing the positions of residues, can be colored, shaded, contoured, and annotated to show physical, sequence, or functional properties such as surface topology, hydrophobicity, or sequence conservation, for example. The automated procedures are useful for surveys of the surface features of proteins sharing similar functional properties.

Computer Graphics↗

CaXML: Chemistry-informed machine learning explains mutual changes between protein conformations and calcium ions in calcium-binding proteins using structural and topological features.

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of CaXML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

Machine Learning↗

Genetic linkage map and expression analysis of genes expressed in the lamellae of the edible basidiomycete Pleurotus ostreatus.

Pleurotus ostreatus is an industrially cultivated basidiomycete with nutritional and environmental applications. Its genome contains 35 Mbp organized in 11 chromosomes. There is currently available a genetic linkage map based predominantly on anonymous molecular markers complemented with the mapping of QTLs controlling growth rate and industrial productivity. To increase the saturation of the existing linkage maps, we have identified and mapped 82 genes expressed in the lamellae. Their manual annotation revealed that 34.1% of the lamellae-expressed and 71.5% of the lamellae-specific genes correspond to previously unknown sequences or to hypothetical proteins without a clearly established function. Furthermore, the expression pattern of some genes provides an experimental basis for studying gene regulation during the change from vegetative to reproductive growth. Finally, the identification of various differentially regulated genes involved in protein metabolism suggests the relevance of these processes in fruit body formation and maturation.

Chromosome Mapping↗

Age-dependent changes of gene expression in the Drosophila head.

Previous gene expression profiling studies in Drosophila have provided clues for understanding the aging process at the gene expression level. For a detailed understanding, studies of specific regions of the body are necessary. We therefore employed microarray analysis to examine gene expression changes in the Drosophila head during aging. Six hundred and eighty-four of the 5405 genes present in the microarray showed significant age-dependent changes as determined by significance analysis of microarray (SAM) (q < 0.05). The biological significance of the changes was analyzed using the gene annotations provided by the Gene Ontology Consortium. Major changes involved genes affecting energy metabolism (proton transport, energy pathways, oxidative phosphorylation) and neuronal function, especially responses to light. Genes involved in protein catabolism and several other metabolic processes also showed age-dependent changes. Most of the changes were reductions in gene expression and occurred before day 13 of adult life. After day 13, the age-dependent gene expression changes were relatively smaller than earlier life. Interestingly, the two biological processes of major gene expression changes are related to the two known environmental changes that increase life span in Drosophila: caloric restriction and light reduction. Our findings suggest that light signaling and energy metabolism may be important biological processes affected by aging and be interesting targets for the further investigation related to the longevity in Drosophila.

Age Factors↗

Assessment of genome-wide protein function classification for Drosophila melanogaster.

The functional classification of genes on a genome-wide scale is now in its infancy, and we make a first attempt to assess existing methods and identify sources of error. To this end, we compared two independent efforts for associating proteins with functions, one implemented by FlyBase and the other by PANTHER at Celera Genomics. Both methods make inferences based on sequence similarity and the available experimental evidence. However, they differ considerably in methodology and process. Overall, assuming that the systematic error across the two methods is relatively small, we find the protein-to-function association error rate of both the FlyBase and PANTHER methods to be <2%. The primary source of error for both methods appears to be simple human error. Although homology-based inference can certainly cause errors in annotation, our analysis indicates that the frequency of such errors is relatively small compared with the number of correct inferences. Moreover, these homology errors can be minimized by careful tree-based inference, such as that implemented in PANTHER. Often, functional associations are made by one method and not the other, indicating that one of the greatest challenges lies in improving the completeness of available ontology associations.

Animals↗

The PredictProtein server.

PredictProtein (http://www.predictprotein.org) is an Internet service for sequence analysis and the prediction of protein structure and function. Users submit protein sequences or alignments; PredictProtein returns multiple sequence alignments, PROSITE sequence motifs, low-complexity regions (SEG), nuclear localization signals, regions lacking regular structure (NORS) and predictions of secondary structure, solvent accessibility, globular regions, transmembrane helices, coiled-coil regions, structural switch regions, disulfide-bonds, sub-cellular localization and functional annotations. Upon request fold recognition by prediction-based threading, CHOP domain assignments, predictions of transmembrane strands and inter-residue contacts are also available. For all services, users can submit their query either by electronic mail or interactively via the World Wide Web.

Internet↗

The ABC transporter gene family of Caenorhabditis elegans has implications for the evolutionary dynamics of multidrug resistance in eukaryotes.

BACKGROUND: Many drugs of natural origin are hydrophobic and can pass through cell membranes. Hydrophobic molecules must be susceptible to active efflux systems if they are to be maintained at lower concentrations in cells than in their environment. Multi-drug resistance (MDR), often mediated by intrinsic membrane proteins that couple energy to drug efflux, provides this function. All eukaryotic genomes encode several gene families capable of encoding MDR functions, among which the ABC transporters are the largest. The number of candidate MDR genes means that study of the drug-resistance properties of an organism cannot be effectively carried out without taking a genomic perspective. RESULTS: We have annotated sequences for all 60 ABC transporters from the Caenorhabditis elegans genome, and performed a phylogenetic analysis of these along with the 49 human, 30 yeast, and 57 fly ABC transporters currently available in GenBank. Classification according to a unified nomenclature is presented. Comparison between genomes reveals much gene duplication and loss, and surprisingly little orthology among analogous genes. Proteins capable of conferring MDR are found in several distinct subfamilies and are likely to have arisen independently multiple times. CONCLUSIONS: ABC transporter evolution fits a pattern expected from a process termed 'dynamic-coherence'. This is an unusual result for such a highly conserved gene family as this one, present in all domains of cellular life. Mechanistically, this may result from the broad substrate specificity of some ABC proteins, which both reduces selection against gene loss, and leads to the facile sorting of functions among paralogs following gene duplication.

ATP-Binding Cassette Transporters↗