Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,315 records · Page 73Linked to original sources

Hunting for genes by functional screens.

Advances in high throughput sequencing technologies have led to an explosion of sequence information available for today's researchers. Efforts in the emerging next phase of the genomic era are focusing on the assignment of function to genes uncovered by genome sequencing programs. The main approaches include high throughput mutagenesis, predictions based on homology in primary sequence, microarray and proteomics. Despite the variety of strategies applied, only 30% of predicted human genes have any function assigned. There is a need, therefore, for additional tools to overcome some of the limitations of existing techniques. In this review we discuss some recent developments and their impact on gene function annotation, especially as they relate to the elucidation of signalling cascades activated by cytokines and growth factors.

Animals↗

Co-clustering and visualization of gene expression data and gene ontology terms for Saccharomyces cerevisiae using self-organizing maps.

We propose a novel co-clustering algorithm that is based on self-organizing maps (SOMs). The method is applied to group yeast (Saccharomyces cerevisiae) genes according to both expression profiles and Gene Ontology (GO) annotations. The combination of multiple databases is supposed to provide a better biological definition and separation of gene clusters. We compare different levels of genome-wide co-clustering by weighting the involved sources of information differently. Clustering quality is determined by both general and SOM-specific validation measures. Co-clustering relies on a sufficient correlation between the different datasets. We investigate in various experiments how much GO information is contained in the applied gene expression dataset and vice versa. The second major contribution is a visualization technique that applies the cluster structure of SOMs for a better biological interpretation of gene (expression) clusterings. Our GO term maps reveal functional neighborhoods between clusters forming biologically meaningful functional SOM regions. To cope with the high variety and specificity of GO terms, gene and cluster annotations are mapped to a reduced vocabulary of more general GO terms. In particular, this advances the ability of SOMs to act as gene function predictors.

Artificial Intelligence↗

The Universal Protein Resource (UniProt).

The Universal Protein Resource (UniProt) provides the scientific community with a single, centralized, authoritative resource for protein sequences and functional information. Formed by uniting the Swiss-Prot, TrEMBL and PIR protein database activities, the UniProt consortium produces three layers of protein sequence databases: the UniProt Archive (UniParc), the UniProt Knowledgebase (UniProt) and the UniProt Reference (UniRef) databases. The UniProt Knowledgebase is a comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase with extensive cross-references. This centrepiece consists of two sections: UniProt/Swiss-Prot, with fully, manually curated entries; and UniProt/TrEMBL, enriched with automated classification and annotation. During 2004, tens of thousands of Knowledgebase records got manually annotated or updated; we introduced a new comment line topic: TOXIC DOSE to store information on the acute toxicity of a toxin; the UniProt keyword list got augmented by additional keywords; we improved the documentation of the keywords and are continuously overhauling and standardizing the annotation of post-translational modifications. Furthermore, we introduced a new documentation file of the strains and their synonyms. Many new database cross-references were introduced and we started to make use of Digital Object Identifiers. We also achieved in collaboration with the Macromolecular Structure Database group at EBI an improved integration with structural databases by residue level mapping of sequences from the Protein Data Bank entries onto corresponding UniProt entries. For convenient sequence searches we provide the UniRef non-redundant sequence databases. The comprehensive UniParc database stores the complete body of publicly available protein sequence data. The UniProt databases can be accessed online (http://www.uniprot.org) or downloaded in several formats (ftp://ftp.uniprot.org/pub). New releases are published every two weeks.

Amino Acid Sequence↗

Age-specific hormonal decline is accompanied by transcriptional changes in human sebocytes in vitro.

The importance of hormones in endogenous aging has been displayed by recent studies performed on animal models and humans. To decipher the molecular mechanisms involved in aging we maintained human sebocytes at defined hormone-substituted conditions that corresponded to average serum levels of females from 20 (f20) to 60 (f60) years of age. The corresponding hormone receptor expression was demonstrated by reverse transcription-polymerase chain reaction (RT-PCR), Western blotting and immunocytochemistry. Cells at f60 produced significantly lower lipids than at f20. Increased mRNA and protein levels of c-Myc and increased protein levels of FN1, which have been associated with aging, were detected in SZ95 sebocytes at f60 compared to those detected at f20 after 5 days of treatment. Expression profiling employing a cDNA microarray composed of 15 529 cDNAs identified 899 genes with altered expression levels at f20 vs. f60. Confirmation of gene regulation was performed by real-time RT-PCR. The functional annotation of these genes according to the Gene Ontology identified pathways related to mitochondrial function, oxidative stress, ubiquitin-mediated proteolysis, cell cycle, immune responses, steroid biosynthesis and phospholipid degradation - all hallmarks of aging. Twenty-five genes in common with those identified in aging kidneys and several genes involved in neurodegenerative diseases were also detected. This is the first report describing the transcriptome of human sebocytes and its modification by a cocktail of hormones administered in age-specific levels and provides an in vitro model system, which approximates some of the hormone-dependent changes in gene transcription that occur during aging in humans.

Aging↗

E-MSD: an integrated data resource for bioinformatics.

The Macromolecular Structure Database (MSD) group (http://www.ebi.ac.uk/msd/) continues to enhance the quality and consistency of macromolecular structure data in the worldwide Protein Data Bank (wwPDB) and to work towards the integration of various bioinformatics data resources. One of the major obstacles to the improved integration of structural databases such as MSD and sequence databases like UniProt is the absence of up to date and well-maintained mapping between corresponding entries. We have worked closely with the UniProt group at the EBI to clean up the taxonomy and sequence cross-reference information in the MSD and UniProt databases. This information is vital for the reliable integration of the sequence family databases such as Pfam and Interpro with the structure-oriented databases of SCOP and CATH. This information has been made available to the eFamily group (http://www.efamily.org.uk/) and now forms the basis of the regular interchange of information between the member databases (MSD, UniProt, Pfam, Interpro, SCOP and CATH). This exchange of annotation information has enriched the structural information in the MSD database with annotation from wider sequence-oriented resources. This work was carried out under the 'Structure Integration with Function, Taxonomy and Sequences (SIFTS)' initiative (http://www.ebi.ac.uk/msd-srv/docs/sifts) in the MSD group.

Amino Acid Sequence↗

ASPIC: a web resource for alternative splicing prediction and transcript isoforms characterization.

Alternative splicing (AS) is now emerging as a major mechanism contributing to the expansion of the transcriptome and proteome complexity of multicellular organisms. The fact that a single gene locus may give rise to multiple mRNAs and protein isoforms, showing both major and subtle structural variations, is an exceptionally versatile tool in the optimization of the coding capacity of the eukaryotic genome. The huge and continuously increasing number of genome and transcript sequences provides an essential information source for the computational detection of genes AS pattern. However, much of this information is not optimally or comprehensively used in gene annotation by current genome annotation pipelines. We present here a web resource implementing the ASPIC algorithm which we developed previously for the investigation of AS of user submitted genes, based on comparative analysis of available transcript and genome data from a variety of species. The ASPIC web resource provides graphical and tabular views of the splicing patterns of all full-length mRNA isoforms compatible with the detected splice sites of genes under investigation as well as relevant structural and functional annotation. The ASPIC web resource-available at http://www.caspur.it/ASPIC/--is dynamically interconnected with the Ensembl and Unigene databases and also implements an upload facility.

Algorithms↗

MultiFun, a multifunctional classification scheme for Escherichia coli K-12 gene products.

An enriched classification system for cellular functions of gene products of Escherichia coli K-12 was developed based on the initial classification by Riley. In the new classification scheme, MultiFun, cellular functions are divided into 10 major categories: Metabolism, Information Transfer, Regulation, Transport, Cell Processes, Cell Structure, Location, Extra-chromosomal Origin, DNA Site, and Cryptic Gene. These major categories are further sub-divided into a hierarchical scheme. Two thousand nine hundred twenty-two gene products of E. coli K-12 were assigned to one or more functions depending on the role they play in the cell. Functional assignments were made to 66% of E. coli gene products, ranging from 1 to 16 assignments per gene product. The expansion of cellular function categories and the assignment to more than one category (multifunction) provides a more complete description of the gene products and their roles and hence better reflects the functional complexity of organisms. We believe this classification system will be useful in the field of genome analysis, both for annotation purposes and for comparative studies. The functional classification scheme and the cellular function assignments made to E. coli gene products can be accessed from the web at the databases GenProtEC (http://genprotec.mbl.edu) and EcoCyc (http://www.ecocyc.org).

Bacterial Proteins↗

Concanavalin A-captured glycoproteins in healthy human urine.

Both the urinary proteome and its glycoproteome can reflect human health status, and more directly, functions of kidney and urinary tracts. Because the high abundance protein albumin is not N-glycosylated, the urine N-glycoprotein enrichment procedure could deplete it, and urine proteome could thus provide a more detailed protein profile in addition to glycosylation information especially when albuminuria occurs in some kidney diseases. In terms of describing the details of urinary proteins, the urine glycoproteome is even a better choice than the proteome itself. Pooled urine samples from healthy volunteers were collected and acetone-precipitated for proteins. N-Linked glycoproteins enriched with concanavalin A affinity purification were separated and analyzed by SDS-PAGE-reverse phase LC/MS/MS or two-dimensional LC/MS/MS. A total of 225 urinary proteins were identified based on two-hit criteria with reliability over 97% for each peptide. Among these proteins, 94 were identified in previous urine proteome works, 150 were annotated as glycoproteins in Swiss-Prot, and 43 were predicted as glycoproteins by NetNGlyc 1.0. A number of known biomarkers and disease-related glycoproteins were identified. Because changes in protein quantity or the glycosylation status can lead to changes in the concanavalin A-captured glycoprotein profile, specific urine glycoproteome patterns might be observed for specific pathological conditions as multiplex urinary biomarkers. Knowledge of the urine glycoproteome is important in understanding kidney and body function.

Concanavalin A↗

Functional metaproteomics for enzyme discovery.

Discovery of microbial biocatalysts traditionally relied on activity screening of isolated bacterial strains. However, since most microorganisms cannot be cultivated in the lab, such an approach leaves the majority of the microbial enzyme diversity untapped. Metagenomic approaches, in which the DNA from a microbial community is directly isolated and then used either for the creation of an expression library or for sequencing and metagenome annotation have alleviated this shortcoming to an extent, but have their own limitations: the generation of large expression libraries is time-consuming and their screening is costly, while metagenome annotation can infer biocatalytic function only from prior knowledge. We have thus developed a functional metaproteomic approach, which combines the immediacy of traditional activity screening with the comprehensiveness of a meta-omics approach. Briefly, the whole metaproteome of an environmental sample is separated on a 2-D gel, biocatalytically active proteins are visualized in-gel through zymography, and those candidate biocatalysts are then identified through mass spectrometry, searching against a metagenome-derived database obtained from the very same environmental sample. Here we explain the process in detail, with a focus on esterases, and give guidelines on how to develop a functional metaproteomic workflow for enzyme discovery.

Proteomics↗

Analysis of the pdx-1 (snz-1/sno-1) region of the Neurospora crassa genome: correlation of pyridoxine-requiring phenotypes with mutations in two structural genes.

We report the analysis of a 36-kbp region of the Neurospora crassa genome, which contains homologs of two closely linked stationary phase genes, SNZ1 and SNO1, from Saccharomyces cerevisiae. Homologs of SNZ1 encode extremely highly conserved proteins that have been implicated in pyridoxine (vitamin B6) metabolism in the filamentous fungi Cercospora nicotianae and in Aspergillus nidulans. In N. crassa, SNZ and SNO homologs map to the region occupied by pdx-1 (pyridoxine requiring), a gene that has been known for several decades, but which was not sequenced previously. In this study, pyridoxine-requiring mutants of N. crassa were found to possess mutations that disrupt conserved regions in either the SNZ or SNO homolog. Previously, nearly all of these mutants were classified as pdx-1. However, one mutant with a disrupted SNO homolog was at one time designated pdx-2. It now appears appropriate to reserve the pdx-1 designation for the N. crassa SNZ homolog and pdx-2 for the SNO homolog. We further report annotation of the entire 36,030-bp region, which contains at least 12 protein coding genes, supporting a previous conclusion of high gene densities (12,000-13,000 total genes) for N. crassa. Among genes in this region other than SNZ and SNO homologs, there was no evidence of shared function. Four of the genes in this region appear to have been lost from the S. cerevisiae lineage.

Cloning, Molecular↗

Bioinformatics: use in bacterial vaccine discovery.

Bioinformatics has now become a common laboratory name for groups studying genomic sequences. It is composed of many different, yet interrelated scientific fields such as genomics, proteomics, and transcriptional profiling. The availability of complete genomic sequences, especially prokaryotic organisms, allows one to rapidly identify, analyze, and clone genes of interest. For bacterial vaccine discovery, one can "mine" the genomic sequence for potential surface targets using various algorithms, characterize these gene targets, and produce primers for cloning, all before one enters the wet laboratory. This review will focus on various genomic mining tools/algorithms available for predicting open reading frames and their associated annotation (if known), physical and functional characterization, and cellular localization. Finally, examples are given of how all of this is being used for the identification of potential bacterial vaccine candidates.

Animals↗

VisANT: an online visualization and analysis tool for biological interaction data.

BACKGROUND: New techniques for determining relationships between biomolecules of all types--genes, proteins, noncoding DNA, metabolites and small molecules--are now making a substantial contribution to the widely discussed explosion of facts about the cell. The data generated by these techniques promote a picture of the cell as an interconnected information network, with molecular components linked with one another in topologies that can encode and represent many features of cellular function. This networked view of biology brings the potential for systematic understanding of living molecular systems. RESULTS: We present VisANT, an application for integrating biomolecular interaction data into a cohesive, graphical interface. This software features a multi-tiered architecture for data flexibility, separating back-end modules for data retrieval from a front-end visualization and analysis package. VisANT is a freely available, open-source tool for researchers, and offers an online interface for a large range of published data sets on biomolecular interactions, including those entered by users. This system is integrated with standard databases for organized annotation, including GenBank, KEGG and SwissProt. VisANT is a Java-based, platform-independent tool suitable for a wide range of biological applications, including studies of pathways, gene regulation and systems biology. CONCLUSION: VisANT has been developed to provide interactive visual mining of biological interaction data sets. The new software provides a general tool for mining and visualizing such data in the context of sequence, pathway, structure, and associated annotations. Interaction and predicted association data can be combined, overlaid, manipulated and analyzed using a variety of built-in functions. VisANT is available at http://visant.bu.edu.

Animals↗

LALNVIEW: a graphical viewer for pairwise sequence alignments.

LALNVIEW is a graphical program for visualising local alignments between two sequences (protein or nucleic acids). Sequences are represented by coloured rectangles to give an overall picture of their similarities. LALNVIEW can display sequence features (exon, intron, active site, domain, propeptide, etc.) along with the alignment. When using LALNVIEW through our Web servers, sequence features are automatically extracted from database annotations (SWISS-PROT, GenBank, EMBL or HOVERGEN) and displayed with the alignment. LALNVIEW is a useful tool for analysing pairwise sequence alignments and for making the link between sequence homology and what is known about the structure or function of sequences. LALNVIEW executables for UNIX, Macintosh and PC computers are freely available from our server (http:// expasy.hcuge.ch/sprot/lalnview.html).

Acyltransferases↗

The X-ray structure of Escherichia coli RraA (MenG), A protein inhibitor of RNA processing.

The Escherichia coli protein regulator of RNase E activity A (RraA) has recently been shown to act as a trans-acting modulator of RNA turnover in bacteria; it binds to the essential endonuclease RNase E and inhibits RNA processing in vivo and in vitro. Here, we report the 2.0A X-ray structure of RraA. The structure reveals a ring-like trimer with a central cavity of approximately 12A in diameter. Based on earlier sequence analysis, RraA had been identified as a putative S-adenosylmethionine:2-demethylmenaquinone and was annotated as MenG. However, an analysis of the RraA structure shows that the protein lacks the structural motifs usually required for methylases. Comparison of the observed fold with that of other proteins (and domains) suggests that the RraA fold is an ancient platform that has been adapted for a wide range of functions. An analysis of the amino acid sequence shows that the E.coli RraA exhibits an ancient relationship to a family of aldolases.

Aldehyde-Lyases↗

Sequence and comparative analysis of the chicken genome provide unique perspectives on vertebrate evolution.

We present here a draft genome sequence of the red jungle fowl, Gallus gallus. Because the chicken is a modern descendant of the dinosaurs and the first non-mammalian amniote to have its genome sequenced, the draft sequence of its genome--composed of approximately one billion base pairs of sequence and an estimated 20,000-23,000 genes--provides a new perspective on vertebrate genome evolution, while also improving the annotation of mammalian genomes. For example, the evolutionary distance between chicken and human provides high specificity in detecting functional elements, both non-coding and coding. Notably, many conserved non-coding sequences are far from genes and cannot be assigned to defined functional classes. In coding regions the evolutionary dynamics of protein domains and orthologous groups illustrate processes that distinguish the lineages leading to birds and mammals. The distinctive properties of avian microchromosomes, together with the inferred patterns of conserved synteny, provide additional insights into vertebrate chromosome architecture.

Animals↗

The Comprehensive Microbial Resource.

One challenge presented by large-scale genome sequencing efforts is effective display of uniform information to the scientific community. The Comprehensive Microbial Resource (CMR) contains robust annotation of all complete microbial genomes and allows for a wide variety of data retrievals. The bacterial information has been placed on the Web at http://www.tigr.org/CMR for retrieval using standard web browsing technology. Retrievals can be based on protein properties such as molecular weight or hydrophobicity, GC-content, functional role assignments and taxonomy. The CMR also has special web-based tools to allow data mining using pre-run homology searches, whole genome dot-plots, batch downloading and traversal across genomes using a variety of datatypes.

Bacteria↗

Sequence annotation of nuclear receptor ligand-binding domains by automated homology modeling.

The quality of three-dimensional homology models derived from protein sequences provides an independent measure of the suitability of a protein sequence for a certain fold. We have used automated homology modeling and model assessment tools to identify putative nuclear hormone receptor ligand-binding domains in the genome of Caenorhabditis elegans. Our results indicate that the availability of multiple crystal structures is crucial to obtaining useful models in this receptor family. The majority of annotated mammalian nuclear hormone receptors could be assigned to a ligand-binding domain fold by using the best model derived from any of four template structures. This strategy also assigned the ligand-binding domain fold to a number of C.elegans. sequences without prior annotation. Interestingly, the retinoic acid receptor crystal structure contributed most to the number of sequences that could be assigned to a ligand-binding domain fold. Several causes for this can be suggested, including the high quality of this protein structure in terms of our assessment tools, similarity between the biological function or ligand of this receptor and the modeled genes and gene duplication in C.elegans.

Animals↗

Genomic signatures of innovation and selection in the extremotolerant yeast Kluyveromyces marxianus.

Extremophiles can be the product of millions of years of evolutionary engineering and refinement. The underlying mechanisms can be quite distinct from the ones operating at earlier stages of trait innovation. In this work, we have developed the compost yeast Kluyveromyces marxianus, which diverged from its closest relative >20 million years ago, as a model for interspecies comparative biology and genomics. We applied a battery of growth assays to species of the Kluyveromyces genus and found that K. marxianus outperformed its relatives in a battery of heat and chemical stress conditions. We then generated and analyzed genomes from across the genus, to find derived genetic features associated with, and potentially causal for, K. marxianus traits. We found robust expansions in gene families in the K. marxianus genome, most notably among genes annotated as transmembrane transporters and in metabolism. In molecular-evolution tests, we identified adaptive protein variants at hundreds of genes, among which plasma membrane transporters were over-represented. Together, these signals enable a model for the molecular mechanisms and evolutionary pressures underlying K. marxianus traits, including gains in transporter function mediating stress resistance, and metabolic variants contributing to its capacity for rapid growth in challenging conditions. Such oligogenic architectures may be the rule rather than the exception in phenotypes that have evolved over long timescales.

Journal Article↗