Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

ESTs from the basidiomycete Schizophyllum commune grown on nitrogen-replete and nitrogen-limited media.

Lambda phage cDNA libraries were constructed using mRNAs from the basidiomycete Schizophyllum commune grown on media with high or low nitrogen concentrations. A total of 440 clones were sequenced, representing 373 distinct transcripts. Of these, 166 showed significant similarity to annotated genes in GenBank. Those that could be tentatively identified using BLAST searches were classified by function using the Gene Ontology (GO) database. Genes with products involved in cell-cycle processes were more frequent in the nitrogen-limited libraries, while genes with products involved in protein biosynthesis were more frequent in the nitrogen-replete library. Overall, clones showed much greater similarity to the one publicly available basidiomycete genome, Phanerochaete chrysosporium, than to any of the ascomycete genomes.

Bacteriophage lambda↗

An integrated data analysis approach to characterize genes highly expressed in hepatocellular carcinoma.

Hepatocellular carcinoma (HCC) is one of the major causes of cancer deaths worldwide. New diagnostic and therapeutic options are needed for more effective and early detection and treatment of this malignancy. We identified 703 genes that are highly expressed in HCC using DNA microarrays, and further characterized them in order to uncover novel tumor markers, oncogenes, and therapeutic targets for HCC. Using Gene Ontology annotations, genes with functions related to cell proliferation and cell cycle, chromatin, repair, and transcription were found to be significantly enriched in this list of highly expressed genes. We also identified a set of genes that encode secreted (e.g. GPC3, LCN2, and DKK1) or membrane-bound proteins (e.g. GPC3, IGSF1, and PSK-1), which may be attractive candidates for the diagnosis of HCC. A significant enrichment of genes highly expressed in HCC was found on chromosomes 1q, 6p, 8q, and 20q, and we also identified chromosomal clusters of genes highly expressed in HCC. The microarray analyses were validated by RT-PCR and PCR. This approach of integrating other biological information with gene expression in the analysis helps select aberrantly expressed genes in HCC that may be further studied for their diagnostic or therapeutic utility.

Biomarkers, Tumor↗

Finding errors in DNA sequences.

An algorithm is described that can detect certain errors within coding regions of DNA sequences. The algorithm is based on the idea that an insertion or deletion error within a coding sequence would interrupt the reading frame and cause the correct translation of a DNA sequence to require one or more frameshifts. If the coding sequence shows similarity to a known protein sequence then such errors can be detected by comparing the conceptual translations of DNA sequences in all six reading frames with every sequence in a protein sequence data base. We have incorporated these ideas into a computer program, called DETECT, that can serve as an aid to the experimentalist who is determining new DNA sequences so that obvious errors may be located and corrected. The program has been tested using raw experimental data and against sequences from the European Molecular Biology Laboratory data base, annotated as containing frameshifts. We have also tested it using unidentified open reading frames that flank known, annotated genes in the GenBank data base. Many potential errors are apparent and in some cases functions can be suggested for the "corrected" versions of these reading frames leading to the identification of new genes. As more sequences are determined the power of this method will increase substantially.

Adenylyl Cyclases↗

Emergence of a subfamily of xylanase inhibitors within glycoside hydrolase family 18.

The xylanase inhibitor protein I (XIP-I), recently identified in wheat, inhibits xylanases belonging to glycoside hydrolase families 10 (GH10) and 11 (GH11). Sequence and structural similarities indicate that XIP-I is related to chitinases of family GH18, despite its lack of enzymatic activity. Here we report the identification and biochemical characterization of a XIP-type inhibitor from rice. Despite its initial classification as a chitinase, the rice inhibitor does not exhibit chitinolytic activity but shows specificities towards fungal GH11 xylanases similar to that of its wheat counterpart. This, together, with an analysis of approximately 150 plant members of glycosidase family GH18 provides compelling evidence that xylanase inhibitors are largely represented in this family, and that this novel function has recently emerged based on a common scaffold. The plurifunctionality of GH18 members has major implications for genomic annotations and predicted gene function. This study provides new information which will lead to a better understanding of the biological significance of a number of GH18 'inactivated' chitinases.

Amino Acid Sequence↗

Identification and functional analysis of 'hypothetical' genes expressed in Haemophilus influenzae.

The progress in genome sequencing has led to a rapid accumulation in GenBank submissions of uncharacterized 'hypothetical' genes. These genes, which have not been experimentally characterized and whose functions cannot be deduced from simple sequence comparisons alone, now comprise a significant fraction of the public databases. Expression analyses of Haemophilus influenzae cells using a combination of transcriptomic and proteomic approaches resulted in confident identification of 54 'hypothetical' genes that were expressed in cells under normal growth conditions. In an attempt to understand the functions of these proteins, we used a variety of publicly available analysis tools. Close homologs in other species were detected for each of the 54 'hypothetical' genes. For 16 of them, exact functional assignments could be found in one or more public databases. Additionally, we were able to suggest general functional characterization for 27 more genes (comprising approximately 80% total). Findings from this analysis include the identification of a pyruvate-formate lyase-like operon, likely to be expressed not only in H.influenzae but also in several other bacteria. Further, we also observed three genes that are likely to participate in the transport and/or metabolism of sialic acid, an important component of the H.influenzae lipo-oligosaccharide. Accurate functional annotation of uncharacterized genes calls for an integrative approach, combining expression studies with extensive computational analysis and curation, followed by eventual experimental verification of the computational predictions.

Amino Acid Sequence↗

Proteome research: complementarity and limitations with respect to the RNA and DNA worlds.

A methodological overview of proteome analysis is provided along with details of efforts to achieve high-throughput screening (HTS) of protein samples derived from two-dimensional electrophoresis gels. For both previously sequenced organisms and those lacking significant DNA sequence information, mass spectrometry has a key role to play in achieving HTS. Prototype robotics designed to conduct appropriate chemistries and deliver 700-1000 protein (genes) per day to batteries of mass spectrometers or liquid chromatography (LC)-based analyses are well advanced, as are efforts to produce high density gridded arrays containing > 1000 proteins on a single matrix assisted laser desorption ionisation/time-of-flight (MALDI-TOF) sample stage. High sensitivity HTS of proteins is proposed by employing principally mass spectrometry in an hierarchical manner: (i) MALDI-TOF-mass spectrometry (MS) on at least 1000 proteins per day; (ii) electrospray ionisation (ESI)/MS/MS for analysis of peptides with respect to predicted fragmentation patterns or by sequence tagging; and (iii) ESI/MS/MS for peptide sequencing. Genomic sequences when complemented with information derived from hybridisation assays and proteome analysis may herald in a new era of holistic cellular biology. The current preoccupation with the absolute quantity of gene-product (RNA and/or protein) should move backstage with respect to more molecularly relevant parameters, such as: molecular half-life; synthesis rate; functional competence (presence or absence of mutations); reaction kinetics; the influence of individual gene-products on biochemical flux; the influence of the environment, cell-cycle, stress and disease on gene-products; and the collective roles of multigenic and epigenetic phenomena governing cellular processes. Proteome analysis is demonstrated as being capable of proceeding independently of DNA sequence information and aiding in genomic annotation. Its ability to confirm the existence of gene-products predicted from DNA sequence is a major contribution to genomic science. The workings of software engines necessary to achieve large-scale proteome analysis are outlined, along with trends towards miniaturisation, analyte concentration and protein detection independent of staining technologies. A challenge for proteome analysis into the future will be to reduce its dependence on two-dimensional (2-D) gel electrophoresis as the preferred method of separating complex mixtures of cellular proteins. Nonetheless, proteome analysis already represents a means of efficiently complementing differential display, high density expression arrays, expressed sequence tags, direct or subtractive hybridisation, chromosomal linkage studies and nucleic acid sequencing as a problem solving tool in molecular biology.

Animals↗

The CATH domain structure database: new protocols and classification levels give a more comprehensive resource for exploring evolution.

We report the latest release (version 3.0) of the CATH protein domain database (http://www.cathdb.info). There has been a 20% increase in the number of structural domains classified in CATH, up to 86 151 domains. Release 3.0 comprises 1110 fold groups and 2147 homologous superfamilies. To cope with the increases in diverse structural homologues being determined by the structural genomics initiatives, more sensitive methods have been developed for identifying boundaries in multi-domain proteins and for recognising homologues. The CATH classification update is now being driven by an integrated pipeline that links these automated procedures with validation steps, that have been made easier by the provision of information rich web pages summarising comparison scores and relevant links to external sites for each domain being classified. An analysis of the population of domains in the CATH hierarchy and several domain characteristics are presented for version 3.0. We also report an update of the CATH Dictionary of homologous structures (CATH-DHS) which now contains multiple structural alignments, consensus information and functional annotations for 1459 well populated superfamilies in CATH. CATH is directly linked to the Gene3D database which is a projection of CATH structural data onto approximately 2 million sequences in completed genomes and UniProt.

Classification↗

The Mouse Genome Database (MGD): updates and enhancements.

The Mouse Genome Database (MGD) integrates genetic and genomic data for the mouse in order to facilitate the use of the mouse as a model system for understanding human biology and disease processes. A core component of the MGD effort is the acquisition and integration of genomic, genetic, functional and phenotypic information about mouse genes and gene products. MGD works within the broader bioinformatics community to define referential and semantic standards to facilitate data exchange between resources including the incorporation of information from the biomedical literature. MGD is also a platform for computational assessment of integrated biological data with the goal of identifying candidate genes associated with complex phenotypes. MGD is web accessible at http://www.informatics.jax.org. Recent improvements in MGD described here include the incorporation of an interactive genome browser, the enhancement of phenotype resources and the further development of functional annotation resources.

Animals↗

High-resolution BAC-based map of the central portion of mouse chromosome 5.

The current strategy for sequencing the mouse genome involves the combination of a whole-genome shotgun approach with clone-based sequencing. High-resolution physical maps will provide a foundation for assembling contiguous segments of sequence. We have established a bacterial artificial chromosome (BAC)-based map of a 5-Mb region on mouse Chromosome 5, encompassing three gene families: receptor tyrosine kinases (PdgfraKit-Kdr), nonreceptor protein-tyrosine type kinases (Tec-Txk), and type-A receptors for the neurotransmitter GABA (Gabra2, Gabrb1, Gabrg1, and Gabra4). The construction of a BAC contig was initiated by hybridization screening the C57BL/6J (RPCI-23) BAC library, using known genes and sequence tagged sites (STSs). Additional overlapping clones were identified by searching the database of available restriction fingerprints for the RPCI-23 and RPCI-24 libraries. This effort resulted in the selection of >600 BAC clones, 251 kb of BAC-end sequences, and the placement of 40 known and/or predicted genes within this 5-Mb region. We use this high-resolution map to illustrate the integration of the BAC fingerprint map with a radiation-hybrid map via assembled expressed sequence tags (ESTs). From annotation of three representative BAC clones we demonstrate that up to 98% of the draft sequence for each contig could be ordered and oriented using known genes, BAC ends, consensus sequences for transcript assemblies, and comparisons with orthologous human sequence. For functional studies, annotation of sequence fragments as they are assembled into 50-200-kb stretches will be remarkably valuable.

Animals↗

Tools for integrated sequence-structure analysis with UCSF Chimera.

BACKGROUND: Comparing related structures and viewing the structures in the context of sequence alignments are important tasks in protein structure-function research. While many programs exist for individual aspects of such work, there is a need for interactive visualization tools that: (a) provide a deep integration of sequence and structure, far beyond mapping where a sequence region falls in the structure and vice versa; (b) facilitate changing data of one type based on the other (for example, using only sequence-conserved residues to match structures, or adjusting a sequence alignment based on spatial fit); (c) can be used with a researcher's own data, including arbitrary sequence alignments and annotations, closely or distantly related sets of proteins, etc.; and (d) interoperate with each other and with a full complement of molecular graphics features. We describe enhancements to UCSF Chimera to achieve these goals. RESULTS: The molecular graphics program UCSF Chimera includes a suite of tools for interactive analyses of sequences and structures. Structures automatically associate with sequences in imported alignments, allowing many kinds of crosstalk. A novel method is provided to superimpose structures in the absence of a pre-existing sequence alignment. The method uses both sequence and secondary structure, and can match even structures with very low sequence identity. Another tool constructs structure-based sequence alignments from superpositions of two or more proteins. Chimera is designed to be extensible, and mechanisms for incorporating user-specific data without Chimera code development are also provided. CONCLUSION: The tools described here apply to many problems involving comparison and analysis of protein structures and their sequences. Chimera includes complete documentation and is intended for use by a wide range of scientists, not just those in the computational disciplines. UCSF Chimera is free for non-commercial use and is available for Microsoft Windows, Apple Mac OS X, Linux, and other platforms from http://www.cgl.ucsf.edu/chimera.

Computer Graphics↗

Exploring trafficking GTPase function by mRNA expression profiling: use of the SymAtlas web-application and the Membrome datasets.

Despite complete sequencing of the human and mouse genomes, functional annotation of novel gene function still remains a major challenge in mammalian biology. Emerging strategies to help elucidate unknown gene function include the analysis of tissue-specific patterns of mRNA expression. A recent study investigated the steady-state mRNA expression profiling of the vast majority of protein-encoding human and mouse genes across a panel of 79 human and 61 mouse nonredundant tissues. The microarray data from this study constitutes the Genomics Institute of Novartis Foundation (GNF) Human and Mouse Gene Atlases and is publicly available for exploration through the SymAtlas web-application (http://symatlas.gnf.org/). We have recently reported the use of these data and hierarchical clustering algorithms to generate a global overview of the distribution of Rabs, SNAREs, and coat machinery components, as well as their respective adaptors, effectors, and regulators. This systems biology approach led us to propose Rab-centric protein activity hubs as a framework for an integrated coding system, the membrome network, which orchestrates the dynamics of specialized membrane architecture of differentiated cells. Here, we describe the use of the SymAtlas web-application and the Membrome datasets to help explore trafficking GTPase function. The human and mouse membrome datasets are available through the Membrome homepage (http://www.membrome.org/) and correspond to subsets of the SymAtlas content restricted to known membrane trafficking components. Considering the fragmentary nature of the current reductionist approaches in elucidating trafficking component functions, the membrome datasets provide a more focused systems biology perspective that not only complements our current understanding of transport in complex tissues but also provides an integrated perspective of Rab activity in controlling membrane architecture.

Animals↗

Comprehensive proteomic analysis of Shigella flexneri 2a membrane proteins.

Shigella flexneri is the causative agent of most shigellosis cases in developing countries. We used different proteolytic enzymes to selectively shave the protruding proteins on the surface of purified bacterial membrane sheets or vesicles, and recovered peptides were subsequently identified using 2-D LC-MS/MS. As a result, a total of 666 proteins were unambiguously assigned, including 159 integral membrane proteins, 35 outer membrane proteins and 114 proteins previously annotated as hypothetical. The former had an average grand average hydrophobicity score of 0.362 and were predicted to separate within a pH range of 4.1-10.6 with molecular mass 8-148 kDa, which represents the largest validated set of integral membrane proteins in this organism to date. A functional classification revealed that a large proportion of the identified proteins were involved in cell envelope biogenesis and energy production and conversion. For the first time, this work provides a global view of the S. flexneri 2a membrane subproteome.

Bacterial Outer Membrane Proteins↗

PRINTS and PRINTS-S shed light on protein ancestry.

The PRINTS database houses a collection of protein fingerprints. These may be used to make family and tentative functional assignments for uncharacterised sequences. The September 2001 release (version 32.0) includes 1600 fingerprints, encoding approximately 10 000 motifs, covering a range of globular and membrane proteins, modular polypeptides and so on. In addition to its continued steady growth, we report here its use as a source of annotation in the InterPro resource, and the use of its relational cousin, PRINTS-S, to model relationships between families, including those beyond the reach of conventional sequence analysis approaches. The database is accessible for BLAST, fingerprint and text searches at http://www.bioinf.man.ac.uk/dbbrowser/PRINTS/.

Amino Acid Motifs↗

Leiomyoma and myometrial gene expression profiles and their responses to gonadotropin-releasing hormone analog therapy.

Gene microarray was used to characterize the molecular environment of leiomyoma and matched myometrium during growth and in response to GnRH analog (GnRHa) therapy as well as GnRHa direct action on primary cultures of leiomyoma and myometrial smooth muscle cells (LSMC and MSMC). Unsupervised and supervised analysis of gene expression values and statistical analysis in R programming with a false discovery rate of P < or = 0.02 resulted in identification of 153 and 122 differentially expressed genes in leiomyoma and myometrium in untreated and GnRHa-treated cohorts, respectively. The expression of 170 and 164 genes was affected by GnRHa therapy in these tissues compared with their respective untreated group. GnRHa (0.1 microm), in a time-dependent manner (2, 6, and 12 h), targeted the expression of 281 genes (P < or = 0.005) in LSMC and MSMC, 48 of which genes were found in common with GnRHa-treated tissues. Functional annotations assigned these genes as key regulators of processes involving transcription, translational, signal transduction, structural activities, and apoptosis. We validated the expression of IL-11, early growth response 3, TGF-beta-induced factor, TGF-beta-inducible early gene response, CITED2 (cAMP response element binding protein-binding protein/p300-interacting transactivator with ED-rich tail), Nur77, growth arrest-specific 1, p27, p57, and G protein-coupled receptor kinase 5, representing cytokine, common transcription factors, cell cycle regulators, and signal transduction, at tissue levels and in LSMC and MSMC in response to GnRHa time-dependent action using real-time PCR, Western blotting, and immunohistochemistry. In conclusion, using different, complementary approaches, we characterized leiomyoma and myometrium molecular fingerprints and identified several previously unrecognized genes as targets of GnRHa action, implying that local expression and activation of these genes may represent features differentiating leiomyoma and myometrial environments during growth and GnRHa-induced regression.

Active Transport, Cell Nucleus↗

The domain architecture of large guanine nucleotide exchange factors for the small GTP-binding protein Arf.

BACKGROUND: Small G proteins, which are essential regulators of multiple cellular functions, are activated by guanine nucleotide exchange factors (GEFs) that stimulate the exchange of the tightly bound GDP nucleotide by GTP. The catalytic domain responsible for nucleotide exchange is in general associated with non-catalytic domains that define the spatio-temporal conditions of activation. In the case of small G proteins of the Arf subfamily, which are major regulators of membrane trafficking, GEFs form a heterogeneous family whose only common characteristic is the well-characterized Sec7 catalytic domain. In contrast, the function of non-catalytic domains and how they regulate/cooperate with the catalytic domain is essentially unknown. RESULTS: Based on Sec7-containing sequences from fully-annotated eukaryotic genomes, including our annotation of these sequences from Paramecium, we have investigated the domain architecture of large ArfGEFs of the BIG and GBF subfamilies, which are involved in Golgi traffic. Multiple sequence alignments combined with the analysis of predicted secondary structures, non-structured regions and splicing patterns, identifies five novel non-catalytic structural domains which are common to both subfamilies, revealing that they share a conserved modular organization. We also report a novel ArfGEF subfamily with a domain organization so far unique to alveolates, which we name TBS (TBC-Sec7). CONCLUSION: Our analysis unifies the BIG and GBF subfamilies into a higher order subfamily, which, together with their being the only subfamilies common to all eukaryotes, suggests that they descend from a common ancestor from which species-specific ArfGEFs have subsequently evolved. Our identification of a conserved modular architecture provides a background for future functional investigation of non-catalytic domains.

ADP-Ribosylation Factors↗

Getting to synaptic complexes through systems biology.

Large numbers of synaptic components have been identified, but the effect so far on our understanding of synaptic function is limited. Now, network maps and annotated functions of individual components have been used in a systems biology approach to analyzing the function of NMDA receptor complexes at synapses, identifying biologically relevant modular networks within the complex.

Animals↗

The functional genomic distribution of protein divergence in two animal phyla: coevolution, genomic conflict, and constraint.

We compare the functional spectrum of protein evolution in two separate animal lineages with respect to two hypotheses: (1) rates of divergence are distributed similarly among functional classes within both lineages, indicating that selective pressure on the proteome is largely independent of organismic-level biological requirements; and (2) rates of divergence are distributed differently among functional classes within each lineage, indicating species-specific selective regimes impact genome-wide substitutional patterns. Integrating comparative genome sequence with data from tissue-specific expressed-sequence-tag (EST) libraries and detailed database annotations, we find a functional genomic signature of rapid evolution and selective constraint shared between mammalian and nematode lineages despite their extensive morphological and ecological differences and distant common ancestry. In both phyla, we find evidence of accelerated evolution among components of molecular systems involved in coevolutionary change. In mammals, lineage-specific fast evolving genes include those involved in reproduction, immunity, and possibly, maternal-fetal conflict. Likelihood ratio tests provide evidence for positive selection in these rapidly evolving functional categories in mammals. In contrast, slowly evolving genes, in terms of amino acid or insertion/deletion (indel) change, in both phyla are involved in core molecular processes such as transcription, translation, and protein transport. Thus, strong purifying selection appears to act on the same core cellular processes in both mammalian and nematode lineages, whereas positive and/or relaxed selection acts on different biological processes in each lineage.

Amino Acid Substitution↗