Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

An infrastructure for comparative genomics to functionally characterize genes and proteins.

Current genome projects are resulting in a flood of sequence data. The interpretation of these sequences is lagging, and optimized data analysis strategies need to be developed. Much can be learned from comparing different genomes, as genomes of distant organisms may still encode proteins with high sequence similarity. The order of genes (co linearity) in genomes may also be conserved to some extend. We have employed both these observations to create a multi-functional, computational analysis system (genomeSCOUT) which allows for rapid identification and functional characterization of genes and proteins through genome comparison. With a number of independent algorithms, information about different levels of protein homology (concerning e.g. paralogs, orthologs and clusters of orthologous groups, COGs) and gene order is collected and stored in several value added databases. These databases are then used for interactive comparison of genomes and subsequent analysis. The application is based on the well established data integration system SRS. This ensures (1) fast handling of large genomic data sets, (2) straightforward access to a multitude of biological databases, (3) unique linking functions between these databases, (4) highly efficient collection of information on genes and proteins, and 5. fully integrated and user friendly graphical representations of search results. This application can be used for projects as diverse as the correct annotation of genomes, the optimization of (micro) organisms for industrial production, or the identification of drug targets.

Computational Biology↗

Interaction networks in yeast define and enumerate the signaling steps of the vertebrate aryl hydrocarbon receptor.

The aryl hydrocarbon receptor (AHR) is a vertebrate protein that mediates the toxic and adaptive responses to dioxins and related environmental pollutants. In an effort to better understand the details of this signal transduction pathway, we employed the yeast S. cerevisiae as a model system. Through the use of arrayed yeast strains harboring ordered deletions of open reading frames, we determined that 54 out of the 4,507 yeast genes examined significantly influence AHR signal transduction. In an effort to describe the relationship between these modifying genes, we constructed a network map based upon their known protein and genetic interactions. Monte Carlo simulations demonstrated that this network represented a description of AHR signaling that was distinct from those generated by random chance. The network map was then explored with a number of computational and experimental annotations. These analyses revealed that the AHR signaling pathway is defined by at least five distinct signaling steps that are regulated by functional modules of interacting modifiers. These modules can be described as mediating receptor folding, nuclear translocation, transcriptional activation, receptor level, and a previously undescribed nuclear step related to the receptor's Per-Arnt-Sim domain.

Active Transport, Cell Nucleus↗

The paradox of functional heterochromatin.

Although heterochromatin has been studied for 80 years, its genetic function and molecular organization have remained elusive. In almost all organisms, heterochromatin has been regarded as genetically inactive chromosome regions. However, from genetic and genomic studies in Drosophila melanogaster and other organisms including humans, it is now clear that transcriptionally active domains are present within constitutive heterochromatin. These domains contain essential coding genes whose expression during development ensures the formation of the proper biochemical and morphological phenotypes, together with several gene models defined by genome annotation whose functions still need to be determined.

Animals↗

Orphan transcripts in Arabidopsis thaliana: identification of several hundred previously unrecognized genes.

Expressed sequence tags (ESTs) represent a huge resource for the discovery of previously unknown genetic information and functional genome assignment. In this study we screened a collection of 178 292 ESTs from Arabidopsis thaliana by testing them against previously annotated genes of the Arabidopsis genome. We identified several hundreds of new transcripts that match the Arabidopsis genome at so far unassigned loci. The transcriptional activity of these loci was independently confirmed by comparison with the Salk Whole Genome Array Data. To a large extent, the newly identified transcriptionally active genomic regions do not encode 'classic' proteins, but instead generate non-coding RNAs and/or small peptide-coding RNAs of presently unknown biological function. More than 560 transcripts identified in this study are not represented by the Affymetrix GeneChip arrays currently widely used for expression profiling in A. thaliana. Our data strongly support the hypothesis that numerous previously unknown genes exist in the Arabidopsis genome.

Arabidopsis↗

TreeDomViewer: a tool for the visualization of phylogeny and protein domain structure.

Phylogenetic analysis and examination of protein domains allow accurate genome annotation and are invaluable to study proteins and protein complex evolution. However, two sequences can be homologous without sharing statistically significant amino acid or nucleotide identity, presenting a challenging bioinformatics problem. We present TreeDomViewer, a visualization tool available as a web-based interface that combines phylogenetic tree description, multiple sequence alignment and InterProScan data of sequences and generates a phylogenetic tree projecting the corresponding protein domain information onto the multiple sequence alignment. Thereby it makes use of existing domain prediction tools such as InterProScan. TreeDomViewer adopts an evolutionary perspective on how domain structure of two or more sequences can be aligned and compared, to subsequently infer the function of an unknown homolog. This provides insight into the function assignment of, in terms of amino acid substitution, very divergent but yet closely related family members. Our tool produces an interactive scalar vector graphics image that provides orthological relationship and domain content of proteins of interest at one glance. In addition, PDF, JPEG or PNG formatted output is also provided. These features make TreeDomViewer a valuable addition to the annotation pipeline of unknown genes or gene products. TreeDomViewer is available at http://www.bioinformatics.nl/tools/treedom/.

Computer Graphics↗

Functional characterization of front-end desaturases from trypanosomatids depicts the first polyunsaturated fatty acid biosynthetic pathway from a parasitic protozoan.

A survey of the three kinetoplastid genome projects revealed the presence of three putative front-end desaturase genes in Leishmania major, one in Trypanosoma brucei and two highly identical ones (98%) in T. cruzi. The encoded gene products were tentatively annotated as Delta8, Delta5 and Delta6 desaturases for L. major, and Delta6 desaturase for both trypanosomes. After phylogenetic and structural analysis of the deduced proteins, we predicted that the putative Delta6 desaturases could have Delta4 desaturase activity, based mainly on the conserved HX(3)HH motif for the second histidine box, when compared with Delta4 desaturases from Thraustochytrium, Euglena gracilis and the microalga, Pavlova lutheri, which are more than 30% identical to the trypanosomatid enzymes. After cloning and expression in Saccharomyces cerevisiae, it was possible to functionally characterize each of the front-end desaturases present in L. major and T. brucei. Our prediction about the presence of Delta4 desaturase activity in the three kinetoplastids was corroborated. In the same way, Delta5 desaturase activity was confirmed to be present in L. major. Interestingly, the putative Delta8 desaturase turned out to be a functional Delta6 desaturase, being 35% and 31% identical to Rhizopus oryzae and Pythium irregulareDelta6 desaturases, respectively. Our results indicate that no conclusive predictions can be made about the function of this class of enzymes merely on the basis of sequence homology. Moreover, they indicate that a complete pathway for very-long-chain polyunsaturated fatty acid biosynthesis is functional in L. major using Delta6, Delta5 and Delta4 desaturases. In trypanosomes, only Delta4 desaturases are present. The putative algal origin of the pathway in kinetoplastids is discussed.

Amino Acid Sequence↗

Earliest changes in the left ventricular transcriptome postmyocardial infarction.

We report a genome-wide survey of early responses of the mouse heart transcriptome to acute myocardial infarction (AMI). For three regions of the left ventricle (LV), namely, ischemic/infarcted tissue (IF), the surviving LV free wall (FW), and the interventricular septum (IVS), 36,899 transcripts were assayed at six time points from 15 min to 48 h post-AMI in both AMI and sham surgery mice. For each transcript, temporal expression patterns were systematically compared between AMI and sham groups, which identified 515 AMI-responsive genes in IF tissue, 35 in the FW, 7 in the IVS, with three genes induced in all three regions. Using the literature, we assigned functional annotations to all 519 nonredundant AMI-induced genes and present two testable models for central signaling pathways induced early post-AMI. First, the early induction of 15 genes involved in assembly and activation of the activator protein-1 (AP-1) family of transcription factors implicates AP-1 as a dominant regulator of earliest post-ischemic molecular events. Second, dramatic increases in transcripts for arginase 1 (ARG1), the enzymes of polyamine biosynthesis, and protein inhibitor of nitric oxide synthase (NOS) activity indicate that NO production may be regulated, in part, by inhibition of NOS and coordinate depletion of the NOS substrate, L: -arginine. ARG1: was the single-most highly induced transcript in the database (121-fold in IF region) and its induction in heart has not been previously reported.

Acute Disease↗

A deep metagenomic atlas of Qinghai-Xizang Plateau lakes reveals their microbial diversity and salinity adaptation mechanisms.

The Qinghai-Xizang Plateau (QXP), harboring the planet's highest density of plateau lakes, offers an exceptional biogeographic environment for studying extremophilic microbial communities and their adaptation to salinity. Through deep metagenomic sequencing, we construct the Qinghai-Xizang Lake Sediment Genome (QXLSG) catalog, a high-resolution genomic catalog comprising 5,866 metagenome-assembled genomes (MAGs), 58.16 million non-redundant protein encoding genes, and 19,008 biosynthetic gene clusters. Notably, 80.78% of the 2,742 species-level MAGs represent undescribed taxa, significantly expanding the known microbial diversity. Salinity emerges as the primary environmental factor influencing microbial community. Functional annotation highlights that the "salt-out" strategy, particularly the uptake of glycine betaine, is the main mechanism for salinity tolerance. This strategy is prevalent in both hypersaline lake communities and the dominant microbial phyla. Overall, this study provides a crucial genetic resource for future bioprospecting and deepens our understanding of the fundamental mechanisms of microbial adaptation to extreme saline environments.

Lakes↗

Proteins with class alpha/beta fold have high-level participation in fusion events.

Now that complete genome sequences are available for a variety of organisms, the elucidation of potential gene products function is a central goal in the post-genome era. Domain fusion analysis has been proposed recently to infer the functional association of the component proteins. Here, we took a new approach to the analysis of the structural features of the proteins involved in fusion events. An exhaustive survey of fusion events within 30 completely sequenced genomes and subsequent structure annotations to the component proteins at a SCOP superfamily level with hidden Markov models was carried out. A domain fusion map was then constructed. The results revealed that proteins with the class alpha/beta fold are frequently involved in fusion events, around 86% of the total 676 assigned single-domain fusion pairs including at least one component protein belonging to the alpha/beta fold class. Moreover, the domain fusion map in our work may offer an attractive framework for designing chimeric enzymes following Nature's lead, and may give useful hints for exploring the evolutionary history of proteins. (c) 2002 Elsevier Science Ltd.

Archaeal Proteins↗

Drug target ontology to classify and integrate drug discovery data.

BACKGROUND: One of the most successful approaches to develop new small molecule therapeutics has been to start from a validated druggable protein target. However, only a small subset of potentially druggable targets has attracted significant research and development resources. The Illuminating the Druggable Genome (IDG) project develops resources to catalyze the development of likely targetable, yet currently understudied prospective drug targets. A central component of the IDG program is a comprehensive knowledge resource of the druggable genome. RESULTS: As part of that effort, we have developed a framework to integrate, navigate, and analyze drug discovery data based on formalized and standardized classifications and annotations of druggable protein targets, the Drug Target Ontology (DTO). DTO was constructed by extensive curation and consolidation of various resources. DTO classifies the four major drug target protein families, GPCRs, kinases, ion channels and nuclear receptors, based on phylogenecity, function, target development level, disease association, tissue expression, chemical ligand and substrate characteristics, and target-family specific characteristics. The formal ontology was built using a new software tool to auto-generate most axioms from a database while supporting manual knowledge acquisition. A modular, hierarchical implementation facilitate ontology development and maintenance and makes use of various external ontologies, thus integrating the DTO into the ecosystem of biomedical ontologies. As a formal OWL-DL ontology, DTO contains asserted and inferred axioms. Modeling data from the Library of Integrated Network-based Cellular Signatures (LINCS) program illustrates the potential of DTO for contextual data integration and nuanced definition of important drug target characteristics. DTO has been implemented in the IDG user interface Portal, Pharos and the TIN-X explorer of protein target disease relationships. CONCLUSIONS: DTO was built based on the need for a formal semantic model for druggable targets including various related information such as protein, gene, protein domain, protein structure, binding site, small molecule drug, mechanism of action, protein tissue localization, disease association, and many other types of information. DTO will further facilitate the otherwise challenging integration and formal linking to biological assays, phenotypes, disease models, drug poly-pharmacology, binding kinetics and many other processes, functions and qualities that are at the core of drug discovery. The first version of DTO is publically available via the website http://drugtargetontology.org/ , Github ( http://github.com/DrugTargetOntology/DTO ), and the NCBO Bioportal ( http://bioportal.bioontology.org/ontologies/DTO ). The long-term goal of DTO is to provide such an integrative framework and to populate the ontology with this information as a community resource.

Biological Ontologies↗

Protein expression, crystallization and preliminary X-ray crystallographic studies on HSCARG from Homo sapiens.

Human HSCARG has been annotated as a possible cancer related protein. Amino acid homology, although at a low percentage, suggested that HSCARG contains NmrA domain and might be a member of short chain dehydrogenase reductase superfamily. In order to investigate its structure and function, HSCARG gene has been successfully expressed and purified in E. coli. HSCARG was crystallized and diffracted to a resolution of 2.4 A on Mar225 CCD Detector at SER-CAT 22BM synchrotron source. The crystals belong to F23 space group, with unit cell parameters a=b=c=223.30A, alpha=beta=gamma=90 degrees . There are two molecules per asymmetry unit.

Amino Acid Sequence↗

B cell pathways implicate shared genetic architecture between schizophrenia and immune-mediated diseases.

BACKGROUND: Schizophrenia and immune-mediated diseases are globally prevalent and highly heritable conditions that frequently co-occur, posing major public health burdens. However, their shared genetic architecture remains poorly understood. METHODS: We applied the bivariate causal mixture model (MiXeR) to investigate the polygenic overlap between schizophrenia and eight common immune-mediated diseases, using genome-wide association study summary statistics comprising 2,489 to 67,323 cases and 9,066 to 497,622 controls. Shared loci were identified through conditional/conjunctional false discovery rate (cond/conjFDR), local genetic correlation (LAVA), and colocalization analyses. Subsequently, gene mapping, functional annotation, expression-trait association, and drug-gene interaction analyses were performed to explore shared genes and enriched pathways, and genetic risk scores (GRS) from the UK Biobank were used to validate the findings. RESULTS: MiXeR estimated substantial polygenic overlap between schizophrenia and immune-mediated diseases, and conjFDR identified 133 shared loci, with eight prioritized through local genetic correlation and colocalization signals. These eight loci were mapped to 85 protein-coding genes enriched in pathways essential for B cell function. Among them, S-PrediXcan analyses identified 14 genes whose expression in brain tissues or blood was associated with both diseases. These genes also interact with immunomodulatory or antihypertensive drugs. Additionally, 11 of the 14 genes were linked to innate immunity and/or cognitive traits. Using UK Biobank data, we further confirmed that overall, shared gene, and B cell activation and receptor signaling pathway–specific genetic risk for schizophrenia is associated with immune-mediated disease susceptibility. CONCLUSIONS: These findings underscore the shared genetic architecture of schizophrenia and immune-mediated diseases, advancing insights at the interface of psychiatric genetics and immunology.

Schizophrenia↗

Gene and pathway analysis of genome-wide genetic associations of bladder cancer.

BACKGROUND: Although genetic variants associated with bladder cancer (BCa) risk have been identified through hypothesis-driven and genome-wide association studies, a systematic understanding of BCa genetic susceptibility at the gene and pathway levels remains to be achieved. MATERIALS AND METHODS: In this 2-stage functional genomics study, we used 5 independent tools for genome-wide gene mapping and ranking based on BCa genome-wide association studies summary statistics, followed by a meta-analysis of gene-level significance p values, to obtain a consensus gene ranking in terms of association with BCa. Subsequently, we performed preranked gene-set enrichment analysis to identify the functional pathways involved in BCa genetic susceptibility. Joint analysis with gene-set enrichment analysis, based on somatic alteration frequency, was performed to explore the pathway-level relationships between genetic susceptibility and somatic alterations in BCa. RESULTS: Other than the well-known BCa genes (such as FGFR3, MYC, TERT, CCNE1, and TP63), we additionally prioritized a set of novel genes likely to be genetically implicated in BCa development, including SETD2, a possible tumor suppressor gene involved in chromatin remodeling. We further demonstrated convergence between genetic associations and somatic alterations at both the gene (eg, FGFR3 and TERT) and pathway levels (eg, cell cycle and chromatin modification), as well as functional ontologies specifically implicated in germline predisposition to BCa (eg, CD8/TCR signaling, immune checkpoints, and cytokine signaling). CONCLUSIONS: We identified several novel genes associated with BCa and demonstrated that genetic variants contribute to the development of BCa by affecting antitumor immunity, response to toxic exposure, and RNA and protein homeostasis and synergizing with somatic alterations in various cancer-related pathways.

Bladder cancer↗

The chromosome-level genome assembly of Prunus cerasifera 'Atropurpurea'.

Prunus cerasifera 'Atropurpurea' (Purpleleaf Plum), known for its unique purple-red foliage, is an important ornamental plant that enhances the aesthetic value of urban greening. To explore the molecular mechanisms underlying leaf color changes, this study assembled the Purpleleaf Plum genome, providing new insights for related research. We used HiFi sequencing data to assemble its genome. After chromosome anchoring, the final genome size was 244.89 Mb, with a contig N50 of 26.60 Mb, and approximately 97.10% of sequences were anchored to 8 chromosomes. Genome annotation identified 28,231 protein-coding genes, with LTR transposons comprising 27.93% of the genome. BUSCO assessment revealed a genome completeness of 98.9%. Telomeric repeat analysis identified 14 telomeres, with six chromosomes capped by double telomeres and two chromosomes containing a single telomere. This high-quality Purpleleaf Plum genome provides a solid foundation for future gene function analysis, cultivar improvement, and genetic research, offering valuable resources for related fields.

Genome, Plant↗

Genome prediction of PhoB regulated promoters in Sinorhizobium meliloti and twelve proteobacteria.

In proteobacteria, genes whose expression is modulated in response to the external concentration of inorganic phosphate are often regulated by the PhoB protein which binds to a conserved motif (Pho box) within their promoter regions. Using a position weight matrix algorithm derived from known Pho box sequences, we identified 96 putative Pho regulon members whose promoter regions contained one or more Pho boxs in the Sinorhizobium meliloti genome. Expression of these genes was examined through assays of reporter gene fusions and through comparison with published microarray data. Of 96 genes, 31 were induced and 3 were repressed by Pi starvation in a PhoB dependent manner. Novel Pho regulon members included several genes of unknown function. Comparative analysis across 12 proteobacterial genomes revealed highly conserved Pho regulon members including genes involved in Pi metabolism (pstS, phnC and ppdK). Genes with no obvious association with Pi metabolism were predicted to be Pho regulon members in S.meliloti and multiple organisms. These included smc01605 and smc04317 which are annotated as substrate binding proteins of iron transporters and katA encoding catalase. This data suggests that the Pho regulon overlaps and interacts with several other control circuits, such as the oxidative stress response and iron homeostasis.

Algorithms↗

A new member of plant CS-lyases. A cystine lyase from Arabidopsis thaliana.

Cystine lyases catalyze the breakdown of l-cystine to thiocysteine, pyruvate, and ammonia. Until now there are no reports of the identification of a plant cystine lyase at a molecular level, and it is not clear what biological role this class of enzymes have in plants. A cystine lyase was isolated from Brassica oleracea (L.), and partial amino acid sequencing allowed the corresponding full-length cDNA (BOCL3) to be cloned. The deduced amino acid sequence of BOCL3 showed highest homology to the deduced amino acid sequences of several Arabidopsis thaliana genes annotated as tyrosine aminotransferase-like, including a coronatine, jasmonic acid, and salt stress-inducible gene, CORI3 (78.8% identity), and the unidentified rooty/superroot1 gene (44.8% identity). A full-length expressed sequence tag clone of CORI3 was obtained and recombinant CORI3 was synthesized in Escherichia coli. Isolated recombinant CORI3 catalyzed a cystine lyase reaction, but no aminotransferase reactions. The present study identifies, for the first time, a cystine lyase from plants at a molecular level and redefines the functional assignment of the only functionally identified member of a group of A. thaliana genes annotated as tyrosine aminotransferase-like.

Amino Acid Sequence↗

Crystallization and preliminary X-ray diffraction data of Mycobacterium tuberculosis FbpC1 (Rv3803c).

The heterotrimeric antigen 85 complex (Ag85) is a major component of the cell wall of Mycobacterium tuberculosis and consists of three abundantly secreted proteins (FbpA, FbpB and FbpC2). These play key roles in the pathogenesis of tuberculosis and in maintaining cell-wall integrity. A homologue of the Ag85 subunits ( approximately 40% identity) was recently annotated in the M. tuberculosis genome as FbpC1. Unlike the Ag85-complex components, FbpC1 lacks mycolyltransferase activity and its function remains to be established. In order to aid functional characterization, FbpC1 has been crystallized. At room temperature, tetragonal crystals of FbpC1 were obtained belonging to space group P4(1)2(1)2 (unit-cell parameters a = b = 109.9, c = 61.8 A), yet when frozen the crystals underwent a phase transition to orthorhombic symmetry, space group P2(1)2(1)2(1) (a = 59.9, b = 108.9, c = 109.9 A). Diffraction data complete to 1.7 A resolution were recorded at 100 K at the synchrotron.

Antigens, Bacterial↗

Protein complex compositions predicted by structural similarity.

Proteins function through interactions with other molecules. Thus, the network of physical interactions among proteins is of great interest to both experimental and computational biologists. Here we present structure-based predictions of 3387 binary and 1234 higher order protein complexes in Saccharomyces cerevisiae involving 924 and 195 proteins, respectively. To generate candidate complexes, comparative models of individual proteins were built and combined together using complexes of known structure as templates. These candidate complexes were then assessed using a statistical potential, derived from binary domain interfaces in PIBASE (http://salilab.org/pibase). The statistical potential discriminated a benchmark set of 100 interface structures from a set of sequence-randomized negative examples with a false positive rate of 3% and a true positive rate of 97%. Moreover, the predicted complexes were also filtered using functional annotation and sub-cellular localization data. The ability of the method to select the correct binding mode among alternates is demonstrated for three camelid VHH domain-porcine alpha-amylase interactions. We also highlight the prediction of co-complexed domain superfamilies that are not present in template complexes. Through integration with MODBASE, the application of the method to proteomes that are less well characterized than that of S.cerevisiae will contribute to expansion of the structural and functional coverage of protein interaction space. The predicted complexes are deposited in MODBASE (http://salilab.org/modbase).

Algorithms↗