Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

Mutations in Mycobacterium tuberculosis Rv0444c, the gene encoding anti-SigK, explain high level expression of MPB70 and MPB83 in Mycobacterium bovis.

It has recently been advanced that Mycobacterium tuberculosis sigma factor K (SigK) positively regulates expression of the antigenic proteins MPB70 and MPB83. As expression of these proteins differs between M. tuberculosis (low) and Mycobacterium bovis (high), this study set out to determine whether M. bovis lacks a functional SigK repressor (anti-SigK). By comparing genes near sigK in M. tuberculosis H37Rv and M. bovis AF2122/97, we observed that Rv0444c, annotated as unknown function, had variable sequence in M. bovis. Analysis of in vitro mpt70/mpt83 expression and Rv0444c sequencing across M. tuberculosis complex (MTC) members revealed that high-level expression was associated with a mutated Rv0444c. Complementation of M. bovis bacillus Calmette-Guerin Russia, a high producer of MPB70/MPB83, with wild-type Rv0444c resulted in a significant decrease in mpb70/mpb83 expression. Conversely, a M. tuberculosis H37Rv mutant which expressed sigK but not Rv0444c manifested the M. bovis phenotype of high-level MPB70/MPB83 expression. Further support that Rv0444c encodes the anti-SigK was obtained by yeast two-hybrid studies, where the N-terminal region of Rv0444c-encoded protein interacted with SigK. Together these findings indicate that Rv0444c encodes the regulator of SigK (RskA) and mutations in this gene explain high-level MPT70/MPT83 expression by certain MTC members.

Antigens, Bacterial↗

Evolution of protein superfamilies and bacterial genome size.

We present the structural annotation of 56 different bacterial species based on the assignment of genes to 816 evolutionary superfamilies in the CATH domain structure database. These assignments have enabled us to analyse the recurrence of specific superfamilies within and across the genomes. We have selected the superfamilies that have a very broad representation and therefore appear to be universally distributed in a significant number of bacterial lineages. Occurrence profiles of these universally distributed superfamilies are compared with genome size in order to estimate the correlation between superfamily duplication and the increase in proteome size. This distinguishes between those size-dependent superfamilies where frequency of occurrence is highly correlated with increase in genome size, and size-independent superfamilies where no correlation is observed. Consideration of the size correlation and the ratio between the mean and the standard deviations for all the superfamily profiles allows more detailed subdivisions and classification of superfamilies. For example, within the size-independent superfamilies, we distinguished a group that are distributed evenly amongst all the genomes. Within the size-dependent superfamilies we differentiated two groups: linearly distributed and non-linearly distributed. Functional annotation using the COG database was performed for all superfamilies in each of these groups, and this revealed significant differences amongst the three sets of superfamilies. Evenly distributed, size-independent domains are shown to be involved primarily in protein translation and biosynthesis. For the size-dependent superfamilies, linearly distributed superfamilies are involved mainly in metabolism, and non-linearly distributed superfamily domains are involved principally in gene regulation.

Databases, Protein↗

Proteomic 2DE database for spot selection, automated annotation, and data analysis.

We present a software solution that enables faster and more accurate data analysis of 2DE/MALDI TOF MS data. The software supports data analysis through a number of automated data selection functions and advanced graphical tools. Once protein identities are determined using MALDI TOF MS, automated data retrieval from online databases provides biological information. The software, called 2DDB, reduces analysis time to a fraction without losing any quality compared to more manual data analysis. The database contains over 100,000 data entries, and selected parts can be reached at http://2ddb.org.

Animals↗

ATP-binding cassette protein E is involved in gene transcription and translation in Caenorhabditis elegans.

ATP-binding cassette protein E (ABCE) gene has been annotated as an RNase L inhibitor in eukaryotes. All eukaryotic species show the ubiquitous presence and high degree of conservation of ABCEs, however, RNase L is present only in mammals. This indicates that ABCEs may function not only as RNase L inhibitors, but also may have other functions that have yet to be determined. As an initial investigation into the novel functions of ABCE, we characterized the gene (Y39E4B.1) in Caenorhabditis elegans by a combination of data mining and functional assays. ABCE promoters drove GFP expressions in hypoderm, pharynx, vulvae, head, and tail neurons at all developmental stages. Three genes, rpl-4, nhr-91, and C07B5.3, were previously found to interact with ABCE. Our expression data showed overlapping expression patterns of ABCE and rpl-4 and nhr-91, but not C07B5.3. RNAi against ABCE resulted in embryonic lethality and slow growth. These data suggest that ABCE protein might be involved in the control of translation and transcription, work as shuttle protein between cytoplasm and nucleus, and possibly as a nucleocytoplasmic transporter. In addition, RNAi data suggest that ABCE and NHR-91 may function in vulvae development and molting pathways in C. elegans. Furthermore, our data suggest that ABCE, along with its interacting components, functions in a well-conserved pathway.

ATP-Binding Cassette Transporters↗

Dynamic Protein Structure Paradox: An Integrative Framework for Endpoint-Conditioned Evidentiary Sufficiency in Structure-to-Function Claims.

Accurate coordinates for a represented protein state do not, by themselves, establish activity or any other condition-specific function. This article defines the Dynamic Protein Structure Paradox (DPSP) as the apparent conflict between structural accuracy and functional underdetermination and develops it as an integrative evidentiary assessment framework rather than a new theory or paradigm. The underlying problem has been longstanding, since structural genomics, function annotation, allostery, and disorder research each established that fold does not determine function and that function does not determine fold. DPSP consolidates those results into one endpoint-conditioned rule. Once a measurable endpoint is defined, it assesses four coupled dimensions: relevant-state completeness, context completeness, ensemble or kinetic dependence, and chemical dependence. A rubric rates each dimension as adequate, uncertain, or missing, and a materiality test determines which gaps influence the stated decision. The outcome is one of three mutually exclusive modes of utilization: geometry-led, conditional, or function-measured. The deliverable is a concise evidence statement delineating what the structure supports, which decisive variable remains unmeasured, and what corroboration is necessary. DPSP complements, rather than replaces, existing structural, ensemble, and computational approaches. The framework remains unvalidated, its thresholds are provisional, and the studies necessary to confirm or refute it are specified.

Proteins↗

An on-line two-dimensional polyacrylamide gel electrophoresis protein database of adult Drosophila melanogaster.

An annotated two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) protein database of adult Drosophila melanogaster has been constructed, based on the protein patterns of heads, thoraces and abdomens of adult male and female Drosophila melanogaster. About 1200 major protein spots are catalogued. Common proteins, found in all body parts, as well as bodypart- and sex-specifically expressed proteins are reported. Of the major proteins, 91, or 7.5%, are differentially expressed in the two sexes or in different body parts, at least in part reflecting specific functional requirements. At the present time 43 proteins, or about 3.5% of the detected proteins, have been identified. These data can be accessed interactively from our World Wide Web (WWW) server through clickable inline gel images and hypertext links. Identified protein spots are cross-referenced, through hypertext links, to the SWISS-PROT annotated database of protein primary sequences and the Fly-Base database of Drosophila genomic data. Our reference gels can be used to gain immediate access to protein spot identify and to the pattern of differentially expressed proteins in Drosophila melanogaster. The work presented in this article ties together information from protein 2-D PAGE, molecular biology and genetics and offers a uniform way to access this large volume of data.

Abdomen↗

Functional convergence of regulatory regions provides vital insights into mammalian gliding adaptation.

Uncovering the key genetic basis of complex phenotypic convergence in distantly related species has been a long-standing focus in evolutionary biology and genetics, and the convergent evolution of gliding in mammals offers a valuable opportunity to address this question. Here, we investigated the genomic basis of convergent evolution of gliding in mammals by analyzing both protein-coding genes and conserved non-coding elements (CNEs). We first de novo assembled and annotated two chromosome-level genomes of gliding mammals, the red and white giant flying squirrel (Petaurista alborufus) and sugar gliders (Petaurus breviceps), and conducted comprehensive comparative genomic analysis combined with another gliding mammal, the Sunda flying lemur (Galeopterus variegatus) and 14 background species. We found that the convergent evolution of protein-coding genes provided relatively limited but functionally relevant evidence linked to gliding phenotypes. By contrast, we found that gliding-accelerated CNEs (GACNEs) cluster near functionally equivalent genes and frequently aggregate into highly diverged yet functionally convergent hotspot regions. Across the three gliding lineages, both GACNEs and hotspot GACNEs show strong convergence in their functional enrichment profiles, suggesting a broad genetic basis underlying the convergent gliding phenotype. Furthermore, we identified 72 core transcription factors underpinning the genetic basis of gliding convergence, including EMX2 and ZFHX3, potentially involved in multiple aspects of gliding adaptation. Our study highlights the role of functional convergence in regulatory regions as a key mechanism in mammalian gliding convergence, offering valuable insights and strategies for uncovering the genetic basis of complex convergent traits, thereby advancing understanding of the molecular basis of convergent traits.

Petaurista alborufus↗

Novel coding regions in four complete archaeal genomes.

In the process of analysing the four available complete archaeal genomes, we have noted that certain regions characterised as 'non-coding' exhibit significant sequence similarity to other protein sequences from Archaea and other species. Using established technology, we have identified a number of potential protein coding regions in these putative 'non-coding' regions. We have detected 524 such cases, of which 113 regions appear to code for proteins present in archaeal or other species, while the remaining 411 regions are mostly start/stop definition conflicts. Of the 113 protein coding regions, only 21 code for proteins with homologues of known function. The number of novel coding sequences identified herein amounts to 1. 5% of the total genome entries, while the conflicting cases represent an additional 5%. The observed differences between the four complete archaeal genomes seem to reflect disparate approaches to genome annotation. Genome sequence collections should be regularly checked to improve gene prediction by sequence similarity and greater effort is required to make gene definitions consistent across related species.

Archaeal Proteins↗

ChloroplastDB: the Chloroplast Genome Database.

The Chloroplast Genome Database (ChloroplastDB) is an interactive, web-based database for fully sequenced plastid genomes, containing genomic, protein, DNA and RNA sequences, gene locations, RNA-editing sites, putative protein families and alignments (http://chloroplast.cbio.psu.edu/). With recent technical advances, the rate of generating new organelle genomes has increased dramatically. However, the established ontology for chloroplast genes and gene features has not been uniformly applied to all chloroplast genomes available in the sequence databases. For example, annotations for some published genome sequences have not evolved with gene naming conventions. ChloroplastDB provides unified annotations, gene name search, BLAST and download functions for chloroplast encoded genes and genomic sequences. A user can retrieve all orthologous sequences with one search regardless of gene names in GenBank. This feature alone greatly facilitates comparative research on sequence evolution including changes in gene content, codon usage, gene structure and post-transcriptional modifications such as RNA editing. Orthologous protein sets are classified by TribeMCL and each set is assigned a standard gene name. Over the next few years, as the number of sequenced chloroplast genomes increases rapidly, the tools available in ChloroplastDB will allow researchers to easily identify and compile target data for comparative analysis of chloroplast genes and genomes.

Chloroplasts↗

Genome-scale functional profiling of the mammalian AP-1 signaling pathway.

Large-scale functional genomics approaches are fundamental to the characterization of mammalian transcriptomes annotated by genome sequencing projects. Although current high-throughput strategies systematically survey either transcriptional or biochemical networks, analogous genome-scale investigations that analyze gene function in mammalian cells have yet to be fully realized. Through transient overexpression analysis, we describe the parallel interrogation of approximately 20,000 sequence annotated genes in cancer-related signaling pathways. For experimental validation of these genome data, we apply an integrative strategy to characterize previously unreported effectors of activator protein-1 (AP-1) mediated growth and mitogenic response pathways. These studies identify the ADP-ribosylation factor GTPase-activating protein Centaurin alpha1 and a Tudor domain-containing hypothetical protein as putative AP-1 regulatory oncogenes. These results provide insight into the composition of the AP-1 signaling machinery and validate this approach as a tractable platform for genome-wide functional analysis.

Animals↗

Evolution of ABCA4 proteins in vertebrates.

The ABCA4 (ABCR) gene encodes a retinal-specific ATP-binding cassette transporter. Mutations in ABCA4 are responsible for several recessive macular dystrophies and susceptibility to age related macular degeneration (AMD). The protein appears to function as a flippase of all-trans-retinaldehyde and/or its derivatives across the membrane of outer segment disks and is a potentially important element in recycling visual cycle metabolites. However, the understanding of ABCA4's role in the visual cycle is limited due to the lack of a direct functional assay. An evolutionary analysis of ABCA4 may aid in the identification of conserved elements, the preservation of which implies functional importance. To date, only human, murine, and bovine ABCA4 genes are described. We have identified ABCA4 genes from African (Xenopus laevis) and Western (Silurana tropicalis) clawed frogs. A comparative analysis describing the evolutionary relationships between the frog ABCA4s, annotated T. rubripes ABCA4, and mammalian ABCA4 proteins was carried out. Several segments are conserved in both intradiscal loop (IL) domains, in addition to the transmembrane and ATP-binding domains. Nonconserved segments were found in the IL and cytoplasmic linker domains. Maximum likelihood analyses of the aligned sequences strongly suggest that ABCA4 was subject to purifying selection. Collectively, these data corroborate the current evolutionary model where two distinct ABCA half-transporter progenitors were combined to form a full ABCA4 progenitor in ancestral chordates. We speculate that evolutionary alterations may increase the retinoid metabolite recycling capacity of ABCA4 and may improve dark adaptation.

ATP-Binding Cassette Transporters↗

Application of a Translational Research Platform to Unveil Efficacy Signals and Mechanisms of Resistance of FGFR Inhibitors in Multiple FGFR-Altered Solid Tumors.

PURPOSE: The predictive value of fibroblast growth factor receptor (FGFR) amplifications (amp) and the role of FGFR mutations (mut) beyond known activating variants remain unclear. We aimed to establish a translational research platform to characterize FGFR alterations (alt) and explore their potential as predictive biomarkers for FGFR-targeted agents. EXPERIMENTAL DESIGN: This ambispective study included a retrospective analysis of patients with FGFR-alt tumors treated with selective FGFR inhibitors (FGFRi) and a prospective collection of longitudinal tumor samples. Patient-derived xenografts (PDX) were generated to investigate FGFRi mechanisms of action and resistance. Molecular characterization included genomic, transcriptomic, proteomic, and functional analyses using the Functional Annotation for Cancer Treatment (FACT) assay. RESULTS: Among 36 retrospectively analyzed patients, clinical benefit from FGFRis was observed in cases with FGFR mRNA overexpression or FGFR2/11q co-amp, but no association was found with the amplification levels. In archival tumor samples, exploratory proteomic analysis showed FGFR1-4 protein expression in 78% of FGFR1/2-amp tumors detected by fluorescence in situ hybridization. RNA sequencing identified a higher prevalence of FGFR mRNA overexpression than proteomic analysis. Among patients harboring FGFR-mut, only one bladder cancer with an FGFR3-mut S249C derived benefit. FACT assay supported the functional activity of selected variants, including FGFR3 T689M, and suggested potential resistance mechanisms involving PI3K/PTEN and MAPK pathway co-alterations. A prospective FGFR-alt PDX biorepository enabled exploratory biomarker analyses, supporting the hypothesis that FGFR1-4 mRNA expression may better reflect FGFR dependency than genomic alterations alone. CONCLUSIONS: These findings highlight the complexity of FGFR-driven oncogenesis and support integrative molecular approaches to refine patient selection for FGFR-targeted therapies.

Humans↗

Virus-PLoc: a fusion classifier for predicting the subcellular localization of viral proteins within host and virus-infected cells.

Viruses can reproduce their progenies only within a host cell, and their actions depend both on its destructive tendencies toward a specific host cell and on environmental conditions. Therefore, knowledge of the subcellular localization of viral proteins in a host cell or virus-infected cell is very useful for in-depth studying of their functions and mechanisms as well as designing antiviral drugs. An analysis on the Swiss-Prot database (version 50.0, released on May 30, 2006) indicates that only 23.5% of viral protein entries are annotated for their subcellular locations in this regard. As for the gene ontology database, the corresponding percentage is 23.8%. Such a gap calls for the development of high throughput tools for timely annotating the localization of viral proteins within host and virus-infected cells. In this article, a predictor called "Virus-PLoc" has been developed that is featured by fusing many basic classifiers with each engineered according to the K-nearest neighbor rule. The overall jackknife success rate obtained by Virus-PLoc in identifying the subcellular compartments of viral proteins was 80% for a benchmark dataset in which none of proteins has more than 25% sequence identity to any other in a same location site. Virus-PLoc will be freely available as a web-server at http://202.120.37.186/bioinf/virus for the public usage. Furthermore, Virus-PLoc has been used to provide large-scale predictions of all viral protein entries in Swiss-Prot database that do not have subcellular location annotations or are annotated as being uncertain. The results thus obtained have been deposited in a downloadable file prepared with Microsoft Excel and named "Tab_Virus-PLoc.xls." This file is available at the same website and will be updated twice a year to include the new entries of viral proteins and reflect the continuous development of Virus-PLoc.

Cells↗

Glucose-6-phosphate isomerase from the hyperthermophilic archaeon Methanococcus jannaschii: characterization of the first archaeal member of the phosphoglucose isomerase superfamily.

ORF MJ1605, previously annotated as pgi and coding for the putative glucose-6-phosphate isomerase (phosphoglucose isomerase, PGI) of the hyperthermophilic archaeon Methanococcus jannaschii, was cloned and functionally expressed in Escherichia coli. The purified 80-kDa protein consisted of a single subunit of 45 kDa, indicating a homodimeric (alpha(2)) structure. The K(m) values for fructose 6-phosphate and glucose 6-phosphate were 0.04 mM and 1 mM, the corresponding V(max) values were 20 U/mg and 9 U/mg, respectively (at 50 degrees C). The enzyme had a temperature optimum at 89 degrees C and showed significant thermostability up to 95 degrees C. The enzyme was inhibited by 6-phosphogluconate and erythrose-4-phosphate. RT-PCR experiments demonstrated in vivo expression of ORF MJ1618 during lithoautotrophic growth of M. jannaschii on H(2)/CO(2). Phylogenetic analyses indicated that M. jannaschii PGI was obtained from bacteria, presumably from the hyperthermophile Thermotoga maritima.

Cloning, Molecular↗

Backbone solution structures of proteins using residual dipolar couplings: application to a novel structural genomics target.

Structural genomics (or proteomics) activities are critically dependent on the availability of high-throughput structure determination methodology. Development of such methodology has been a particular challenge for NMR based structure determination because of the demands for isotopic labeling of proteins and the requirements for very long data acquisition times. We present here a methodology that gains efficiency from a focus on determination of backbone structures of proteins as opposed to full structures with all sidechains in place. This focus is appropriate given the presumption that many protein structures in the future will be built using computational methods that start from representative fold family structures and replace as many as 70% of the sidechains in the course of structure determination. The methodology we present is based primarily on residual dipolar couplings (RDCs), readily accessible NMR observables that constrain the orientation of backbone fragments irrespective of separation in space. A new software tool is described for the assembly of backbone fragments under RDC constraints and an application to a structural genomics target is presented. The target is an 8.7 kDa protein from Pyrococcus furiosus, PF1061, that was previously not well annotated, and had a nearest structurally characterized neighbor with only 33% sequence identity. The structure produced shows structural similarity to this sequence homologue, but also shows similarity to other proteins, which suggests a functional role in sulfur transfer. Given the backbone structure and a possible functional link this should be an ideal target for development of modeling methods.

Amino Acid Sequence↗

Gene expression profiles of human proximal tubular epithelial cells in proteinuric nephropathies.

In kidney disease renal proximal tubular epithelial cells (RPTEC) actively contribute to the progression of tubulointerstitial fibrosis by mediating both an inflammatory response and via epithelial-to-mesenchymal transition. Using laser capture microdissection we specifically isolated RPTEC from cryosections of the healthy parts of kidneys removed owing to renal cell carcinoma and from kidney biopsies from patients with proteinuric nephropathies. RNA was extracted and hybridized to complementary DNA microarrays after linear RNA amplification. Statistical analysis identified 168 unique genes with known gene ontology association, which separated patients from controls. Besides distinct alterations in signal-transduction pathways (e.g. Wnt signalling), functional annotation revealed a significant upregulation of genes involved in cell proliferation and cell cycle control (like insulin-like growth factor 1 or cell division cycle 34), cell differentiation (e.g. bone morphogenetic protein 7), immune response, intracellular transport and metabolism in RPTEC from patients. On the contrary we found differential expression of a number of genes responsible for cell adhesion (like BH-protocadherin) with a marked downregulation of most of these transcripts. In summary, our results obtained from RPTEC revealed a differential regulation of genes, which are likely to be involved in either pro-fibrotic or tubulo-protective mechanisms in proteinuric patients at an early stage of kidney disease.

Aged↗

UniProt: the Universal Protein knowledgebase.

To provide the scientific community with a single, centralized, authoritative resource for protein sequences and functional information, the Swiss-Prot, TrEMBL and PIR protein database activities have united to form the Universal Protein Knowledgebase (UniProt) consortium. Our mission is to provide a comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase, with extensive cross-references and query interfaces. The central database will have two sections, corresponding to the familiar Swiss-Prot (fully manually curated entries) and TrEMBL (enriched with automated classification, annotation and extensive cross-references). For convenient sequence searches, UniProt also provides several non-redundant sequence databases. The UniProt NREF (UniRef) databases provide representative subsets of the knowledgebase suitable for efficient searching. The comprehensive UniProt Archive (UniParc) is updated daily from many public source databases. The UniProt databases can be accessed online (http://www.uniprot.org) or downloaded in several formats (ftp://ftp.uniprot.org/pub). The scientific community is encouraged to submit data for inclusion in UniProt.

Animals↗

NPRD: Nucleosome Positioning Region Database.

Nucleosome Positioning Region Database (NPRD), which is compiling the available experimental data on locations and characteristics of nucleosome formation sites (NFSs), is the first curated NFS-oriented database. The object of the database is a single NFS described in an individual entry. When annotating results of NFS experimental mapping, we pay special attention to several important functional characteristics, such as the relationship between type of gene activity and nucleosome positioning, the influence of non-histone proteins on nucleosome formation, type of the variant of nucleosome positioning (translational or rotational), indication of tissue types and states of cell activity, description of experimental methods used and accuracy of nucleosome position determination, and the results of applying theoretical and computer methods to the analysis of contextual and conformational DNA properties. At present, the NPRD database contains 438 entries and integrates the data described in 124 original papers. The database URL: http://srs6.bionet.nsc.ru/srs6/. Then click the button 'Databank' and open the link NUCLEOSOME.

Chromosome Mapping↗