Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Using optimized evidence-theoretic K-nearest neighbor classifier and pseudo-amino acid composition to predict membrane protein types.

Knowledge of membrane protein type often provides crucial hints toward determining the function of an uncharacterized membrane protein. With the avalanche of new protein sequences emerging during the post-genomic era, it is highly desirable to develop an automated method that can serve as a high throughput tool in identifying the types of newly found membrane proteins according to their primary sequences, so as to timely make the relevant annotations on them for the reference usage in both basic research and drug discovery. Based on the concept of pseudo-amino acid composition [K.C. Chou, Proteins: Struct. Funct. Genet. 43 (2001) 246-255; Erratum: Proteins: Struct. Funct. Genet. 44 (2001) 60] that has made it possible to incorporate a considerable amount of sequence-order effects by representing a protein sample in terms of a set of discrete numbers, a novel predictor, the so-called "optimized evidence-theoretic K-nearest neighbor" or "OET-KNN" classifier, was proposed. It was demonstrated via the self-consistency test, jackknife test, and independent dataset test that the new predictor, compared with many previous ones, yielded higher success rates in most cases. The new predictor can also be used to improve the prediction quality for, among many other protein attributes, structural class, subcellular localization, enzyme family class, and G-protein coupled receptor type. The OET-KNN classifier will be available as a web-server at http://www.pami.sjtu.edu.cn/kcchou.

Algorithms↗

High-Quality Genome Assembly, Metabolome, Pangenome, and Metabolic Models of Megasphaera hexanoica KCCM 43214T.

Megasphaera hexanoica KCCM 43214T, isolated from cow rumen, is capable of producing medium-chain carboxylic acids such as hexanoate and octanoate. In this study, we present a high-quality genome assembly, along with intracellular metabolomic profiling and pangenomic analysis. Illumina sequencing generated 2.3 Gbp from 15,293,634 reads with a GC content of 49.5%, while PacBio HiFi sequencing produced 331.5 Mbp across 45,266 reads, with an average read length of 7,323 bp and a HiFi read N50 of 8,214 bp. Hybrid assembly of short and long reads resulted in a single 2.88 Mbp contig, containing 2,835 protein-coding genes. Genome-scale metabolic models were constructed to evaluate its metabolic capabilities under specific growth conditions. Intracellular metabolomic analysis of cells grown in medium containing fructose and lactate revealed key metabolic activities associated with chain elongation. Pangenomic analysis across nine annotated genomes identified 6,721 orthologous genes using OrthoMCL, emphasizing the genetic and functional diversity within the Megasphaera genus. This dataset offers valuable insights into the metabolism and biotechnological potential of M. hexanoica KCCM 43214T.

Metabolome↗

Localization, annotation, and comparison of the Escherichia coli K-12 proteome under two states of growth.

Here we describe a proteomic analysis of Escherichia coli in which 3,199 protein forms were detected, and of those 2,160 were annotated and assigned to the cytosol, periplasm, inner membrane, and outer membrane by biochemical fractionation followed by two-dimensional gel electrophoresis and tandem mass spectrometry. Represented within this inventory were unique and modified forms corresponding to 575 different ORFs that included 151 proteins whose existence had been predicted from hypothetical ORFs, 76 proteins of completely unknown function, and 222 proteins currently without location assignments in the Swiss-Prot Database. Of the 575 unique proteins identified, 42% were found to exist in multiple forms. Using DIGE, we also examined the relative changes in protein expression when cells were grown in the presence and absence of amino acids. A total of 23 different proteins were identified whose abundance changed significantly between the two conditions. Most of these changes were found to be associated with proteins involved in carbon and amino acid metabolism, transport, and chemotaxis. Detailed information related to all 2,160 protein forms (protein and gene names, accession numbers, subcellular locations, relative abundances, sequence coverage, molecular masses, and isoelectric points) can be obtained upon request in either tabular form or as interactive gel images.

Amino Acids↗

Gene and alternative splicing annotation with AIR.

Designing effective and accurate tools for identifying the functional and structural elements in a genome remains at the frontier of genome annotation owing to incompleteness and inaccuracy of the data, limitations in the computational models, and shifting paradigms in genomics, such as alternative splicing. We present a methodology for the automated annotation of genes and their alternatively spliced mRNA transcripts based on existing cDNA and protein sequence evidence from the same species or projected from a related species using syntenic mapping information. At the core of the method is the splice graph, a compact representation of a gene, its exons, introns, and alternatively spliced isoforms. The putative transcripts are enumerated from the graph and assigned confidence scores based on the strength of sequence evidence, and a subset of the high-scoring candidates are selected and promoted into the annotation. The method is highly selective, eliminating the unlikely candidates while retaining 98% of the high-quality mRNA evidence in well-formed transcripts, and produces annotation that is measurably more accurate than some evidence-based gene sets. The process is fast, accurate, and fully automated, and combines the traditionally distinct gene annotation and alternative splicing detection processes in a comprehensive and systematic way, thus considerably aiding in the ensuing manual curation efforts.

Alternative Splicing↗

Dehydrogenases from all three domains of life cleave RNA.

Specific interactions of glyceraldehyde-3-phosphate dehydrogenase (GAPDH) with RNA have been reported both in vitro and in vivo. We show that eukaryotic and bacterial GAPDH and two proteins from the hyperthermophilic archaeon Sulfolobus solfataricus, which are annotated as dehydrogenases, cleave RNA producing similar degradation patterns. RNA cleavage is most efficient at 60 degrees C, at MgCl(2) concentrations up to 5 mm, and takes place between pyrimidine and adenosine. The RNase active center of the putative aspartate semialdehyde dehydrogenase from S. solfataricus is located within the N-terminal 73 amino acids, which comprise the first mononucleotide-binding site of the predicted Rossmann fold. Thus, RNA cleavage has to be taken into account in the ongoing discussion of the possible biological function of RNA binding by dehydrogenases.

Archaeal Proteins↗

Genome-Wide Characterization of β-Glucosidase (TaBGLU) Genes in Bread Wheat and Their Expression Under Drought, Cold, and Combined Stress.

Glycoside hydrolase 1 (GH1) β-glucosidases were known to activate hormone conjugates and defense metabolites, yet their genomic organization and stress-response dynamics in wheat remained incompletely defined. We therefore performed an integrated characterization of TaBGLUs spanning phylogeny, gene structure and conserved motifs, subcellular localization, promoter cis-elements, Gene Ontology enrichment, protein-protein interaction networks, and targeted expression profiling. Wheat TaBGLUs partitioned into well-supported clades that shared canonical GH1 catalytic residues and a largely conserved motif scaffold. Subcellular localization predictions indicated predominant nuclear and chloroplast targeting, with a smaller cohort directed to secretory or endomembrane compartments. Promoters were enriched for light-responsive, hormone-related (ABA, JA/SA, auxin, GA) and stress-associated (MYB/WRKY, heat, low temperature) cis-elements, and functional annotations were consistent with roles in carbohydrate and cell-wall metabolism, hormone homeostasis, and defense. Network analysis revealed a densely connected TaBGLU submodule embedded within broader carbohydrate and defense interaction networks, suggesting coordinated or cooperative functions. Expression profiling under cold, drought, and combined drought and cold demonstrated broad stress inducibility, with early activation detected by 6 h, cold-responsive maxima typically at 12 h, drought-responsive peaks predominating at 24 h, and combined stress eliciting both earlier and more sustained expression maxima between 12-24 h. Representative strongly responsive genes included TaBGLU20, TaBGLU44, TaBGLU6, and TaBGLU23, which showed pronounced late induction under combined stress, TaBGLU30, which exhibited an earlier combined-stress peak, and TaBGLU12, which displayed a marked late drought-specific response. Taken together, this integrated genomic, regulatory, and expression atlas refined the wheat BGLU repertoire relative to previous gene model inventories, highlighted candidate TaBGLUs with central network positions and strong stress inducibility, and provided concrete entry points for functional validation and breeding for improved stress resilience.

Triticum↗

An enzyme that regulates ether lipid signaling pathways in cancer annotated by multidimensional profiling.

Hundreds, if not thousands, of uncharacterized enzymes currently populate the human proteome. Assembly of these proteins into the metabolic and signaling pathways that govern cell physiology and pathology constitutes a grand experimental challenge. Here, we address this problem by using a multidimensional profiling strategy that combines activity-based proteomics and metabolomics. This approach determined that KIAA1363, an uncharacterized enzyme highly elevated in aggressive cancer cells, serves as a central node in an ether lipid signaling network that bridges platelet-activating factor and lysophosphatidic acid. Biochemical studies confirmed that KIAA1363 regulates this pathway by hydrolyzing the metabolic intermediate 2-acetyl monoalkylglycerol. Inactivation of KIAA1363 disrupted ether lipid metabolism in cancer cells and impaired cell migration and tumor growth in vivo. The integrated molecular profiling method described herein should facilitate the functional annotation of metabolic enzymes in any living system.

Carbamates↗

Fulfilling the promise: drug discovery in the post-genomic era.

The genomic era has brought with it a basic change in experimentation, enabling researchers to look more comprehensively at biological systems. The sequencing of the human genome coupled with advances in automation and parallelization technologies have afforded a fundamental transformation in the drug target discovery paradigm, towards systematic whole genome and proteome analyses. In conjunction with novel proteomic techniques, genome-wide annotation of function in cellular models is possible. Overlaying data derived from whole genome sequence, expression and functional analysis will facilitate the identification of causal genes in disease and significantly streamline the target validation process. Moreover, several parallel technological advances in small molecule screening have resulted in the development of expeditious and powerful platforms for elucidating inhibitors of protein or pathway function. Conversely, high-throughput and automated systems are currently being used to identify targets of orphan small molecules. The consolidation of these emerging functional genomics and drug discovery technologies promises to reap the fruits of the genomic revolution.

Animals↗

Focusing of gene expression as the basis of stem cell differentiation.

In a prior report (Stem Cells Dev 14(4):354-366, 2005), we employed two-dimensional gel electrophoresis followed by advanced proteomics and the Database for Annotation, Visualization and Integrated Discovery (DAVID) to compare the protein expression profiles of mesenchymal stem cells to that of fully differentiated osteoblasts. These data were reported to advance technical approaches to define the basis of differentiation, but also led us to suggest that osteogenic differentiation of stem cells may result from the focusing of gene expression in functional clusters (e.g., calcium-regulated signaling proteins or adherence proteins) rather than simply from the induced expression of new genes, as many have assumed. Here, we have employed these analytical techniques to compare protein expression by mesenchymal stem cells directly with that of cells derived from them after induced osteogenic differentiation. Our results support the concept of gene focusing as the basis of differentiation. Specifically, induced differentiation results in a decrease in the number of mesenchymal cell markers and calcium-mediated signaling molecules expressed by their differentiated progeny. This effect was seen in parallel to increased expression of specific extracellular matrix (ECM) molecules and their receptors. These results strongly imply that changes in the ECM have a direct impact on stem cell differentiation, and that osteogenic differentiation of stem cells directed by matrix clues results from focusing of the expression of genes involved in Ca2+-dependent signaling pathways.

Calcium Signaling↗

GeConT: gene context analysis.

SUMMARY: The fact that adjacent genes in bacteria are often functionally related is widely known. GeConT (Gene Context Tool) is a web interface designed to visualize genome context of a gene or a group of genes and their orthologs in all the completely sequenced genomes. The graphical information of GeConT can be used to analyze genome annotation, functional ortholog identification or to verify the genomic context congruence of any set of genes that share a common property. AVAILABILITY: http://www.ibt.unam.mx/biocomputo/gecont.html

Chromosome Mapping↗

circASbase: A Comprehensive Database of Alternative Splicing Events in circRNAs.

Although extensive evidence has underscored the critical role of alternative splicing (AS) in generating mature circular RNA (circRNA) isoforms and augmenting their functional diversity, a significant gap remains in the availability of specialized databases housing circRNA AS events. To bridge this gap, we develop circASbase, a pioneering and comprehensive database that catalogs 452,129 AS events in 884,047 full-length circRNAs from 581 samples across 13 species, and provides rich annotations to facilitate understanding the splicing regulation of circRNA. Our findings reveal substantial differences between circRNAs and linear transcripts regarding the distribution and occurrence of AS events, highlighting the unique regulatory landscape of circRNAs. These special splicing events result in functional differences of circRNAs by affecting internal ribosome entry sites, N6-methyladenosine sites, open reading frames, protein features, microRNA targets, and more. In summary, circASbase not only meets the urgent need of the research community for data repositories, but also represents a significant advancement in our understanding of circRNA biology. With its user-friendly interfaces and web-based visualization tools, circASbase is poised to become an indispensable resource for researchers exploring the regulatory mechanisms and functional roles of AS events in circRNAs. This database will continuously drive new insights and discoveries in the field, setting the stage for further advancements in circRNA research. circASbase is freely available at http://reprod.njmu.edu.cn/cgi-bin/circASbase/.

Alternative Splicing↗

Development of human protein reference database as an initial platform for approaching systems biology in humans.

Human Protein Reference Database (HPRD) is an object database that integrates a wealth of information relevant to the function of human proteins in health and disease. Data pertaining to thousands of protein-protein interactions, posttranslational modifications, enzyme/substrate relationships, disease associations, tissue expression, and subcellular localization were extracted from the literature for a nonredundant set of 2750 human proteins. Almost all the information was obtained manually by biologists who read and interpreted >300,000 published articles during the annotation process. This database, which has an intuitive query interface allowing easy access to all the features of proteins, was built by using open source technologies and will be freely available at http://www.hprd.org to the academic community. This unified bioinformatics platform will be useful in cataloging and mining the large number of proteomic interactions and alterations that will be discovered in the postgenomic era.

BRCA1 Protein↗

Molecular cloning and characterization of Tap, a putative multidrug efflux pump present in Mycobacterium fortuitum and Mycobacterium tuberculosis.

A recombinant plasmid isolated from a Mycobacterium fortuitum genomic library by selection for gentamicin and 2-N'-ethylnetilmicin resistance conferred low-level aminoglycoside and tetracycline resistance when introduced into M. smegmatis. Further characterization of this plasmid allowed the identification of the M. fortuitum tap gene. A homologous gene in the M. tuberculosis H37Rv genome has been identified. The M. tuberculosis tap gene (Rv1258 in the annotated sequence of the M. tuberculosis genome) was cloned and conferred low-level resistance to tetracycline when introduced into M. smegmatis. The sequences of the putative Tap proteins showed 20 to 30% amino acid identity to membrane efflux pumps of the major facilitator superfamily (MFS), mainly tetracycline and macrolide efflux pumps, and to other proteins of unknown function but with similar antibiotic resistance patterns. Approximately 12 transmembrane regions and different sequence motifs characteristic of the MFS proteins also were detected. In the presence of the protonophore carbonyl cyanide m-chlorophenylhydrazone (CCCP), the levels of resistance to antibiotics conferred by plasmids containing the tap genes were decreased. When tetracycline accumulation experiments were carried out with the M. fortuitum tap gene, the level of tetracycline accumulation was lower than that in control cells but was independent of the presence of CCCP. We conclude that the Tap proteins of the opportunistic organism M. fortuitum and the important pathogen M. tuberculosis are probably proton-dependent efflux pumps, although we cannot exclude the possibility that they act as regulatory proteins.

Amino Acid Sequence↗

The Drosophila melanogaster genome.

Drosophila's importance as a model organism made it an obvious choice to be among the first genomes sequenced, and the Release 1 sequence of the euchromatic portion of the genome was published in March 2000. This accomplishment demonstrated that a whole genome shotgun (WGS) strategy could produce a reliable metazoan genome sequence. Despite the attention to sequencing methods, the nucleotide sequence is just the starting point for genome-wide analyses; at a minimum, the genome sequence must be interpreted using expressed sequence tag (EST) and complementary DNA (cDNA) evidence and computational tools to identify genes and predict the structures of their RNA and protein products. The functions of these products and the manner in which their expression and activities are controlled must then be assessed-a much more challenging task with no clear endpoint that requires a wide variety of experimental and computational methods. We first review the current state of the Drosophila melanogaster genome sequence and its structural annotation and then briefly summarize some promising approaches that are being taken to achieve an initial functional annotation.

Animals↗

Validation and functional annotation of expression-based clusters based on gene ontology.

BACKGROUND: The biological interpretation of large-scale gene expression data is one of the paramount challenges in current bioinformatics. In particular, placing the results in the context of other available functional genomics data, such as existing bio-ontologies, has already provided substantial improvement for detecting and categorizing genes of interest. One common approach is to look for functional annotations that are significantly enriched within a group or cluster of genes, as compared to a reference group. RESULTS: In this work, we suggest the information-theoretic concept of mutual information to investigate the relationship between groups of genes, as given by data-driven clustering, and their respective functional categories. Drawing upon related approaches (Gibbons and Roth, Genome Research 12:1574-1581, 2002), we seek to quantify to what extent individual attributes are sufficient to characterize a given group or cluster of genes. CONCLUSION: We show that the mutual information provides a systematic framework to assess the relationship between groups or clusters of genes and their functional annotations in a quantitative way. Within this framework, the mutual information allows us to address and incorporate several important issues, such as the interdependence of functional annotations and combinatorial combinations of attributes. It thus supplements and extends the conventional search for overrepresented attributes within a group or cluster of genes. In particular taking combinations of attributes into account, the mutual information opens the way to uncover specific functional descriptions of a group of genes or clustering result. All datasets and functional annotations used in this study are publicly available. All scripts used in the analysis are provided as additional files.

Algorithms↗

Genome sequences and great expectations.

To assess how automatic function assignment will contribute to genome annotation in the next five years, we have performed an analysis of 31 available genome sequences. An emerging pattern is that function can be predicted for almost two-thirds of the 73,500 genes that were analyzed. Despite progress in computational biology, there will always be a great need for large-scale experimental determination of protein function.

Animals↗

Intraspecies sequence comparisons for annotating genomes.

Analysis of sequence variation among members of a single species offers a potential approach to identify functional DNA elements responsible for biological features unique to that species. Due to its high rate of allelic polymorphism and ease of genetic manipulability, we chose the sea squirt, Ciona intestinalis, to explore intraspecies sequence comparisons for genome annotation. A large number of C. intestinalis specimens were collected from four continents, and a set of genomic intervals were amplified, resequenced, and analyzed to determine the mutation rates at each nucleotide in the sequence. We found that regions with low mutation rates efficiently demarcated functionally constrained sequences: these include a set of noncoding elements, which we showed in C. intestinalis transgenic assays to act as tissue-specific enhancers, as well as the location of coding sequences. This illustrates that comparisons of multiple members of a species can be used for genome annotation, suggesting a path for the annotation of the sequenced genomes of organisms occupying uncharacterized phylogenetic branches of the animal kingdom. It also raises the possibility that the resequencing of a large number of Homo sapiens individuals might be used to annotate the human genome and identify sequences defining traits unique to our species.

Animals↗

Assessing Metal Ion Assignment Accuracy in Protein Data Bank Models via Elemental Spectroscopy.

Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.

Databases, Protein↗