Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Helicobacter pylori FlhB function: the FlhB C-terminal homologue HP1575 acts as a "spare part" to permit flagellar export when the HP0770 FlhBCC domain is deleted.

In Helicobacter pylori 26695, a gene annotated HP1575 encodes a putative protein of unknown function which shows significant similarity to part of the C-terminal domain of the flagellar export protein FlhB. In Salmonella enterica, this part (FlhB(CC)) is proteolytically cleaved from the full-length FlhB, a processing event that is required for flagellar protein export and, thus, motility. The role of FlhB (HP0770) and its C-terminal homologue HP1575 was studied in H. pylori using a range of nonpolar deletion mutants defective in HP1575, HP0770, and the CC domain of HP0770 (HP0770(CC)). Deletion of HP0770 abolished swimming motility, whereas mutants carrying a deletion of either HP1575 or HP0770(CC) retained their ability to swim. An H. pylori strain containing deletions in both HP1575 and HP0770(CC) was nonmotile and did not produce flagella, suggesting that at least one of the two proteins had to be present for flagellar assembly to occur. Indeed, motility was restored when HP1575 was reintroduced into this strain immediately downstream of, but not fused to, the truncated HP0770 gene. Thus, HP1575 can functionally replace HP0770(CC) in this background. Like FlhB in S. enterica, HP0770 appeared to be proteolytically processed at a conserved NPTH processing site. However, mutation of the proline contained within the NPTH site of HP0770 did not affect motility and flagellar assembly, although it clearly interfered with processing when the protein was heterologously produced in Escherichia coli.

Amino Acid Sequence↗

pdbFun: mass selection and fast comparison of annotated PDB residues.

pdbFun (http://pdbfun.uniroma2.it) is a web server for structural and functional analysis of proteins at the residue level. pdbFun gives fast access to the whole Protein Data Bank (PDB) organized as a database of annotated residues. The available data (features) range from solvent exposure to ligand binding ability, location in a protein cavity, secondary structure, residue type, sequence functional pattern, protein domain and catalytic activity. Users can select any residue subset (even including any number of PDB structures) by combining the available features. Selections can be used as probe and target in multiple structure comparison searches. For example a search could involve, as a query, all solvent-exposed, hydrophylic residues that are not in alpha-helices and are involved in nucleotide binding. Possible examples of targets are represented by another selection, a single structure or a dataset composed of many structures. The output is a list of aligned structural matches offered in tabular and also graphical format.

Algorithms↗

A gap-free, telomere-to-telomere chromosome-scale genome assembly of the mangrove red snapper, Lutjanus argentimaculatus.

The mangrove red snapper (Lutjanus argentimaculatus) is a commercially important marine fish species in the Indo-Pacific region. Despite its significant economic value for aquaculture, existing genomic resources remain fragmented, limiting the advancement of molecular breeding and functional genomic studies. Here, we present a gap-free, telomere-to-telomere (T2T) genome assembly of L. argentimaculatus, generated using a hybrid approach combining PacBio HiFi, Oxford Nanopore ultra-long reads and Hi-C technology. The resulting assembly comprises exactly 24 scaffolds spanning 1.03 Gb, perfectly matching the haploid chromosome number with a contig N50 of 46.17 Mb. Notably, this assembly resolves all physical gaps present in previous versions, achieving a BUSCO completeness score of 98.2%. Comprehensive genome annotation successfully predicted 23,167 protein-coding genes. Among these, 22,067 genes (95.25%) were functionally annotated across major public databases, including eggNOG, InterPro, and Swiss-Prot. Furthermore, structural analysis successfully identified 19 telomeres and 20 centromeres, validating the chromosomal integrity. This high-fidelity, gap-free reference genome provides a robust foundation for comparative genomics, population genetics, and the genetic improvement of Lutjanidae species.

Animals↗

Long homopurine*homopyrimidine sequences are characteristic of genes expressed in brain and the pseudoautosomal region.

Homo(purine*pyrimidine) sequences (R*Y tracts) with mirror repeat symmetries form stable triplexes that block replication and transcription and promote genetic rearrangements. A systematic search was conducted to map the location of the longest R*Y tracts in the human genome in order to assess their potential function(s). The 814 R*Y tracts with > or =250 uninterrupted base pairs were preferentially clustered in the pseudoautosomal region of the sex chromosomes and located in the introns of 228 annotated genes whose protein products were associated with functions at the cell membrane. These genes were highly expressed in the brain and particularly in genes associated with susceptibility to mental disorders, such as schizophrenia. The set of 1957 genes harboring the 2886 R*Y tracts with > or =100 uninterrupted base pairs was additionally enriched in proteins associated with phosphorylation, signal transduction, development and morphogenesis. Comparisons of the > or =250 bp R*Y tracts in the mouse and chimpanzee genomes indicated that these sequences have mutated faster than the surrounding regions and are longer in humans than in chimpanzees. These results support a role for long R*Y tracts in promoting recombination and genome diversity during evolution through destabilization of chromosomal DNA, thereby inducing repair and mutation.

Animals↗

The Lipase Engineering Database: a navigation and analysis tool for protein families.

The Lipase Engineering Database (LED) (http://www.led.uni-stuttgart.de) integrates information on sequence, structure, and function of lipases, esterases, and related proteins. Sequence data on 806 protein entries are assigned to 38 homologous families, which are grouped into 16 superfamilies with no global sequence similarity between each other. For each family, multisequence alignments are provided with functionally relevant residues annotated. Pre-calculated phylogenetic trees allow navigation inside superfamilies. Experimental structures of 45 proteins are superposed and consistently annotated. The LED has been applied to systematically analyze sequence-structure-function relationships of this vast and diverse enzyme class. It is a useful tool to identify functionally relevant residues apart from the active site residues, and to design mutants with desired substrate specificity.

Amino Acid Sequence↗

Functional clues for hypothetical proteins based on genomic context analysis in prokaryotes.

Three integrated genomic context methods were used to annotate uncharacterized proteins in 102 bacterial genomes. Of 7853 orthologous groups with unknown function containing 45,110 proteins, 1738 groups could be linked to functionally associated partners. In many cases, those partners are uncharacterized themselves (hinting at newly identified modules) or have been described in general terms only. However, we were able to assign pathways, cellular processes or physical complexes for 273 groups (encompassing 3624 previously functionally uncharacterized proteins).

Bacterial Proteins↗

Automatic discovery of cross-family sequence features associated with protein function.

BACKGROUND: Methods for predicting protein function directly from amino acid sequences are useful tools in the study of uncharacterized protein families and in comparative genomics. Until now, this problem has been approached using machine learning techniques that attempt to predict membership, or otherwise, to predefined functional categories or subcellular locations. A potential drawback of this approach is that the human-designated functional classes may not accurately reflect the underlying biology, and consequently important sequence-to-function relationships may be missed. RESULTS: We show that a self-supervised data mining approach is able to find relationships between sequence features and functional annotations. No preconceived ideas about functional categories are required, and the training data is simply a set of protein sequences and their UniProt/Swiss-Prot annotations. The main technical aspect of the approach is the co-evolution of amino acid-based regular expressions and keyword-based logical expressions with genetic programming. Our experiments on a strictly non-redundant set of eukaryotic proteins reveal that the strongest and most easily detected sequence-to-function relationships are concerned with targeting to various cellular compartments, which is an area already well studied both experimentally and computationally. Of more interest are a number of broad functional roles which can also be correlated with sequence features. These include inhibition, biosynthesis, transcription and defence against bacteria. Despite substantial overlaps between these functions and their corresponding cellular compartments, we find clear differences in the sequence motifs used to predict some of these functions. For example, the presence of polyglutamine repeats appears to be linked more strongly to the "transcription" function than to the general "nuclear" function/location. CONCLUSION: We have developed a novel and useful approach for knowledge discovery in annotated sequence data. The technique is able to identify functionally important sequence features and does not require expert knowledge. By viewing protein function from a sequence perspective, the approach is also suitable for discovering unexpected links between biological processes, such as the recently discovered role of ubiquitination in transcription.

Algorithms↗

A functional annotation of subproteomes in human plasma.

The data collected by Human Proteome Organization's Plasma Proteome Pilot project phase was analyzed by members of our working group. Accordingly, a functional annotation of the human plasma proteome was carried out. Here, we report the findings of our analyses. First, bioinformatic analyses were undertaken to determine the likely sources of plasma proteins and to develop a protein interaction network of proteins identified in this project. Second, annotation of these proteins was performed in the context of functional subproteomes involved in the coagulation pathway, the mononuclear phagocytic system, the inflammation pathway, the cardiovascular system, and the liver; as well as the subset of proteins associated with DNA binding activities. Our analyses contributed to the Plasma Proteome Database (http://www.plasmaproteomedatabase.org), an annotated database of plasma proteins identified by HPPP as well as from other published studies. In addition, we address several methodological considerations including the selective enrichment of post-translationally modified proteins by the use of multi-lectin chromatography as well as the use of peptidomic techniques to characterize the low molecular weight proteins in plasma. Furthermore, we have performed additional analyses of peptide identification data to annotate cleavage of signal peptides, sites of intra-membrane proteolysis and post-translational modifications. The HPPP-organized, multi-laboratory effort, as described herein, resulted in much synergy and was essential to the success of this project.

Blood Coagulation↗

A mathematical and computational framework for quantitative comparison and integration of large-scale gene expression data.

Analysis of large-scale gene expression studies usually begins with gene clustering. A ubiquitous problem is that different algorithms applied to the same data inevitably give different results, and the differences are often substantial, involving a quarter or more of the genes analyzed. This raises a series of important but nettlesome questions: How are different clustering results related to each other and to the underlying data structure? Is one clustering objectively superior to another? Which differences, if any, are likely candidates to be biologically important? A systematic and quantitative way to address these questions is needed, together with an effective way to integrate and leverage expression results with other kinds of large-scale data and annotations. We developed a mathematical and computational framework to help quantify, compare, visualize and interactively mine clusterings. We show that by coupling confusion matrices with appropriate metrics (linear assignment and normalized mutual information scores), one can quantify and map differences between clusterings. A version of receiver operator characteristic analysis proved effective for quantifying and visualizing cluster quality and overlap. These methods, plus a flexible library of clustering algorithms, can be called from a new expandable set of software tools called CompClust 1.0 (http://woldlab.caltech.edu/compClust/). CompClust also makes it possible to relate expression clustering patterns to DNA sequence motif occurrences, protein-DNA interaction measurements and various kinds of functional annotations. Test analyses used yeast cell cycle data and revealed data structure not obvious under all algorithms. These results were then integrated with transcription motif and global protein-DNA interaction data to identify G1 regulatory modules.

Algorithms↗

The SWISS-PROT protein sequence data bank and its new supplement TREMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domain structure, post-translational modifications, variants, etc), a minimal level of redundancy and a high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to seven additional databases; a variety of new documentation files; the creation of TREMBL, and unannotated supplement to SWISS-PROT. This supplement consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except CDS already included in SWISS-PROT.

Amino Acid Sequence↗

Genome-scale gene function prediction using multiple sources of high-throughput data in yeast Saccharomyces cerevisiae.

Characterizing gene function is one of the major challenging tasks in the post-genomic era. To address this challenge, we have developed GeneFAS (Gene Function Annotation System), a new integrated probabilistic method for cellular function prediction by combining information from protein-protein interactions, protein complexes, microarray gene expression profiles, and annotations of known proteins through an integrative statistical model. Our approach is based on a novel assessment for the relationship between (1) the interaction/correlation of two proteins' high-throughput data and (2) their functional relationship in terms of their Gene Ontology (GO) hierarchy. We have developed a Web server for the predictions. We have applied our method to yeast Saccharomyces cerevisiae and predicted functions for 1548 out of 2472 unannotated proteins.

Automation↗

Cluster analysis of protein array results via similarity of Gene Ontology annotation.

BACKGROUND: With the advent of high-throughput proteomic experiments such as arrays of purified proteins comes the need to analyse sets of proteins as an ensemble, as opposed to the traditional one-protein-at-a-time approach. Although there are several publicly available tools that facilitate the analysis of protein sets, they do not display integrated results in an easily-interpreted image or do not allow the user to specify the proteins to be analysed. RESULTS: We developed a novel computational approach to analyse the annotation of sets of molecules. As proof of principle, we analysed two sets of proteins identified in published protein array screens. The distance between any two proteins was measured as the graph similarity between their Gene Ontology (GO) annotations. These distances were then clustered to highlight subsets of proteins sharing related GO annotation. In the first set of proteins found to bind small molecule inhibitors of rapamycin, we identified three subsets containing four or five proteins each that may help to elucidate how rapamycin affects cell growth whereas the original authors chose only one novel protein from the array results for further study. In a set of phosphoinositide-binding proteins, we identified subsets of proteins associated with different intracellular structures that were not highlighted by the analysis performed in the original publication. CONCLUSION: By determining the distances between annotations, our methodology reveals trends and enrichment of proteins of particular functions within high-throughput datasets at a higher sensitivity than perusal of end-point annotations. In an era of increasingly complex datasets, such tools will help in the formulation of new, testable hypotheses from high-throughput experimental data.

Algorithms↗

PIRSF: family classification system at the Protein Information Resource.

The Protein Information Resource (PIR) is an integrated public resource of protein informatics. To facilitate the sensible propagation and standardization of protein annotation and the systematic detection of annotation errors, PIR has extended its superfamily concept and developed the SuperFamily (PIRSF) classification system. Based on the evolutionary relationships of whole proteins, this classification system allows annotation of both specific biological and generic biochemical functions. The system adopts a network structure for protein classification from superfamily to subfamily levels. Protein family members are homologous (sharing common ancestry) and homeomorphic (sharing full-length sequence similarity with common domain architecture). The PIRSF database consists of two data sets, preliminary clusters and curated families. The curated families include family name, protein membership, parent-child relationship, domain architecture, and optional description and bibliography. PIRSF is accessible from the website at http://pir.georgetown.edu/pirsf/ for report retrieval and sequence classification. The report presents family annotation, membership statistics, cross-references to other databases, graphical display of domain architecture, and links to multiple sequence alignments and phylogenetic trees for curated families. PIRSF can be utilized to analyze phylogenetic profiles, to reveal functional convergence and divergence, and to identify interesting relationships between homeomorphic families, domains and structural classes.

Amino Acid Motifs↗

Mass spectrometric analysis of the editosome and other multiprotein complexes in Trypanosoma brucei.

The composition of the editosome, a multi-protein complex that catalyzes uridine insertion and deletion RNA editing to produce mature mitochondrial mRNAs in trypanosomes, was analyzed by mass spectrometry. The editosomes were isolated by column chromatography, glycerol gradient sedimentation, and monoclonal antibody affinity purifications. At least 16 proteins form the catalytic core of the editosome, and additional associated proteins were identified. Analyses of mitochondrial fractions identified several non-editosome proteins and multi-protein complexes. These studies contribute to the functional annotation of T. brucei genome.

Amino Acid Sequence↗

The proteome of Mannheimia succiniciproducens, a capnophilic rumen bacterium.

Mannheimia succiniciproducens MBEL55E isolated from bovine rumen is an industrially important bacterium as an efficient succinic acid producer. Recently, its full genome sequence was determined. In the present study, we analyzed the M. succiniciproducens proteome based on the genome information using 2-DE and MS. We established proteome reference map of M. succiniciproducens by analyzing whole cellular proteins, membrane proteins, and secreted proteins. More than 200 proteins were identified and characterized by MS/MS supported by various bioinformatic tools. The presence of proteins previously annotated as hypothetical proteins or proteins having putative functions were also confirmed. Based on the proteome reference map, cells in the different growth phases were analyzed at the proteome level. Comparative proteome profiling revealed valuable information to understand physiological changes during growth, and subsequently suggested target genes to be manipulated for the strain improvement.

Animals↗

Transcriptional profiling of wheat caryopsis development using cDNA microarrays.

The expression of 7,835 genes in developing wheat caryopses was analyzed using cDNA arrays. Using a mixed model analysis of variance (ANOVA) method, 29% (2,237) of the genes on the array were identified to be differentially expressed at the 6 different time-points examined, which covers the developmental stages from coenocytic endosperm to physiological maturity. Comparison of genes differentially expressed between two time-points revealed a dynamic transcript accumulation profile with major re-programming events that occur at 3-7, 7-14 and 21-28 DPA. A k-means clustering algorithm grouped the differentially expressed genes into 10 clusters, revealing co-expression of genes involved in the same pathway such as carbohydrate and protein synthesis or preparation for desiccation. Functional annotation of genes that show peak expression at specific time-points correlated with the developmental events associated with the respective stages. Results provide information on the temporal expression during caryopsis development for a significant number of differentially expressed genes with unknown function.

DNA, Complementary↗

Chromosomal level genome assembly of medicinal plant Chrysosplenium macrophyllum.

Chrysosplenium macrophyllum Oliv., a perennial herb native to China, is widely used in traditional medicine for its notable therapeutic properties. However, the absence of a reference genome has constrained its full potential for research and application. This study presents the first chromosome-level de novo genome assembly of C. macrophyllum, constructed by integrating long reads from Oxford Nanopore Technologies (ONT), short reads from BGI, and Hi-C data. The final assembly spans 2.55 Gb, with a scaffold N50 of 93.38 Mb, and 83.70% of the genome has been assigned to 22 chromosomes. The mapping rate of the BGI short reads to the genome is approximately 97.94%, and BUSCO analysis reveals that 97.94% of the predicted genes are complete. A total of 62,921 protein-coding genes were predicted, with functional annotations for 93.67% of them. This chromosome-level genome assembly represents an important resource for expanding our understanding of Chrysosplenium species and supports future genomic studies and applications.

Genome, Plant↗

The SBASE domain library: a collection of annotated protein segments.

SBASE is a database of annotated protein domain sequences representing various structural, functional, ligand binding and topogenic segments of proteins. The current release of SBASE contains 27,211 entries which are provided with standardized names in order to facilitate retrieval. SBASE is cross-referenced to the major protein and nucleic acid databanks as well as to the PROSITE catalog of protein sequence patterns [Bairoch, A. (1992) Nucleic Acids Res., 20, Suppl., 2013-2118]. SBASE can be used to establish domain homologies through database search using programs such as FASTA [Lipman and Pearson (1985) Science, 227, 1436-1441], FASTDB [Brutlag et al. (1990) Comp. Appl. Biosci., 6, 237-245] or BLAST3 [Altschul and Lipman (1990) Proc. Natl. Acad. Sci. USA, 87, 5509-5513], which is especially useful in the case of loosely defined domain types for which efficient consensus patterns cannot be established. The use of SBASE is illustrated on the DNA binding protein Brain-4. The database and a set of search and retrieval tools are freely available on request to the authors or by anonymous 'ftp' file transfer from < ftp.icgeb.trieste.it >.

Amino Acid Sequence↗