Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Genome-Resolved Functional Profiling of Osteoporosis-Associated Gut Bacteria Highlights Putative Metabolic and Immunogenic Signatures of the Gut-Bone Axis.

The gut microbiota has emerged as a potential regulator of bone metabolism, but the genome-encoded functional repertoire of osteoporosis-associated gut bacteria remains insufficiently characterized. This study performed in silico functional profiling of gut bacterial taxa associated with osteoporosis, low bone mineral density, or comparator bone-related phenotypes. Twenty candidate taxa were selected from evidence in the human microbiome and represented by 26 curated bacterial reference genomes. Genome-wide annotations were used to map predicted gut-bone axis signatures, carbohydrate-active enzyme (CAZyme) repertoires, selected Kyoto Encyclopedia of Genes and Genomes pathways, and gutSMASH-predicted metabolic gene clusters. Functional burdens were normalized as hits per 1000 annotated proteins and integrated into metabolic, immunogenic, CAZyme, KEGG, and metabolic gene cluster profiles. Twelve predicted gut-bone axis signatures were identified, comprising 3337 primary candidate protein hits and a strict high-confidence subset of 2497 hits. Dominant signatures included vitamin B12/cobalamin metabolism, folate/one-carbon metabolism, peptidoglycan/cell-wall biosynthesis, and short-chain fatty acid-related functions. Dialister invisus, Dialister succinatiphilus, Megamonas funiformis, and Megamonas hypermegale showed the strongest normalized predicted gut-bone axis signal. These hypothesis-generating findings prioritize microbial metabolic and immunogenic features for future metagenomic, metabolomic, and experimental validation studies.

Osteoporosis↗

Global analysis of bacterial transcription factors to predict cellular target processes.

Whole-genome sequences are now available for >100 bacterial species, giving unprecedented power to comparative genomics approaches. We have applied genome-context methods to predict target processes that are regulated by transcription factors (TFs). Of 128 orthologous groups of proteins annotated as TFs, to date, 36 are functionally uncharacterized; in our analysis we predict a probable cellular target process or biochemical pathway for half of these functionally uncharacterized TFs.

Bacteria↗

Chromosome-level genome assembly of Elaeocarpus petiolatus (Elaeocarpaceae).

Elaeocarpus petiolatus is an ecologically and economically important species in tropical and subtropical forests. Despite its significance, the lack of genomic resources has hindered research on the genetic diversity and adaptive traits of E. petiolatus. To address this gap, we present a comprehensive chromosome-level genome assembly of E. petiolatus generated using advanced PacBio high-fidelity (HiFi) long-read sequencing and Hi-C technology. The assembly spans 322.45 Mb, with a scaffold N50 of 20.58 Mb, indicating that 37.11% of the genome is composed of repetitive elements. We identified 25,295 protein-coding genes, of which 96.74% were functionally annotated. This high-quality genome provides a critical resource for understanding the genetic mechanisms underlying environmental adaptability and biosynthesis of bioactive compounds in E. petiolatus, thereby supporting conservation efforts and sustainable forest management. The assembled genome and associated sequencing data are publicly available, facilitating further evolutionary and functional studies on the Elaeocarpaceae family.

Chromosomes, Plant↗

Chromosome-level genome assembly of the hemiparasitic Taxillus sutchuenensis (Loranthaceae).

Taxillus sutchuenensis, an ecologically and medicinally important hemiparasitic plant that parasitizes diverse woody hosts, was sequenced to generate a high-quality chromosome-level genome assembly. PacBio HiFi long reads, RNA-seq transcriptome data, and Hi-C data were used to assemble a 406.32 Mb genome anchored onto nine pseudo-chromosomes, with a scaffold N50 of 45.59 Mb. The assembly showed high completeness and accuracy, supported by BUSCO (93.6%) and Merqury QV (70.6) assessments. The LTR Assembly Index (LAI) of 13.98 indicated excellent continuity. A total of 21,795 protein-coding genes were predicted, with 94.46% functionally annotated. Repetitive sequences accounted for 50.05% of the genome, primarily LTR retrotransposons. This genome provides a valuable resource for investigating the evolution, functional genomics, and parasitic mechanisms of hemiparasitic plants.

Genome, Plant↗

The SBASE protein domain library, Release 4.0: a collection of annotated protein sequence segments.

SBASE 4.0 is the fourth release of SBASE, a collection of annotated protein domain sequences that represent various structural, functional, ligand binding and topogenic segments of proteins. SBASE was designed to facilitate the detection of functional homologies and can be searched with standard database search tools, such as FASTA and BLAST3. The present release contains 61 137 entries provided with standardized names and cross-referenced to all major protein, nucleic acid and sequence pattern collections. The entries are clustered into 13 155 groups in order to facilitate detection of distant similarities. SBASE 4.0 is freely available by anonymous ftp file transfer from ftp.icgeb.trieste.it. Individual records can be retrieved with the gopher server at icgeb.trieste.it and with a World Wide Web server at http://www.icgeb.trieste.it. Automated searching of SBASE with BLAST can be carried out with the electronic mail server sbase@icgeb.trieste.it, which now also provides a graphic representation of the homologies. A related mail server, domain@hubi.abc.hu, assigns SBASE domain homologies on the basis of SWISS-PROT searches.

Amino Acid Sequence↗

The SBASE protein domain library, release 5.0: a collection of annotated protein sequence segments.

SBASE 5.0 is the fifth release of SBASE, a collection of annotated protein domain sequences that represent various structural, functional, ligand-binding and topogenic segments of proteins. SBASE was designed to facilitate the detection of functional homologies and can be searched with standard database-search programs. The present release contains over 79863 entries provided with standardized names and is cross-referenced to all major sequence databases and sequence pattern collections. The information is assigned to individual domains rather than to entire protein sequences, thus SBASE contains substantially more cross-references and links than do the protein sequence databases. The entries are clustered into >16 000 groups in order to facilitate the detection of distant similarities. SBASE 5.0 is freely available by anonymous 'ftp' file transfer from . Automated searching of SBASE with BLAST can be carried out with the WWW-server . and with the electronic mail server which now also provides a graphic representation of the homologies. A related WWW-server and e-mail server predicts SBASE domain homologies on the basis of SWISS-PROT searches.

Amino Acid Sequence↗

The molecular structure of Rv1873, a conserved hypothetical protein from Mycobacterium tuberculosis, at 1.38 A resolution.

The X-ray crystal structure of the gene product encoded by open reading frame Rv1873 of Mycobacterium tuberculosis has been determined by single isomorphous replacement with anomalous scattering (SIRAS) phasing techniques at 1.38 A resolution from monoclinic crystals with unit-cell parameters a = 33.44, b = 31.63, c = 53.19 A, beta = 90.8 degrees. The 16.2 kDa Rv1873 is a monomer that adopts a primarily alpha-helical fold with limited structural similarity to previously determined tertiary structures. It has been annotated as a conserved hypothetical protein of unknown function and is classified by the Clusters of Orthologous Groups (COG) database as belonging to COG5579. The three-dimensional structure of the Rv1873 gene product reveals limited similarity to a repeated motif that is found in a variety of other proteins. While not a novel fold, it serves as a model for orthologues predicted to be related by sequence and it is hoped that knowledge of the structure of Rv1873 will aid in determining a possible function for this protein.

Amino Acid Sequence↗

Integration with the human genome of peptide sequences obtained by high-throughput mass spectrometry.

A crucial aim upon the completion of the human genome is the verification and functional annotation of all predicted genes and their protein products. Here we describe the mapping of peptides derived from accurate interpretations of protein tandem mass spectrometry (MS) data to eukaryotic genomes and the generation of an expandable resource for integration of data from many diverse proteomics experiments. Furthermore, we demonstrate that peptide identifications obtained from high-throughput proteomics can be integrated on a large scale with the human genome. This resource could serve as an expandable repository for MS-derived proteome information.

Amino Acid Sequence↗

Rapid proteome analysis of bronchoalveolar lavage samples of lifelong smokers and never-smokers by micro-scale liquid chromatography and mass spectrometry.

BACKGROUND: The aim of this study was to determine whether relative qualitative and quantitative differences in protein expression could be related to smoke exposure or smoke-induced airway inflammation. We therefore explored and characterized the protein components found in bronchoalveolar lavage (BAL) fluid sampled from either lifelong smokers or never-smokers. METHODS: BAL fluid samples obtained by bronchoscopy from 60-year-old healthy never-smokers (n = 18) and asymptomatic smokers (n = 30) were analyzed in either pooled or individual form. Initial global proteomic analysis used shotgun digestion approaches on unfractionated BAL fluid samples (after minimal sample preparation) and separation of peptides by gradient (90-min) liquid chromatography (LC) coupled with on-line linear ion trap quadropole mass spectrometry (LTQ MS) for identification and analysis. RESULTS: LTQ MS identified 481 high- to low-abundance proteins. Relative differences in patterns of BAL fluid proteins in smokers compared with never-smokers were observed in pooled and individual samples as well as by 2-dimensional gel analysis. Gene ontology categorization of all annotated proteins showed a wide spectrum of molecular functions and biological processes. CONCLUSIONS: The described method provides comprehensive qualitative proteomic analysis of BAL fluid protein expression from never-smokers and from smokers at risk of developing chronic obstructive pulmonary disease. Many of the proteins identified had not been detected in previous studies of BAL fluid; thus, the use of LC-tandem MS with LTQ may provide new information regarding potentially important patterns of protein expression associated with lifelong smoking.

Bronchoalveolar Lavage Fluid↗

Application of string kernels in protein sequence classification.

INTRODUCTION: The production of biological information has become much greater than its consumption. The key issue now is how to organise and manage the huge amount of novel information to facilitate access to this useful and important biological information. One core problem in classifying biological information is the annotation of new protein sequences with structural and functional features. METHOD: This article introduces the application of string kernels in classifying protein sequences into homogeneous families. A string kernel approach used in conjunction with support vector machines has been shown to achieve good performance in text categorisation tasks. We evaluated and analysed the performance of this approach, and we present experimental results on three selected families from the SCOP (Structural Classification of Proteins) database. We then compared the overall performance of this method with the existing protein classification methods on benchmark SCOP datasets. RESULTS: According to the F1 performance measure and the rate of false positive (RFP) measure, the string kernel method performs well in classifying protein sequences. The method outperformed all the generative-based methods and is comparable with the SVM-Fisher method. DISCUSSION: Although the string kernel approach makes no use of prior biological knowledge, it still captures sufficient biological information to enable it to outperform some of the state-of-the-art methods.

Algorithms↗

Analysis of expressed sequence tags from Brassica rapa L. ssp. pekinensis.

Non-redundant expressed sequence tags (ESTs) were generated from six different organs at various developmental stages of Chinese cabbage, Brassica rapa L. ssp. pekinensis. Of the 1,295 ESTs, 915 (71%) showed significantly high homology in nucleotide or deduced amino acid sequences with other sequences deposited in databases, while 380 did not show similarity to any sequences. Briefly, 598 ESTs matched with proteins of identified biological function, 177 with hypothetical proteins or non-annotated Arabidopsis genome sequences, and 140 with other ESTs. About 82% of the top-scored matching sequences were from Arabidopsis or Brassica, but overall 558 (43%) ESTs matched with Arabidopsis ESTs at the nucleotide sequence level. This observation strongly supports the idea that gene-expression profiles of Chinese cabbage differ from that of Arabidopsis, despite their genome structures being similar to each other. Moreover, sequence analyses of 21 Brassica ESTs revealed that their primary structure is different from those of corresponding annotated sequences of Arabidopsis genes. Our data suggest that direct prediction of Brassica gene expression pattern based on the information from Arabidopsis genome research has some limitations. Thus, information obtained from the Brassica EST study is useful not only for understanding of unique developmental processes of the plant, but also for the study of Arabidopsis genome structure.

Arabidopsis↗

Visualization of biochemical networks in living cells.

Functional annotation of novel genes can be achieved by detection of interactions of their encoded proteins with known proteins followed by assays to validate that the gene participates in a specific cellular function. We report an experimental strategy that allows for detection of protein interactions and functional assays with a single reporter system. Interactions among biochemical network component proteins are detected and probed with stimulators and inhibitors of the network. In addition, the cellular location of the interacting proteins is determined. We used this strategy to map a signal transduction network that controls initiation of translation in eukaryotes. We analyzed 35 different pairs of full-length proteins and identified 14 interactions, of which five have not been observed previously, suggesting that the organization of the pathway is more ramified and integrated than previously shown. Our results demonstrate the feasibility of using this strategy in efforts of genomewide functional annotation.

Animals↗

A categorization approach to automated ontological function annotation.

Automated function prediction (AFP) methods increasingly use knowledge discovery algorithms to map sequence, structure, literature, and/or pathway information about proteins whose functions are unknown into functional ontologies, typically (a portion of) the Gene Ontology (GO). While there are a growing number of methods within this paradigm, the general problem of assessing the accuracy of such prediction algorithms has not been seriously addressed. We present first an application for function prediction from protein sequences using the POSet Ontology Categorizer (POSOC) to produce new annotations by analyzing collections of GO nodes derived from annotations of protein BLAST neighborhoods. We then also present hierarchical precision and hierarchical recall as new evaluation metrics for assessing the accuracy of any predictions in hierarchical ontologies, and discuss results on a test set of protein sequences. We show that our method provides substantially improved hierarchical precision (measure of predictions made that are correct) when applied to the nearest BLAST neighbors of target proteins, as compared with simply imputing that neighborhood's annotations to the target. Moreover, when our method is applied to a broader BLAST neighborhood, hierarchical precision is enhanced even further. In all cases, such increased hierarchical precision performance is purchased at a modest expense of hierarchical recall (measure of all annotations that get predicted at all).

Computational Biology↗

LigProf: a simple tool for in silico prediction of ligand-binding sites.

With the increasing amount of data provided by both high-throughput sequencing and structural genomics studies, there is a growing need for tools to augment functional predictions for protein sequences. Broad descriptions of function can be provided by establishing the presence of protein domains associated with a particular function. To extend the domain-based annotation, LigProf provides predictions of potential ligands that bind to a protein, as well as critical residues that stabilize ligands. A P-value statistic for estimating the significance of motif occurrence is provided for all sites. Although the usefulness of the method will rise with increasing numbers of crystallographically solved molecules deposited in the PDB database, we show that it can already be applied successfully to the highly represented ligand-bound protein kinase domains of viral and human origin. The LigProf webserver is freely available at: http://www.cropnet.pl/ligprof . At present, LigProf descriptors annotate and extend major protein families from the PfamA database.

Binding Sites↗

A putative novel alpha/beta hydrolase ORFan family in Bacillus.

A large number of sequences in each newly sequenced genome correspond to lineage and species-specific proteins, also known as ORFans. Amongst these ORFans, a large number are sequences with unknown structures and functions. We have identified a family of sequences, annotated as hypothetical proteins, which are specific to Bacillus and have carried out a computational study aimed at characterizing this family. Fold-recognition methods predict that these sequences belong to the alpha/beta hydrolase fold. We suggest possible catalytic triads for the ORFans and propose a hypothesis regarding the possible families within the alpha/beta hydrolase superfamily to which they may belong.

Amino Acid Sequence↗

Methods to map protein interactions in mammalian cells: different tools to address different questions.

In the post-genome era, functional annotation of the predicted gene-sets will be one of the most important upcoming challenges. So-called interactome analysis positions a protein in its subcellular environment by mapping its interaction partners. Such interaction maps are essential for an accurate insight into protein function since many cellular processes are organised to operate in protein complexes. These assemblies have dynamic structures and can interact with each other, two properties which are often controlled by regulated protein expression and modification. Various methods exist to unravel protein interaction circuitries, which can be roughly divided into biochemical and genetic strategies. In this review we focus on the different strategies to study protein-protein interactions in living mammalian cells. Recently developed analytical and screening methods are also addressed.

Animals↗

Use of search algorithms to define specificity in Rab GTPase domain function.

The continuing explosion of sequencing data has inspired a corresponding effort in the annotation and classification of protein families. Within a particular protein family, however, individual members may have distinct functions, although they share a common fold and broadly defined physiological role. Rab GTPases are the largest subfamily of the Ras superfamily, yet from early in their discovery, it was apparent that each Rab protein has a unique subcellular localization and regulates a particular stage(s) membrane traffic. To gain insight into the contribution of individual residues to unique protein functions a general strategy is outlined. This method should allow the cell and molecular biologist with no specialist expertise to implement an algorithm that makes use of a combination of experimental and phylogenetic data. The algorithm is applicable to the analysis of any protein domain and here is illustrated with the analysis of residues contributing to the individual functions of a pair of Rab GTPases.

Algorithms↗

LocustDB: a relational database for the transcriptome and biology of the migratory locust (Locusta migratoria).

BACKGROUND: The migratory locust (Locusta migratoria) is an orthopteran pest and a representative member of hemimetabolous insects for biological studies. Its transcriptomic data provide invaluable information for molecular entomology and pave a way for the comparative research of other medically, agronomically, and ecologically relevant insects. We developed the first transcriptomic database of the locust (LocustDB), building necessary infrastructures to integrate, organize, and retrieve data that are either currently available or to be acquired in the future. DESCRIPTION: LocustDB currently hosts 45,474 high-quality EST sequences from the locust, which were assembled into 12,161 unigenes. It, through user-friendly web interfaces, allows investigators to freely access sequence data, including homologous/orthologous sequences, functional annotations, and pathway analysis, based on conserved orthologous groups (COG), gene ontology (GO), protein domain (InterPro), and functional pathways (KEGG). It also provides information from comparative analysis based on data from the migratory locust and five other invertebrate species, including the silkworm, the honeybee, the fruitfly, the mosquito and the nematode. The website address of LocustDB is http://locustdb.genomics.org.cn/. CONCLUSION: LocustDB starts with the first transcriptome information for an orthopteran and hemimetabolous insect and will be extended to provide a framework for incorporating in-coming genomic data of relevant insect groups and a workbench for cross-species comparative studies.

Animals↗