Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Inferring sub-cellular localization through automated lexical analysis.

MOTIVATION: The SWISS-PROT sequence database contains keywords of functional annotations for many proteins. In contrast, information about the sub-cellular localization is available for only a few proteins. Experts can often infer localization from keywords describing protein function. We developed LOCkey, a fully automated method for lexical analysis of SWISS-PROT keywords that assigns sub-cellular localization. With the rapid growth in sequence data, the biochemical characterisation of sequences has been falling behind. Our method may be a useful tool for supplementing functional information already automatically available. RESULTS: The method reached a level of more than 82% accuracy in a full cross-validation test. Due to a lack of functional annotations, we could infer localization for fewer than half of all proteins in SWISS-PROT. We applied LOCkey to annotate five entirely sequenced proteomes, namely Saccharomyces cerevisiae (yeast), Caenorhabditis elegans (worm), Drosophila melanogaster (fly), Arabidopsis thaliana (plant) and a subset of all human proteins. LOCkey found about 8000 new annotations of sub-cellular localization for these eukaryotes.

Abstracting and Indexing↗

Genome-wide association study of estimated glomerular filtration rate using repeated measurements in the Taiwan Biobank.

BACKGROUND: Chronic kidney disease (CKD) is a major global public health issue, with genetic factors playing a significant role in kidney function. Although genome-wide association studies (GWAS) have identified numerous loci associated with estimated glomerular filtration rate (eGFR), most studies relied on a single time-point measurement, which limits the capacity to account for within-individual measurement variability. METHODS: We performed a repeated-measurement GWAS in the prospective Taiwan Biobank (Taiwanese ancestry; n = 25,004) using two repeated creatinine-based eGFR measurements. Repeated eGFR values were analyzed using a linear mixed-effects model with a subject-specific random intercept and time-varying covariates, providing a more precise estimate of eGFR level. Identified loci underwent functional annotation (expression quantitative trait locus, deleteriousness prediction, and epigenetic markers) and were compared with results from a single-measurement GWAS. RESULTS: Six loci associated with eGFR were identified, including four previously reported regions (1q22, 4q21.1, 11p14.1, and 17q21.2) and two additional loci (6p21.32 and 15q24.2). Functional annotation implicated several candidate genes-such as MUC1/EFNA1, SHROOM3, HLA-DQB1, MPPED2, NRG4, and PGAP3/FBXL20-in the regulation of kidney function. CONCLUSION: Incorporating repeated eGFR measurements into GWAS may improve phenotypic precision for identifying genetic associations with kidney function. This study identified eGFR-associated loci and biologically plausible candidate genes in a Taiwanese population, which require further replication and functional validation.

Chronic kidney disease↗

Visualization of biochemical networks in living cells.

Functional annotation of novel genes can be achieved by detection of interactions of their encoded proteins with known proteins followed by assays to validate that the gene participates in a specific cellular function. We report an experimental strategy that allows for detection of protein interactions and functional assays with a single reporter system. Interactions among biochemical network component proteins are detected and probed with stimulators and inhibitors of the network. In addition, the cellular location of the interacting proteins is determined. We used this strategy to map a signal transduction network that controls initiation of translation in eukaryotes. We analyzed 35 different pairs of full-length proteins and identified 14 interactions, of which five have not been observed previously, suggesting that the organization of the pathway is more ramified and integrated than previously shown. Our results demonstrate the feasibility of using this strategy in efforts of genomewide functional annotation.

Animals↗

SCORPION, a molecular database of scorpion toxins.

Increasing interest in the studies of toxins and the requirements for better structural and functional annotations have created a need for improved data management in the field of toxins. The molecular database, SCORPION, contains more than 200 entries of fully referenced scorpion toxin data including primary sequences, three-dimensional structures, structural and functional annotations of scorpion toxins along with relevant literature references. SCORPION has a set of search tools that allow users to extract data and perform specific queries. These entries have been compiled from public databases and literature, cleaned of errors and enriched with additional structural and functional information. The grouping of scorpion toxins provides a basis for extending and clarifying the existing structural and functional classifications. The bioinformatics modules in SCORPION facilitate analyses aimed at classification of scorpion toxins and identification of sequence patterns associated with specific structural or functional properties of scorpion toxins. The SCORPION database is accessible via the Internet at sdmc.krdl.org.sg:8080/scorpion.

Animals↗

Correlation between gene expression profiles and protein-protein interactions within and across genomes.

MOTIVATION: Function annotation of an unclassified protein on the basis of its interaction partners is well documented in the literature. Reliable predictions of interactions from other data sources such as gene expression measurements would provide a useful route to function annotation. We investigate the global relationship of protein-protein interactions with gene expression. This relationship is studied in four evolutionarily diverse species, for which substantial information regarding their interactions and expression is available: human, mouse, yeast and Escherichia coli. RESULTS: In E.coli the expression of interacting pairs is highly correlated in comparison to random pairs, while in the other three species, the correlation of expression of interacting pairs is only slightly stronger than that of random pairs. To strengthen the correlation, we developed a protocol to integrate ortholog information into the interaction and expression datasets. In all four genomes, the likelihood of predicting protein interactions from highly correlated expression data is increased using our protocol. In yeast, for example, the likelihood of predicting a true interaction, when the correlation is > 0.9, increases from 1.4 to 9.4. The improvement demonstrates that protein interactions are reflected in gene expression and the correlation between the two is strengthened by evolution information. The results establish that co-expression of interacting protein pairs is more conserved than that of random ones.

Animals↗

Protein interaction mapping in C. elegans using proteins involved in vulval development.

Protein interaction mapping using large-scale two-hybrid analysis has been proposed as a way to functionally annotate large numbers of uncharacterized proteins predicted by complete genome sequences. This approach was examined in Caenorhabditis elegans, starting with 27 proteins involved in vulval development. The resulting map reveals both known and new potential interactions and provides a functional annotation for approximately 100 uncharacterized gene products. A protein interaction mapping project is now feasible for C. elegans on a genome-wide scale and should contribute to the understanding of molecular mechanisms in this organism and in human diseases.

Animals↗

Functional discrimination of gene expression patterns in terms of the gene ontology.

The ever-growing amount of experimental data in molecular biology and genetics requires its automated analysis, by employing sophisticated knowledge discovery tools. We use an Inductive Logic Programming (ILP) learner to induce functional discrimination rules between genes studied using microarrays and found to be differentially expressed in three recently discovered subtypes of adenocarcinoma of the lung. The discrimination rules involve functional annotations from the Proteome HumanPSD database in terms of the Gene Ontology, whose hierarchical structure is essential for this task. While most of the lower levels of gene expression data (pre)processing have been automated, our work can be seen as a step toward automating the higher level functional analysis of the data. We view our application not just as a prototypical example of applying more sophisticated machine learning techniques to the functional analysis of genes, but also as an incentive for developing increasingly more sophisticated functional annotations and ontologies, that can be automatically processed by such learning algorithms.

Adenocarcinoma↗

Characterization of a draft chromosome-scale genome assembly for the mutton snapper, Lutjanus analis.

BACKGROUND: The mutton snapper (Lutjanus analis) is a reef fish commonly found in tropical waters of the Western Atlantic Ocean. Genomic studies of this species are needed to support conservation efforts and breeding programs. OBJECTIVE: Here, we report the development of a chromosome-scale reference assembly for the mutton snapper and conduct an initial comparative genomic analysis with other lutjanids. METHODS: The genome of one mutton snapper specimen was sequenced using PAC-Bio HiFi long reads and Illumina short reads. Contigs and scaffolds were assembled in the Flye pipeline and anchored using Hi-C proximity guided assembly. Gene prediction and functional annotations were obtained in AUGUSTUS and eggNOG-mapper, respectively. The mutton snapper genome was compared to those of other lutjanids to infer gene family evolution and chromosome synteny conservation. RESULTS: Assembly and polishing yielded 946 contigs and 926 scaffolds (N50 of 3.16 Mb, complete BUSCO score 98.1%) that were anchored using Hi-C scaffolding in 24 draft chromosomes. The anchored assembly featured a N50 of 42.47 Mb and contained 97.6% of the unanchored assembly length. The 24 mutton snapper chromosomes showed a one-to-one syntenic relationship with their counterparts in medaka, and other Lutjanids. AUGUSTUS predicted 29,023 genes, 24,335 of which (83.85%) could be functionally annotated. Gene family evolution analysis revealed 1,014 significantly expanded or contracted hierarchical ortholog groups in mutton snapper. Expansions and contractions were linked to several biological functions including growth, oocyte maturation, and response to exogenous stressors. CONCLUSION: The draft genome will be a valuable tool for forthcoming applied genomic studies of mutton snapper.

Animals↗

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals↗

Ulmus minor response to Dutch elm disease: de novo transcriptome assembly and annotation.

Dutch elm disease (DED), caused by Ophiostoma novo-ulmi (ONU), has devastated elm populations across Europe and North America since the 20th century. In this work, a de novo transcriptome assembly of Ulmus minor in response to ONU is presented. We used two DED-resistant genotypes, MDV2.3 and VAD2, and one DED-susceptible genotype, MDV1, to capture responses to ONU at four time points post-inoculation (6, 24, 72, and 144 hours). RNA from collected samples was isolated and sequenced producing 60.88 M 100 bp paired-end reads per sample. We performed a de novo transcriptome assembly combining data from the three genotypes. The assembly was functionally annotated and validated through differential gene expression analysis of the response. This dataset provides a valuable resource for studying molecular mechanisms of DED resistance in elms, contributing to broadening our understanding of tree immunity and facilitating potential applications in functional annotation of future genome assemblies.

Transcriptome↗

The cadherin superfamily database.

The cadherin superfamily is a large protein family with diverse structures and functions. Because of this diversity and the growing biological interest in cell adhesion and signaling processes, in which many members of the cadherin superfamily play a crucial role, it is becoming increasingly important to develop tools to manage, distribute and analyze sequences in this protein family. Current profile and motif databases classify protein sequences into a broad spectrum of protein superfamilies, however to provide a more specific functional annotation, the next step should include classification of subfamilies of these protein superfamilies. Here, we present a tool that classified greater than 90% of the proteins belonging to the cadherin superfamily found in the SWISS PROT database. Therefore, for most members of the cadherin superfamily, this tool can assist in adding more specific functional annotations than can be achieved with current profile and motif databases. Finally, the classification tool and the results of our analysis were integrated into a web-accessible database (http://calcium.uhnres. utoronto.ca/cadherin).

Amino Acid Motifs↗

A cross-species analysis of the rodent uterotrophic program: elucidation of conserved responses and targets of estrogen signaling.

Physiological, morphological, and transcriptional alterations elicited by ethynyl estradiol in the uteri of Sprague-Dawley rats and C57BL/6 mice were assessed using comparable study designs, microarray platforms, and analysis methods to identify conserved estrogen signaling networks. Comparative analysis identified 153 orthologous gene pairs that were positively correlated, suggesting conserved transcriptional targets important in uterine proliferation. Functional annotation for these responses were associated with angiogenesis, water and solute transport, cell cycle control, redox control, DNA replication, protein synthesis and transport, xenobiotic metabolism, cell-cell communication, energetics, and cholesterol and fatty acid regulation. The identification of conserved temporal expression patterns of these orthologs provides experimental support for the transfer of functional annotation from mouse orthologs to 44 previously unannotated rat expressed sequence tags based on their homology and co-expression patterns. The identification of comparable temporal phenotypic responses linked to related gene expression profiles demonstrates the ability of systematic comparative genomic assessments to elucidate important conserved mechanisms in rodent estrogen signaling during uterine proliferation.

Animals↗

Phage bioinformatics tools: a review of computational approaches for bacteriophage research.

Rising clinical interest in phage therapy and the exponential growth of metagenomic sequence catalogues have driven a rapid expansion of bacteriophage bioinformatics. More than 80 dedicated tools, mostly published since 2020, now span identification, assembly, annotation, taxonomy, lifestyle prediction, defence-system detection, and host prediction. Aimed at experienced practitioners and developers, this review synthesizes the field through the lens of three successive computational paradigms: sequence homology, bounded by database completeness; machine learning, constrained by labelled training data; and foundation models, which now achieve Matthews correlation coefficients above 0.95 in identification tasks and, through structure-informed prediction, raise functional annotation to over half of phage genes. Furthermore, we map the upstream components, namely, gene callers, homology engines, protein language models, and structural search tools, that underpin most downstream pipelines, exposing shared infrastructure and ecosystem-level fragility when dependencies change. To translate this into practice, we propose web-based and command-line reference workflows calibrated to user expertise and sample types. Finally, we set an agenda for the next wave of tool development. Roughly half of phage genes still resist functional annotation despite structural methods; no broadly generalizable strain-level host predictor exists for phage therapy; varying true-positive rates (0%-97%) underscore the absence of standardized community benchmarks analogous to Critical Assessment of Structure Prediction or Critical Assessment of Metagenome Interpretation. As generative genome models begin designing synthetic phages, progress will depend less on producing standalone tools than on rigorous evaluation, interoperable infrastructure, and clinically meaningful prediction targets.

Computational Biology↗

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Artificial Intelligence↗

Molecular dynamics of the Shewanella oneidensis response to chromate stress.

Temporal genomic profiling and whole-cell proteomic analyses were performed to characterize the dynamic molecular response of the metal-reducing bacterium Shewanella oneidensis MR-1 to an acute chromate shock. The complex dynamics of cellular processes demand the integration of methodologies that describe biological systems at the levels of regulation, gene and protein expression, and metabolite production. Genomic microarray analysis of the transcriptome dynamics of midexponential phase cells subjected to 1 mm potassium chromate (K(2)CrO(4)) at exposure time intervals of 5, 30, 60, and 90 min revealed 910 genes that were differentially expressed at one or more time points. Strongly induced genes included those encoding components of a TonB1 iron transport system (tonB1-exbB1-exbD1), hemin ATP-binding cassette transporters (hmuTUV), TonB-dependent receptors as well as sulfate transporters (cysP, cysW-2, and cysA-2), and enzymes involved in assimilative sulfur metabolism (cysC, cysN, cysD, cysH, cysI, and cysJ). Transcript levels for genes with annotated functions in DNA repair (lexA, recX, recA, recN, dinP, and umuD), cellular detoxification (so1756, so3585, and so3586), and two-component signal transduction systems (so2426) were also significantly up-regulated (p < 0.05) in Cr(VI)-exposed cells relative to untreated cells. By contrast, genes with functions linked to energy metabolism, particularly electron transport (e.g. so0902-03-04, mtrA, omcA, and omcB), showed dramatic temporal alterations in expression with the majority exhibiting repression. Differential proteomics based on multidimensional HPLC-MS/MS was used to complement the transcriptome data, resulting in comparable induction and repression patterns for a subset of corresponding proteins. In total, expression of 2,370 proteins were confidently verified with 624 (26%) of these annotated as hypothetical or conserved hypothetical proteins. The initial response of S. oneidensis to chromate shock appears to require a combination of different regulatory networks that involve genes with annotated functions in oxidative stress protection, detoxification, protein stress protection, iron and sulfur acquisition, and SOS-controlled DNA repair mechanisms.

Anion Transport Proteins↗

Overview of BioCreAtIvE: critical assessment of information extraction for biology.

BACKGROUND: The goal of the first BioCreAtIvE challenge (Critical Assessment of Information Extraction in Biology) was to provide a set of common evaluation tasks to assess the state of the art for text mining applied to biological problems. The results were presented in a workshop held in Granada, Spain March 28-31, 2004. The articles collected in this BMC Bioinformatics supplement entitled "A critical assessment of text mining methods in molecular biology" describe the BioCreAtIvE tasks, systems, results and their independent evaluation. RESULTS: BioCreAtIvE focused on two tasks. The first dealt with extraction of gene or protein names from text, and their mapping into standardized gene identifiers for three model organism databases (fly, mouse, yeast). The second task addressed issues of functional annotation, requiring systems to identify specific text passages that supported Gene Ontology annotations for specific proteins, given full text articles. CONCLUSION: The first BioCreAtIvE assessment achieved a high level of international participation (27 groups from 10 countries). The assessment provided state-of-the-art performance results for a basic task (gene name finding and normalization), where the best systems achieved a balanced 80% precision / recall or better, which potentially makes them suitable for real applications in biology. The results for the advanced task (functional annotation from free text) were significantly lower, demonstrating the current limitations of text-mining approaches where knowledge extrapolation and interpretation are required. In addition, an important contribution of BioCreAtIvE has been the creation and release of training and test data sets for both tasks. There are 22 articles in this special issue, including six that provide analyses of results or data quality for the data sets, including a novel inter-annotator consistency assessment for the test set used in task 2.

Computational Biology↗

Metaproteomic Analysis to Assess the Impact of Storage Media on Human Gut Microbiome in Fecal Samples.

The human gut microbiome is a diverse community of microorganisms residing in the gastrointestinal tract. The storage condition of fecal samples may impact the taxonomic and protein compositions of microbiomes in these samples. Here, we performed a mass spectrometry-based metaproteomic study to assess the impact of storage media on human gut microbiome in fecal samples. We evaluated FDA-authorized OMNIgene&#xb7;GUT (OG), phosphate-buffered saline (PBS), and RNALater (RNAL) buffers and identified 38,185 microbial peptides corresponding to 7348 microbial proteins, which matched 16 phyla, 20 classes, 50 orders, 104 families, 332 genera, and 453 species. We found a high similarity among the fecal microbiomes preserved in OG, PBS, and RNAL in terms of the identification of proteins, taxa, and functional annotations. Both alpha and beta diversity suggested the high similarity among samples stored in the three media. Nonetheless, we also found some notable differences among buffers regarding the abundances of a few taxon groups. A partial human proteome (over 400 proteins) was identified in the fecal samples, with most of these proteins associated with the membrane and extracellular regions. The findings indicate the similarity among microbiomes in the fecal samples stored in OG, PBS, and RNAL regarding proteome profile, taxa, and functional capacity. SUMMARY: This study thoroughly analyzed and compared the metaproteomes of fecal samples preserved at -80&#xb0;C in PBS, RNALater, and OMNIgene&#xb7;GUT Dx buffers, offering novel insights into the effectiveness of these buffers in maintaining the stability and composition of the human gut microbiome. We found a high similarity in the identification and quantification of proteins, taxa, and functional annotations across the three buffers, with notable quantitative differences highlighting subtle yet important variations in preservation efficacy. The unique datasets and findings could offer valuable revelations into the impact of fecal sample preservation on translational and clinical analyses of the human gut microbiome.

Humans↗

Struct2net: integrating structure into protein-protein interaction prediction.

UNLABELLED: This paper presents a framework for predicting protein-protein interactions (PPI) that integrates structure-based information with other functional annotations, e.g. GO, co-expression and co-localization, etc., Given two protein sequences, the structure-based interaction prediction technique threads these two sequences to all the protein complexes in the PDB and then chooses the best potential match. Based on this match, structural information is incorporated into logistic regression to evaluate the probability of these two proteins interacting. This paper also describes a random forest classifier which can effectively combine the structure-based prediction results and other functional annotations together to predict protein interactions. Experimental results indicate that the predictive power of the structure-based method is better than many other information sources. Also, combining the structure-based method with other information sources allows us to achieve a better performance than when structure information is not used. We also tested our method on a set of approximately 1000 yeast genes and, interestingly, the predicted interaction network is a scale-free network. Our method predicted some potential interactions involving yeast homologs of human disease-related proteins. SUPPLEMENTARY INFORMATION: http://theory.csail.mit.edu/struct2net

Algorithms↗