Search PubMedSearch

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Chromosome-level assembly and annotation of the yellow-shelled fish (Barbodes Wynaadensis).

Barbodes wynaadensis, a unique cyprinid species native to Yunnan Province in China, stands out as an allotetraploid (AABB) fish with a complex evolutionary history. Leveraging a multi-platform sequencing strategy combining MGI short-read, PacBio long-read, and Hi-C scaffolding technologies, we assembled the first chromosome-level genome for B. wynaadensis. The final assembled genome spans 1.76 Gb in length with a contig N50 of 33.53 Mb, demonstrating high assembly continuity. Hi-C scaffolding enabled the reconstruction of 50 pseudochromosomes, representing 99.94% of the total genome assembly. Genome annotation identified 46,121 protein-coding genes, with a functional annotation rate of 99.76%. Repetitive elements constituted 48.26% of the genomic sequences, including lineage-specific expansions of DNA transposons (29.26%) and LTRs (6.36%). This high-quality assembly resolves challenges in polyploid genome reconstruction and provides a critical resource for investigating Cyprinidae evolution, particularly subgenome divergence and adaptation. The dataset also enables practical applications, such as molecular marker development for population monitoring, supporting conservation efforts for this threatened endemic species amid habitat degradation in the Nujiang River basin.

Animals

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats

On the state of protein function prediction: a report on the fourth CAFA challenge.

BACKGROUND: The Critical Assessment of Functional Annotation (CAFA) is a community effort held to understand the field of computational protein function prediction. Every three years, since 2010, the organizers initiate an experiment to collect function predictions on a large set of proteins and then evaluate the performance of predicting methods on a subset of proteins that have accumulated experimental annotations between the submission deadline and the evaluation time. CAFA provides an independent and rigorous assessment of the current state of the art, thus leveling the playing field, highlighting successes, revealing bottlenecks, and offering a forum for the exchange of ideas in protein science. Here, we report the results of the fourth CAFA experiment (CAFA4). RESULTS: CAFA4 featured the participation of 148 methods from 70 research groups on a total of 46,205 unique proteins over a 5-year annotation accumulation phase, the longest in any CAFA. In a comparison across CAFA2-CAFA4 methods, the prediction of Gene Ontology (GO) terms has clearly improved across all three GO aspects and traditional evaluation settings. While not achieving the first rank, several CAFA2 and CAFA3 methods featured in the top ten methods in many evaluations, suggesting that earlier methods still hold relevance. The performance is weaker in the newly introduced "partial knowledge" evaluation category (proteins with experimental annotations before submission deadline that gained additional annotations in the same GO aspect during the annotation accumulation phase), highlighting the need for a new class of methods. The rankings of the methods were stable over the years in traditional evaluation settings, but less so in the new partial knowledge evaluation. Overall, the field continues to progress with some influx of new participants. Sustained efforts will be necessary to substantially advance it.

Journal Article

Chromosome-Level Genome Assembly of Solanum carolinense.

Horsenettle (Solanum carolinense L.) is a noxious weed widely distributed across North America and increasingly invasive in other regions. Its strong environmental adaptability, complex defense strategies, and distinctive reproductive traits make it an important model for studying plant-herbivore coevolution. However, the absence of high-quality genomic resources has limited deeper investigation into its adaptive evolutionary mechanisms. In this study, we generated a chromosome-level reference genome assembly for S. carolinense using an integrated approach combining PacBio HiFi long-read sequencing, Illumina second-generation sequencing, and Hi-C chromatin interaction scaffolding. The final genome assembly had a total length of 915.40 Mb, with a contig N50 of 51.06 Mb and a scaffold N50 of 73.17 Mb; 96.05% of the sequences were successfully anchored onto 12 pseudochromosomes. The genome was characterized by a high proportion of repetitive sequences (73.64%) and substantial heterozygosity (1.13%), consistent with a highly repetitive and moderately high heterozygous genome. BUSCO analysis indicated that the chromosome-level genome assembly of S. carolinense reached a completeness score of 94.8%. A total of 32,206 protein-coding genes were annotated, of which 97.95% received functional annotations. The evaluation of the annotated protein-coding gene set returned a completeness value of 94.9%. This reference genome provides a valuable resource for advancing research on the adaptive evolution of weedy Solanaceae species, supports the development of more effective management strategies for this troublesome species, and offers a technical reference for assembling other highly heterozygous weed genomes.

Solanum carolinense

Enhanced identification of key bacterial motility genes via a cross-species genomic hybrid feature machine learning approach.

Efficient and accurate identification of functional genes is critical to biological research, yet traditional single-species approaches are often limited by low efficiency. Previously, we established a novel method for identifying key genes using cross-species protein domain features and machine learning. However, the high multiplicity of gene members associated with specific domains creates a substantial workload for subsequent experimental validation. To address this, this study proposes an enhanced approach that integrates EggNOG-based protein sequence annotation with domain analysis. Unannotated sequences are subsequently analyzed for protein domains, generating a comprehensive "direct gene annotation plus domain" hybrid feature matrix. While the hybrid matrix model yielded comparable predictive accuracy, it significantly enhanced feature resolution: the top 50 predicted features were all known motility-related genes or domains. Furthermore, among the top 100 ranked features, 58 are confirmed to be directly related to motility based on experimental evidence. Although strict genus-level control still yielded 51 confirmed features, excessive taxonomic restriction drastically reduces the number of training genomes, which may paradoxically impair identification efficiency. These results demonstrate that the new method effectively reduces the subsequent experimental workload and enables high-throughput identification of functional genes in a single analysis. With accuracy and efficiency far exceeding those of existing single-species identification methods, it provides a highly efficient solution for mining key genes underlying other complex bacterial phenotypes.

Machine Learning

Selection system for genes encoding nuclear-targeted proteins.

Nuclear proteins have essential roles in cell proliferation and differentiation. We have developed a yeast selection system-the nuclear transportation trap (NTT)-to identify genes encoding nuclear transport signals. Both unknown and previously identified nuclear localization signals were identified from a human fetal brain cDNA library. The majority (75%) of the unknown proteins examined were exclusively localized to the nucleus in COS-7 cells. We propose that NTT is an efficient method for isolating cDNAs that encode nuclear targeted proteins that can be applied to the retrieval of novel nuclear proteins and to annotate gene function.

Amino Acid Sequence

Genomic and transcriptomic characterization of genes expressed at 20 MPa by the marine actinobacterium Kocuria flava.

A marine hydrocarbonoclastic actinobacterium Kocuria flava IOS11 was isolated from 3500 m deep-sea water of the Indian Ocean. The isolate efficiently degraded phenanthrene (250 mg/L) achieving 82 and 98% of degradation at 0.1 MPa and 20 MPa, respectively within a period of 5 days. Whole genome, transcriptomee and metabolomic analysis elucidated its phenanthrene biodegradation efficiency under in situ deep-sea conditions. The genome sequence comprises 3.47 Mb distributed across 88 scaffolds with a high GC content of 74.30%. The genome analysis encoded 3126 genes including 3052 protein coding sequences with functional annotation identifying a broad array of genes associated with PAHs degradation, environmental stress adaptation, biosurfactant and siderophore synthesis. Transcriptome profiling under 0.1 and 20 MPa conditions with phenanthrene as a sole carbon source revealed enhanced expression of hydrocarbon degrading genes, transporters, biosurfactant associated enzymes and stress responsive genes including integrases, DNA repair protein Rad, alanine ligase, heat and cold shock proteins under high pressure conditions underscoring the deep-sea adaptation capabilities of the strain. The degradation pathway of phenanthrene was proposed through integrated genome, transcriptome and metabolomic analysis. These studies provided K. flava IOS11 as a metabolically versatile and pressure adapted bacterium with promising potential for bioremediation application in extreme marine environment.

Transcriptome

Large language models improve annotation of prokaryotic viral proteins.

Viral genomes are poorly annotated in metagenomic samples, representing an obstacle to understanding viral diversity and function. Current annotation approaches rely on alignment-based sequence homology methods, which are limited by the paucity of characterized viral proteins and divergence among viral sequences. Here we show that protein language models can capture prokaryotic viral protein function, enabling new portions of viral sequence space to be assigned biologically meaningful labels. When applied to global ocean virome data, our classifier expanded the annotated fraction of viral protein families by 29%. Among previously unannotated sequences, we highlight the identification of an integrase defining a mobile element in marine picocyanobacteria and a capsid protein that anchors globally widespread viral elements. Furthermore, improved high-level functional annotation provides a means to characterize similarities in genomic organization among diverse viral sequences. Protein language models thus enhance remote homology detection of viral proteins, serving as a useful complement to existing approaches.

Viral Proteins

Large-scale functional annotation establishes a reference framework for human LRRK2 variants.

Pathogenic variants in leucine-rich repeat kinase 2 (LRRK2)1are among the most frequent monogenic causes of Parkinson's disease (PD)2 and act through a gain-of-function mechanism of increased kinase activity. LRRK2-targeted therapies are in clinical development, but interpretation of the rapidly expanding catalogue of rare LRRK2 variants remains a barrier to translation. Here, we present functionally annotated data on >350 LRRK2 coding variants using a standardized cellular assay with Rab10 phosphorylation as a readout of kinase activity and integrated these data with curated genetic and clinical annotations from the Movement Disorders Society Genetic Mutation Database (MDSGene). Variants differed in activation magnitude, ranging from modest increases (e.g., p.G2019S) to strongly activating substitutions such as p.Y1699C or p.L1795F. Activating variants occurred across the full length of LRRK2, although the largest effects clustered within the ROC-COR regulatory hub, where structural analysis identified subdomains forming an allosteric scaffold controlling kinase output. All known/established pathogenic variants showed increased activity, whereas benign and likely benign variants remained within the wild-type range. Functional effect sizes correlated with pathway activation in patient-derived immune cells, altogether providing a framework for ACMG-based variant interpretation in which kinase activation can support PS3 functional evidence for reclassification of variants.

Protein phosphorylation

Systematic Proteome Profiling of Maternal Plasma for Development of Preeclampsia Biomarkers.

Preeclampsia (PE) is a hypertensive disorder of pregnancy with various clinical symptoms. However, traditional markers for the disease including high blood pressure and proteinuria are poor indicators of the related adverse outcomes. Here, we performed systematic proteome profiling of plasma samples obtained from pregnant women with PE to identify clinically effective diagnostic biomarkers. Proteome profiling was performed using TMT-based liquid chromatography-mass spectrometry (LC-MS/MS) followed by subsequent verification by multiple reaction monitoring (MRM) analysis on normal and PE maternal plasma samples. Functional annotations of differentially expressed proteins (DEPs) in PE were predicted using bioinformatic tools. The diagnostic accuracies of the biomarkers for PE were estimated according to the area under the receiver-operating characteristics curve (AUC). A total of 1307 proteins were identified, and 870 proteins of them were quantified from plasma samples. Significant differences were evident in 138 DEPs, including 71 upregulated DEPs and 67 downregulated DEPs in the PE group, compared with those in the control group. Upregulated proteins were significantly associated with biological processes including platelet degranulation, proteolysis, lipoprotein metabolism, and cholesterol efflux. Biological processes including blood coagulation and acute-phase response were enriched for down-regulated proteins. Of these, 40 proteins were subsequently validated in an independent cohort of 26 PE patients and 29 healthy controls. APOM, LCN2, and QSOX1 showed high diagnostic accuracies for PE detection (AUC >0.9 and p&#xa0;<&#xa0;0.001, for all) as validated by MRM and ELISA. Our data demonstrate that three plasma biomarkers, identified by systematic proteomic profiling, present a possibility for the assessment of PE, independent of the clinical characteristics of pregnant women.

Humans

Chromosome-level genome assembly with telomeric repeats at scaffold ends for Rhabdosargus sarba.

Rhabdosargus sarba, the goldlined seabream, is a euryhaline marine fish of great aquaculture potential. Genome sequencing and assembly of R. sarba was carried utilizing a multi-platform sequencing strategy that included long-read sequencing (PacBio HiFi), short-read sequencing (Illumina), and chromatin interaction mapping (Hi-C). The final genome assembly size after scaffolding was 764.59&#x2009;Mb in 31 scaffolds with an N50 length of 33.98&#x2009;Mb. Repeat profiling of primary assembly showed that 28.71% of the genome comprises of repeat elements. Gene prediction utilising the evidence from ab initio prediction and transcriptome data revealed 26,913 protein encoding genes and functional annotation and pathway analysis showed their participation in 332 pathways. This genome is an excellent resource for future research on genetic improvement and molecular breeding programmes for R. sarba.

Animals

Glycosyltransferases in SWISS-PROT.

SWISS-PROT is a curated protein sequence database with a high level of annotation (such as description of the function of a protein, its domain structure, post-translational modification, variants, etc), a minimal level of redundancy and a high level of integration with other databases. An ongoing project is to maintain the glycosyltransferase family of enzymes with comprehensive annotation and documentation in the SWISS-PROT database and to represent the most recent research developments.

Amino Acid Sequence

Assessment of the impact of manual curation in BioCyc.

INTRODUCTION: BioCyc is an extensive collection of databases of genomic and pathway information for microorganisms and model eukaryotes. These organismal databases integrate diverse biological data by combining computationally inferred information, data imported from other databases, and, for selected organisms, literature-based manual curation. This study investigates the magnitude and significance of annotation changes performed during the curation of 10 prokaryotic genomes to better understand the rate of erroneous annotations and the value of BioCyc curation. METHODS: We identified curation changes by finding cases where the annotation of the protein at the start of the curation process differed from its annotation at the end of the process. RESULTS: We found that across a sample of curated databases (n = 10), the annotation of 6,753, or 25.6% of the proteins in the pooled protein dataset (n = 26,126) were modified. Assessment of considerable sampling fractions of these proteins found that a median of 62% (mean of 52.9%) represented functionally informative name changes, rather than stylistic annotation changes. These results were then extrapolated to total proteins with name changes with uncertainty quantified via finite population correction, indicating that most Tier 2 Biocyc PGDBs received hundreds of functionally informative name changes during manual curation. On average 363, or13% (&#xb1;5.4% SD) of the proteins encoded in each genome received functionally informative annotation changes, ranging from 5.3% (Streptococcus pneumoniae D39V) to 22.7% (Staphylococcus aureus NCTC 8325). DISCUSSION: These findings demonstrate a substantial improvement in the accuracy of manually curated BioCyc databases compared with automated annotation pipelines. This result is particularly impactful as the rate of downstream propagation of erroneous annotations across biological databases can significantly compromise scientific discovery.

annotation errors

Serum Proteomic Profiling Reveals Renin-Associated Immune and Cytoskeletal Dysregulation in Post-COVID-19 Condition Patients with Secondary Adrenal Insufficiency.

Post-COVID-19 condition (PCC) with secondary adrenal insufficiency (SAI) involves multiorgan dysfunction, potentially linked to renin-angiotensin-aldosterone system dysregulation. The molecular basis of renin-associated pathology remains unclear. Here, PCC+SAI patients were stratified by upright renin into low- (<38.8&#x202f;pg/mL) and high-renin (&#x2265;38.8&#x202f;pg/mL) groups. Clinical, endocrine, and proteomic analyses were performed. We found that high-renin patients showed increased BMI, lipids, renin, and aldosterone, but reduced aldosterone-to-renin ratio. Proteomic annalysis identified 20 differentially expressed proteins (DEPs), including 17 upregulated and 3 downregulated proteins in Ren-H patients. Functional annotation revealed that 15 DEPs were immune-related (e.g., APOC4, APOE, C4BPA, CFAH, CFHR3, PF4V, PLF4), while FLNA and COF1 represented cytoskeletal proteins. These DEPs were primarily involved in immune response, complement and coagulation cascades, and MAPK signaling pathways. Correlation analyses indicated that upright renin was positively correlated with complement-related proteins and platelet-derived immune factors, while cytoskeletal proteins (FLNA, COF1) showed positive associations with serum Na+ levels. Additionally, white blood cell and platelet counts were positively correlated with the majority of DEPs. In conclusion, exploratory proteomic analyses suggest that elevated upright renin in PCC+SAI may be associated with immune dysregulation, complement activation, and cytoskeletal remodeling, offering novel insights into the endocrine-immune interactions driving postviral sequelae.

Humans

Chromosome-Level Genome Assembly and Annotation of the Chinese Lizard Gudgeon (Saurogobio dabryi).

The Chinese lizard gudgeon (Saurogobio dabryi) is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in China, the lack of genomic resources has rendered the genetic breeding and conservation research. Here, we present the first chromosome-level genome assembly of S. dabryi using PacBio HiFi long reads, short reads, and Hi-C sequencing data. The final assembly reaches a total size of 1.09 Gb and Hi-C scaffolding anchors 99.55% of the assembled contigs onto 25 chromosomes, with a scaffold N50 reaching 43.15 Mb. The final genome assembly shows a BUSCO completeness of 98.39%. We annotated 659.55 Mb repetitive sequences and 26,036 protein-coding genes, 99.47% of which are functionally annotated. Comparative phylogenomic analysis clarifies the phylogenetic position of Saurogobio within Gobioninae. This high-quality genome provides a critical genetic basis for exploring cyprinid phylogeny, benthic adaptive evolution, genetic improvement, and conservation efforts of S. dabryi.

Saurogobio dabryi

Rare variant contribution to the heritability of coronary artery disease.

Whole genome sequences (WGS) enable discovery of rare variants which may contribute to missing heritability of coronary artery disease (CAD). To measure their contribution, we apply the GREML-LDMS-I approach to WGS of 4949 cases and 17,494 controls of European ancestry from the NHLBI TOPMed program. We estimate CAD heritability at 34.3% assuming a prevalence of 8.2%. Ultra-rare (minor allele frequency &#x2264;&#x2009;0.1%) variants with low linkage disequilibrium (LD) score contribute ~50% of the heritability. We also investigate CAD heritability enrichment using a diverse set of functional annotations: i) constraint; ii) predicted protein-altering impact; iii) cis-regulatory elements from a cell-specific chromatin atlas of the human coronary; and iv) annotation principal components representing a wide range of functional processes. We observe marked enrichment of CAD heritability for most functional annotations. These results reveal the predominant role of ultra-rare variants in low LD on the heritability of CAD. Moreover, they highlight several functional processes including cell type-specific regulatory mechanisms as key drivers of CAD genetic risk.

Humans

Shared genetic architecture and therapeutic targets across paediatric immune-mediated diseases.

OBJECTIVES: Paediatric-onset immune-mediated inflammatory diseases (IMIDs), including juvenile idiopathic arthritis and related rheumatic diseases, remain genetically undercharacterised. We aimed to define shared and category-specific genetic architecture across paediatric IMIDs, compare signals with adult IMIDs, and identify therapeutic opportunities. METHODS: We analysed 24 paediatric IMIDs classified as autoimmune, polygenic-autoinflammatory, mixed-pattern, or allergic. Genome-wide association analyses included 18,086 cases and 131,019 controls of European ancestry. We estimated single nucleotide polymorphism (SNP)-based heritability, genetic correlations, and polygenic overlap; performed subset-based meta-analysis; and conducted functional annotation, gene prioritisation, pathway and protein network analyses, adult-IMID comparison, and drug-target prioritisation. RESULTS: SNP-based heritability ranged from 28.9% for allergic IMIDs to 61.9% for autoimmune IMIDs. Genetic correlation and polygenic modelling supported partial sharing across categories with category-specific components. Meta-analysis identified 39 genome-wide significant loci outside the Major Histocompatibility Complex (MHC) region, including 15 previously unreported loci; 19 loci were shared between categories. Gene-prioritisation and protein interaction analyses identified a core MHC-centred antigen-presentation network, with category-enriched modules involving complement, innate/barrier pathways, epithelial biology, and type 2 immunity. Enriched pathways included nuclear factor &#x3ba;B signalling, T helper 17 related pathways, Janus kinase-signal transducer and activator of transcription signalling, programmed cell death protein 1/programmed death&#x2011;ligand 1, cytotoxic T&#x2011;lymphocyte associated protein 4 regulation, and osteoclast differentiation, several of which are relevant to rheumatic diseases. Paediatric IMIDs shared broad polygenic architecture with adult IMIDs, whereas top-ranked genes converged strongly with adult rheumatic diseases. Priority Index analysis identified 178 high-scoring genes, including 43 approved or investigational IMID drug targets. CONCLUSIONS: Paediatric-onset IMIDs share core pathways with adult forms but exhibit distinct genetic architecture shaped by age-specific immune and neurodevelopmental biology. These findings provide a genomic framework for paediatric precision medicine, guiding classification, risk prediction, and therapeutic development.

Humans

A conserved distal-tail helical extension defines a tailspike attachment architecture in Gram-negative siphophages.

Rapid growth of bacteriophage genome collections has outpaced functional annotation of tail-tip proteins, limiting comparative analysis of host-recognition structures. Starting from a shared distal-tail gene organization in the Salmonella phages 9NA and Jersey, I developed a morphogenetic bioinformatic framework integrating gene synteny, sequence comparison, profile hidden Markov model (HMM) screening, structural evidence, structure-aware searching, and AlphaFold modeling. Comparison with the experimentally characterized lambda and Sf11 tail assemblies identified a predominantly alpha-helical C-terminal extension of the distal-tail (DT) protein associated with tailspike attachment, termed the distal-tail helical extension (DT-helix). Screening 541,986 proteins from 5167 complete NCBI RefSeq tailed-phage genomes, followed by evidence-based evaluation of sequence, genomic context, and structural architecture, identified 165 curated DT-helical-extension-associated phages. Their DT proteins segregated into six sequence groups. In the four principal multi-member groups, cognate tailspikes showed group-specific conservation in proximal N-terminal regions but substantially greater downstream diversity, consistent with sequence constraint at the DT-tailspike attachment boundary. A complementary ProstT5/Foldseek search supported the established groups but revealed no convincing additional highly divergent family. Together with the experimentally characterized Sf11 attachment interface, these findings define a recurrent morphogenetic architecture linking conserved distal-tail scaffolds to more variable receptor-binding proteins across siphophages infecting Gram-negative bacteria. Although universal exchangeability is not established, the identified scaffold-receptor-binding boundaries provide a framework for molecular characterization and rational phage engineering. Accession-level information for the 165 curated phages is available through PhageTailDB.

Viral Tail Proteins