Search PubMedSearch

SEARCH · Search PubMed

Results for “Protein function prediction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

A chromosome-level assembly of the alpine snow alga Chloromonas typhlos.

Chloromonas typhlos is a cosmopolitan alpine snow alga distributed across continents, and its blooming accelerates snow melting by decreasing the amount of snow albedo. To elucidate the genetic traits underlying the adaptation of C. typhlos to the alpine habitat, we combined PacBio sequencing and Hi-C to generate a high-quality chromosome-level genome assembly (contig N50: 1.29 Mb; scaffold N50: 7.23 Mb) with 31 chromosomes and a genome size of 200.86 Mb. Repetitive elements constituted 11.05% of the genome, and 16,133 protein-coding genes were predicted, of which 82% were functionally annotated. This study provides a set of omics resources both for snow algae and the genus Chloromonas.

Snow

The Biobank Rare Variant consortium powers the discovery of rare genetic associations through global collaboration.

Rare coding variants can have large effects on disease risk and provide direct routes from human genetics to disease mechanisms and therapeutic targets, but their discovery is constrained by sample size, particularly for low-prevalence diseases. Here we establish the Biobank Rare Variant Analysis (BRaVa) consortium, a global rare variant association resource that integrates sequencing and linked health-record data from ten biobanks and cohorts comprising over 1.2 million individuals across diverse ancestries. We performed gene-based meta-analyses of rare coding variation across 33 clinical endpoints and 11 quantitative traits. Aggregating evidence across biobanks and ancestries identified 514 gene-trait associations, including 31 not previously reported in prior studies or curated association resources following systematic literature review. Notably, 36.1% of gene-level associations were undetectable in any individual biobank, and 91 emerged only through cross-ancestry meta-analysis, demonstrating that federated integration enables discovery beyond the reach of single cohorts. Similar gains were observed at the variant level, where 25.0% of phenotype-locus associations were detectable only through meta-analysis. Effect size estimates were correlated across ancestries with concordant directions of effect, supporting the generalizability of rare variant associations. The identified signals implicate pathways involved in transcriptional and epigenetic regulation, metabolism, vascular and epithelial biology, and immune function, highlighting rare coding variation as an engine for biological discovery across medical record phenotypes. For example, damaging variation in ANKRD12 implicates inflammatory transcriptional dysregulation in asthma and chronic obstructive pulmonary disease, and ultra-rare predicted loss-of-function variants in NAA15 link protein acetylation processes to type 2 diabetes risk. BRaVa establishes a scalable framework and freely available community resource for rare variant meta-analysis across global biobanks. Public release of gene- and variant-level association summary statistics provides a reference map of rare coding variant associations to support disease gene discovery, biological interpretation, and therapeutic target prioritization as sequencing-linked health-record resources continue to expand.

Journal Article

AnoEST: toward A. gambiae functional genomics.

Here, we present an analysis of 215,634 EST and cDNA sequences of a major vector of human malaria Anopheles gambiae structured into the AnoEST database. The expressed sequences are grouped into clusters using genomic sequence as template and associated with inferred functional annotation, including the following: corresponding Ensembl gene prediction, putative orthologous genes in other species, homology to known proteins, protein domains, associated Gene Ontology terms, and corresponding classification into broad GO-slim functional groups. AnoEST is a vital resource for interpretation of expression profiles derived using recently developed A. gambiae cDNA microarrays. Using these cDNA microarrays, we have experimentally confirmed the expression of 7961 clusters during mosquito development. Of these, 3100 are not associated with currently predicted genes. Moreover, we found that clusters with confirmed expression are nonbiased with respect to the current gene annotation or homology to known proteins. Consequently, we expect that many as yet unconfirmed clusters are likely to be actual A. gambiae genes. [AnoEST is publicly available at http://komar.embl.de, and is also accessible as a Distributed Annotation Service (DAS).].

Animals

Simulation of CRISPR/Cas9-mediated gene editing for the Vitellogenin gene in Apis mellifera.

CRISPR/Cas9 genome editing provides a powerful framework for interrogating gene function in Apis mellifera. Yet, empirical application remains challenging due to biological constraints, including haplodiploid genetics, narrow embryonic injection window, and the social rearing requirements that complicate functional validation. These constraints necessitate in silico pre-screening to maximize editing success before resource-intensive wet-lab implementation. Within the omnigenic framework, which distinguishes core regulatory genes from peripheral loci buffered by network effects, vitellogenin (Vg) represents an optimal target which is ancestrally dedicated to yolk provisioning; it has been co-opted to orchestrate diverse non-reproductive functions including longevity, stress resistance, immunity, and social behavior. We developed a computational pipeline to design a list of 57 and 56 candidate guide RNAs (gRNA) for targeted Vg knockout, evaluating candidate sites in both functional exons 2 and 3 based on structural accessibility and frameshift efficiency. Comparative analysis revealed complementary strengths in two top-best candidates from initial target pool of predicted gRNAs. The gRNA targeting exon 2 exhibits weaker secondary structure (ΔG = -0.25 kcal/mol versus -2.10 kcal/mol for exon 3), aligning with empirical evidence that sites with ΔG > -1.0 kcal/mol achieve 2-5 × higher Cas9 binding efficiency. This site yielded moderate frameshift frequency (77.8%; 61.9 percentile). Conversely, the predicted editing outcome for the gRNA targeting exon 3, despite stronger structural constraints, demonstrated superior functional disruption metrics demonstrating very high frameshift frequency (88.3%; 95.2 percentile), high in silico editing precision, minimal microhomology-mediated repair bias, and reproducible outcomes wherein nearly all predicted indels disrupt the coding sequence. Protein structure and domain analyses further predict that frameshift edits will generate a truncated protein missing all downstream functional domains. We recommend parallel empirical validation of both exon 2 and exon 3 targets to resolve the trade-off between structural accessibility (favoring higher editing rates) and frameshift efficacy (favoring complete loss-of-function). This dual-target strategy accommodates uncertainty in in vivo performance while maximizing the probability of generating informative phenotypes. Our in silico framework enables rational CRISPR design in non-model organisms by computationally balancing biophysical accessibility with functional impact, accelerating functional genomics in species where empirical optimization faces substantial biological constraints.

Animals

De novo chromatin remodelling variants in sporadic Chiari 1 malformation.

Chiari 1 malformation (CM1) is the most common congenital malformation of the human hindbrain. Although prior studies have implicated chromatin-remodeling genes in CM1, the de novo genetic architecture and underlying neurodevelopmental mechanisms remain incompletely defined. To investigate the molecular genetics of a novel familial form of CM1 linked with syringomyelia and tethered cord and determine whether rare, damaging de novo variants (DNVs) contribute to sporadic CM1 risk with gene- and pathway-level resolution, we performed whole-exome sequencing in an ultra-rare multigenerational family with CM1 and associated spinal pathology, and in the largest assembled trio-based cohort to date, comprising 1,585 proband-parent trios with sporadic, idiopathic CM1 (2017-2025). The comparison cohort included 1,798 unaffected control siblings. Clinical phenotyping was by systematic medical record review. Structural domain mapping, in silico modeling, and integration with single-cell transcriptomic data from developing human cerebellum was conducted to assess biological plausibility. A heterozygous loss-of-function variant in CHD3 segregated with CM1 and syringomyelia in a multigenerational family. In the trio-based cohort, rare protein-altering DNVs were significantly enriched across multiple chromodomain helicase DNA-binding (CHD) genes, including CHD1, CHD3, CHD4, and CHD8, exceeding gene-specific mutation expectations (protein-damaging variants: P = 1.3 × 10-9; predicted loss-of-function variants: P = 8.6 × 10-5). CHD1 contained two pathogenic DNVs (p.A999D and p.E984K). CHD4 (p.D744N, p.T1813P, and p.I1102T) and CHD8 (p.R1402X, p.R1472X, and p.R2035X) each contained three new DNVs. Variants clustered within conserved ATPase, helicase, and chromodomain regions essential for chromatin remodeling, and these patients frequently had comorbid developmental delay and related neurodevelopmental features. Single-cell transcriptomic analyses demonstrated enrichment in Purkinje cells and inhibitory neurons of midgestational cerebellum, where CHD gene products form a coherent chromatin-regulatory network. Rare, large-effect DNVs that disrupt chromatin-remodeling programs contribute to sporadic CM1, implicating genetically encoded dysregulation of cerebellar development as a central disease mechanism. Exome sequencing may complement surgical evaluation of children with sporadic CM1, particularly when accompanied by neurodevelopmental concerns, informing prognosis and family counseling.

de novo variants

A recurrent CCDC82 frameshift variant associated with syndromic neurodevelopmental disorder in a consanguineous Pakistani family.

BACKGROUND: Intellectual disabilities (IDs) are part of neurodevelopmental disorders (NDDs) and are genetically heterogeneous conditions characterized by impairments in cognition, learning, and adaptive functioning. Despite advances in gene discovery, many individuals, particularly those from understudied populations, remain without a molecular diagnosis. Recent reports implicate CCDC82 (HGNC: 26282) as an autosomal recessive ID gene, although the phenotypic spectrum and biological context remain incompletely defined. METHODS: Exome sequencing (ES) was performed in a consanguineous Pakistani family (PKMR06A) with four affected individuals presenting with moderate to severe ID. Variant segregation was confirmed by Sanger sequencing. In silico analyses, including pathogenicity prediction, protein structural modeling, and domain intolerance assessment, were used to evaluate the functional consequences of the identified variant. Spatiotemporal gene expression patterns were examined using bulk and single-cell human brain transcriptomic datasets. RESULTS: Clinically, affected individuals of family PKMR06A presented with early childhood global developmental delay, speech delay, hypotonia, gait abnormalities, spasticity, and mild facial dysmorphism. Genetic screening revealed a recurrent rare homozygous frameshift variant in CCDC82 (NM_024725.4): c.373del; p.(Asp125Ilefs*6), segregating with disease in all available affected individuals of the family. The identified c.373del variant was absent from the gnomAD database and was classified as pathogenic (PVS1, PM2, and PP1) based on ACMG/AMP criteria. The c.373del variant is predicted to introduce a premature termination codon, p.(Asp125Ilefs*6), leading to deletion of essential coiled-coil domains from the encoded protein, supporting a loss-of-function mechanism. In silico, transcriptomic analyses demonstrated preferential CCDC82 expression during prenatal human brain development, providing developmental context for the neurodevelopmental phenotype associated with the identified truncating variant. CONCLUSIONS: This study expands the mutational landscape of CCDC82 and provides additional clinical and molecular evidence supporting its role in autosomal recessive NDD. The findings reinforce the importance of CCDC82 in human neurodevelopment and highlight the value of genomic investigation in underrepresented populations.

Autosomal recessive

A machine learning-based predictive model for radiosensitivity in nasopharyngeal carcinoma utilizing serum proteomics.

BACKGROUND: Nasopharyngeal carcinoma (NPC) remains highly sensitive to radiotherapy; however, radioresistance in a subset of patients leads to local recurrence and distant metastasis. Serum proteomics provides a minimally invasive approach to capturing dynamic physiological changes, and machine learning enables efficient construction of predictive models. This study aimed to develop and validate a serum proteomics–based machine-learning model for predicting radiotherapy sensitivity in nasopharyngeal carcinoma (NPC). METHODS: Pretreatment serum samples from newly diagnosed NPC patients were analyzed using SELDI-TOF-MS. Differentially expressed proteins between radiosensitive and radioresistant groups were identified using limma. GO and KEGG analyses were performed to explore functional enrichment. Twelve machine-learning algorithms were used to construct predictive models, and the top-performing models were optimized through feature selection. A Random Forest model with seven features was identified as the optimal model. External validation was performed using an independent cohort with ELISA-quantified protein levels. Model performance was assessed using Receiver operating characteristic curve (ROC), calibration analysis, decision curve analysis (DCA), and 10-fold cross-validation. SHapley Additive exPlanations (SHAP) analysis was applied for model interpretability, and the final model was deployed via a ShinyAPP. RESULTS: A total of 96 differentially expressed proteins were identified, which involved multiple function and signaling pathways. The Random Forest model demonstrated the best predictive performance, achieving an area under the curve (AUC) of 0.963 in the training set and 0.975 in the validation set. Cross-validation yielded an average AUC of 0.965. DCA indicated high clinical utility across a broad threshold range, and calibration curves showed good model agreement. Seven proteins (PLXND1, GSR, PGD, PTPRC, OR2T29, ACTG2, CHAD) were selected as final features. SHAP analysis provided global and individual-level interpretability. A web-based tool was developed to facilitate clinical application. CONCLUSION: This study establishes a robust serum proteomics–based machine-learning model capable of accurately predicting radiotherapy sensitivity in NPC. The model offers clinical interpretability and practical implementation, supporting personalized radiotherapy decision-making.

Humans

Genomic insights of first varicella zoster clade9 strain: a potential silent surge in Pakistan.

The study presents the first-time detection of one of the rare clades (clade9 strain) of varicella zoster virus (VZV) from Pakistan. The next-generation sequencing confirmed wild-type clade9 strain through clade-specific markers at C5827A, T33722C, T33725C, T33728C, T38055C, G69424A, C87841T and T95241C and restriction profile of PstI+BgII+SmaI-. The rarely reported SNPs (22/134) were detected along with 12/42 rare amino-acid mutations. However, the mutations at C77Y, Q43H, D613E and A2V were predicted to be not-tolerated hence might affect protein function. The VZV (PV934234) strain clustered with clade9 strains upon phylogenetics. Thus, the first-time detection of clade9 raises concern of limited genomic surveillance of VZV in Pakistan. This necessitates genomic surveillance and continuous clinical vigilance in Pakistan to avoid any potential silent surge in the country.

Clade 9

Systematic Proteome Profiling of Maternal Plasma for Development of Preeclampsia Biomarkers.

Preeclampsia (PE) is a hypertensive disorder of pregnancy with various clinical symptoms. However, traditional markers for the disease including high blood pressure and proteinuria are poor indicators of the related adverse outcomes. Here, we performed systematic proteome profiling of plasma samples obtained from pregnant women with PE to identify clinically effective diagnostic biomarkers. Proteome profiling was performed using TMT-based liquid chromatography-mass spectrometry (LC-MS/MS) followed by subsequent verification by multiple reaction monitoring (MRM) analysis on normal and PE maternal plasma samples. Functional annotations of differentially expressed proteins (DEPs) in PE were predicted using bioinformatic tools. The diagnostic accuracies of the biomarkers for PE were estimated according to the area under the receiver-operating characteristics curve (AUC). A total of 1307 proteins were identified, and 870 proteins of them were quantified from plasma samples. Significant differences were evident in 138 DEPs, including 71 upregulated DEPs and 67 downregulated DEPs in the PE group, compared with those in the control group. Upregulated proteins were significantly associated with biological processes including platelet degranulation, proteolysis, lipoprotein metabolism, and cholesterol efflux. Biological processes including blood coagulation and acute-phase response were enriched for down-regulated proteins. Of these, 40 proteins were subsequently validated in an independent cohort of 26 PE patients and 29 healthy controls. APOM, LCN2, and QSOX1 showed high diagnostic accuracies for PE detection (AUC >0.9 and p&#xa0;<&#xa0;0.001, for all) as validated by MRM and ELISA. Our data demonstrate that three plasma biomarkers, identified by systematic proteomic profiling, present a possibility for the assessment of PE, independent of the clinical characteristics of pregnant women.

Humans

Identification of novel cytoskeleton protein involved in spermatogenic cells and sertoli cells of non-obstructive azoospermia based on microarray and bioinformatics analysis.

BACKGROUND: During mammalian spermatogenesis, the cytoskeleton system plays a significant role in morphological changes. Male infertility such as non-obstructive azoospermia (NOA) might be explained by studies of the cytoskeletal system during spermatogenesis. METHODS: The cytoskeleton, scaffold, and actin-binding genes were analyzed by microarray and bioinformatics (771 spermatogenic cellsgenes and 774 Sertoli cell genes). To validate these findings, we cross-referenced our results with data from a single-cell genomics database. RESULTS: In the microarray analyses of three human cases with different NOA spermatogenic cells, the expression of TBL3, MAGEA8, KRTAP3-2, KRT35, VCAN, MYO19, FBLN2, SH3RF1, ACTR3B, STRC, THBS4, and CTNND2 were upregulated, while expression of NTN1, ITGA1, GJB1, CAPZA1, SEPTIN8, and GOLGA6L6 were downregulated. There was an increase in KIRREL3, TTLL9, GJA1, ASB1, and RGPD5 expression in the Sertoli cells of three human cases with NOA, whereas expression of DES, EPB41L2, KCTD13, KLHL8, TRIOBP, ECM2, DVL3, ARMC10, KIF23, SNX4, KLHL12, PACSIN2, ANLN, WDR90, STMN1, CYTSA, and LTBP3 were downregulated. A combined analysis of Gene Ontology (GO) and STRING, were used to predict proteins' molecular interactions and then to recognize master pathways. Functional enrichment analysis showed that the biological process (BP) mitotic cytokinesis, cytoskeleton-dependent cytokinesis, and positive regulation of cell-substrate adhesion were significantly associated with differentially expressed genes (DEGs) in spermatogenic cells. Moleculare function (MF) of DEGs that were up/down regulated, it was found that tubulin bindings, gap junction channels, and tripeptide transmembrane transport were more significant in our analysis. An analysis of GO enrichment findings of Sertoli cells showed BP and MF to be common DEGs. Cell-cell junction assembly, cell-matrix adhesion, and regulation of SNARE complex assembly were significantly correlated with common DEGs for BP. In the study of MF, U3 snoRNA binding, and cadherin binding were significantly associated with common DEGs. CONCLUSION: Our analysis, leveraging single-cell data, substantiated our findings, demonstrating significant alterations in gene expression patterns.

Male

Molecular mechanisms for proton transport in membranes.

Likely mechanisms for proton transport through biomembranes are explored. The fundamental structural element is assumed to be continuous chains of hydrogen bonds formed from the protein side groups, and a molecular example is presented. From studies in ice, such chains are predicted to have low impedance and can function as proton wires. In addition, conformational changes in the protein may be linked to the proton conduction. If this possibility is allowed, a simple proton pump can be described that can be reversed into a molecular motor driven by an electrochemical potential across the membrane.

Biological Transport, Active

shinyDeepGxP: a user-friendly R shiny app for predicting surface protein abundance from scRNA-seq expression using deep learning in blood cells.

MOTIVATION: Understanding accurate immune cell heterogeneity and function in single-cell datasets requires access to protein-level information, which is often unavailable due to experimental limitations. RESULTS: We present shinyDeepGxP, an interactive web application featuring our deep learning model, DeepGxP, for predicting surface protein abundance from single-cell RNA-sequencing (scRNA-seq) data. This platform makes DeepGxP accessible to researchers without programming skills. Users can upload scRNA-seq count matrices and use "Predict Protein" to predict the abundance of 224 biologically relevant surface proteins. shinyDeepGxP provides visualizations to help identify distinct cell populations based on predicted protein profiles. Moreover, users can choose "Explore Model" to reveal key RNA predictors and their associated biological pathways for each protein. Overall, shinyDeepGxP is a user-friendly, freely available web tool that provides protein-level detail for RNA-only single-cell datasets, enabling multimodal discovery without additional experiments. AVAILABILITY AND IMPLEMENTATION: shinyDeepGxP can be launched on https://shiny.crc.pitt.edu/deepgxp/.

Journal Article

Freely available genomic datasets for atrial fibrillation research: current resources and analytical pipeline.

Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia, characterized by clinical and genetic heterogeneity. Increasing use of genomics and other omics approaches has driven reliance on publicly available AF datasets to advance biological discovery. Thus, this systematic review aimed to identify freely available genomic AF datasets through Mendeley Data and its interconnected repositories, and to characterize the most common analyses performed on these data. The search was conducted in adherence to the PRISMA 2020 guideline. Nineteen freely available genomic AF datasets were identified: Summary statistics for 'Biobank-driven genomic discovery yields new insight into atrial fibrillation biology', hum0014.v8.58qt.v1, AF GWAS in UK Biobank, UK Biobank (Publication 9659), GWAS summary statistics from a 2025 multi-ancestry AF meta-analysis, GSE115574, GSE128188, GSE14975, GSE2240, GSE238242, GSE254133, GSE261170, GSE271748, GSE271839, GSE293813, GSE294456, GSE31821, GSE41177, and GSE79768. The GEO datasets were further examined using differential gene expression, functional enrichment, protein-protein interaction networks, hub gene analysis, microRNA target prediction, and gene clustering, as well as, for the more recently deposited datasets, eQTL colocalization, single-cell/single-nucleus clustering, cell-cell communication analysis, and gene-dosage-dependent transcriptional and electrophysiological profiling. These analyses show some consistency but also considerable heterogeneity in initial conditions, data normalization, and analytical methodological settings. In conclusion, only a limited number of datasets are freely available, so additional, well-characterized and standardized datasets are needed to provide a complete picture of the AF pathology.

Mendeley Data

From pan-life phase insights to PhaseHub: Analyzing protein condensate complexity.

Intracellular biomolecular condensation forms multicomponent signaling hubs that regulate development, stress responses, and environmental adaptation. While the molecular grammar encoded within scaffold proteins defines the basal associative features driving condensation, heterotypic condensates are intrinsically dynamic, multicomponent, and far-from-equilibrium systems. Consequently, how condensates organize component composition, stoichiometry, and functional specificity in space and time under physiological conditions remains poorly understood. Addressing this challenge requires integrative frameworks that combine predictive biophysical features with experimental information on protein abundance, interaction networks, subcellular localization, and evolutionary conservation. In this study, we first analyzed phase separation (PS) proteins across the tree of life in 1106 species, revealing a stark contrast in computationally predicted PS propensity between eukaryotes and prokaryotes, with genome size as a key determinant. Through a broad analysis of amino acid homorepeat-containing proteins (HRPs) across all species, we uncovered how PS evolves via a balance between functional condensation and avoidance of harmful, aggregation-prone sequences. We further identified potential signaling hubs and components across kingdoms by integrating PS-positive proteins with experimentally derived abundance and interactome data from four model eukaryotic species. Using Arabidopsis as a model, we dissected the relationships among PS propensity, condensation hub prediction, HRPs, subcellular localization, and structural conservation. Finally, we developed PhaseHub, a user-friendly interface for exploring scaffold-client dynamics, PS components, sequence signatures within each PS protein, and hubs. Collectively, our work provides an evolutionary framework for understanding multicomponent PS hubs by integrating molecular grammar with physiological context, thereby facilitating hypothesis generation and rational design.

Phase Separation

The signed two-space proximity model for learning representations in protein-protein interaction networks.

MOTIVATION: Accurately predicting complex protein-protein interactions (PPIs) is crucial for decoding biological processes, from cellular functioning to disease mechanisms. However, experimental methods for determining PPIs are computationally expensive. Thus, attention has been recently drawn to machine learning approaches. Furthermore, insufficient effort has been made toward analyzing signed PPI networks, which capture both activating (positive) and inhibitory (negative) interactions. To accurately represent biological relationships, we present the Signed Two-Space Proximity Model (S2-SPM) for signed PPI networks, which explicitly incorporates both types of interactions, reflecting the complex regulatory mechanisms within biological systems. This is achieved by leveraging two independent latent spaces to differentiate between positive and negative interactions while representing protein similarity through proximity in these spaces. Our approach also enables the identification of archetypes representing extreme protein profiles. RESULTS: S2-SPM's superior performance in predicting the presence and sign of interactions in SPPI networks is demonstrated in link prediction tasks against relevant baseline methods. Additionally, the biological prevalence of the identified archetypes is confirmed by an enrichment analysis of Gene Ontology (GO) terms, which reveals that distinct biological tasks are associated with archetypal groups formed by both interactions. This study is also validated regarding statistical significance and sensitivity analysis, providing insights into the functional roles of different interaction types. Finally, the robustness and consistency of the extracted archetype structures are confirmed using the Bayesian Normalized Mutual Information (BNMI) metric, proving the model's reliability in capturing meaningful SPPI patterns. AVAILABILITY: S2-SPM is implemented and freely available under the MIT license at https://github.com/Nicknakis/S2SPM.

Protein Interaction Mapping

Unravelling the genomic potential of sponge-associated Streptomyces sp. BLC 17-3 from Indonesia for mannooligosaccharide production.

This research aims to show the promising capacity of Streptomyces sp. BLC 17-3 to produce high &#x3b2;-mannanase enzymes and generate mannooligosaccharide (MOS) such as mannobiose, mannotriose, mannotetraose and mannopentaose when exposed to mannan polymers. Streptomyces sp. BLC 17-3 was isolated from the sponge (Rhabdastrella globostellata) Put4 obtained from the marine waters of Putus Island in Bitung, North Sulawesi, Indonesia. The characterization results showed that the peak enzyme activity was achieved at 50&#xa0;mM sodium acetate, 6.0 pH, and 60&#xa0;&#xb0;C temperature on the seventh day of production with a value of 155.77&#xa0;&#xb1;&#xa0;3.21&#xa0;U/mL. The SDS-PAGE and zymograms also showed that the size of the enzyme molecule was approximately &#xb1;34.8-49.1&#xa0;kDa. Moreover, whole-genome sequencing was conducted to identify the genetic basis of MOS-synthesizing capabilities in the selected strain, followed by functional annotation of genes encoding mannan degradation and associated functions. The results showed an 8,248,862&#xa0;Mb complete draft genome of the strain which comprised 111 predicted gene models. Gene annotation also provided important information about the location and function of protein-encoding genes. A total of 6 mannan degradation-related genes encoding mannanase-related metabolism were identified and the three-dimensional structures were predicted using AlphaFold 3. This characterization and modeling further enhanced the bioprospecting and development of this strain which exhibited efficient mannose metabolism. The results showed Streptomyces sp. BLC 17-3 as a promising microorganism for the future bioproduction of MOS which were discovered to have the capability of serving as a potential prebiotic substance to enhance digestion and promote health.

Bioprospecting

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Artificial Intelligence

Safe and Stable Germline Transmission of MSTN Mutations in Cattle.

With the global population expected to reach 10 billion by 2050, sustainable livestock production is critical. Gene editing of the myostatin (MSTN) gene represents a promising strategy to enhance muscle growth in cattle. In this study, MSTN-mutated founder (F0) cows were used to generate F1 offspring via ovum pick-up, in&#xa0;vitro fertilization, and embryo transfer. Four F1 calves were born, all confirmed to be heterozygous for the MSTN mutation. Long-term monitoring showed normal growth and no visible health abnormalities. Whole-genome sequencing identified SNPs, INDELs, and structural variants, most with minimal predicted functional effects. Proteomic profiling of Longissimus dorsi muscle quantified 2947 proteins, revealing only subtle expression differences between MSTN-mutated and wild-type cattle. These results demonstrate stable inheritance and confirm that MSTN editing does not disrupt genome integrity or protein expression. Overall, our findings support the safety and utility of MSTN gene editing to improve livestock productivity for future food security.

Animals