Search PubMedSearch

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 × homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering ∼18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals

dcHiChIP: a comprehensive Nextflow-based pipeline for multiscale analysis of chromatin architecture from HiChIP data.

MOTIVATION: Despite the growing use of HiChIP to investigate protein-directed chromatin architecture, a comprehensive and reproducible pipeline for analysing these datasets-from raw reads to multiscale 3D genome features-remains lacking. Existing tools often focus on isolated components, such as loop calling or matrix generation, but fall short in integrating structural annotation, functional enrichment, and spatial modeling within a unified framework. To address this gap, we developed dcHiChIP, a modular, scalable Nextflow-based workflow that streamlines the analysis of HiChIP data, enabling both routine processing and in-depth exploration of chromatin organization and regulatory interactions. RESULTS: dcHiChIP enables robust and reproducible analysis of HiChIP datasets across multiple scales of chromatin architecture. It accepts raw sequencing data as input and generates high-quality loop calls, domain annotations, and 3D genome models. It also performs functional annotation and motif enrichment analyses. Applied to benchmark CTCF HiChIP datasets, dcHiChIP identifies major chromatin architectural features such as TADs/CCDs, A/B compartments, and chromatin stripes, and offers efficient, end-to-end execution with support for batch processing and workflow resumability. AVAILABILITY: dcHiChIP is publicly available on GitHub at https://github.com/SFGLab/dcHiChIP, with documentation at https://sfglab.github.io/dcHiChIP/. The software version used in this study is archived at Zenodo: https://doi.org/10.5281/zenodo.22030542.

Chromatin

Genome-wide annotation of human multi-nucleotide variants reveals widespread functional differences from single nucleotide variants.

Multi-nucleotide variants (MNVs) represent a crucial yet underexplored category of genetic variation. Despite previous studies highlighting the prevalence and potential biological impact of MNVs in populations, comprehensive identification and detailed functional annotation of MNVs remain challenging. Here, we develop MNVAnno, a toolbox for rapid identification and annotation of complex MNVs, and utilize it to identify 3,984,258 MNVs from 700,134 human samples, expanding the human MNV list to 8,199,654. Our analysis reveals that MNVs can not only lead to distinct amino acid changes from their constituent single-nucleotide variants, but also significantly impact the function of non-coding regions. Furthermore, through genome-wide association studies, we identify some MNVs associated with multiple cancers, and establish the Human MNV Database to facilitate MNV research. Our study emphasizes the importance of MNV annotation, broadens the human MNV landscape, and opens avenues for exploring genetic variation in phenotypes and diseases.

Humans

Genome-wide association study of estimated glomerular filtration rate using repeated measurements in the Taiwan Biobank.

BACKGROUND: Chronic kidney disease (CKD) is a major global public health issue, with genetic factors playing a significant role in kidney function. Although genome-wide association studies (GWAS) have identified numerous loci associated with estimated glomerular filtration rate (eGFR), most studies relied on a single time-point measurement, which limits the capacity to account for within-individual measurement variability. METHODS: We performed a repeated-measurement GWAS in the prospective Taiwan Biobank (Taiwanese ancestry; n = 25,004) using two repeated creatinine-based eGFR measurements. Repeated eGFR values were analyzed using a linear mixed-effects model with a subject-specific random intercept and time-varying covariates, providing a more precise estimate of eGFR level. Identified loci underwent functional annotation (expression quantitative trait locus, deleteriousness prediction, and epigenetic markers) and were compared with results from a single-measurement GWAS. RESULTS: Six loci associated with eGFR were identified, including four previously reported regions (1q22, 4q21.1, 11p14.1, and 17q21.2) and two additional loci (6p21.32 and 15q24.2). Functional annotation implicated several candidate genes-such as MUC1/EFNA1, SHROOM3, HLA-DQB1, MPPED2, NRG4, and PGAP3/FBXL20-in the regulation of kidney function. CONCLUSION: Incorporating repeated eGFR measurements into GWAS may improve phenotypic precision for identifying genetic associations with kidney function. This study identified eGFR-associated loci and biologically plausible candidate genes in a Taiwanese population, which require further replication and functional validation.

Chronic kidney disease

Characterization of a draft chromosome-scale genome assembly for the mutton snapper, Lutjanus analis.

BACKGROUND: The mutton snapper (Lutjanus analis) is a reef fish commonly found in tropical waters of the Western Atlantic Ocean. Genomic studies of this species are needed to support conservation efforts and breeding programs. OBJECTIVE: Here, we report the development of a chromosome-scale reference assembly for the mutton snapper and conduct an initial comparative genomic analysis with other lutjanids. METHODS: The genome of one mutton snapper specimen was sequenced using PAC-Bio HiFi long reads and Illumina short reads. Contigs and scaffolds were assembled in the Flye pipeline and anchored using Hi-C proximity guided assembly. Gene prediction and functional annotations were obtained in AUGUSTUS and eggNOG-mapper, respectively. The mutton snapper genome was compared to those of other lutjanids to infer gene family evolution and chromosome synteny conservation. RESULTS: Assembly and polishing yielded 946 contigs and 926 scaffolds (N50 of 3.16 Mb, complete BUSCO score 98.1%) that were anchored using Hi-C scaffolding in 24 draft chromosomes. The anchored assembly featured a N50 of 42.47 Mb and contained 97.6% of the unanchored assembly length. The 24 mutton snapper chromosomes showed a one-to-one syntenic relationship with their counterparts in medaka, and other Lutjanids. AUGUSTUS predicted 29,023 genes, 24,335 of which (83.85%) could be functionally annotated. Gene family evolution analysis revealed 1,014 significantly expanded or contracted hierarchical ortholog groups in mutton snapper. Expansions and contractions were linked to several biological functions including growth, oocyte maturation, and response to exogenous stressors. CONCLUSION: The draft genome will be a valuable tool for forthcoming applied genomic studies of mutton snapper.

Animals

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals

Ulmus minor response to Dutch elm disease: de novo transcriptome assembly and annotation.

Dutch elm disease (DED), caused by Ophiostoma novo-ulmi (ONU), has devastated elm populations across Europe and North America since the 20th century. In this work, a de novo transcriptome assembly of Ulmus minor in response to ONU is presented. We used two DED-resistant genotypes, MDV2.3 and VAD2, and one DED-susceptible genotype, MDV1, to capture responses to ONU at four time points post-inoculation (6, 24, 72, and 144 hours). RNA from collected samples was isolated and sequenced producing 60.88 M 100 bp paired-end reads per sample. We performed a de novo transcriptome assembly combining data from the three genotypes. The assembly was functionally annotated and validated through differential gene expression analysis of the response. This dataset provides a valuable resource for studying molecular mechanisms of DED resistance in elms, contributing to broadening our understanding of tree immunity and facilitating potential applications in functional annotation of future genome assemblies.

Transcriptome

Phage bioinformatics tools: a review of computational approaches for bacteriophage research.

Rising clinical interest in phage therapy and the exponential growth of metagenomic sequence catalogues have driven a rapid expansion of bacteriophage bioinformatics. More than 80 dedicated tools, mostly published since 2020, now span identification, assembly, annotation, taxonomy, lifestyle prediction, defence-system detection, and host prediction. Aimed at experienced practitioners and developers, this review synthesizes the field through the lens of three successive computational paradigms: sequence homology, bounded by database completeness; machine learning, constrained by labelled training data; and foundation models, which now achieve Matthews correlation coefficients above 0.95 in identification tasks and, through structure-informed prediction, raise functional annotation to over half of phage genes. Furthermore, we map the upstream components, namely, gene callers, homology engines, protein language models, and structural search tools, that underpin most downstream pipelines, exposing shared infrastructure and ecosystem-level fragility when dependencies change. To translate this into practice, we propose web-based and command-line reference workflows calibrated to user expertise and sample types. Finally, we set an agenda for the next wave of tool development. Roughly half of phage genes still resist functional annotation despite structural methods; no broadly generalizable strain-level host predictor exists for phage therapy; varying true-positive rates (0%-97%) underscore the absence of standardized community benchmarks analogous to Critical Assessment of Structure Prediction or Critical Assessment of Metagenome Interpretation. As generative genome models begin designing synthetic phages, progress will depend less on producing standalone tools than on rigorous evaluation, interoperable infrastructure, and clinically meaningful prediction targets.

Computational Biology

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Artificial Intelligence

Metaproteomic Analysis to Assess the Impact of Storage Media on Human Gut Microbiome in Fecal Samples.

The human gut microbiome is a diverse community of microorganisms residing in the gastrointestinal tract. The storage condition of fecal samples may impact the taxonomic and protein compositions of microbiomes in these samples. Here, we performed a mass spectrometry-based metaproteomic study to assess the impact of storage media on human gut microbiome in fecal samples. We evaluated FDA-authorized OMNIgene·GUT (OG), phosphate-buffered saline (PBS), and RNALater (RNAL) buffers and identified 38,185 microbial peptides corresponding to 7348 microbial proteins, which matched 16 phyla, 20 classes, 50 orders, 104 families, 332 genera, and 453 species. We found a high similarity among the fecal microbiomes preserved in OG, PBS, and RNAL in terms of the identification of proteins, taxa, and functional annotations. Both alpha and beta diversity suggested the high similarity among samples stored in the three media. Nonetheless, we also found some notable differences among buffers regarding the abundances of a few taxon groups. A partial human proteome (over 400 proteins) was identified in the fecal samples, with most of these proteins associated with the membrane and extracellular regions. The findings indicate the similarity among microbiomes in the fecal samples stored in OG, PBS, and RNAL regarding proteome profile, taxa, and functional capacity. SUMMARY: This study thoroughly analyzed and compared the metaproteomes of fecal samples preserved at -80°C in PBS, RNALater, and OMNIgene·GUT Dx buffers, offering novel insights into the effectiveness of these buffers in maintaining the stability and composition of the human gut microbiome. We found a high similarity in the identification and quantification of proteins, taxa, and functional annotations across the three buffers, with notable quantitative differences highlighting subtle yet important variations in preservation efficacy. The unique datasets and findings could offer valuable revelations into the impact of fecal sample preservation on translational and clinical analyses of the human gut microbiome.

Humans

Common genetic variants associated with urinary phthalate levels in children: A genome-wide study.

INTRODUCTION: Phthalates, or dieters of phthalic acid, are a ubiquitous type of plasticizer used in a variety of common consumer and industrial products. They act as endocrine disruptors and are associated with increased risk for several diseases. Once in the body, phthalates are metabolized through partially known mechanisms, involving phase I and phase II enzymes. OBJECTIVE: In this study we aimed to identify common single nucleotide polymorphisms (SNPs) and copy number variants (CNVs) associated with the metabolism of phthalate compounds in children through genome-wide association studies (GWAS). METHODS: The study used data from 1,044 children with European ancestry from the Human Early Life Exposome (HELIX) cohort. Ten phthalate metabolites were assessed in a two-void pooled urine collected at the mean age of 8&#xa0;years. Six ratios between secondary and primary phthalate metabolites were calculated. Genome-wide genotyping was done with the Infinium Global Screening Array (GSA) and imputation with the Haplotype Reference Consortium (HRC) panel. PennCNV was used to estimate copy number variants (CNVs) and CNVRanger to identify consensus regions. GWAS of SNPs and CNVs were conducted using PLINK and SNPassoc, respectively. Subsequently, functional annotation of suggestive SNPs (p-value&#xa0;<&#xa0;1E-05) was done with the FUMA web-tool. RESULTS: We identified four genome-wide significant (p-value&#xa0;<&#xa0;5E-08) loci at chromosome (chr) 3 (FECHP1 for oxo-MiNP_oh-MiNP ratio), chr6 (SLC17A1 for MECPP_MEHHP ratio), chr9 (RAPGEF1 for MBzP), and chr10 (CYP2C9 for MECPP_MEHHP ratio). Moreover, 115 additional loci were found at suggestive significance (p-value&#xa0;<&#xa0;1E-05). Two CNVs located at chr11 (MRGPRX1 for oh-MiNP and SLC35F2 for MEP) were also identified. Functional annotation pointed to genes involved in phase I and phase II detoxification, molecular transfer across membranes, and renal excretion. CONCLUSION: Through genome-wide screenings we identified known and novel loci implicated in phthalate metabolism in children. Genes annotated to these loci participate in detoxification, transmembrane transfer, and renal excretion.

Humans

Shared genetic architecture between ADHD and intelligence varies across ADHD subtypes.

BACKGROUND: Attention-deficit/hyperactivity disorder (ADHD) is a heterogeneous neurodevelopmental condition frequently accompanied by cognitive difficulties. Although previous genetic studies have demonstrated substantial overlap between ADHD and intelligence, most have treated ADHD as a single phenotype. However, whether this shared genetic architecture differs across ADHD subtypes remains unclear. METHODS: We conducted a genome-wide cross-trait analysis integrating large-scale genome-wide association study (GWAS) datasets of overall ADHD, its subtypes-childhood ADHD, persistent ADHD, and late-diagnosed ADHD-and intelligence (total N&#x2009;>&#x2009;300,000). Genome-wide genetic correlations, polygenic overlap, local genetic correlations, and variant-level associations between ADHD phenotypes and intelligence were evaluated to characterize their shared genetic architecture. Shared variants were identified through cross-trait enrichment analyses and subsequently mapped to genes for functional annotation and gene-set enrichment. Bidirectional associations were evaluated using two-sample Mendelian randomization with sensitivity analyses. Additional GWAS datasets were used to validate the robustness of shared loci by assessing the consistency of effect directions. RESULTS: All ADHD phenotypes showed significant negative genetic correlations with intelligence (rg ranging from -0.3442 to -0.4205). Despite these modest genome-wide correlations, cross-trait analyses revealed substantial genetic overlap, including polygenic overlap, local genetic correlations, and variant-level associations. We identified 184 loci jointly associated with ADHD traits and intelligence, including 64 novel loci, whereas no shared loci were detected for persistent ADHD under the current analysis. Functional annotation revealed biologically distinct enrichment patterns across subtypes: childhood ADHD loci were linked to early neurodevelopmental processes, while late-diagnosed ADHD loci were enriched in synapse-related and neuronal signaling pathways. Mendelian randomization analyses suggested bidirectional associations, with stronger evidence supporting a directional association from intelligence to ADHD risk. Furthermore, these shared loci showed largely consistent effect directions across additional GWAS datasets, providing support for the robustness of the findings. CONCLUSIONS: The shared genetic architecture between ADHD and intelligence varies across ADHD subtypes, highlighting distinct biological pathways underlying cognitive heterogeneity in ADHD. These findings suggest that the relationship between ADHD liability and general cognitive ability is not uniform across ADHD subtypes and may inform future research on risk stratification and early identification in child and adolescent psychiatry.

Humans

Structural insights into adeno-associated virus serotype 5.

The adeno-associated viruses (AAVs) display differential cell binding, transduction, and antigenic characteristics specified by their capsid viral protein (VP) composition. Toward structure-function annotation, the crystal structure of AAV5, one of the most sequence diverse AAV serotypes, was determined to 3.45-&#xc5; resolution. The AAV5 VP and capsid conserve topological features previously described for other AAVs but uniquely differ in the surface-exposed HI loop between &#x3b2;H and &#x3b2;I of the core &#x3b2;-barrel motif and have pronounced conformational differences in two of the AAV surface variable regions (VRs), VR-IV and VR-VII. The HI loop is structurally conserved in other AAVs despite amino acid differences but is smaller in AAV5 due to an amino acid deletion. This HI loop is adjacent to VR-VII, which is largest in AAV5. The VR-IV, which forms the larger outermost finger-like loop contributing to the protrusions surrounding the icosahedral 3-fold axes of the AAVs, is shorter in AAV5, creating a smoother capsid surface topology. The HI loop plays a role in AAV capsid assembly and genome packaging, and VR-IV and VR-VII are associated with transduction and antigenic differences, respectively, between the AAVs. A comparison of interior capsid surface charge and volume of AAV5 to AAV2 and AAV4 showed a higher propensity of acidic residues but similar volumes, consistent with comparable DNA packaging capacities. This structure provided a three-dimensional (3D) template for functional annotation of the AAV5 capsid with respect to regions that confer assembly efficiency, dictate cellular transduction phenotypes, and control antigenicity.

Capsid Proteins

The extracellular vesicle transcriptome provides tissue-specific functional genomic annotation relevant to disease susceptibility in obesity.

We characterized circulating extracellular vesicles (EVs) in obese and lean humans, identifying transcriptional cargo differentially expressed in obesity (277 unique genes; false discovery rate < 10%). Since circulating EVs may have broad origin, we compared this obesity EV transcriptome with expression from human visceral-adipose-tissue-derived EVs from freshly collected and cultured biopsies from the same obese individuals, observing high concordance. Using a comprehensive set of adipose-specific epigenomic and chromatin conformation assays, we found that the differentially expressed transcripts from the EVs were those regulated in adipose by body mass index-associated SNPs (p < 5 &#xd7; 10-8) from a large-scale genome-wide association study (GWAS). Using a phenome-wide association study of the regulatory SNPs for the EV-derived transcripts, we identified a substantial enrichment for inflammatory phenotypes, including type 2 diabetes. Collectively, these findings represent the convergence of the GWAS (genetics), epigenomics (transcript regulation), and EV (liquid biopsy) fields, enabling powerful future genomic studies of complex diseases.

Humans

Chromosome-Level Genome Assembly of Solanum carolinense.

Horsenettle (Solanum carolinense L.) is a noxious weed widely distributed across North America and increasingly invasive in other regions. Its strong environmental adaptability, complex defense strategies, and distinctive reproductive traits make it an important model for studying plant-herbivore coevolution. However, the absence of high-quality genomic resources has limited deeper investigation into its adaptive evolutionary mechanisms. In this study, we generated a chromosome-level reference genome assembly for S. carolinense using an integrated approach combining PacBio HiFi long-read sequencing, Illumina second-generation sequencing, and Hi-C chromatin interaction scaffolding. The final genome assembly had a total length of 915.40 Mb, with a contig N50 of 51.06 Mb and a scaffold N50 of 73.17 Mb; 96.05% of the sequences were successfully anchored onto 12 pseudochromosomes. The genome was characterized by a high proportion of repetitive sequences (73.64%) and substantial heterozygosity (1.13%), consistent with a highly repetitive and moderately high heterozygous genome. BUSCO analysis indicated that the chromosome-level genome assembly of S. carolinense reached a completeness score of 94.8%. A total of 32,206 protein-coding genes were annotated, of which 97.95% received functional annotations. The evaluation of the annotated protein-coding gene set returned a completeness value of 94.9%. This reference genome provides a valuable resource for advancing research on the adaptive evolution of weedy Solanaceae species, supports the development of more effective management strategies for this troublesome species, and offers a technical reference for assembling other highly heterozygous weed genomes.

Solanum carolinense

16S rRNA and Metagenomic Datasets of Gastrointestinal Microbiota in Fetal and 7-Day-Old Goat Kids.

The perinatal period (from late gestation to the neonatal stage) in ruminants is a critical phase for fetal organ maturation, where ecological succession of gastrointestinal microbial communities significantly impacts livestock production efficiency. However, research remains insufficient regarding the distribution patterns and functional annotation of microbial communities across different gastrointestinal compartments during this period. This study characterized early microbiota dynamics in Hutianshi Goats using 16S rRNA sequencing (4 fetal goats at 90&#x2009;&#xb1;&#x2009;10 gestational days) and metagenomics (3 7-day-old goat&#xa0;kids). The fetal goat group generated 852,694 valid reads, yielding 688,277 high-quality reads after chimera removal for downstream analysis. The 7-day-old goat kids group produced 1,081,588,182 final valid reads, after data processing and assembly, 8,561,345 contigs were generated. Gene prediction identified 6,095,352 genes. Multi-database annotations (NR, KEGG, CAZy, etc.) revealed functional potential and antimicrobial resistance traits. The public release of this dataset facilitates academic understanding of microbial community dynamics and host-microbe interactions during this developmental stage, providing both theoretical foundations and data resources for ruminant developmental biology and precision breeding regulation.

Animals

GUANinE v1.1 reveals complementarity of supervised and genomic language models.

There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness predictions to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models, and we conclude that moderate-context hybrid or post-trained language models may define the next era of machine learning in genomics.

Genomics

Unravelling the genomic potential of sponge-associated Streptomyces sp. BLC 17-3 from Indonesia for mannooligosaccharide production.

This research aims to show the promising capacity of Streptomyces sp. BLC 17-3 to produce high &#x3b2;-mannanase enzymes and generate mannooligosaccharide (MOS) such as mannobiose, mannotriose, mannotetraose and mannopentaose when exposed to mannan polymers. Streptomyces sp. BLC 17-3 was isolated from the sponge (Rhabdastrella globostellata) Put4 obtained from the marine waters of Putus Island in Bitung, North Sulawesi, Indonesia. The characterization results showed that the peak enzyme activity was achieved at 50&#xa0;mM sodium acetate, 6.0 pH, and 60&#xa0;&#xb0;C temperature on the seventh day of production with a value of 155.77&#xa0;&#xb1;&#xa0;3.21&#xa0;U/mL. The SDS-PAGE and zymograms also showed that the size of the enzyme molecule was approximately &#xb1;34.8-49.1&#xa0;kDa. Moreover, whole-genome sequencing was conducted to identify the genetic basis of MOS-synthesizing capabilities in the selected strain, followed by functional annotation of genes encoding mannan degradation and associated functions. The results showed an 8,248,862&#xa0;Mb complete draft genome of the strain which comprised 111 predicted gene models. Gene annotation also provided important information about the location and function of protein-encoding genes. A total of 6 mannan degradation-related genes encoding mannanase-related metabolism were identified and the three-dimensional structures were predicted using AlphaFold 3. This characterization and modeling further enhanced the bioprospecting and development of this strain which exhibited efficient mannose metabolism. The results showed Streptomyces sp. BLC 17-3 as a promising microorganism for the future bioproduction of MOS which were discovered to have the capability of serving as a potential prebiotic substance to enhance digestion and promote health.

Bioprospecting