Search PubMedSearch

SEARCH · Search PubMed

Results for “Gene functional annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Genome-wide association study reveals candidate genes associated with body weight and wool traits in Ordos fine-wool sheep.

BACKGROUND: The Ordos fine-wool sheep is a high-quality fine-wool breed in China, renowned for its excellent wool quality, meat production, and adaptability to the arid and semi-arid regions of Inner Mongolia. Body weight and wool traits are important economic characteristics in sheep breeding. This study aimed to identify genetic loci associated with body weight (BW), wool length (WL), and wool fineness (WF) in Ordos fine-wool sheep. METHODS: A genome-wide association study (GWAS) was conducted in 388 Ordos fine-wool sheep genotyped using the GenoBaits® Ovine 40K SNP panel. Single nucleotide polymorphisms (SNPs) associated with BW, WL, and WF were identified, and candidate genes located near the SNPs reaching the suggestive threshold were subjected to functional annotation and enrichment analysis. RESULTS: A total of 22 SNPs were identified as potentially associated with BW, WL, and WF traits, corresponding to 27 annotated genes. Functional annotation highlighted six potential candidate genes, including LAMA2, ARHGAP18, IGFBP2, IGFBP5, CA10, and AXIN1, which may play important roles in regulating body weight and wool growth in sheep. CONCLUSIONS: The identified genes provide valuable candidate loci for BW, WL, and WF traits in Ordos fine-wool sheep. The results of this study provide preliminary references for further exploration of the genetic mechanisms of wool traits in Ordos fine-wool sheep and the development of molecular breeding markers.

GWAS

Molecular Characterization of Listeria monocytogenes Isolated from Retail Yak Meat in Nyingchi, Xizang, China.

Listeria monocytogenes is a Gram-positive zoonotic pathogen responsible for listeriosis, a severe foodborne disease with high mortality in humans and animals. This study aimed to investigate the molecular epidemiology and genomic characteristics of L. monocytogenes isolated from raw yak meat in Nyingchi, Xizang, China. A total of 231 yak-related samples were collected in Nyingchi, consisting of 214 retail raw yak meat samples, 14 farm environmental samples, and 3 nearby water source samples. L. monocytogenes isolates were identified and characterized using culture-based methods, PCR serotyping, and whole-genome sequencing (WGS). Bioinformatic analyses were performed for virulence, antimicrobial resistance, and functional gene annotation using KEGG and COG databases. The overall contamination rate of Lm was 13.08% (28/214) for retail raw yak meat samples, whereas no isolates were recovered from 14 farm environmental samples (0.00%, 0/14) and 3 nearby water source samples (0.00%, 0/3). The serotypes of isolates were 1/2a (9/28, 32.14%), 1/2b (7/28, 25.00%), and 1/2c (12/28, 42.86%). These 28 isolates exhibited varied antimicrobial resistance profiles, with universal resistance to trimethoprim-sulfamethoxazole, high resistance to erythromycin and clindamycin, and low resistance to vancomycin. MLST analysis revealed seven sequence types (STs): ST9 (12/28, 42.86%), ST619 (6/28, 21.43%), ST8 (6/28, 21.43%), ST7 (1/28, 3.57%), ST87 (1/28, 3.57%), ST121 (1/28, 3.57%), ST297 (1/28, 3.57%). ST619 isolates harbored multiple virulence genes, including those located on Listeria pathogenicity islands LIPI-1, LIPI-3, and LIPI-4, indicating high genomic potential for virulence. Representative isolate Y2 (ST619) possessed a 3,009,858 bp genome with 3036 coding genes, four genomic islands, and two prophages. Functional annotation revealed enrichment of genes involved in carbohydrate transport and metabolism and amino acid biosynthesis pathways. Our findings provide the first genomic insight into L. monocytogenes contamination in yak meat from Nyingchi, Xizang, China, highlighting the urgent need to strengthen food safety monitoring and hygiene management in this region.

Listeria monocytogenes

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 × homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering ∼18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals

Whole-Genome Analysis of Bacillus Licheniformis Ali5 and Synthesis of Lichenysin via Genome Shuffling.

Whole-genome sequencing of Bacillus licheniformis Ali5 was performed via MGI-seq PE150 and Nanopore single-molecule real-time sequencing. The strain has a 4,114,664 bp circular genome encoding 4030 protein-coding genes. Functional annotation across NR, COG, GO, KEGG, CARD, BacMet, and CAZy databases identified 4025, 2812, 988, 1242, 72, 69, and 94 corresponding genes, respectively, and antiSMASH 6.0 revealed multiple antimicrobial biosynthetic gene clusters, including intact lichenysin and lichenicidin VK21 A1/A2 gene clusters. Three rounds of recursive protoplast fusion-based genome shuffling, paired with a dual-index screening system, significantly improved strain growth and lichenysin biosynthesis. Recombinants exhibited shortened lag phase, enhanced proliferation, improved stationary-phase stability, and higher diauxic peak biomass. PP3-176 and PP3-186 showed 4.6%-8.1% higher 12-h shake-flask titer and 3.1%-4.0% higher maximum titer than the parental average, with excellent fermentation stability. 1-L bioreactor validation confirmed strong scale-up potential. PP3-186 achieved 27.2% and 31.6% titer increases at 12 h and 20 h, while PP3-176 yielded 20.4% and 14.6% improvements with robust metabolic performance. This study validates genome shuffling as an effective strategy for enhancing lichenysin production, providing candidate strains and technical support for industrial application.

Bacillus licheniformis

Genetic overlap between depression and C-reactive protein levels: Evidence from a cross-trait analysis.

Inflammation and depression have been consistently associated, with elevated C-reactive protein (CRP) levels observed in a significant subset of affected individuals. However, the genetic mechanisms underlying this association remain poorly understood. We integrated results from large-scale genome-wide association studies (GWAS) of depression and CRP levels in a cross-trait analysis specifically focusing on identifying horizontally pleiotropic loci. Identified variants were stratified as concordant versus discordant based on their direction of effects on the two traits and followed up using functional annotation, gene set enrichment, and colocalization analyses. We also explored causal relationships using Mendelian Randomization (MR) analysis with extensive sensitivity analyses, including adjustment for body mass index (BMI). We identified 9 novel loci. Functional analyses revealed that concordant loci were enriched in genes linked to immune and inflammatory processes, while discordant loci mostly mapped to metabolic pathways, including lipid regulation. MR provided strong evidence for body mass index driving a causal relationship between the genetic liability of depression on CRP levels. Our findings suggest that the association between depression and CRP levels is partly driven by shared genetic influences, pointing to different biological pathways depending on whether genetic effects are concordant or discordant. These results underscore the importance of considering effect direction when assessing the genetic overlap between depression and inflammatory processes. In addition, they highlight BMI as a key factor in the causal relationship between depression and systemic inflammation.

C-Reactive Protein

Comprehensive profiling of antibiotic resistance genes and functional clusters of orthologous groups annotation of gut microbiota in Indonesian Kedu chickens.

Antibiotic resistance is a growing global health concern, with poultry systems acting as important reservoirs of antibiotic resistance genes (ARGs). However, resistome and functional profiles of indigenous chickens raised under traditional systems remain underexplored. This study aimed to characterize the antibiotic resistome, virulence factor genes, and metabolic potential of gut microbiota in Indonesian Kedu chickens using a shotgun metagenomic approach. Digesta samples from five gastrointestinal segments of 21 healthy adult chickens were analyzed through high-throughput sequencing. ARGs were identified using the Comprehensive Antibiotic Resistance Database (CARD) and Antibiotic Resistance Genes Databases (ARDB), while virulence factors and functional genes were annotated using Virulence Factor Database (VFDB), Clusters of Orthologous Groups (COG), and Carbohydrate-Active EnZymes (CAZy) databases. Results revealed a diverse resistome dominated by multidrug resistance and efflux pump mechanisms, with prominent genes associated with fluoroquinolone, tetracycline, β-lactam, and glycopeptide resistance. The detection of clinically relevant ARGs suggests that genetic determinants associated with antimicrobial resistance are present in the gut microbiota of traditionally raised Kedu chickens, although metagenomic data alone cannot determine whether these genes are actively expressed or confer phenotypic resistance. Virulence factor analysis showed functions related to adherence, immune evasion, iron acquisition, quorum sensing, and efflux activity, reflecting strong microbial adaptability. Functional profiling demonstrated enrichment in translation, carbohydrate and amino acid metabolism, genome maintenance, and cell envelope biogenesis. Additionally, CAZyme analysis indicated a high capacity for complex polysaccharide degradation, supporting efficient utilization of fiber-rich traditional diets. In conclusion, this study provides a comprehensive metagenomic overview of antibiotic resistance and functional potential in Kedu chicken gut microbiota, emphasizing the importance of incorporating indigenous poultry into antimicrobial resistance surveillance within a One Health framework.

Antibiotic resistance genes

Bioinformatics analysis of miR-2861 and miR-5011-5p that function as potential tumor suppressors in colorectal carcinogenesis.

BACKGROUND: The study aimed to was to investigate the relationship between miR-2861, miR-5011-5p, and colorectal carcinogenesis. METHOD: In the present study, it was isolated RNA from both the tumor and non-tumor tissue of a total of 80 CRC patients and after synthesizing the cDNA, it was performed qRT-PCR to determine the expression levels of miR‑2861 and miR‑5011-5p. In addition, it was predicted that dysregulated miRNAs targets, pathways and functional gene annotations that may be important in colorectal carcinogenesis using KEGG pathway and GO analysis. RESULTS: The resulting data revealed that both expression levels of miR-2861 and miR-5011-5p were significantly decreased in tumor tissues compared with non-tumor tissues of CRC patients. The GO and KEGG pathway analysis showed that miR-2861 and miR-5011-5p may participate in multiple the biological process, cellular components, and molecular function subcategories such as mitotic cell cycle, regulation of small GTPase mediated signal transduction, cell death, and acid binding transcription factor activity. It was also revealed that target genes of miRNAs can be found in signaling pathways such as TGF-beta, Rap1, Ras, cAMP, Wnt, mTOR and, PI3K-Akt signaling pathways. CONCLUSION: These findings imply that miR-2861 and miR-5011-5p might function as tumor suppressors in the development of CRC.

MicroRNAs

Conventional and Shared Genetic Association Analysis Between Diabetes Mellitus and Sensorineural Hearing Loss.

PURPOSE: This study aims to investigate the epidemiological and genetic associations between diabetes mellitus (DM) and sensorineural hearing loss (SNHL) across different subtypes. METHODS: We analyzed 502,490 participants from the UK Biobank using multivariate logistic regression to examine the association between DM and SNHL, considering gender, age, and HbA1c levels. Genetic correlations and causality were examined by linkage disequilibrium score regression and bidirectional Mendelian randomization. Cross-trait meta-analyses identified shared loci between DM and SNHL, followed by gene annotation, functional analysis, and drug candidate exploration for the shared traits. RESULTS: Observational analysis revealed significant associations between DM and SNHL, consistent in subgroups based on age, sex, and certain HbA1c levels. A positive genetic correlation was found between type 2 diabetes mellitus (T2D) and SNHL (Rg = 0.0982, p = 0.0095) between T2D and SNHL, and four loci were identified, with ARHGEF28 and TCF7L2 prioritized as credible pleiotropic genes. Enrichment was indicated in glucose metabolism and organogenesis, with shared heritability in metabolic tissues and outer hair cells. Metformin was identified as potential drug candidates for the T2D-SNHL comorbidity. CONCLUSION: These findings progress our understanding of the epidemiological association, shared genetic basis, and potential therapeutic targets between T2D and SNHL, which might contribute to the management of their comorbidity.

Humans

Transcriptomic insights into the coordinated regulation of signaling, apoptosis, immunity, and metabolism during Sinonovacula constricta larval metamorphosis.

Metamorphosis is a critical ontogenetic transition for marine bivalves, marking the shift from planktonic to benthic lifestyles, where successful transformation dictates survival. The razor clam Sinonovacula constricta is economically important; however, low larval metamorphosis rates remain a major bottleneck in seedling production. To elucidate the mechanisms governing this process, we performed a comparative transcriptome analysis of S. constricta larvae at pre- and post-metamorphosis stages using Illumina sequencing. A total of 3701 differentially expressed genes (DEGs) were identified, including 3254 up-regulated and 447 down-regulated genes. Functional annotation of the respective top 20 significantly up-regulated and down-regulated DEGs indicated their potential pivotal roles in signal transduction (e.g., up-regulated: CAV1, CHRNA2; down-regulated: APP, NOTCH1), cellular proliferation and differentiation (e.g., up-regulated: TUBA, EGF1; down-regulated: KIF23, TTC25), transcriptional and epigenetic regulation (e.g., up-regulated: NFIL3; down-regulated: OVO, HMX1), substance transport (e.g., up-regulated: LRP2, LRP1B; down-regulated: SLC51A, Slc33a1), substance metabolism (e.g., up-regulated: CPK3, CYP26A1; down-regulated: RDMT1, ADAC), immunomodulation (e.g., up-regulated: CPN2, CRISP2), and protein homeostasis (e.g., up-regulated: HSP27, NAS-27). Functional enrichment analysis further revealed that DEGs were significantly enriched in pathways related to signal transduction and developmental regulation (e.g., Ras, TNF), cell death and homeostasis (e.g., apoptosis), immune responses (e.g., Toll-like receptor), energy metabolism (e.g., lipid), cardiovascular related (e.g., Fluid shear stress), cell junction and architecture (e.g., Tight junction), and infectious disease (e.g., measles). These results suggest a synergistic interplay between signaling, apoptosis, immunity, and metabolism during S. constricta metamorphosis. This study advances our understanding of marine bivalve metamorphosis and offers candidate genes for further mechanistic studies.

Animals

Shared genetic architecture and therapeutic targets across paediatric immune-mediated diseases.

OBJECTIVES: Paediatric-onset immune-mediated inflammatory diseases (IMIDs), including juvenile idiopathic arthritis and related rheumatic diseases, remain genetically undercharacterised. We aimed to define shared and category-specific genetic architecture across paediatric IMIDs, compare signals with adult IMIDs, and identify therapeutic opportunities. METHODS: We analysed 24 paediatric IMIDs classified as autoimmune, polygenic-autoinflammatory, mixed-pattern, or allergic. Genome-wide association analyses included 18,086 cases and 131,019 controls of European ancestry. We estimated single nucleotide polymorphism (SNP)-based heritability, genetic correlations, and polygenic overlap; performed subset-based meta-analysis; and conducted functional annotation, gene prioritisation, pathway and protein network analyses, adult-IMID comparison, and drug-target prioritisation. RESULTS: SNP-based heritability ranged from 28.9% for allergic IMIDs to 61.9% for autoimmune IMIDs. Genetic correlation and polygenic modelling supported partial sharing across categories with category-specific components. Meta-analysis identified 39 genome-wide significant loci outside the Major Histocompatibility Complex (MHC) region, including 15 previously unreported loci; 19 loci were shared between categories. Gene-prioritisation and protein interaction analyses identified a core MHC-centred antigen-presentation network, with category-enriched modules involving complement, innate/barrier pathways, epithelial biology, and type 2 immunity. Enriched pathways included nuclear factor κB signalling, T helper 17 related pathways, Janus kinase-signal transducer and activator of transcription signalling, programmed cell death protein 1/programmed death‑ligand 1, cytotoxic T‑lymphocyte associated protein 4 regulation, and osteoclast differentiation, several of which are relevant to rheumatic diseases. Paediatric IMIDs shared broad polygenic architecture with adult IMIDs, whereas top-ranked genes converged strongly with adult rheumatic diseases. Priority Index analysis identified 178 high-scoring genes, including 43 approved or investigational IMID drug targets. CONCLUSIONS: Paediatric-onset IMIDs share core pathways with adult forms but exhibit distinct genetic architecture shaped by age-specific immune and neurodevelopmental biology. These findings provide a genomic framework for paediatric precision medicine, guiding classification, risk prediction, and therapeutic development.

Humans

VicMAG, an open-source tool for visualizing circular metagenome-assembled genomes highlighting bacterial virulence and antimicrobial resistance.

Bacterial pathogens spread in clinical and environmental settings, and mobile genetic elements (MGEs), such as plasmids and phages, mediate the transfer of virulence factor genes (VFGs) and antimicrobial resistance genes (ARGs) among bacterial communities. Metagenomic analysis of environmental and wastewater samples using highly accurate long-read sequencing technologies, such as Pacific Biosciences (PacBio) HiFi sequencing, provides valuable insights into monitoring the regional spread of VFGs and ARGs, including dissemination mediated by MGEs. No visualization tool is currently available for the comprehensive display of numerous resulting circular metagenome-assembled genomes (cMAGs) with functional gene annotations. Here, we developed visualization of circular metagenome-assembled genome (VicMAG), a visualization tool for highly complex cMAGs derived from long-read metagenome assemblies annotated using updated databases of VFGs, ARGs, and MGEs. Using 353 cMAGs from PacBio HiFi sequencing of a wastewater sample, we demonstrated the utility of VicMAG for metagenome visualization. VicMAG provides comprehensive, size-aware visualization of cMAGs representing bacterial chromosomes and plasmids, annotated with VFGs, ARGs, and phages. By simultaneously visualizing all cMAGs in a framework, VicMAG facilitates a holistic understanding of the distribution and genomic context of VFGs and ARGs across complex microbial communities. This tool supports integrated surveillance of bacteria associated with virulence and antimicrobial resistance across clinical, environmental, and One Health contexts.

Metagenome

Genetic Contributors to Postoperative Delirium and Their Implications for Dementia Outcomes.

BACKGROUND: Postoperative delirium (POD) is a perioperative neurocognitive disorder that substantially impairs patient recovery. Unfortunately, its genetic risk profile and relationship with subsequent dementia remain unclear. This study aimed to elucidate genetic contributors to POD identified via Hospital Episode Statistics codes and to examine its association with subsequent dementia. METHODS: The study included 230,179 noncardiac and 21,254 cardiac surgery subjects from the UK Biobank, defining POD using delirium codes from the International Classification of Diseases (10th revision) recorded within the first 7 postoperative days. Genome-wide association studies were performed in the noncardiac and cardiac cohorts and their prespecified subgroups, followed by functional annotation, gene prioritization and drug-target analyses. Associations between POD and subsequent dementia were estimated using Cox models. RESULTS: In the noncardiac cohort, one genome-wide significant locus was identified at the APOE region, with rs429358 as the lead variant ( P = 5.00 × 10 -28 ). Integrative gene prioritization analyses highlighted multiple genes within this locus. Exploratory drug-target analyses suggested potential subgroup-specific drug-target enrichment. In the cardiac cohort, no genome-wide significant signals were detected. POD was associated with all-cause dementia after both noncardiac (hazard ratio, 6.45; 95% CI, 5.45 to 7.63) and cardiac (hazard ratio, 2.95; 95% CI, 1.71 to 5.08) surgeries. CONCLUSIONS: This study demonstrates APOE as a genetic risk locus for International Classification of Diseases-coded POD in the noncardiac surgery setting and confirms an association between POD and subsequent dementia.

Humans

Multi-omics analysis identifies key genes and functional loci affecting teat number in American Large White and Landrace pigs and their application in optimizing genomic selection models.

BACKGROUND: Teat number is a crucial economic trait in pigs. It directly affects the ability of sows to lactate, which in turn influences the survival and health of piglets. The teat number of French Large White pigs is close to 16, while the teat number of American Large White and Landrace pigs is about 14. In order to improve the teat number of American Landrace and Large White pigs through molecular approaches and precise breeding techniques, we genotyped 2,131 American Landrace and 4,564 American Large White with teat number phenotype using a 50 K SNP chip. Then, the SNP-chip data was imputed to the level of whole-genome sequencing (iWGS). Based on iWGS data, we conducted GWAS to identify novel, significant SNPs associated with teat number and to incorporate them into genomic selection. RESULTS: In Landrace pigs, significant SNPs for TTN mapped to SSC2, SSC7, SSC8, and SSC14; the SSC8 and SSC14 effects are novel. LTN mapped to SSC7, RTN to SSC7 and SSC8. The lead SSC7 SNP explained 2.60% of TTN phenotypic variance. In Large White pigs, significant SNPs were detected on SSC7 and SSC10 for TTN; SSC7, SSC10, and SSC12 for LTN; and SSC7 and SSC10 for RTN. The most significant locus on SSC7 accounted for 2.99% of the phenotypic variance in TTN. Additionally, a multi-population meta-analysis detected significant novel SNPs for LTN on SSC1 and SSC8. By utilizing Bayesian fine mapping, the most precise QTL confidence interval on SSC7 for both TTN and RTN in Large White pigs was reduced to 40 kb. By integrating functional gene annotation with RNA-seq and ATAC-seq data from Erhualian and Bamaxiang pigs mammary placodes at embryonic day 26, we prioritized PTPN13, TRPV3, ZDHHC13, and BRD2 as novel candidate genes for teat number. We then incorporated the significant SNPs to GBLUP and benchmarked genomic-selection accuracy. In both breeds, fitting the top SNP as fixed maximized prediction for TTN and RTN, whereas treating all significant loci as an additional random effect optimized LTN. CONCLUSIONS: Our findings provide a theoretical basis for dissecting new key genes affecting teat number and for advancing molecular breeding of teat number in pigs.

Animals

GOtcha: a new method for prediction of protein function assessed by the annotation of seven genomes.

BACKGROUND: The function of a novel gene product is typically predicted by transitive assignment of annotation from similar sequences. We describe a novel method, GOtcha, for predicting gene product function by annotation with Gene Ontology (GO) terms. GOtcha predicts GO term associations with term-specific probability (P-score) measures of confidence. Term-specific probabilities are a novel feature of GOtcha and allow the identification of conflicts or uncertainty in annotation. RESULTS: The GOtcha method was applied to the recently sequenced genome for Plasmodium falciparum and six other genomes. GOtcha was compared quantitatively for retrieval of assigned GO terms against direct transitive assignment from the highest scoring annotated BLAST search hit (TOPBLAST). GOtcha exploits information deep into the 'twilight zone' of similarity search matches, making use of much information that is otherwise discarded by more simplistic approaches. At a P-score cutoff of 50%, GOtcha provided 60% better recovery of annotation terms and 20% higher selectivity than annotation with TOPBLAST at an E-value cutoff of 10(-4). CONCLUSIONS: The GOtcha method is a useful tool for genome annotators. It has identified both errors and omissions in the original Plasmodium falciparum annotation and is being adopted by many other genome sequencing projects.

Animals

Unravelling the genomic potential of sponge-associated Streptomyces sp. BLC 17-3 from Indonesia for mannooligosaccharide production.

This research aims to show the promising capacity of Streptomyces sp. BLC 17-3 to produce high β-mannanase enzymes and generate mannooligosaccharide (MOS) such as mannobiose, mannotriose, mannotetraose and mannopentaose when exposed to mannan polymers. Streptomyces sp. BLC 17-3 was isolated from the sponge (Rhabdastrella globostellata) Put4 obtained from the marine waters of Putus Island in Bitung, North Sulawesi, Indonesia. The characterization results showed that the peak enzyme activity was achieved at 50 mM sodium acetate, 6.0 pH, and 60 °C temperature on the seventh day of production with a value of 155.77 ± 3.21 U/mL. The SDS-PAGE and zymograms also showed that the size of the enzyme molecule was approximately ±34.8-49.1 kDa. Moreover, whole-genome sequencing was conducted to identify the genetic basis of MOS-synthesizing capabilities in the selected strain, followed by functional annotation of genes encoding mannan degradation and associated functions. The results showed an 8,248,862 Mb complete draft genome of the strain which comprised 111 predicted gene models. Gene annotation also provided important information about the location and function of protein-encoding genes. A total of 6 mannan degradation-related genes encoding mannanase-related metabolism were identified and the three-dimensional structures were predicted using AlphaFold 3. This characterization and modeling further enhanced the bioprospecting and development of this strain which exhibited efficient mannose metabolism. The results showed Streptomyces sp. BLC 17-3 as a promising microorganism for the future bioproduction of MOS which were discovered to have the capability of serving as a potential prebiotic substance to enhance digestion and promote health.

Bioprospecting

Large-scale benchmarking of prokaryotic annotation tools across thousands of species.

BACKGROUND: Genome annotation is an important step in deriving functional meaning from prokaryotic sequencing data, yet systematic evaluations guiding tool selection are lacking. We present the first large-scale investigation of four prominent open-source annotation tools (Prokka, Bakta, EggNOG-mapper, and PGAP) across 156,033 diverse genomes. This includes Escherichia coli strains for baseline performance, thousands of archaea and bacteria genomes, as well as frameshifted and metagenome-assembled genomes. RESULTS: Bakta excels in annotating high-quality bacterial genomes, while PGAP was better for archaeal genomes and challenging bacterial assemblies, including metagenome-assembled, fragmented, or contaminated samples. For Gene Ontology annotation, PGAP consistently provides broader term coverage, whereas EggNOG-mapper offers more terms per feature. CONCLUSIONS: Our findings highlight tool-specific strengths crucial for selecting optimal solutions based on genome quality, taxonomy, and origin (e.g. MAGs). This study provides an evidence-based guide for users and informs future tool development.

Molecular Sequence Annotation

Shared genetic architecture between ADHD and intelligence varies across ADHD subtypes.

BACKGROUND: Attention-deficit/hyperactivity disorder (ADHD) is a heterogeneous neurodevelopmental condition frequently accompanied by cognitive difficulties. Although previous genetic studies have demonstrated substantial overlap between ADHD and intelligence, most have treated ADHD as a single phenotype. However, whether this shared genetic architecture differs across ADHD subtypes remains unclear. METHODS: We conducted a genome-wide cross-trait analysis integrating large-scale genome-wide association study (GWAS) datasets of overall ADHD, its subtypes-childhood ADHD, persistent ADHD, and late-diagnosed ADHD-and intelligence (total N > 300,000). Genome-wide genetic correlations, polygenic overlap, local genetic correlations, and variant-level associations between ADHD phenotypes and intelligence were evaluated to characterize their shared genetic architecture. Shared variants were identified through cross-trait enrichment analyses and subsequently mapped to genes for functional annotation and gene-set enrichment. Bidirectional associations were evaluated using two-sample Mendelian randomization with sensitivity analyses. Additional GWAS datasets were used to validate the robustness of shared loci by assessing the consistency of effect directions. RESULTS: All ADHD phenotypes showed significant negative genetic correlations with intelligence (rg ranging from -0.3442 to -0.4205). Despite these modest genome-wide correlations, cross-trait analyses revealed substantial genetic overlap, including polygenic overlap, local genetic correlations, and variant-level associations. We identified 184 loci jointly associated with ADHD traits and intelligence, including 64 novel loci, whereas no shared loci were detected for persistent ADHD under the current analysis. Functional annotation revealed biologically distinct enrichment patterns across subtypes: childhood ADHD loci were linked to early neurodevelopmental processes, while late-diagnosed ADHD loci were enriched in synapse-related and neuronal signaling pathways. Mendelian randomization analyses suggested bidirectional associations, with stronger evidence supporting a directional association from intelligence to ADHD risk. Furthermore, these shared loci showed largely consistent effect directions across additional GWAS datasets, providing support for the robustness of the findings. CONCLUSIONS: The shared genetic architecture between ADHD and intelligence varies across ADHD subtypes, highlighting distinct biological pathways underlying cognitive heterogeneity in ADHD. These findings suggest that the relationship between ADHD liability and general cognitive ability is not uniform across ADHD subtypes and may inform future research on risk stratification and early identification in child and adolescent psychiatry.

Humans

Whole-genome sequencing and analysis of the endophytic fungus Alternaria alternata Y-2 from Leymus chinensis.

To explore the genetic basis and functional potential of beneficial symbiosis between the endophytic fungus Alternaria alternata Y-2 and its host Leymus chinensis, we performed Illumina-based draft whole-genome sequencing and systematic bioinformatic analysis. Although this assembly does not reach telomere-to-telomere completeness, it provides high-quality gene-level information for gene prediction, functional annotation, carbohydrate-active enzyme (CAZyme) identification, and secondary metabolite biosynthetic gene cluster analysis. The final genome size of A. alternata Y-2 was 34,383,676 bp with a GC content of 51.0%, containing 12,724 predicted protein-coding genes, 90 tRNAs, and 12 rRNAs. BUSCO assessment showed 98.9% completeness, supporting the high quality of this draft genome. A total of 12,627 genes were successfully annotated in the NCBI NR database, and 17,183 genes were functionally categorized using GO terms. In total, 448 CAZyme genes and 21 secondary metabolite biosynthetic gene clusters were identified, which are potentially involved in lignocellulose degradation, cellular redox homeostasis and biosynthesis of bioactive metabolites. Based on ITS sequence alignment, NR annotation, and phylogenetic analysis of single-copy orthologous genes, the strain was confidently identified as A. alternata. This study firstly reports the draft genome of an endophytic A. alternata strain derived from L. chinensis and provides valuable genetic resources for exploring the endophytic lifestyle, stress tolerance, and bioactive metabolite potential of this fungus.

Alternaria