Search PubMedSearch

SEARCH · Search PubMed

Results for “gene function”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

XsiAMT1.1a was identified as a novel ammonium uptake functional gene and its overexpression combined with GA4 application significantly increased yield in Arabidopsis thaliana.

Nitrogen (N) is a key limiting factor for plant yield. Ammonium is one of the main N forms absorbed by plants. Overexpression of ammonium uptake functional genes, such as ammonium transporter (AMT), can increase yield. However, the AMTs reported to enhance yield significantly is still limited. No researches have focused on the effect of overexpressing AMT combined with hormone application on yield improvement. In this study, we first investigated the role of XsiAMT1.1a, a potential ammonium uptake functional gene in an ammonium preference plant Xanthium sibiricum, in ammonium uptake by the analysis of bioinformatics, gene expression and subcellular localization, and the determination of ammonium uptake rate in endogenous silencing and heterologous overexpression plants. Subsequently, the effect of XsiAMT1.1a overexpression combined with hormone application on yield increase was further investigated in model plant Arabidopsis thaliana. Our results showed that XsiAMT1.1a shared the same conserved domains with AtAMT1 subfamily members and localized on the plasma membrane. XsiAMT1.1a was induced by N deficiency and highly expressed during the reproductive period. XsiAMT1.1a endogenous silencing and heterologous overexpression significantly decreased and increased ammonium uptake rates in X. sibiricum and A. thaliana, respectively. Overexpression of XsiAMT1.1a significantly improved total N accumulation, biomass and yield in A. thaliana, while XsiAMT1.1a overexpression combined with GA4 application had a stronger promoting effect on the above indicators. Our research identified a novel ammonium uptake functional gene, XsiAMT1.1a, and provided a new yield-increasing strategy which was verified in A. thaliana.

Arabidopsis

Generation of spCAS9 expressing human mesenchymal stem cell line to study gene function during osteoblast differentiation.

Human bone marrow-derived stromal cells (hMSCs) are a great resource for studying how genes influence cell fate and differentiation into various cell types like osteoblasts, adipocytes, and chondrocytes, among other cell types. However, genetic manipulation of primary hMSCs has been challenging due to their short lifespan and cellular senescence after limited passaging. Their low and unstable transfection efficiency also complicates gene delivery or inactivation, hindering long-term functional studies. The limited lifespan has been effectively solved by immortalizing hMSCs with telomerase reverse transcriptase (hMSCs-TERT). The use of these cells is ideal for functional studies of osteoblast and adipocyte differentiation through genetic manipulation, providing a stable and reliable model. Here, we have engineered a stable CAS9 expressing hMSC-TERT cell line (hMSC-TERTCAS9) via lentiviral transduction. The constitutive expression of spCas9 enables efficient and reproducible gene editing. We demonstrate the potential of these hMSC-TERTCAS9 cells for generating gene disruptions using plasmid delivery of guide RNAs as a fast and efficient strategy for targeted genome editing. The edited cells can be sorted and expanded as single cells to obtain homogenous clonal cell lines with mono- as well as bi-allelic gene deletions, a crucial step for producing reliable experimental results. We further validate this cell line as a powerful tool for studying gene function during hMSC proliferation and differentiation, providing 3 distinct examples of its utility. Through the generation of indels, single-cell sorting, and clonal selection, we have efficiently inactivated the vitamin D receptor and created both larger (256 nucleotides) gene disruptions in Forkhead box protein O1 and precise removals of a small genomic sequence (73 nucleotides) coding for microRNA MIR675. This novel hMSC-TERTCAS9 cell line represents a significant advancement, offering a stable, efficient, and versatile platform for advanced genetic studies, high-throughput screening, and the creation of reliable cellular disease models.

CRISPR-Cas9

CRISPR as a Tool to Uncover Gene Function in Polycystic Ovary Syndrome: A Literature Review of Experimental Models Targeting Ovarian and Metabolic Genes.

Polycystic ovary syndrome (PCOS) is a complex disorder characterized by reproductive abnormalities such as hyperandrogenism, ovulatory dysfunction, and polycystic ovarian morphology, and is frequently accompanied by metabolic disturbances such as insulin resistance, obesity and dyslipidemia. Genome-wide association studies (GWASs) have identified several susceptibility loci, yet little is known about their functional implications. Clustered regularly interspaced short palindromic repeats (CRISPR)/CRISPR-associated protein 9 (CRISPR/Cas9) has emerged as a powerful gene editing tool in bridging this gap by allowing researchers to directly target candidate genes in ovarian and metabolic pathways. For instance, experimental models have highlighted the role of CYP17A1 and DENND1A.V2 in androgen excess, anti-Müllerian hormone (AMH) in follicular arrest, and insulin receptor substrate 1 (IRS1) and PPARγ in insulin signaling and adipogenesis. To highlight the multifactorial nature of PCOS, animal models, including zebrafish and rodents, have been used to reveal interactions between reproductive and metabolic phenotypes. Nevertheless, most studies remain restricted to single-gene models, and dual-gene models or combined gene editing and hormonal induction models remain underexplored. Future research integrating precision editing, multi-omic platforms, and patient-derived organoids may provide more accurate disease models and novel therapeutic strategies.

Polycystic Ovary Syndrome

Large-scale pleiotropic analysis across cancers reveals shared genetic mechanisms and identifies novel functional genes.

Pleiotropic genetic loci have been increasingly reported in cancer, and identifying genetic variants with pleiotropic associations can reveal shared biological pathways influencing multiple cancers. Using summary statistics from genome-wide association studies for 37 cancer types (N = 433 836), we identified extensive genome-wide and local genetic correlations among cancers. Through pairwise pleiotropic analysis, we identified 75 243 significant pleiotropic single nucleotide polymorphisms (SNPs) across 372 cancer pairs, among which 3472 were lead SNPs with potential regulatory functions. Using FUMA and MAGMA, we identified 2527 pleiotropic risk loci and 4272 candidate pleiotropic genes. Notably, genes such as TERT (5p15.33), POU5F1B (8q24.21), and FANCA (16q24.3) exhibited widespread pleiotropy across multiple cancer types. Pathway enrichment analysis highlighted the critical roles of pigment synthesis, metabolism, and apoptosis in skin-related cancers, while cross-cancer enrichment analysis emphasized pathways related to apoptosis, chromatin structure, and intermediate filaments. We also identified 33 novel functional genes harboring previously unreported cancer risk variants. Drug-gene interaction analysis revealed several repositionable FDA-approved drugs. Importantly, drug sensitivity assays demonstrated that bosutinib and cobimetinib exhibited promising therapeutic potential in breast cancer cell lines. Finally, we developed the PleioCancer database (https://gonglab.hzau.edu.cn/PleioCancer/), providing a comprehensive resource for cancer pleiotropy research. These findings have important implications for carcinogenesis cancer, prevention and treatment.

Humans

From Gene Function to Precision Intervention: CRISPR/Cas9 and Stem Cell-Based Strategies as Emerging Disease-Modifying Approaches in PMOS.

Polyendocrine metabolic ovarian syndrome (PMOS) is a complex endocrine-metabolic disorder affecting up to 18% of women worldwide and remains the leading cause of anovulatory infertility. Despite extensive research, current treatments primarily target symptoms, including menstrual irregularities, hyperandrogenism, and metabolic dysfunction, without addressing the underlying molecular and tissue-level disturbances. Advances in multi‑omic profiling have identified disruptions across neuroendocrine, metabolic, inflammatory, and extracellular matrix pathways, alongside genetic susceptibility at loci such as DENND1A, CYP17A1, LHCGR, FSHR, IRS1, and PPARG. However, the functional roles of many variants remain unresolved. CRISPR/Cas9 gene editing enables precise interrogation of these pathways, while stem cell-based platforms, including mesenchymal stem cells (MSCs), exosomes, and gene-edited induced pluripotent stem cells (iPSCs), may serve as complementary platforms for regeneration and disease modeling. Preclinical studies demonstrate that MSCs and their derivatives modulate inflammation, restore ovarian structure, and improve metabolic parameters, while iPSC-based models enable patient-specific investigation of steroidogenic and metabolic abnormalities. Translational challenges remain, including targeted delivery, off-target effects, phenotypic heterogeneity, and regulatory considerations. Integrating CRISPR‑based functional genomics with stem cell research may shift PMOS management from symptom‑focused care to targeted, mechanism‑driven interventions that could modify the course of PMOS (Graphical Abstract).

Humans

A User-Friendly Protocol for Microinjection into Teleost Embryos to Study Gene Function.

Zebrafish (Danio rerio) and medaka (Oryzias latipes) are popular teleost models used in developmental biology and functional genomics. To achieve high-quality and reproducible microinjections, it is essential to have robust protocols for breeding, egg collection, and the precise delivery of genetic material. In this protocol, we present a comprehensive and optimized methodology for setting up breeding tanks under controlled photoperiod conditions to maximize egg yield while minimizing contamination. We provide detailed procedures for sex identification, pair selection, the use of grated breeding inserts, and methods to increase egg collection efficiency. We outline procedures for making injection gel beds, pulling needles, and calibration using one-microliter microcapillaries to achieve consistent nanoliter-scale injections. Our protocol outlines settings for the pico-liter injector that are optimized to deliver a precise amount per pulse with minimal variability. Finally, we demonstrate the application of these methods for gene knockdown using morpholino antisense oligonucleotides, gene knockout using CRISPR-Cas9, and gain-of-function mRNA overexpression experiments. Phenotypic assessments conducted at various developmental stages to evaluate gene-specific effects reveal consistent phenotypic outcomes between the morpholino and CRISPR-Cas9 approaches. This easy and comprehensive protocol enables efficient, precise, and scalable genetic manipulation of zebrafish and medaka embryos, thereby supporting advanced functional studies in developmental biology and disease modeling. To our knowledge, this is the first unified protocol for both zebrafish and medaka microinjection systems achieving 97.7% phenotype penetrance in CRISPR-Cas9 knockouts with precision together with a triple validation approach that confirms gene function across multiple techniques.

Animals

Role of functional genes for seed vigor related traits through genome-wide association mapping in finger millet (Eleusine coracana L. Gaertn.).

Finger millet (Eleusine coracana (L.) Gaertn.) is a calcium-rich, nutritious and resilient crop that thrives even in harsh environmental conditions. In such ecologies, seed longevity and seedling vigor are crucial for sustainable crop production amid climate change. The current study explores the genetics of accelerated aging on seed longevity traits across 221 diverse accessions of finger millet through genome-wide association approach (GWAS). A significant variation was identified in germination percentage, germination rate indices, mean germination time, seedling vigor indices and dry weight upon aging treatment. GWAS model from 11,832 high-quality SNPs identified through Genotyping-by-Sequencing (GBS) approach produced 491 marker-trait associations (MTAs) for 27 traits, of which 54 were FDR-corrected. A pleiotropic SNP, FM_SNP_9478 identified on chromosome 7B was associated with the traits viz., germination after aging, germination index after aging and their relative measures. Functional annotation revealed DET1 and expansin-A2 influenced seed coat integrity, critical for germination and aging resilience. Probable protein phosphatase 2C3 and piezo-type ion channels contributed to mechanical sensing and stress adaptation in seeds. Beta-amylase and acetyl-CoA carboxylase 2 were identified for seed metabolism and stress response. These insights lay the framework for targeted breeding efforts to improve seed quality and resilience under diverse production conditions.

Eleusine

Characterization of the gut phageome and functional genes carried by phages in laying hens with fatty liver hemorrhagic syndrome.

BACKGROUND: The gut microbiota is closely associated with the development of fatty liver hemorrhagic syndrome (FLHS); however, the function of its viral component, particularly bacteriophages, remains poorly understood. This study compared clinical parameters and the cecal phageome between 30-week-old (W30) and 50-week-old (W50) laying hens to characterize gut phages in the context of this metabolic disorder. RESULTS: Clinical analysis revealed that the W50 group exhibited typical FLHS, accompanied by elevated serum liver function and lipid markers (P&#x2009;<&#x2009;0.05). Functional prediction of the gut microbiota suggested a reduced lipid-metabolic capacity in W50 compared to the W30 group. A total of 20,274 phage genomes were identified from the two groups. These phages were primarily classified into 67 viral families, including Salasmaviridae, Herelleviridae, Suoliviridae, Peduoviridae, Crevaviridae, and Casjensviridae. The families Druskaviridae, Felixviridae, and Stanwilliamsviridae were uniquely detected in the W50 group. The phage community structure differed significantly between groups, with both phage diversity and richness markedly lower in W50 (P&#x2009;<&#x2009;0.05). LEfSe analysis revealed that phage taxa such as Stegnyidae, Herpelidae, and Chasovidae were significantly enriched in the W50 group, whereas Crewdviridae, Salasmaviridae, and Castroviridae were predominantly enriched in the W30 group. Functional annotation showed that these phages encode numerous metabolism-related genes and carry antimicrobial resistance genes (ARGs) as well as virulence factor genes. Notably, the diversity of ARGs carried by W50 phages was significantly higher (P&#x2009;<&#x2009;0.05), and ARG-rank analysis indicated a greater potential risk to human health. CONCLUSIONS: This study provides the first characterization of the gut phageome associated with FLHS in laying hens and confirms that gut phages constitute an important reservoir of ARGs. These findings offer a new perspective for understanding the pathogenesis of this disease and its associated public health risks. Video Abstract.

Animals

Boolean matrix logic programming for active learning of gene functions in genome-scale metabolic network models.

Reasoning about hypotheses and updating knowledge through empirical observations are central to scientific discovery. In this work, we applied logic-based machine learning methods to drive biological discovery by guiding experimentation. Genome-scale metabolic network models (GEMs) - comprehensive representations of metabolic genes and reactions - are widely used to evaluate genetic engineering of biological systems. However, GEMs often fail to accurately predict the behaviour of genetically engineered cells, primarily due to incomplete annotations of gene interactions. The task of learning the intricate genetic interactions within GEMs presents computational and empirical challenges. To efficiently predict using GEM, we describe a novel approach called Boolean Matrix Logic Programming (BMLP) by leveraging Boolean matrices to evaluate large logic programs. We developed a new system, [Formula: see text], which guides cost-effective experimentation and uses interpretable logic programs to encode a state-of-the-art GEM of a model bacterial organism. Notably, [Formula: see text] successfully learned the interaction between a gene pair with fewer training examples than random experimentation, overcoming the increase in experimental design space. [Formula: see text] enables rapid optimisation of metabolic models to reliably engineer biological systems for producing useful compounds. It offers a realistic approach to creating a self-driving lab for biological discovery, which would then facilitate microbial engineering for practical applications.

Active learning

Construction of a prognostic model for gastric cancer based on immune infiltration and microenvironment, and exploration of MEF2C gene function.

BACKGROUND: Advanced gastric cancer (GC) exhibits a high recurrence rate and a dismal prognosis. Myocyte enhancer factor 2c (MEF2C) was found to contribute to the development of various types of cancer. Therefore, our aim is to develop a prognostic model that predicts the prognosis of GC patients and initially explore the role of MEF2C in immunotherapy for GC. METHODS: Transcriptome sequence data of GC was obtained from The Cancer Genome Atlas (TCGA), the Gene Expression Omnibus (GEO) and PRJEB25780 cohort for subsequent immune infiltration analysis, immune microenvironment analysis, consensus clustering analysis and feature selection for definition and classification of gene M and N. Principal component analysis (PCA) modeling was performed based on gene M and N for the calculation of immune checkpoint inhibitor (ICI) Score. Then, a Nomogram was constructed and evaluated for predicting the prognosis of GC patients, based on univariate and multivariate Cox regression. Functional enrichment analysis was performed to initially investigate the potential biological mechanisms. Through Genomics of Drug Sensitivity in Cancer (GDSC) dataset, the estimated IC50 values of several chemotherapeutic drugs were calculated. Tumor-related transcription factors (TFs) were retrieved from the Cistrome Cancer database and utilized our model to screen these TFs, and weighted correlation network analysis (WGCNA) was performed to identify transcription factors strongly associated with immunotherapy in GC. Finally, 10 patients with advanced GC were enrolled from Sun Yat-sen University Cancer Center, including paired tumor tissues, paracancerous tissues and peritoneal metastases, for preparing sequencing library, in order to perform external validation. RESULTS: Lower ICI Score was correlated with improved prognosis in both the training and validation cohorts. First, lower mutant-allele tumor heterogeneity (MATH) was associated with lower ICI Score, and those GC patients with lower MATH and lower ICI Score had the best prognosis. Second, regardless of the T or N staging, the low ICI Score group had significantly higher overall survival (OS) compared to the high ICI Score group. For its mechanisms, consistently, for Camptothecin, Doxorubicin, Mitomycin, Docetaxel, Cisplatin, Vinblastine, Sorafenib and Paclitaxel, all of the IC50 values were significantly lower in the low ICI Score group compared to the high ICI Score group. As a result, based on univariate and multivariate Cox regression, ICI Score was considered to be an independent prognostic factor for GC. And our Nomogram showed good agreement between predicted and actual probabilities. Based on CIBERSORT deconvolution analysis, there was difference of immune cell composition found between high and low ICI Score groups, probably affecting the efficacy of immunotherapy. Then, MEF2C, a tumor-related transcription factor, was screened out by WGCNA analysis. Higher MEF2C expression is significantly correlated with a worse OS. Moreover, its higher expression is also negatively correlated with tumor mutation burden (TMB) and microsatellite instability (MSI), but positively correlated with several immunosuppressive molecules, indicating MEF2C may exert its influence on tumor development by upregulating immunosuppressive molecules. Finally, based on transcriptome sequencing data on 10 paired tumor tissues from Sun Yat-sen University Cancer Center, MEF2C expression was significantly lower in paracancerous tissues compared to tumor tissues and peritoneal metastases, and it was also lower in tumor tissues compared to peritoneal metastases, indicating a potential positive association between MEF2C expression and tumor invasiveness. CONCLUSIONS: Our prognostic model can effectively predict outcomes and facilitate stratification GC patients, offering valuable insights for clinical decision-making. The identified transcription factor MEF2C can serve as a biomarker for assessing the efficacy of immunotherapy for GC.

Humans

Exploring diagnostic m6A regulators in primary open-angle glaucoma: insight from gene signature and possible mechanisms by which key genes function.

PURPOSE: The purpose of this study was to interrogate the potential role of N6-methyladenosine (m6A) regulators in the process of trabecular meshwork (TM) tissue damage in patients with primary open-angle glaucoma (POAG). METHODS: Firstly, the expression profile of m6A regulators in TM tissues of POAG patients was comprehensively analyzed by bioinformatics analysis; Plasmid transfection and siRNA gene interference were used to enhance or weaken the expression levels of YTHDC2 in human trabecular meshwork cells (HTMCs); Cell migration ability was detected by transwell chamber assay; Immunofluorescence staining assay was used to evaluate the expression of extracellular matrix (ECM) related proteins. RESULTS: Through the analysis of GSE27276 database, 5 m6A regulators with different expression in POAG were screened out. The results of random forest model showed that these 5 m6A regulators exhibited diagnostic potential and were characteristic genes of POAG. All POAG samples could be effectively divided into two groups based on the expression levels of these 5 hub m6A regulators. Immune cell infiltration analysis indicated that the levels of activated CD8+ T cells and regulatory T cells were different in the two subtypes. HTMC oxidative stress cell model and TGF-&#x3b2;2 stimulation cell model were further constructed to verify the expression of the aforementioned hub m6A regulators, and it was found that YTHDC2 mRNA showed the same expression trend in both models. The silencing of YTHDC2 enhanced the migration ability of HTMCs and increased the synthesis ability of ECM. However, when YTHDC2&#x394;YTH, which lacks the YTH domain, is overexpressed in HTMCs, there is no significant change in the ECM synthesis ability. CONCLUSIONS: The differentially expressed m6A regulators in TM tissues may serve as potential diagnostic biomarkers for POAG. And, in HTMCs, the expression level of YTHDC2 mRNA was changed under oxidative stress or TGF-&#x3b2;2 intervention, and then exerted its regulation on cell migration and ECM synthesis capability through m6A modification, which may be an important part of the disease process of POAG.

Humans

Mechanism exploration of divergent partial denitrification performance under tetracycline stress: Insights from functional gene, electron transport and molecular docking.

Nitrates and antibiotics like tetracycline (TC) coexist in wastewater and inhibit nitrite (NO&#x2082;--N) accumulation during partial denitrification (PD), restricting anammox coupling. A moving bed biofilm reactor (PD-MBBR) and a sequencing batch reactor (PD-SBR) were compared under TC stress (0-8&#x202f;mg/L). The PD-MBBR proved more robust, sustaining a high nitrate transformation ratio (NTR) of 95.11% and &#x223c;53% TC removal. Metagenomic sequencing, quantitative polymerase chain reaction (qPCR), and molecular docking revealed this tolerance stemmed from physical shielding and metabolic compensation. Carrier-attached growth promoted extracellular polymeric substances (EPS) overproduction, forming a dense barrier preventing TC from binding to key denitrifying enzymes. The biofilm maintained stable nitrate reductase (NAR) activity via high narG and napA gene abundances, while nitrite reductase (NIR) was inhibited, ensuring efficient NO&#x2082;--N accumulation. This was supported by hyperactivated electron transport chain components, with complex III relative abundance increasing 15.08% and peak enzymatic activity reaching 149.02%. While IntI1-mediated horizontal gene transfer fortified community defense, concentrated antibiotic resistance genes (ARGs) within the biofilm pose a secondary dissemination risk. Thus, PD-MBBR provides an efficient pretreatment strategy for anammox, though downstream ARGs management is warranted.

Denitrification

Differential expression of neuronal function genes follows a tissue-specific temporal dynamic during Deformed Wing Virus infection in honey bees.

Deformed Wing Virus type A (DWV-A) is one of the primary threats to honeybees (Apis mellifera), significantly impacting their nervous system physiology and behavior. While its neurotropic nature is well-recognized, the temporal dynamics of the neuronal transcriptomic response following oral infection, the natural transmission route, remains poorly understood. In this study, we analyzed gene expression in the heads of worker bees orally inoculated with DWV-A over a 16-day time course (1, 4, 7, 10, 13, and 16 days post-inoculation). RNA-seq analysis at day 10 identified 147 differentially expressed genes associated with different biological processes that are critical to the organism, including cellular metabolism and neuronal activity. RT-qPCR validation revealed a persistent downregulation of key genes related to glutamatergic system (eaat-2, neto, and kainate) and sensory perception-related genes in the antennae. Notably, the simultaneous co-expression of nurse-associated and forager-associated marker genes suggests that DWV-A infection induces an asynchrony in behavioral maturation. Our findings demonstrate that DWV-A disrupts neuronal homeostasis and peripheral sensory perception in a tissue-specific and time-dependent manner, providing a molecular framework to understand the behavioral impairment and the loss of coordination at the colony level.

Animals

Multi-omics analysis identifies key genes and functional loci affecting teat number in American Large White and Landrace pigs and their application in optimizing genomic selection models.

BACKGROUND: Teat number is a crucial economic trait in pigs. It directly affects the ability of sows to lactate, which in turn influences the survival and health of piglets. The teat number of French Large White pigs is close to 16, while the teat number of American Large White and Landrace pigs is about 14. In order to improve the teat number of American Landrace and Large White pigs through molecular approaches and precise breeding techniques, we genotyped 2,131 American Landrace and 4,564 American Large White with teat number phenotype using a 50&#xa0;K SNP chip. Then, the SNP-chip data was imputed to the level of whole-genome sequencing (iWGS). Based on iWGS data, we conducted GWAS to identify novel, significant SNPs associated with teat number and to incorporate them into genomic selection. RESULTS: In Landrace pigs, significant SNPs for TTN mapped to SSC2, SSC7, SSC8, and SSC14; the SSC8 and SSC14 effects are novel. LTN mapped to SSC7, RTN to SSC7 and SSC8. The lead SSC7 SNP explained 2.60% of TTN phenotypic variance. In Large White pigs, significant SNPs were detected on SSC7 and SSC10 for TTN; SSC7, SSC10, and SSC12 for LTN; and SSC7 and SSC10 for RTN. The most significant locus on SSC7 accounted for 2.99% of the phenotypic variance in TTN. Additionally, a multi-population meta-analysis detected significant novel SNPs for LTN on SSC1 and SSC8. By utilizing Bayesian fine mapping, the most precise QTL confidence interval on SSC7 for both TTN and RTN in Large White pigs was reduced to 40&#xa0;kb. By integrating functional gene annotation with RNA-seq and ATAC-seq data from Erhualian and Bamaxiang pigs mammary placodes at embryonic day 26, we prioritized PTPN13, TRPV3, ZDHHC13, and BRD2 as novel candidate genes for teat number. We then incorporated the significant SNPs to GBLUP and benchmarked genomic-selection accuracy. In both breeds, fitting the top SNP as fixed maximized prediction for TTN and RTN, whereas treating all significant loci as an additional random effect optimized LTN. CONCLUSIONS: Our findings provide a theoretical basis for dissecting new key genes affecting teat number and for advancing molecular breeding of teat number in pigs.

Animals

An embedding-based framework enables statistical testing of gene-set function hypotheses inferred by large language models.

Emerging large language models (LLMs) can infer gene functions directly from gene lists, enabling hypothesis generation without predefined gene sets. However, these LLM-derived predictions are qualitative, and principled statistical validation is lacking. Here, we develop an embedding-based statistical framework that transforms gene and function descriptions into vector representations, enabling statistical testing of gene-gene and gene-function relationships and quantitative prioritization of de novo functional hypotheses inferred by LLMs. We benchmark seven state-of-the-art embedding models using curated and retrieval-augmented literature-derived gene descriptions across diverse biological contexts. OpenAI's text-embedding-3-large and Google's gemini-embedding-001 perform best, capturing gene-gene functional relationships in 88.7-92.5% of Gene Ontology biological processes and approximately 98.6% of canonical pathways. In gene-function association analyses, these models achieve high sensitivity (95.2-98.4%) and specificity (72.7-84.3%). Through contamination analysis and evaluation using experimentally informed protein assembly gene sets, our framework distinguishes biologically meaningful LLM-inferred hypotheses from noise, outperforming confidence-based inference and conventional enrichment analysis. We further develop the open-source R package DEGEmbedR and demonstrate its utility for interpreting a drug perturbation-derived differentially expressed gene (DEG) signature lacking significant conventional enrichment results. Together, these results establish LLM-derived embeddings as a quantitative foundation for functional genomics and the statistical validation of LLM-based gene function inference.

Large Language Models

Emerging Principles in Spatial Functional Genomics.

Spatial transcriptomic and proteomic atlases have enabled mapping of gene programs within intact tissues, but these measurements remain largely descriptive and do not define the mechanisms controlling tissue biology. Pooled CRISPR screening provides scalable causal interrogation of gene function but remains largely confined to dissociated systems that lack spatial context. In vivo spatial functional genomics (SFG) bridges these approaches by integrating genetic perturbations with in situ transcriptomic and proteomic readouts to measure gene function within intact tissue ecosystems. By preserving spatial organization, SFG enables interpretation of perturbations through effects on cell-cell interactions, diffusible signals, multicellular niches, and tissue architecture. Here, we outline key design axes of SFG: perturbation strategy, barcoding strategy, and phenotypic readout. We discuss computational challenges, including spatial autocorrelation, neighborhood dependence, and context-aware null modeling, and highlight how SFG reveals non-cell-autonomous, architecture-dependent mechanisms of gene function, advancing toward predictive models of tissue organization and gene function.

Genomics

High-throughput recovery of integron cassettes for gene discovery screens.

Integrons capture functional genes in mobile genetic elements called integron cassettes, which represent an untapped source of genes of biotechnological interest. Here we present two tools, cassette gatherer and cassette hunter, that enable high-throughput establishment of gene libraries either from genetically tractable strains or directly from DNA. We re-engineered a class 1 integron into counterselection markers on a plasmid or on the chromosome of a naturally competent Vibrio cholerae, which enabled capture of single cassettes in a sequence- and function-independent manner. When applied to Vibrio strains and genomic libraries, our tools recovered hundreds of single cassettes per assay with more than 99% specificity. We further subjected the library of cassettes generated by the hunter and gatherer tools to screens against phages ICP2 and T4, and identified nine phage-defence systems, including five previously undescribed. These tools enable rapid and large-scale recovery of integron cassettes that could be leveraged for functional gene discovery.

Journal Article