Search PubMedSearch

SEARCH · Search PubMed

Results for “Candidate gene”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

GWAS-based identification of a candidate gene and development of a predictive KASP marker for seed protein and oil contents in soybean.

BACKGROUND: Soybean [Glycine max (L.) Merrill] is one of the most widely cultivated crops worldwide. Its seeds contain about 40% protein and 20% oil, serving as essential nutrient sources for humans. Given the nutritional importance of seed protein and oil, identifying genes that regulate their levels is crucial for improving soybean seed quality. OBJECTIVE: This study aimed to identify genetic factors associated with seed protein and oil content using a genome-wide association study (GWAS). METHODS: Seed protein and oil contents were quantified in 192 soybean mutant accessions in a mutant diversity pool (MDP), and GWAS was conducted using 17,631 SNPs filtered from genotyping-by-sequencing. Expression of a candidate gene was examined across seed developmental stages (R5 to R7), and a significant SNP was converted into a Kompetitive Allele-Specific PCR (KASP) marker for validation. RESULTS: GWAS detected significant SNPs associated with seed protein and oil content. Chr20_7635098 was identified as a nonsynonymous SNP located in the exon of Glyma.20g042400. This gene showed differential expression across seed developmental stages between mutant accessions with contrasting protein and oil contents. The KASP marker for Chr20_7635098 was validated using the MDP and six domestic soybean cultivars showing predictive accuracies of ≥ 80.50% for protein content and ≥ 61.18% for oil content. CONCLUSION: Overall, this study identified a candidate gene linked to both seed protein and oil content, providing valuable insights for molecular breeding strategies aimed at efficiently improving these nutritional traits.

Glycine max

Fourier-transform infrared-based genome-wide association study identifies candidate genes and variants for sow colostrum composition.

Sows with high prolificacy and better lactation traits are beneficial for weaned piglet number. Because the genetic basis of sow lactation traits remains elusive, genetic improvement for lactation traits lags behind that for litter traits, constraining the full realisation of genetic potential for large litters. Here, we measured 1&#xa0;060 Fourier-transform infrared (FTIR) wavenumbers and five predicted colostrum composition traits from sow colostrum samples. Heritability estimates for both the FTIR spectra and predicted traits ranged from moderate to high. Correlation analysis revealed that lactose percentage was negatively genetically correlated with the other four predicted traits (fat percentage, protein percentage, total solid content, and urea nitrogen content) and with 24&#xa0;h litter weight, which was positively genetically correlated with both protein and total solid content. Genome-wide association studies on the FTIR spectra and predicted traits identified 134 significant single-nucleotide polymorphisms (SNPs) (False discovery rate < 0.05), with most clustering on Sus scrofa chromosomes (SSC) 5 and 7. Among the candidate genes, two expressed in lactating mammary tissue have established roles in milk trait determination: (1) LALBA, which encodes a major colostrum protein and is responsible for lactose synthesis, and (2) BTN1A1, which mediates milk fat secretion. Additionally, the study detected two important candidate variants on SSC7: (1) rs691487382, which was colocalised with the expression quantitative trait locus signal for TRIM26 in the liver, a key metabolic organ supporting lactation, and (2) rs327923027, a missense variant located in a phylogenetically conserved domain of TRIM26, an E3 ubiquitin ligase implicated in liver homeostasis. Taken together, this study identifies, for the first time, candidate genes and variants for sow colostrum, providing genomic markers useful for genetic improvement.

Association analysis

Comparative Transcriptomic Analyses Identify Candidate Genes for Convergent Reproductive Shifts in a Bimodal Viviparous Amphibian.

Shifts in reproductive mode represent key evolutionary innovations that shape species' life histories and evolutionary trajectories. Species showing bimodal reproductive strategies with multiple independent origins offer a rare opportunity to gain insights into the adaptive processes and mechanisms underlying convergent traits. The fire salamander, Salamandra salamandra, is the only amphibian exhibiting intraspecific variation in reproductive mode across multiple independent reproductive shifts, enabling investigation of the transition between larviparity (females give birth to aquatic larvae) and pueriparity (females give birth to fully developed terrestrial juveniles) within a single species and across different timescales. Pueriparity is an adaptive innovation that skips the aquatic larval stage, allowing individuals to exploit habitats with no available water bodies. The fire salamander is larviparous across most of its range, but pueriparity has evolved independently at least three times: once in the early Pleistocene within S. s. bernardezi in the mountains of northern Spain, and more recently on two land-bridge islands (NW Spain) inhabited by S. s. gallaica. To identify candidate genes associated with these distinct reproductive modes, we compared gene expression profiles of the uterus and oviduct of pregnant females across two independent evolutionary transitions using RNA-sequencing. We detected shared changes in maternal gene expression among pueriparous S. s. bernardezi and S. s. gallaica relative to their larviparous counterparts, in addition to differences unique to each independent evolutionary transition. Functional enrichment analyses indicated that differentially expressed genes were associated with reproductive timing, angiogenesis, and maternal signalling, consistent with the phenotypic differences observed in the uterine environment and embryonic development between the two reproductive modes. This study represents an important first step towards understanding the genomic basis of the evolution of pueriparity in a remarkable bimodal reproductive system, and provides transcriptomic resources and candidate genes for future research into the genomic architecture underlying this poorly understood adaptive trait.

Animals

Genome-wide scans reveal candidate genes associated with wing morph differentiation in Tetrix japonica.

Wing dimorphism is an important dispersal-related trait in insects, but its genomic basis remains poorly understood in pygmy grasshoppers. Here, we integrated genome-wide single-nucleotide polymorphism (SNP) analyses, population structure inference, selection scans, and functional annotation to investigate genomic differentiation between long- and short-winged Tetrix japonica. Principal component analysis (PCA), ADMIXTURE, and phylogenetic analyses revealed weak genome-wide separation between morphs, indicating differentiation on a largely shared genetic background. Genome-wide scans based on the fixation index (FST), nucleotide diversity ratios, and Tajima's D, using 50-kb non-overlapping windows and empirical top-5% outlier thresholds, identified multiple candidate regions across seven chromosomes. The broader long- and short-winged candidate sets spanned 9.35&#xa0;Mb and 9.37&#xa0;Mb and directly overlapped 82 and 77 genes, respectively. Candidate genes were associated with signaling/hormone regulation, membrane transport, metabolism, cytoskeletal organization, extracellular matrix structure, and development. Short-winged candidate genes were significantly enriched for ABC-type transporter activity and ATP hydrolysis activity. Because all individuals originated from a single laboratory-maintained population with weak genome-wide structure, these regions should be regarded as candidate loci from a screening-stage analysis that require validation in independent populations and by functional assays, rather than as confirmed targets of selection.

Animals

Whole-genome safety assessment of Loigolactobacillus coryniformis WBB05 and identification of a candidate gene for aerobic reuterin production.

This study reports on the safety profile of Loigolactobacillus coryniformis WBB05 for food industry applications and identifies glycerol-3-phosphate oxidase (GlpO) as a candidate gene associated with aerobic reuterin production. The safety of L. coryniformis WBB05 was evaluated through whole-genome sequencing, phenotypic analysis of haemolytic activity and determination of minimum inhibitory concentrations (MICs) of antibiotics. Comparative genomic analysis was performed to identify candidate genetic determinants for aerobic reuterin production. The draft genome (2.83 Mb, 179 contigs) harboured no known virulence factors, acquired antimicrobial resistance (AMR) genes or biogenic amine biosynthetic genes. Prophage analysis identified only one incomplete prophage region, and four CRISPR-Cas systems (212 spacers) were consistent with phage defence capacity. Secondary metabolite analysis revealed biosynthetic gene clusters encoding a coagulin-like bacteriocin. No &#x3b2;-haemolytic activity was observed. The MICs of all antibiotics tested were below the European Food Safety Authority cut-off values except for kanamycin (128&#xa0;mg/L), although no acquired AMR genes were detected. Comparative genomic analysis revealed that L. coryniformis WBB05 possesses two putative copies of GlpO, a gene not detected in publicly available genomes of Limosilactobacillus reuteri, which produces reuterin only under anaerobic conditions. These findings support the use of L. coryniformis WBB05 as a safe adjunct culture for dairy applications and highlight GlpO as a candidate determinant of aerobic reuterin production. Further studies comparing GlpO-positive and GlpO-negative strains under aerobic and anaerobic conditions are warranted to confirm the role of GlpO.

Loigolactobacillus coryniformis

Identification of ultrasound-associated gene candidates in myeloid cells and construction of a prognostic risk model for acute myeloid leukemia.

BACKGROUND: Incorporating ultrasound (US) treatment sensitivity analysis may improve the treatment of acute myeloid leukemia (AML). METHODS: This study integrated single-cell and bulk datasets for analysis. Differential expression analysis between US-treated and control samples was performed using limma package. The AUCell package was used to calculate US-associated scores in the single-cell dataset. Differentially expressed genes (DEGs) between the specific groups were identified, followed by intersection analysis with previously identified DEGs. Univariate regression, Least Absolute Shrinkage and Selection Operator (LASSO) analysis (using the glmnet package), and stepwise multivariate regression (using the MASS package) were used to refine the candidate genes and to construct a risk model. The model genes were validated using in vitro experiments. Enrichment analysis was conducted using gene set enrichment analysis (GSEA), and immune infiltration was evaluate by single-sample GSEA (ssGSEA) and ESTIMATE algorithms. The correlations between RiskScores and drug sensitivity were analyzed by oncoPredict package. Finally, tumor mutational burden (TMB) and genomic mutations were compared between the risk groups. RESULTS: Nine prognostic signatures (SPINK2, HNRNPAB, SH3BGRL3, CLEC11A, ITGA4, RPL39L, MX1, HEXIM1, and MAP4K4) were identified. Particularly, low expression of SPINK2 attenuated the activity and invasion of AML cells. High-risk group had higher immune cell infiltration. Eight drugs were predicted to be correlated with the RiskScore model. DNMT3A and RUNX1 showed higher mutation frequencies in the high-risk group, whereas KIT and MUC16 showed higher mutation frequencies in the low-risk group. CONCLUSION: The RiskScore model established in this study provides a theoretical basis for clinically screening responsive populations and optimizing treatment strategies.

Humans

Identification of candidate genes and regulatory mechanisms for thick shank phenotype of Dong Tao chickens.

The Dong Tao Chicken has garnered widespread attention owing to its hallmark trait of remarkably thick shanks. However, the genetic mechanisms and molecular basis underlying this unique phenotype remain largely elusive to date. We carried out crossbreeding trials, performed genome-wide association study (GWAS), and completed transcriptome sequencing of leg tissues. Crossbreeding experiments confirmed that the thick-leg trait of Dong Tao chickens is polygenically controlled. Furthermore, GWAS analysis on the F&#x2082; segregating population screened NAKIN3 as a key candidate gene associated with this trait. Transcriptomic results indicated that tarsometatarsal skin acts as the primary tissue regulating shank circumference growth. Moreover, ACTB was recognized as a core gene driving dermal thickening of the tarsometatarsus. By interacting with IRF family members, TLR4, SOS1, IFNG, STAT family members, EGF and other molecules, ACTB participates in multiple Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways including the Toll-like receptor signaling pathway and cytokine-cytokine receptor interaction pathway to coordinately modulate dermal thickening in the tarsometatarsal skin of Dong Tao chickens. This study identifies key loci and genes governing shank development, offering theoretical basis and genetic resources for dissecting molecular mechanisms and selective breeding of Dong Tao chickens. Nevertheless, these preliminary findings need to be further verified by expanding the experimental population and sample size, together with additional independent validation experiments.

Dong Tao chicken

Integrating GWAS and Transcriptome Analysis Identifies Candidate Genes for Kernel Starch Quality Traits in Maize.

Maize (Zea mays L.) starch quality is a complex trait with significant implications for grain processing and industrial applications. However, the genetic basis underlying starch quality, particularly for gelatinization and thermodynamic properties, remains poorly understood. In this study, we evaluated 12 starch quality traits, including seven gelatinization characteristics, four thermodynamic traits, and kernel starch content (KSC) in a diverse panel of 335 maize inbred lines. Considerable phenotypic variation was observed for all traits. A total of 228 quantitative trait loci (QTLs) were significantly associated with 12 starch quality traits through genome-wide association studies (GWAS). By integrating a dynamic transcriptome analysis of two maize inbred lines with contrasting starch quality, we identified 60 candidate genes. One gene, waxy1, encoding a starch synthase, was found to be associated with enthalpy of gelatinization (&#x394;Hgel) and pasting temperature (Ptemp). Six variants in waxy1 contributed to natural variation in &#x394;Hgel and Ptemp, and a cost-effective InDel and two PARMS-based molecular markers were developed and validated in 144 maize inbred lines, enabling efficient marker-assisted selection. Our findings provide key genes and molecular markers for high-quality maize breeding with improved starch properties.

Zea mays

Genome-wide characterization of the sugar transporter protein family identifies candidate genes for bacterial wilt resistance breeding in tobacco.

Sugar transporter proteins (STPs) play pivotal roles in hexose allocation and plant stress responses. However, systematic characterization of the STP family in tobacco (Nicotiana tabacum) and its involvement in Ralstonia solanacearum resistance remains unclear. In this study, 37 NtSTP genes were identified and classified into six groups, with Group VI being the most conserved and Group V exhibiting dicot-specific expansion. Gene structure and conserved motif analyses revealed that most NtSTP members possess the typical MFS_STP domain, although variations in exon-intron organization and motif composition suggested functional divergence. Tandem duplication (TD) served as the primary driver of NtSTP family expansion, and Ka/Ks values of all paralogous pairs were less than 1, indicative of purifying selection. Promoter cis-element analysis revealed a complex regulatory network involving hormone signaling (ABA, JA, SA, GA, ET), stress responses, and light signaling. RT-qPCR expression profiling revealed that ten NtSTP genes (NtSTP1, 5, 7, 21, 22, 24, 26, 27, 28, and 29) exhibited significant transcriptional upregulation upon R. solanacearum infection. Specifically, NtSTP5, NtSTP7, NtSTP21, NtSTP22, NtSTP24, NtSTP26, and NtSTP27 peaked at 12&#xa0;h post-inoculation (hpi), whereas NtSTP1, NtSTP28, and NtSTP29 reached their highest expression levels at 24 hpi. By contrast, NtSTP6, NtSTP13, and NtSTP30 displayed reduced expression upon R. solanacearum infection. These expression patterns indicate functional diversification within the NtSTP family and imply that these members may be transcriptionally modulated during plant responses to R. solanacearum. The present work provides preliminary and valuable candidate gene resources that may facilitate future disease resistance breeding programs in tobacco.

NtSTP gene family

Prioritizing Parkinson's disease risk-associated mitochondrial candidate genes via multi-omics integrative analysis.

BACKGROUND: Mitochondrial dysfunction has been implicated in Parkinson's disease (PD), but the genetically regulated mitochondrial genes associated with PD risk remain incompletely defined. METHODS: We conducted a summary-data-based genetic epidemiology study integrating summary-based Mendelian randomization (SMR), Heterogeneity in dependent instruments (HEIDI) filtering, and Bayesian colocalization to prioritize mitochondrial-related molecular features associated with PD risk. Mitochondrial-related genes were defined using MitoCarta3.0. Genetically predicted gene expression and plasma protein abundance were evaluated using expression quantitative trait loci (eQTL) data from eQTLGen and GTEx v8, and protein quantitative trait loci (pQTL) data was assessed using International Parkinson's Disease Genomics Consortium (IPDGC) as the discovery genome-wide association study (GWAS) and FinnGen as the replication dataset. Prespecified QTL analyses were interpreted using FDR correction, HEIDI filtering, and colocalization support. DNA methylation QTL analysis, mitochondrial phenotype MR, and single-nucleus RNA-seq analysis were performed as complementary analyses. RESULTS: In the primary eQTL analysis, higher genetically predicted TTC19 expression was associated with lower PD risk (OR = 0.80, 95% CI: 0.74-0.87, PPH4&#x202f;= 0.80), whereas higher MALSU1 expression was associated with increased PD risk (OR = 2.21, 95% CI: 1.59-3.06, PPH4&#x202f;= 0.96). Both associations survived FDR correction, passed HEIDI filtering, and showed colocalization support. GTEx whole-blood data supported the direction of the TTC19 association. No mitochondrial protein reached significance after FDR correction and colocalization filtering in the primary pQTL analysis. Complementary methylation analysis highlighted cg06270993 as an exploratory regulatory signal for MALSU1. CONCLUSIONS: This MR-colocalization study prioritizes TTC19 and MALSU1 as genetically supported mitochondrial-related candidate genes associated with PD risk. Further validation is required to define their functional roles in PD pathogenesis.

Humans

Combining QTL mapping and RNA-Seq reveals candidate genes controlling flag leaf width in foxtail millet.

BACKGROUND: The flag leaf, a crucial component of plant architecture, significantly influences final grain yield in crops, including foxtail millet (Setaria italica L.). Optimizing flag leaf size is considered an effective strategy for enhancing grain yield potential under higher planting densities. However, the genetic mechanism underlying flag leaf size, particularly flag leaf width (FLW), remains largely unknown under varying planting densities in foxtail millet. RESULTS: An FLW phenotype variation analysis was conducted across multiple planting densities using a recombinant inbred line (RIL) population derived from Heizhigu (narrow leaf) and Changnong 35 (wide leaf). Based on a high-density genetic map with 3795 Bin markers, 11 flag leaf width (FLW) QTLs were identified on chromosomes 3, 5, and 6, explaining 2.35%-36.06%. Among these, qFLW5-2 was a major QTL, detected consistently across 3 environments and explaining a large proportion of FLW variation. The QTL was further validated with 9 InDel markers with its candidate region across different planting densities. Moreover, RNA-seq revealed 2,293 and 2,338 differentially expressed genes (DEGs) between biparents at heading stage and grain filling stage, respectively. There were 11 and 9 DEGs within the location range of qFLW5-2 among 2 comparison groups (HZG-H_vs_CN35-H and HZG-G_vs_CN35-G). Combining QTL mapping and RNA-seq, we speculated that Seita.5g134600 (encoding an auxin responsive protein Aux/IAA) and Seita.5G123900 (encoding a cytochrome P450 family protein) as key candidate genes for qFLW5-2. Furthermore, variation analysis confirmed that the lines or germplasm with Seita.5G1346005UTR277+ allele, both within the RIL population and natural populations, exhibited significantly wider leaves than those with Seita.5G1346005UTR277- allele. These findings advance our understanding of the genetic and molecular regulatory mechanisms governing flag leaf growth. CONCLUSIONS: This study elucidates genetic and molecular mechanism regulating flag leaf growth and development in foxtail millet. The results provide a theoretical foundation for improving plant architecture and facilitating molecular marker-assisted breeding in this crop.

Quantitative Trait Loci

Automating candidate gene prioritization with large language models: from naive scoring to literature-grounded validation.

MOTIVATION: Identifying promising therapeutic targets from thousands of genes in transcriptomic studies remains a major bottleneck in biomedical research. While large language models (LLMs) show potential for gene prioritization, they suffer from hallucination and lack systematic validation against expert knowledge. RESULTS: The framework identified 609 sepsis-relevant genes with >94% filtering efficiency, demonstrating strong enrichment for inflammatory pathways including TNF-&#x3b1; signaling, complement activation, and interferon responses. Literature validation yielded 30 ultra-high confidence therapeutic candidates, including both established sepsis genes (IL10, TREM1, S100A9, NLRP3) and novel targets warranting investigation. Benchmark validation against expert-curated databases achieved 71.2% recall, with systematic correlation between computational confidence and evidence quality. The final candidate set balanced discovery (11 novel genes) with validation (19 known genes), maintaining biological coherence throughout the filtering process. This framework demonstrates that rigorous methodology can transform unreliable LLM outputs into systematically validated biological insights. By combining computational efficiency with literature grounding, the approach provides a practical tool for prioritizing experimental validation efforts. The modular design enables adaptation to other diseases through knowledge base substitution, offering a systematic approach to literature-guided biomarker discovery. AVAILABILITY AND IMPLEMENTATION: We developed a two-stage computational framework that combines LLM-based screening with literature validation for systematic gene prioritization. Starting with 10&#xa0;824 genes from the BloodGen3 repertoire, we applied multi-criteria evaluation for sepsis relevance, followed by retrieval-augmented generation using 6346 curated sepsis publications. A novel faithfulness evaluation system verified that LLM predictions aligned with retrieved literature evidence. Source code and implementation details are available at https://github.com/taushifkhan/llm-geneprioritization-framework, vector database at https://doi.org/10.5281/zenodo.15802241, and Interactive demonstration at https://llm-geneprioritization.streamlit.app/.

Humans

Multidimensional GWAS analyses on longitudinal phenotypes reveal candidate genes regulating multi-stage egg production traits in Wannan yellow chicken.

Egg production performance directly determines the economic viability of indigenous chicken breeding. However, the genetic regulation of multi-stage egg production traits remains difficult to characterize due to their complex and dynamic nature. Here, we integrated a multidimensional GWAS framework, including single-trait GWAS, multi-trait GWAS (MTAG), and longitudinal trajectory-based GWAS (TrajGWAS), to identify stage-specific and shared genetic effects underlying egg production traits in Wannan yellow chickens (WNY). Whole-genome sequencing of 354 WNY hens (10&#xd7; depth) and quality control yielded 14,253,816 SNPs for analysis. Selective sweep analyses comparing red jungle fowl, commercial layers, and WNY identified a genomic region containing IGF1 under significant selection pressure. Single-trait GWAS identified SNPs 4_57990480 (BMPR1B) and 17_370912 (LOC112531479) associated with egg production across three laying stages (21-30, 31-40, and 21-40 weeks). MTAG further identified loci 8_4336468 (FASLG) and 21_654726 (CHD5) with shared effects across the laying period, whereas TrajGWAS revealed longitudinal associations involving PRKG1 and identified dynamic loci associated with clutch traits, including GRID1. For clutch traits, stage-specific loci were detected for average clutch size (ACS) and maximum clutch size (MCS), including SNP 8_8542036 at 21-30 weeks, PROK1 at 31-40 weeks, and CUL5, ALKBH8 across the entire laying period. These results demonstrate that integrating complementary GWAS strategies improves the resolution of genetic architecture underlying egg production traits by capturing trait-specific, shared, and stage-dependent genetic effects. The identified GWAS loci and selective-sweep candidate regions provide insights into the genetic architecture of egg production traits and breed differentiation.

Egg production

De novo transcriptome meta-analysis reveals candidate genes involved in life-stage transitions for RNAi-mediated management of the citrus root weevil (Diaprepes abbreviatus).

BACKGROUND: The citrus root weevil, Diaprepes abbreviatus, is a destructive agricultural pest for which molecular control options remain limited due to historically sparse genomic resources. Leveraging a comprehensive de novo transcriptome, we investigated developmental gene regulation across larval, pupal, and adult stages and identified essential targets for RNA interference (RNAi)-based intervention. RESULTS: Stage-resolved transcriptomic analyses revealed extensive transcriptional reprogramming associated with metabolism, detoxification, cuticle biosynthesis, endocrine signaling, and sensory perception. Among these, chitin synthase (DaCHS) emerged as a critical developmental gene, exhibiting pronounced up-regulation during late larval and pupal stages corresponding to intensive cuticle synthesis. Phylogenetic and structural analyses demonstrated that DaCHS is highly conserved among insects and retains canonical catalytic domains and transmembrane topology. Alpha Fold-based structural modeling and molecular docking confirmed stable interaction of DaCHS with its substrate, N-acetylglucosamine, supporting functional conservation of enzymatic activity. Oral delivery of DaCHS double-stranded RNA induced robust transcript suppression, leading to significant mortality and severe developmental defects, including larval and pupal abnormalities, and adults with disrupted wing and abdominal morphogenesis. CONCLUSION: These findings establish DaCHS as an indispensable gene for D. abbreviates development and validate transcriptome-guided RNAi as a powerful framework for target discovery. This work provides a strong molecular foundation for developing RNAi-based strategies that can be integrated into sustainable management programs for citrus root weevil control. &#xa9; 2026 Society of Chemical Industry.

Animals

Functional characterization of the 9q34.13 locus identifies RAPGEF1 as a candidate gene modulating risk for melanoma and nevi via RAS activation.

Genome-wide association studies identified a melanoma- and nevus count-associated locus on chromosome band 9q34.13. Fine-mapping and melanocyte expression data collectively suggest two potential risk genes with opposite associations with risk: higher levels of Rap guanine nucleotide exchange factor 1 (RAPGEF1) and lower levels of uridine-cytidine kinase 1 (UCK1). Colocalization analyses and conditional transcriptome-wide association studies (TWASs) suggest multiple causal cis-regulatory sequence variants in partial linkage disequilibrium (LD) to each other. Melanocyte capture-HiC and CRISPR inhibition demonstrated regulatory interactions between fine-mapped variants and the RAPGEF1 and UCK1 promoters. Focusing on RAPGEF1, we demonstrate that RAPGEF1 expression promotes melanocyte growth and drives colony formation of human immortalized melanocytes. Following treatment with human epidermal growth factor (EGF), RAPGEF1 overexpression activated both RAP1 and RAS. Further, we show that RAPGEF1 expression is significantly enriched in melanomas that lack strongly activating RAS-MAPK pathway mutations, which suggests that RAPGEF1 may promote oncogenic RAS-MAPK pathway signaling in melanomas. Furthermore, in these tumors, we provide preliminary evidence to support the prognostic relevance of RAPGEF1 expression in individuals whose melanomas lack RAS or BRAF mutations. Together with other recent studies, these data suggest that germline variation influencing RAS activation may play a key role in nevus development and melanoma risk.

GWAS

Cr3a, a candidate gene conferring fruit cracking resistance, was fine-mapped in an introgression line of Solanum lycopersicum L.

In the cultivation and production of tomato (Solanum lycopersicum L.), fruit cracking is a prevalent and detrimental issue that significantly impacts the esthetic quality and commercial value of the fruit. The complexity of the trait has resulted in a slow advancement in research aimed at identifying genes that influence tomato fruit cracking and the underlying regulatory mechanisms. In this study, a sub-introgression population for tomato crack-resistant fruit has been constructed from the cross between S. lycopersicum 1052 and Solanum pennellii LA0716, followed by 11 generations of selfing. Utilizing specifically designed InDel markers, the tomato crack-resistant gene, Cr3a, was fine-mapped, cloned, and its functionality was confirmed through transgenic and gene-knockout approaches. The precise localization of Cr3a was delineated to a 30&#x2009;kb genomic region on chromosome 3, corresponding to the gene Sopen03g034650 in S. pennellii and Solyc03g115660.3 in the Heinz1706 variety. An integrated transcriptomic and metabolomic analysis of fruits with and without the Cr3a gene was finally conducted to elucidate the intricate regulatory mechanisms associated with Cr3a. The findings revealed a molecular regulatory network for tomato fruit crack resistance, characterized by 7 key metabolites, 13 pivotal genes, and 4 critical pathways: the phenylpropanoid biosynthesis pathway, the phenylalanine, tyrosine, and tryptophan biosynthesis pathway, the linolenic acid metabolism pathway, and the cysteine and methionine metabolism pathway. In summary, this research provides novel insights into the molecular underpinnings of tomato fruit crack resistance and holds substantial promise for accelerating the molecular breeding of tomatoes with enhanced fruit crack resistance.

Solanum lycopersicum

Gut microbiota-derived metabolites target C5AR1/KDM2A/HCAR3 axis in inflammatory bowel disease: a multi-machine learning algorithms and molecular docking study.

BACKGROUND: Inflammatory bowel disease (IBD) is a chronic recurrent disorder. Gut microbiota-derived metabolites regulate intestinal homeostasis, but their molecular mechanisms in IBD remain unclear. Current studies lack systematic "microbiota-metabolite-target" network mining with multi-method validation. This study integrates network pharmacology, three machine learning algorithms, and molecular docking to construct this regulatory network in IBD. METHODS: Transcriptome data were obtained from the Gene Expression Omnibus (GEO) database. Differentially expressed genes (DEGs) were identified using limma (p < 0.05, |log2FC| > 0.5). Weighted gene co-expression network analysis (WGCNA) with an optimal soft threshold of &#x3b2; = 7 was performed to identify key module genes. Candidate genes were obtained by intersecting DEGs, gut microbiota-associated genes from the gutMGene database, and WGCNA module genes. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were conducted to explore the functional roles of candidate genes. Core genes were identified using three machine learning algorithms (LASSO, Boruta, and SVM-RFE), followed by protein-protein interaction (PPI) network analysis. Molecular docking was performed to assess the binding affinities between hub proteins and gut microbiota-derived metabolites. RESULTS: A total of 885 DEGs were identified between the IBD and control groups, including 463 upregulated and 422 downregulated genes. WGCNA identified 280 key module genes from the purple and yellow modules. The intersection of DEGs, gut microbiota-associated genes, and WGCNA module genes yielded 19 core candidate genes. PPI network analysis combined with three machine learning algorithms jointly identified C5AR1, KDM2A, and HCAR3 as core hub genes. ROC curve analysis demonstrated that all three hub genes achieved AUC values greater than 0.7 in both the training and validation sets, indicating excellent diagnostic performance for IBD. Enrichment analysis revealed significant associations with the TNF, NF-&#x3ba;B, and IL-17 signaling pathways. Molecular docking confirmed stable binding of C5AR1 with 1,3-Diphenylpropan-2-Ol (-7.87 &#xb1; 0.83 kcal&#xb7;mol-&#xb9;) and HCAR3 with 3-Indolepropionic Acid (-6.35 &#xb1; 0.70 kcal&#xb7;mol-&#xb9;), both below -5.0 kcal&#xb7;mol-&#xb9;. CONCLUSION: This study first constructs a "gut microbiota-metabolite-hub gene" axis in IBD, providing a computational framework for microbiota-targeted precision therapy, and identifying C5AR1/KDM2A/HCAR3 as computationally predicted diagnostic biomarkers and 1,3-Diphenylpropan-2-Ol/3-Indolepropionic Acid as candidate intervention molecules that warrant further experimental validation.

Molecular Docking Simulation