Search PubMedSearch

SEARCH · Search PubMed

Results for “Genes, Overlapping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Localization of Escherichia coli RNA polymerase binding sites on bacteriophage S13 and phiX174 DNA: alignment with restriction enzyme maps.

Escherichia coli RNA polymerase has been shown to bind to a limited number of Hind and Hae III restriction enzyme fragments. On S13 replicative form DNA there are three major binding sites, and the locations correlate with promoter sites at the beginning of genes a and B and a site overlapping gene D and the beginning of gene E. Two less definite binding sites have been localized, one in gene F and one at the gene G-H junction. In phiX174 replicative form DNA, five sites, each with apparently similar binding properties, have been located, four of which correspond exactly to binding sites in S13. One site, at the beginning of the B gene, could not be assigned to exactly the same location found in S13. This was due in part to differences in the restriction cleavage maps in the area of the DNA and possibly to the higher background of nonspecific binding in the phiX174 experiments. The location of two of the phiX174 sites at the beginning of genes A and D-E corresponds very well with transcription data, but the site at the start of the B gene indicates the promoter site is closer to the initiation sequence of the B protein than was previously suggested on the basis of transcription data.

Binding Sites

Genome-wide cis-expression Quantitative Trait Loci (eQTL) and transcriptomic signals reveal distinct molecular regulation across correlated feed efficiency traits.

INTRODUCTION: Feed efficiency (FE) is a complex trait which determines livestock production profitability, yet the molecular mechanisms behind it remain unclear. This study investigated the blood transcriptomic profile of lambs, alongside genotype data with the aim to uncover the genetic basis of FE traits such as absolute dry matter intake (DMIabsolute), DMI adjusted for body size (DMIadjusted), average daily live weight gain (ADG), and residual feed intake (RFI). MATERIALS AND METHODS: Bulk RNA-Seq and genotype data were analysed using three complementary approaches: differential gene expression (DGE) analysis, weighted gene co-expression network analysis (WGCNA), and cis-expression Quantitative Trait Loci (cis-eQTL) mapping. These methods were used independently to identify genes and regulatory networks associated with FE traits and to investigate evidence supporting multi-trait candidate gene selection. RESULTS: DGE analysis revealed 2, 24, 85 and 4 differentially expressed genes for DMIabsolute, DMIadjusted, ADG, and RFI (Padjusted < 0.05), functionally enriched in sensory perception, ATP-dependent chromatin remodeling, Notch signaling and immune response pathways. 9 gene modules significantly associated with the FE traits (P &#x2264; 0.05) with correlations ranging from r = -0.56 to 0.49, were identified using WGCNA. Single nucleotide polymorphism (SNP)-level cis-eQTL analysis identified 93 eSNPs associated with 74 genes (false discovery rate (FDR) < 0.05), while permutation-derived gene level analysis identified 280 eGenes (FDR < 0.2, empirical P < 0.03). Across the three analyses, applying thresholds of DGE (Padjusted < 0.05), WGCNA (correlation, P &#x2264; 0.05), and cis-eQTL gene-level significance (empirical P < 0.05), multiple overlapping genes were identified including DNMT3A, KANSL1, NCOR1 for DMIadjusted, ACOX2, FANCF, CIMIP2B, LOC101115106, ARMH2, LOC132657496 for ADG, and LOC114114576 for RFI representing regulators of variations in FE. DISCUSSION: The integration of DGE, WGCNA, and cis-eQTL analyses identified key genes and regulatory mechanisms associated with variation in FE traits. These results highlight that integrated multi-trait candidate gene identification approaches can reveal key genes that lower feed intake while maintaining animal growth, supporting breeding strategies aimed at improving efficiency and long-term economic sustainability in sheep.

average daily gain (ADG)

Genome-wide Association Studies of the Pathogenic Sphingosine-1-Phosphate Gene in Ulcerative Colitis.

BACKGROUND: Ulcerative colitis (UC) is a chronic inflammatory bowel disease that can lead to malignancies over time. Sphingosine-1-phosphate (S1P) receptor signaling affects lymphocyte trafficking and vascular integrity, influencing intestinal inflammation. This study aimed to identify S1P-related key genes in UC. METHODS: Differentially expressed genes (DEGs) between the UC and control groups were analyzed in the GSE87473 (training) dataset. Genes overlapping between the DEGs and S1P-related genes were considered candidate genes. These genes were incorporated into machine learning algorithms and subjected to expression analysis to identify key genes. Gene functions were determined through a gene&#x2013;gene interaction network, enrichment analysis, and immune cell infiltration analysis. In addition, transcription factor&#x2013;mRNA and mRNA&#x2013;miRNA&#x2013;lncRNA networks were constructed. Finally, reverse transcription&#x2013;quantitative polymerase chain reaction (RT-qPCR) was performed to evaluate the expression of key candidate genes in UC and control tissues. RESULTS: This study identified two key genes (SPHK2 and SPNS2) associated with UC. Notably, SPHK2 expression was lower and SPNS2 expression was higher in the UC group in both training and validation datasets and in clinical UC tissues (RT-qPCR). The area under the curve values of SPHK2 and SPNS2 exceeded 0.7 in both datasets, indicating that the genes had good diagnostic efficacy for UC. Consistently, the nomogram showed that the two genes had promising diagnostic value in UC. SPHK2 and SPNS2 were found to be localized to the plasma membrane. The correlations of the two genes with different immune cells showed significantly opposite trends. In particular, SPHK2 had the strongest positive correlation with M2 macrophages (r = 0.6) and the strongest negative correlation with neutrophils. Moreover, mRNA&#x2013;miRNA&#x2013;lncRNA and transcription factor&#x2013; mRNA networks of the key genes were constructed. CONCLUSION: This study suggests that SPHK2 and SPNS2 are key genes associated with UC, highlighting their potential as effective diagnostic biomarkers.

Humans

CTSG Suppresses Breast Cancer Progression by Inhibiting the EGFR/ERK Signaling Pathway and Enhancing CD8&#x207a; T Cell Activation.

BACKGROUND: Breast cancer (BC), the most common female malignancy, has metastasis as its main cause of mortality. Cathepsin G (CTSG) is involved in tumorigenesis and immunity. This study explores the role of CTSG in BC progression and CD8 + T cell regulation. METHODS: Differentially expressed genes and proteins (DEGs/DEPs) were analyzed using Limma, and core genes were screened using Random Forest (RF) and Least absolute shrinkage and selection operator (LASSO). CTSG expression was analyzed using GSE36295, the Cancer Genome Atlas (TCGA), reverse transcription-quantitative polymerase chain reaction (RT-qPCR), and western blot. Cell viability, proliferation, cell cycle, migration, and invasion were detected using Cell Counting Kit-8 (CCK8), 5&#x2011;Ethynyl&#x2011;2'&#x2011;deoxyuridine (EdU), flow cytometry, and Transwell assays, respectively. Sphere diameter was analyzed via sphere formation assay. Downstream mechanisms were examined using western blot, CCK8, flow cytometry, and Transwell assays. CD8 + T cell activity was examined using EdU, western blot, and flow cytometry. RESULTS: A total of 177 genes overlapped between GSE36295 DEGs and PDC000173 DEPs. CTSG was the hub gene identified by RF and LASSO. CTSG expression was significantly reduced in BC (P < 0.01). CTSG overexpression suppressed cell viability, proliferation, migration, invasion, sphere formation, and CD44 and CD133 expression (P < 0.01). CTSG up-regulation inhibited epidermal growth factor receptor (EGFR)/extracellular signal-regulated kinase (ERK) signaling axis and reduced cancer cell malignancy (P < 0.01). CTSG overexpression activated CD8 + T cells via EGFR/ERK inhibition, enhancing their cytotoxic effect on cancer cells (P < 0.01). CONCLUSION: CTSG inhibits BC malignancy and enhances CD8 + T cell function via EGFR/ERK inhibition.

Humans

Integrative ATAC-seq and RNA-seq analysis reveals lactation performance between Sewa sheep and East Friesian sheep.

Lactation performance is a pivotal economic trait in sheep production, yet its underlying epigenetic regulatory mechanisms remain poorly understood. In the present study, we integrated ATAC-seq and RNA-seq to compare chromatin accessibility landscapes and transcriptomic in mammary gland tissues from Sewa sheep (SWS) and East Friesian sheep (EFS). Histological characterization revealed that SWS exhibited significantly smaller mammary acini area, smaller lipid droplet area, and reduced lipid droplet diameter compared to EFS. ATAC-seq analysis identified 15,902 differentially accessible regions (DARs) between the two breeds, with motif enrichment analysis uncovering key transcription factors potentially governing lactation traits. RNA-seq analysis revealed 1,163 differentially expressed genes (DEGs), which were involved in lactation regulation. Integrated analysis identified 441 overlapping genes, and enriched in glycolysis/gluconeogenesis (e.g., PGAM1, ENO1) and pyruvate metabolism (e.g., ACACA, ACSS1, ACYP1). Collectively, our study provides new insights into the epigenetic regulatory mechanisms underlying lactation performance differences in sheep.

Animals

Lambda transducing phages derived from a FinO- R100::lambda cointegrate plasmid: proteins encoded by the R100 replication/incompatibility region and the antibiotic resistance determinant.

Three lambda transducing phages have been isolated from pEDR20, an R100::lambda cointegrate plasmid in which the lambda insertion inactivated the R100 finO gene. Physical analysis of the three phages showed that the lambda is inserted at kilobase coordinate 81.3 of R100. All three phages carry different amounts of R100 DNA in the left arm of lambda. Each pahge contains ISlb, the mer genes and the region between coordinate 81.3 and 88.6; thus, all contain the genes necessary for R100 replication. One phage, VA lambda 73, contains the entire r-determination of R100 in addition to the above DNA. Five proteins coded by the region between 81.3 and 88.6 were detected. These had subunit molecular weights of 10,400; 12,200; 16,200; 19,600; and 38,300. The first was made constitutively and the other four only from a lambda promoter. Other constitutive proteins were one from the cml fus region with a molecular weight of 22,400 (cml) and two from the str sul region with molecular weights of 31,500 (str?) and 30,100 (sul?). Mercuric ion induced synthesis of at least 10 proteins. Six of these were known from earlier work. The total size of the proteins which appear to derive from the mer genes exceeds by a factor of 1.5, the coding capacity of this region without overlapping genes. Some, or all of these extra proteins may be chromosomal in origin, possibly derepressed in response to mercury gene products.

Bacteriophage lambda

Prognostic value of genes associated with metastasis and propionate metabolism in rectal cancer.

BACKGROUND: Research indicates that alterations in propionate metabolic pathways play a critical role in cancer development and invasion. Postoperative metastatic recurrence remains a major cause of mortality in patients with rectal cancer. However, propionate metabolism-related genes (PMRGs) in rectal cancer remain insufficiently characterized. Therefore, this study aimed to identify prognostic biomarkers associated with lymph node metastasis and propionate metabolism and construct a risk&#x2011;prediction model for rectal cancer via bioinformatic analyses. METHODS: The Cancer Genome Atlas-Rectum Adenocarcinoma (TCGA-READ) and GSE87211 datasets, together with a curated PMRGs gene set, were used in this study. Pearson correlation analysis was performed to assess associations between overlapping genes (differentially expressed genes between READ and normal tissues, as well as between N0 and N1-N2 stages) and PMRGs, leading to the identification of candidate genes. Functional enrichment analyses were subsequently conducted to characterize the biological roles of these candidates. Prognostic biomarkers were identified using univariate Cox regression combined with least absolute shrinkage and selection operator (LASSO) regression, and a prognostic model was constructed accordingly. Independent prognostic validation was then performed. In addition, immune checkpoint profiling and immunotherapy response analyses were conducted across risk subgroups. Single-gene Gene Set Enrichment Analysis (GSEA) was applied to elucidate the pathways associated with the identified biomarkers. Finally, drug sensitivity analyses were performed. RESULTS: A total of 157 candidate genes were identified through the analytical pipeline. Functional enrichment analysis indicated that these genes were primarily involved in inflammatory response regulation and tumor necrosis factor (TNF) signaling pathways. Five prognostic biomarkers were subsequently identified and incorporated into a predictive model. External validation using the GSE87211 cohort confirmed the robustness of the model. Risk score and disease status were identified as independent prognostic factors. Six immune checkpoint molecules exhibited differential expression between risk groups. Correlation analyses revealed that the risk score was positively associated with most immune checkpoint genes. Single-gene GSEA demonstrated that the biomarkers were mainly enriched in ribosomal biogenesis and cell adhesion molecule-related pathways. Furthermore, 51 therapeutic agents exhibited significantly different half-maximal inhibitory concentration (IC50) values between risk subgroups. CONCLUSIONS: This study identified five biomarkers (CCL24, IGFBP3, ODC1, PYGM, and VKORC1) associated with lymph node metastasis and propionate metabolism pathways, providing a potential foundation for prognostic prediction in patients with rectal cancer.

Rectal cancer

Differential gene expression study in whole blood identifies candidate genes for psychosis in African American individuals.

Genome-wide association has identified regions of the genome that mediate risk for psychosis. It is possible that variants in these regions confer risk by altering gene expression. This work has predominantly been conducted in individuals of European descent and has focused narrowly on schizophrenia rather than psychosis as a syndrome. In the present study we investigated alterations in gene expression in African American individuals with a range of psychotic diagnoses to increase understanding of the etiology in an underserved population. We performed RNA-seq in whole bloody to survey the transcriptome in 126 patients with a psychosis-spectrum disorder and 217 healthy controls and applied differential gene expression analyses across the genome while controlling for age, sex, population stratification and batch. We found 18 differentially expressed genes (DEGs), some of the locations of the corresponding genes overlap with previously implicated regions for psychosis, but many of which were novel associations. Enrichment analysis of nominally significant genes (p&#xa0;<&#xa0;0.05) revealed overrepresentation of biological processes relating to platelet, immune and cellular function, and sensory perception. Weighted gene co-expression network analysis, applied to identify modules of co-expressed genes associated with psychosis, revealed 10 modules, one of which was significantly associated with psychosis. This module was significantly enriched for DEGs, and for platelet function. These results support the potential role of immune function in the etiology of psychosis, identify novel candidate gene expression phenotypes that correspond to both established and new genomic regions, in individuals of African American ancestry.

Humans

Population Genomics of Almond (Prunus dulcis) Reveals Region-Specific Selection and a Complex History of Domestication.

The domestication of perennial crops in the Mediterranean Basin remains unclear, particularly regarding the genomic consequences of human-mediated demographic shifts and selection. We analysed 8.1 million single nucleotide polymorphisms from 96 cultivated almond (Prunus dulcis) accessions from Europe, North America, Central Asia, and New Zealand, alongside four wild relatives. Population structure analyses revealed four geographically differentiated cultivated groups (Central Asian, North American, and two European) and three wild populations (P. spinosissima, P. orientalis, and P. fenzliana). Cultivated almonds retained high genetic diversity, consistent with weak domestication bottlenecks typical of outcrossing perennials. Elevated diversity and private allele counts in Central Asian cultivars, together with limited evidence of crop-wild gene flow, support Central Asia as an important reservoir of ancestral cultivated diversity that may have played a major role during the early stages of almond domestication. In contrast, allele sharing consistent with historical wild-to-crop introgression-especially involving P. orientalis-has contributed to the genomic composition of European and North American almonds. Genome-wide scans for selective sweeps showed most genes overlapping candidate sweep regions were population-specific, though often associated with similar biological functions, including stress responses and agronomic traits. This suggests repeated targeting of comparable pathways during and post-domestication, despite distinct selection histories. Notably, a subset of candidate genes detected in cultivated populations also occurs in wild relatives, particularly P. orientalis. This overlap is consistent with shared ancestral variation, introgression/gene flow between wild and cultivated lineages, and/or parallel adaptation. Altogether, our results support a complex domestication and diversification history for almonds, shaped by geographic expansion, gene flow with wild relatives, and recurrent selection acting in different regions. This study highlights wild relatives as important reservoirs of genetic diversity and emphasises the need for broader geographic sampling to clarify their contributions to almond domestication and adaptation.

Prunus dulcis

Transcriptome Analysis and Experimental Validation of Palmitoylation- Related Biomarkers in Atherosclerosis.

INTRODUCTION: Protein palmitoylation contributes to membrane localisation, signal transduction, and cell-fate regulation. It is closely associated with lipid metabolic dysfunction, immune inflammation, and vascular remodelling in atherosclerosis (AS). However, key palmitoylation-related transcriptomic markers and their potential causal associations with AS remain incompletely defined. METHODS: The Gene Expression Omnibus (GEO) dataset GSE100927 was used as the training cohort, and GSE43292 was used as an external validation cohort. Differentially expressed genes were identified using limma and intersected with palmitoylation-related genes to obtain palmitoylation-related differentially expressed genes (PRDEGs). Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were then performed using clusterProfiler. Two-sample Mendelian randomisation was used to evaluate potential causal relationships between characteristic genes and AS. Feature selection was conducted using random forest and support vector machine recursive feature elimination (SVM-RFE), and the overlapping genes selected by both methods were retained. Receiver operating characteristic (ROC) curves were used to assess diagnostic performance. A five-gene nomogram was constructed, and its clinical utility was evaluated using calibration curves and decision curve analysis (DCA). Gene set variation analysis (GSVA) was applied to compare pathway activity between high- and low-expression groups for each core gene. Single-cell analysis using Seurat and expression-based cell-cell communication analysis using CellChat were conducted with GSE159677, and upstream transcription factors were predicted using NetworkAnalyst. For in vivo validation, an AS model was established in ApoE&#x2078;/&#x2078; mice fed a high-fat diet, and aortic gene and protein expression were assessed by RT-qPCR and western blotting. RESULTS: In GSE100927, 51 PRDEGs were identified. GO and KEGG enrichment analyses highlighted pathways associated with regulation of monoatomic ion transport, sarcomere and myofibril organisation, and immune inflammation. Mendelian randomisation suggested a potential protective causal association between SLC7A7 and AS. By integrating MR with random forest and SVM-RFE feature selection, we prioritised five core genes: PLCB2, GMIP, NEXN, PLN, and SLC7A7. These genes showed good diagnostic performance in GSE43292. The resulting nomogram was well calibrated and demonstrated stable net benefit in decision curve and clinical impact curve analyses. Single-gene GSVA identified consistently activated pathways across multiple genes, including innate and adaptive immune recognition, calcium signalling and myocardial contraction/cardiomyopathy, extracellular matrix-receptor interaction, cell junction pathways, autophagy-lysosome pathways, and several metabolic programmes. At the single-cell level, PLCB2 and GMIP were predominantly expressed in T cells and macrophages, NEXN and PLN were enriched in vascular smooth muscle cells, and SLC7A7 was mainly expressed in macrophages. CellChat analysis indicated increased signals for immune-related ligand-receptor interactions. In ApoE&#x2078;/&#x2078; mice fed a high-fat diet, PLCB2, GMIP, and SLC7A7 were upregulated, whereas NEXN and PLN were downregulated; protein-level changes were concordant with the transcriptomic trends. DISCUSSION: These findings indicate that palmitoylation-related dysregulation in AS converges on immune inflammation, calcium signalling/contractile programmes, ECM remodelling, and autophagy-linked metabolism. The five-gene panel is supported by external validation, single-cell localisation to immune and vascular compartments, and concordant results in ApoE&#x2078;/&#x2078; mice. CONCLUSION: This study identified and validated five palmitoylation-related genes associated with AS. SLC7A7 showed a potential protective causal signal in MR analysis. The enriched pathway patterns linked these genes to immune inflammation, calcium signalling-contraction coupling, ECM remodelling, cell adhesion, and autophagy- associated metabolic reprogramming. The five-gene nomogram showed potential utility for diagnostic classification and decision support, nominating candidate biomarkers and pathway targets for AS molecular subtyping, diagnosis, and mechanistic investigation.

Atherosclerosis (AS)

NR3C1 Modulates Wnt Signalling to Influence the Invasiveness and Immune Features of Nonfunctioning Invasive Pituitary Adenomas.

Pituitary adenomas (PAs) are common intracranial tumours, and invasiveness in nonfunctioning invasive pituitary adenomas (NIPAs) predicts poor prognosis. The molecular mechanisms driving this phenotype remain unclear. This study explored the role of nuclear receptor subfamily 3 group C member 1 (NR3C1) in NIPA invasiveness and its regulation of Wnt signalling. mRNA expression profiles of 32 PA samples were generated by RNA-seq, and proteomic data from 19 samples were obtained by mass spectrometry. Immune-related differentially expressed genes (DEGs) were retrieved from GeneCards. Weighted gene coexpression network analysis identified modules and hub genes linked to invasiveness, while machine learning methods (support vector machine, LASSO, random forest) prioritised key genes. Gene set enrichment analysis (GSEA) assessed pathways associated with candidate gene expression. NR3C1 expression and function were validated by immunohistochemistry, Western blotting and invasion assays. Integration of transcriptomic, proteomic and immune-related datasets yielded 11 overlapping genes, with NR3C1 emerging as the top candidate. NR3C1 was significantly upregulated in NIPAs and demonstrated good discriminatory power by ROC analysis. GSEA associated high NR3C1 expression with Wnt pathway activation. Functional experiments confirmed that NR3C1 overexpression enhances the invasive capacity of PA cells. NR3C1 promotes the invasive phenotype of NIPAs by activating Wnt signalling. These findings suggest NR3C1 as a potential biomarker and therapeutic target for invasive pituitary adenomas.

Humans

Multiple Forms and Functions of Premature Termination by RNA Polymerase II.

Eukaryotic genomes are widely transcribed by RNA polymerase II (pol II) both within genes and in intergenic regions. POL II elongation complexes comprising the polymerase, the DNA template and nascent RNA transcript must be extremely processive in order to transcribe the longest genes which are over 1 megabase long and take many hours to traverse. Dedicated termination mechanisms are required to disrupt these highly stable complexes. Transcription termination occurs not only at the 3' ends of genes once a full length transcript has been made, but also within genes and in promiscuously transcribed intergenic regions. Termination at these latter positions is termed "premature" because it is not triggered in response to a specific signal that marks the 3' end of a gene, like a polyA site. One purpose of premature termination is to remove polymerases from intergenic regions where they are "not wanted" because they may interfere with transcription of overlapping genes or the progress of replication forks. Premature termination has recently been appreciated to occur at surprisingly high rates within genes where it is speculated to serve regulatory or quality control functions. In this review I summarize current understanding of the different mechanisms of premature termination and its potential functions.

RNA Polymerase II

Strong but diffuse genetic divergence underlies differentiation in an incipient species of marine stickleback.

Understanding how lineages proceed along the "speciation continuum" and how species boundaries are maintained over time remain central questions in evolutionary biology. Populations early in the speciation process can give us detailed insight into the reproductive barriers that first initiate speciation. In this study, we explore the nature of genomic divergence between two sympatric marine stickleback ecotypes from Atlantic Canada, "whites" and "commons". Males of each ecotype exhibit distinct nuptial colorations, nesting habits, and parental care strategies. Using population genomic analyses of SNPs and copy number variants (CNVs; deletions and duplications) we show that whites and commons consistently form distinct populations. We uncover genomic differentiation in the white ecotype characteristic of an incipient species, showing extremely low genome-wide differentiation (FST) and very recent divergence (~1 kya). Demographic analysis detected very low levels of ongoing gene flow between populations. Our results and prior genomic studies suggest that reproductive isolation is being maintained between ecotypes despite recent evidence that hybridization in nature does occur. Contrary to other systems, we found many small, but dispersed regions of high differentiation throughout the genome rather than explicitly within chromosomal inversions or the sex chromosomes. On chromosomes VII and XVI, we identified CNVs overlapping genes enriched for olfaction, which may play a role in differences in reproductive strategies between ecotypes. Ultimately, our results demonstrate that genome-wide rather than localized differences can underlie the early stages of divergence, and that this pattern is corroborated by both SNPs and CNVs.

Copy Number Variation

Multi-omics identification of therapeutic targets of compound sappan decoction in hepatocellular carcinoma.

BACKGROUND: Compound sappan decoction (CSD) is a multi-herbal traditional Chinese medicine formulation with clinical relevance in hepatocellular carcinoma (HCC). However, its therapeutic mechanisms remain unclear. METHODS: Bioactive compounds of CSD were identified and standardized using pharmacological and chemical databases. Potential targets were predicted via multiple target inference platforms. HCC-related genes were curated from comprehensive disease databases. Summary-data-based Mendelian randomization (SMR) was conducted to infer causal relationships between compound targets and HCC risk using large-scale quantitative trait loci (QTL) datasets and HCC genome-wide association study data. Colocalization analysis, protein-protein interaction (PPI) network construction, and GO/KEGG enrichment were performed on SMR-identified targets. Molecular docking evaluated binding affinities of representative compounds to prioritized targets. RESULTS: A total of 784 overlapping genes between predicted CSD targets and HCC-related genes were subjected to SMR analysis. Among these, 22 targets were significantly associated with HCC risk based on transcriptomic or proteomic QTLs and showed colocalization evidence. Notably, four targets (ADRB2, APOE, SYK, and PGF) were supported by both replication in an independent cohort and strong colocalization. These 22 targets were enriched in apoptosis, PI3K-Akt signaling, redox metabolism, and detoxification pathways. PPI analysis revealed central hubs including MMP9, BCL2, CASP1, and MCL1. Molecular docking demonstrated strong binding of APOE to quercetin, PGF to luteolin-7-olate, and SYK to kaempferol. CONCLUSIONS: CSD may exert therapeutic effects on HCC through modulation of genetically validated targets involved in tumor progression, inflammation, and metabolic reprogramming, supporting its potential clinical utility as an adjunctive treatment strategy. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s12672-026-04740-8.

Caesalpinia

Big data analytics for CLEC5A dynamics based on single cell genomics and proteomics reveal its diverse functions in human diseases.

BACKGROUND: CLEC5A (C-type lectin domain family 5 member A) is an innate immune receptor implicated in inflammatory signaling, contributing to hyperinflammatory responses in infections and sterile inflammation. However, CLEC5A dynamics in human diseases remain to be identified. Here, we systematically characterized CLEC5A dynamics in humans across cells, tissues, and disease states, and to explore the functional significance of CLEC5A in macrophage activation based on single-cell genomics. METHODS: With multi-omics (scRNA-seq, proteomics and big data analytics), we analyzed extensive human transcriptomic datasets (>42,000 samples) to profile CLEC5A expression by cell type, tissue, and disease. Single-nucleus RNA-seq (snRNA-seq) from pediatric congenital heart disease and a virtual CLEC5A gene knockout were also performed to characterize CLEC5A dynamics in humans. RESULTS: CLEC5A is highly enriched in innate immune cells, particularly in macrophages and neutrophils. Baseline CLEC5A in most tissues is low, but it is markedly upregulated in inflammatory and infectious diseases. CLEC5A expression has sex-specific differences in certain organs. Single-cell analysis showed that CLEC5A can be considered novel marker of proinflammatory macrophages with elevated cytokine production, antigen presentation, and impaired phagocytosis. Virtual CLEC5A knockout analysis identified coordinated perturbation of immune-regulatory pathways and overlapping genes linking CLEC5A to macrophage activation networks. CONCLUSION: CLEC5A is predominantly expressed in myeloid cells and acts as a key amplifier of inflammation in human diseases. Our findings highlight CLEC5A as a potential biomarker and therapeutic target in myeloid-driven hyperinflammatory conditions, warranting further experimental and translational validation.

Humans

Segment 8 of the influenza virus genome is unique in coding for two polypeptides.

In previous studies we showed that a ninth polypeptide with a molecular weight of approximately 11,000 (NS2) found in influenza virus-infected cells was unique, that it could be synthesized in vitro, and that its expression in vivo required early protein synthesis. On the basis of these results we suggested that one of the eight genome RNA segments of influenza virus codes for two polypeptides [Lamb, R.A., Etkind, P.R. & Choppin, P.W. (1978) Virology 91, 60-78]. We describe here differences in the electrophoretic mobility of the NS2 polypeptides of different strains of influenza A virus. These results provided further evidence that NS2 is virus coded and also made possible genetic studies using recombinants between two virus strains (HK and PR8) whose NS2 polypeptides differ. These studies showed that the gene for NS2 reassorts with that of the nonstructural polypeptide NS1, which is coded by genome segment 8. A mRNA for NS2 has been separated from that of NS1 and the other viral polypeptides by centrifugation and has been translated in vitro. Hybridization of genome segment 8 to the total mRNAs from infected cells specifically prevented the synthesis of NS2 and NS1. These results indicate that influenza virus genome segment 8 is transcribed into two separate mRNAs that code for two polypeptides, NS1 and NS2. Possible mechanisms for the transcription of the two mRNAs from either contiguous or overlapping genes are discussed.

DNA, Viral

A Computational Workflow for Prioritizing Microbial Metabolite-Associated Host Genes in Constipation-Predominant Irritable Bowel Syndrome.

No standardized computational pipeline exists for systematically prioritizing microbial metabolite-associated host genes and protein-ligand complexes from publicly available chemical, genomic, and structural databases. This article describes an eight-stage workflow that accepts a user-defined set of gut microbiota-derived metabolites and produces a ranked shortlist of candidate metabolite-associated host genes, enriched biological pathways, and structurally prioritized protein-ligand complexes for experimental follow-up. The pipeline integrates (i) chemoinformatic metabolite profiling; (ii) multi-database candidate target prediction using protein-chemical interaction and ligand-based target-prediction tool and a molecular docking program; (iii) differential gene expression analysis of publicly available transcriptomic data; (iv) target-differentially expressed gene overlap; (v) protein-protein interaction network construction and pathway enrichment; (vi) molecular docking with a molecular docking program; (vii) 200 ns molecular dynamics simulation using a molecular dynamics engine with a protein force field used for molecular dynamics simulations; and (viii) MM-PBSA binding free-energy estimation. As a worked example, nine gut microbiota-derived or microbiota-modified metabolites representing short-chain fatty acids, bile acids, tryptophan-derived metabolites, and urolithin A were processed using the public IBS-C rectal mucosal transcriptomic dataset GSE36701. The workflow ranked 17 unique predicted metabolite-associated genes that were differentially expressed in this dataset. Docking, molecular dynamics simulation, and MM-PBSA analyses structurally prioritized five metabolite-protein complexes: lithocholic acid-VDR, lithocholic acid-NR1H4/FXR, ursodeoxycholic acid-NR1H4/FXR, tryptamine-HTR2A (simulated in an explicit 1-Palmitoyl-2-oleoyl-sn-glycero-3-phosphocholine (POPC) lipid bilayer), and urolithin A-CASP3. The protocol is designed to be adaptable to other metabolite sets, disease transcriptomic datasets, and target classes; all outputs are hypothesis-generating computational predictions that require independent transcriptomic replication, protein-level validation, and functional ligand-response assays before causal or therapeutic conclusions can be drawn.

Irritable Bowel Syndrome

The value of structural variants to conservation genomics in the pangenome era.

Structural variants (SVs) comprise an axis of genetic diversity with strong consequences for phenotype and fitness, making them a potentially important target for conservation genomics. Here, we review how and why SVs can play a role in conservation genomics; the different types of SVs and how they can affect phenotype; and how pangenomes and long-read sequencing are illuminating their evolution in populations, including small populations and those of conservation concern. SVs comprise multinucleotide mutations including insertions, deletions, transpositions, inversions, and other multinucleotide mutations, often overlapping genes and other functional genome regions. As a result, SVs often play important roles in phenotypic evolution and local adaptation and can contribute substantially to genetic load in inbred populations. However, our understanding of the factors influencing SV diversity in populations is still in its infancy and is complicated by the vast range of sizes, effects, and mechanisms of formation of these mutations. We argue that SVs are an important axis of genetic diversity which should be characterized alongside more traditional metrics of genetic diversity in conservation contexts. There are a number of analytical challenges to detecting and studying SVs, but analyses aimed at understanding the role of SVs in inbreeding load and population health are rapidly becoming realizable goals, accelerated by new technologies and analytical approaches. New tools, including population-scale long-read sequencing and pangenome approaches, are beginning to make SVs accessible in ways which can be readily applied in conservation settings.

Genomic Structural Variation