Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Proteomics-Based Identification of the Pyroptosis-Related Biomarker PCSK9 and Its Association With the Pathogenesis of Rheumatoid Arthritis.

Rheumatoid arthritis (RA) is a common autoimmune disease, and early diagnosis is critical for effective treatment. This study aims to identify potential biomarkers related to pyroptosis through serum proteomics analysis, offering new insights for the early diagnosis of RA. We enrolled 100 participants, including 50 patients with RA and 50 healthy controls. Serum samples were collected and analyzed using high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) for proteomics profiling. Differential protein expression analysis and functional annotation revealed significant upregulation of pyroptosis-related proteins in the serum of patients with RA. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses, along with protein-protein interaction (PPI) network analysis, showed that these proteins are involved in inflammation and immune pathways, particularly the activation of the NOD-like receptor protein 3 (NLRP3) inflammasome. Enzyme-linked immunosorbent assay (ELISA) validation confirmed a significant increase in PCSK9 levels in patients with RA, suggesting that PCSK9 may play a key role in the pathogenesis of RA. This study provides new directions for biomarker research in RA, particularly regarding the potential involvement of the pyroptosis pathway, with significant clinical application prospects.

Humans↗

Network based approach identifies miR-145-3p as a central regulatory hub associated to the progression from localized to metastatic medullary thyroid carcinoma.

Medullary thyroid carcinoma (MTC) is a neuroendocrine tumor originating from calcitonin producing C-cells and accounts for 1-5% of thyroid cancers. Total thyroidectomy is curative in localized disease (N0), whereas lymph node metastases (N1) are associated with poorer prognosis. However, the molecular mechanisms driving the metastatic shift remain poorly understood. This study aimed to identify miRNA features linked to metastatic spread in MTC, focusing on the transition from N0 to N1. Co-expression networks were constructed for N0 and N1 tumors, and differential connectivity analysis was used to identify key miRNAs acting as regulatory hubs. Functional annotation of their target genes was performed using the Kyoto Encyclopedia of Genes and Genomes (KEGG), Gene Ontology (GO), and Reactome pathway analyses. Validation experiments were carried out in MTC cells to evaluate the effects of selected miRNAs on cell proliferation, survival, and MAPK pathway activation. Network analysis revealed distinct miRNA co-expression patterns between N0 and N1 tumors. Differential network analysis highlighted miR-145-3p as a central regulatory hub, exhibiting 29 altered co-expression changes and a marked loss of connectivity in N1. Target enrichment identified 59 validated genes, including key oncogenic drivers such as MYC, PTEN, BCL2, PIK3CA, AKT1, and MAPK7. In MTC cells, simultaneous inhibition of miR-145-3p together with its top co-expressed miRNAs increased proliferation and survival, and enhanced ERK phosphorylation, indicating MAPK pathway activation and a shift toward a more aggressive phenotype. In conclusion, this study identifies a miRNA regulatory hub centered on miR-145-3p that is associated with metastatic progression and highlights the value of network-based approaches in uncovering mechanisms of cancer dissemination. © 2026 The Author(s). The Journal of Pathology published by John Wiley & Sons Ltd on behalf of The Pathological Society of Great Britain and Ireland.

MAPK signaling↗

A proteomic analysis of salivary glands of female Anopheles gambiae mosquito.

Understanding the development of the malaria parasite within the mosquito vector at the molecular level should provide novel targets for interrupting parasitic life cycle and subsequent transmission. Availability of the complete genomic sequence of the major African malaria vector, Anopheles gambiae, allows discovery of such targets through experimental as well as computational methods. In the female mosquito, the salivary gland tissue plays an important role in the maturation of the infective form of the malaria parasite. Therefore, we carried out a proteomic analysis of salivary glands from female An. gambiae mosquitoes. Salivary gland extracts were digested with trypsin using two complementary approaches and analyzed by LC-MS/MS. This led to identification of 69 unique proteins, 57 of which were novel. We carried out a functional annotation of all proteins identified in this study through a detailed bioinformatics analysis. Even though a number of cDNA and Edman degradation-based approaches to catalog transcripts and proteins from salivary glands of mosquitoes have been published previously, this is the first report describing the application of MS for characterization of the salivary gland proteome. Our approach should prove valuable for characterizing proteomes of parasites and vectors with sequenced genomes as well as those whose genomes are yet to be fully sequenced.

Animals↗

Plasma Proteome Database as a resource for proteomics research.

Plasma is one of the best studied compartments in the human body and serves as an ideal body fluid for the diagnosis of diseases. This report provides a detailed functional annotation of all the plasma proteins identified to date. In all, gene products encoded by 3778 distinct genes were annotated based on proteins previously published in the literature as plasma proteins and the identification of multiple peptides from proteins under HUPO's Plasma Proteome Project. Our analysis revealed that 51% of these genes encoded more than one protein isoform. All single nucleotide polymorphisms involving protein-coding regions were mapped onto the protein sequences. We found a number of examples of isoform-specific subcellular localization as well as tissue expression. This database is an attempt at comprehensive annotation of a complex subproteome and is available on the web at http://www.plasmaproteomedatabase.org.

Amino Acid Motifs↗

Modeling a whole organ using proteomics: the avian bursa of Fabricius.

While advances in proteomics have improved proteome coverage and enhanced biological modeling, modeling function in multicellular organisms requires understanding how cells interact. Here we used the chicken bursa of Fabricius, a common experimental system for B cell function, to model organ function from proteomics data. The bursa has two major functional cell types: B cells and the supporting stromal cells. We used differential detergent fractionation-multidimensional protein identification technology (DDF-MudPIT) to identify 5198 proteins from all cellular compartments. Of these, 1753 were B cell specific, 1972 were stroma specific and 1473 were shared between the two. By modeling programmed cell death (PCD), cell differentiation and proliferation, and transcriptional activation, we have improved functional annotation of chicken proteins and placed chicken-specific death receptors into the PCD process using phylogenetics. We have identified 114 transcription factors (TFs); 42 of the bursal B cell TFs have not been reported before in any B cells. We have also improved the structural annotation of a newly sequenced genome by confirming the in vivo expression of 4006 "predicted", and 6623 ab initio, ORFs. Finally, we have developed a novel method for facilitating structural annotation, "expressed peptide sequence tags" (ePSTs) and demonstrate its utility by identifying 521 potential novel proteins from the chicken "unassigned chromosome".

Amino Acid Sequence↗

Accurate prediction of solvent accessibility using neural networks-based regression.

Accurate prediction of relative solvent accessibilities (RSAs) of amino acid residues in proteins may be used to facilitate protein structure prediction and functional annotation. Toward that goal we developed a novel method for improved prediction of RSAs. Contrary to other machine learning-based methods from the literature, we do not impose a classification problem with arbitrary boundaries between the classes. Instead, we seek a continuous approximation of the real-value RSA using nonlinear regression, with several feed forward and recurrent neural networks, which are then combined into a consensus predictor. A set of 860 protein structures derived from the PFAM database was used for training, whereas validation of the results was carefully performed on several nonredundant control sets comprising a total of 603 structures derived from new Protein Data Bank structures and had no homology to proteins included in the training. Two classes of alternative predictors were developed for comparison with the regression-based approach: one based on the standard classification approach and the other based on a semicontinuous approximation with the so-called thermometer encoding. Furthermore, a weighted approximation, with errors being scaled by the observed levels of variability in RSA for equivalent residues in families of homologous structures, was applied in order to improve the results. The effects of including evolutionary profiles and the growth of sequence databases were assessed. In accord with the observed levels of variability in RSA for different ranges of RSA values, the regression accuracy is higher for buried than for exposed residues, with overall 15.3-15.8% mean absolute errors and correlation coefficients between the predicted and experimental values of 0.64-0.67 on different control sets. The new method outperforms classification-based algorithms when the real value predictions are projected onto two-class classification problems with several commonly used thresholds to separate exposed and buried residues. For example, classification accuracy of about 77% is consistently achieved on all control sets with a threshold of 25% RSA. A web server that enables RSA prediction using the new method and provides customizable graphical representation of the results is available at http://sable.cchmc.org.

Artificial Intelligence↗

Similarity networks of protein binding sites.

An increasing attention has been dedicated to the characterization of complex networks within the protein world. This work is reporting how we uncovered networked structures that reflected the structural similarities among protein binding sites. First, a 211 binding sites dataset has been compiled by removing the redundant proteins in the Protein Ligand Database (PLD) (http://www-mitchell.ch.cam.ac.uk/pld/). Using a clique detection algorithm we have performed all-against-all binding site comparisons among the 211 available ones. Within the set of nodes representing each binding site an edge was added whenever a pair of binding sites had a similarity higher than a threshold value. The generated similarity networks revealed that many nodes had few links and only few were highly connected, but due to the limited data available it was not possible to definitively prove a scale-free architecture. Within the same dataset, the binding site similarity networks were compared with the networks of sequence and fold similarity networks. In the protein world, indications were found that structure is better conserved than sequence, but on its own, sequence was better conserved than the subset of functional residues forming the binding site. Because a binding site is strongly linked with protein function, the identification of protein binding site similarity networks could accelerate the functional annotation of newly identified genes. In view of this we have discussed several potential applications of binding site similarity networks, such as the construction of novel binding site classification databases, as well as the implications for protein molecular design in general and computational chemogenomics in particular.

Animals↗

Using evolutionary and structural information to predict DNA-binding sites on DNA-binding proteins.

Proteins that interact with DNA are involved in a number of fundamental biological activities such as DNA replication, transcription, and repair. A reliable identification of DNA-binding sites in DNA-binding proteins is important for functional annotation, site-directed mutagenesis, and modeling protein-DNA interactions. We apply Support Vector Machine (SVM), a supervised pattern recognition method, to predict DNA-binding sites in DNA-binding proteins using the following features: amino acid sequence, profile of evolutionary conservation of sequence positions, and low-resolution structural information. We use a rigorous statistical approach to study the performance of predictors that utilize different combinations of features and how this performance is affected by structural and sequence properties of proteins. Our results indicate that an SVM predictor based on a properly scaled profile of evolutionary conservation in the form of a position specific scoring matrix (PSSM) significantly outperforms a PSSM-based neural network predictor. The highest accuracy is achieved by SVM predictor that combines the profile of evolutionary conservation with low-resolution structural information. Our results also show that knowledge-based predictors of DNA-binding sites perform significantly better on proteins from mainly-alpha structural class and that the performance of these predictors is significantly correlated with certain structural and sequence properties of proteins. These observations suggest that it may be possible to assign a reliability index to the overall accuracy of the prediction of DNA-binding sites in any given protein using its sequence and structural properties. A web-server implementation of the predictors is freely available online at http://lcg.rit.albany.edu/dp-bind/.

Amino Acid Sequence↗

The relationship between protein structure and function: a comprehensive survey with application to the yeast genome.

For most proteins in the genome databases, function is predicted via sequence comparison. In spite of the popularity of this approach, the extent to which it can be reliably applied is unknown. We address this issue by systematically investigating the relationship between protein function and structure. We focus initially on enzymes functionally classified by the Enzyme Commission (EC) and relate these to by structurally classified domains the SCOP database. We find that the major SCOP fold classes have different propensities to carry out certain broad categories of functions. For instance, alpha/beta folds are disproportionately associated with enzymes, especially transferases and hydrolases, and all-alpha and small folds with non-enzymes, while alpha+beta folds have an equal tendency either way. These observations for the database overall are largely true for specific genomes. We focus, in particular, on yeast, analyzing it with many classifications in addition to SCOP and EC (i.e. COGs, CATH, MIPS), and find clear tendencies for fold-function association, across a broad spectrum of functions. Analysis with the COGs scheme also suggests that the functions of the most ancient proteins are more evenly distributed among different structural classes than those of more modern ones. For the database overall, we identify the most versatile functions, i.e. those that are associated with the most folds, and the most versatile folds, associated with the most functions. The two most versatile enzymatic functions (hydro-lyases and O-glycosyl glucosidases) are associated with seven folds each. The five most versatile folds (TIM-barrel, Rossmann, ferredoxin, alpha-beta hydrolase, and P-loop NTP hydrolase) are all mixed alpha-beta structures. They stand out as generic scaffolds, accommodating from six to as many as 16 functions (for the exceptional TIM-barrel). At the conclusion of our analysis we are able to construct a graph giving the chance that a functional annotation can be reliably transferred at different degrees of sequence and structural similarity. Supplemental information is available from http://bioinfo.mbb.yale.edu/genome/foldfunc++ +.

Enzymes↗

N-terminal N-myristoylation of proteins: refinement of the sequence motif and its taxon-specific differences.

N-terminal N-myristoylation is a lipid anchor modification of eukaryotic and viral proteins targeting them to membrane locations, thus changing the cellular function of modified proteins. Protein myristoylation is critical in many pathways; e.g. in signal transduction, apoptosis, or alternative extracellular protein export. The myristoyl-CoA:protein N-myristoyltransferase (NMT) recognizes the sequence motif of appropriate substrate proteins at the N terminus and attaches the lipid moiety to the absolutely required N-terminal glycine residue. Reliable recognition of capacity for N-terminal myristoylation from the substrate protein sequence alone is desirable for proteome-wide function annotation projects but the existing PROSITE motif is not practical, since it produces huge numbers of false positive and even some false negative predictions. As a first step towards a new prediction method, it is necessary to refine the sequence motif coding for N-terminal N-myristoylation. Relying on the in-depth study of the amino acid sequence variability of substrate proteins, on binding site analyses in X-ray structures or 3D homology models for NMTs from various taxa, and on consideration of biochemical data extracted from the scientific literature, we found indications that, at least within a complete substrate protein, the N-terminal 17 protein residues experience different types of variability restrictions. We identified three motif regions: region 1 (positions 1-6) fitting the binding pocket; region 2 (positions 7-10) interacting with the NMT's surface at the mouth of the catalytic cavity; and region 3 (positions 11-17) comprising a hydrophilic linker. Each region was characterized by physical requirements to single sequence positions or groups of positions regarding volume, polarity, backbone flexibility and other typical properties of amino acids (http://mendel.imp.univie.ac.at/myristate/). These specificity differences are confined partly to taxonomic ranges and are proposed for the design of NMT inhibitors in pathogenic fungal and protozoan systems including Aspergillus fumigatus, Leishmania major, Trypanosoma cruzi, Trypanosoma brucei, Giardia intestinalis, Entamoeba histolytica, Pneumocystis carinii, Strongyloides stercoralis and Schistosoma mansoni. An exhaustive search for NMT-homologues led to the discovery of two putative entomopoxviral NMTs.

Acyltransferases↗

High-throughput yeast two-hybrid assays for large-scale protein interaction mapping.

Protein-protein interactions play fundamental roles in many biological processes. Hence, protein interaction mapping is becoming a well-established functional genomics approach to generate functional annotations for predicted proteins that so far have remained uncharacterized. The yeast two-hybrid system is currently one of the most standardized protein interaction mapping techniques. Here, we describe the protocols for a semiautomated, high-throughput, Gal4-based yeast two-hybrid system.

Culture Media↗

The anthracnose resistance locus Co-4 of common bean is located on chromosome 3 and contains putative disease resistance-related genes.

The broadest based resistance to anthracnose of common bean ( Phaseolus vulgaris L.) is conferred by the Co-4 locus. We sequenced a bacterial artificial chromosome clone harboring part of the Co-4 locus of the bean genotype Sprite and assembled a single contig of 106.5 kb for functional annotation. This region contained five copies of the COK-4 gene that encodes for a serine threonine kinase protein previously mapped to the Co-4 locus and 19 novel genes with no similarity to any previously identified genes of common bean. Several putative genes of the Co-4 locus seemed to be expressed as they matched perfectly with bean expressed sequence tags. The expression of the COK-4 genes was assessed by reverse transcription (RT)-PCR, and a single 850-bp cDNA fragment was sequenced and compared with the genomic sequences of the COK-4 homologs. Although the COK-4 cDNA was isolated from a different bean cultivar, it showed high similarity (95%) to the exons of genes BA17 and BA21, suggesting that they were expressed. In a phylogenetic tree including all currently available Pto-like sequences from Phaseolus species, the COK-4 homologs formed a single cluster with the Pto gene, whereas two sequences from P. coccineus and all sequences of P. vulgaris formed two closely related clusters. The Co-4 locus was physically mapped to the short arm of bean chromosome 3, which corresponds to linkage group B8. This study represents a first step in gaining an understanding of the genomic organization of an anthracnose resistance locus of common bean and provides molecular data for comparative analysis with other plant species.

Ascomycota↗

Genomic and Structural Analysis of Gamete Recognition Proteins in a Broadcast Spawning Echinoderm Mesocentrotus franciscanus.

Gamete recognition proteins are expressed on the surfaces of sperm and eggs, where they mediate interactions between gametes. The genetic basis for gamete recognition proteins, as well as their structure and interactions, have yet to be fully resolved. Using a new high-quality de novo genome assembly for the sea urchin Mesocentrotus franciscanus, we investigated the genomic structure, expression, and protein forms of several gamete recognition proteins: sperm bindin, egg receptor for sperm (HSP110), and egg bindin receptor (EBR1), as well as the receptor for egg jelly (REJ) and its paralogs. To inform future population genetic and evolutionary studies, we resolve the genomic structure of the large EBR1 protein, identifying fewer tandem CUB-TSP1 repeats in EBR1 compared to the initial characterization of this protein. As expected for an egg receptor for sperm, EBR1 is highly expressed in female reproductive tissues (eggs and female gonad), compared to other tissues. In contrast, HSP110 shows similar levels of expression across male and female reproductive tissues, as well as across non-reproductive tissues and development stages. HSP110 might be a pleiotropic gene that in part influences fertilization. Using protein structural modeling and functional domain predictions, we propose hypotheses about potential interactions among EBR1, bindin, and HSP110 proteins that may provide insight into sperm-egg interactions in sea urchins. Resolving the genomic structure of genes encoding gamete recognition proteins, in combination with functional annotations and protein structural modeling, enables deeper investigation into the consequences of variation in gamete recognition proteins and the evolution of reproductive isolation.

Mesocentrotus franciscanus↗

Hybrid genome assembly and phenotypic assays reveal carbohydrate metabolism diversity in Lacticaseibacillus strains.

Investigation of carbohydrate metabolism in lactic acid bacteria is essential for the rational selection of strains for fermentation processes, particularly in emerging applications involving non-conventional substrates or building of synthetic microbial consortia. However, establishing robust genotype-phenotype relationships remains challenging, as gene presence alone often fails to explain observed metabolic traits without considering the genomic context and regulatory architecture. In the present study, we combined hybrid genome assembly (Illumina and Oxford Nanopore) with high-throughput phenotype profiling (Biolog GENIII and PM2A) to investigate carbohydrate utilization in five Lacticaseibacillus strains. Phenotypic assays revealed clear intra- and inter-specific variability in substrate utilization. We therefore investigated whether such differences could be attributed to the organization and regulatory context of carbohydrate-associated loci, rather than to gene presence alone. Functional annotation based on COG and CAZyme databases revealed candidate genomic regions potentially involved in carbohydrate metabolism. Comparative analysis between predicted and experimentally observed substrate usage highlighted specific loci associated with carbohydrate utilization profile. The trehalose (tre) operon was conserved across all strains, while at least two distinct cellobiose-associated loci were detected in each genome. Despite the presence of these loci, L. paracasei strains were unable to metabolize cellobiose, a phenotype likely linked to the presence of a downstream TetR-type transcriptional repressor within the cellobiose (cel) operon. Additionally, a genomic region uniquely found in L. rhamnosus strains was associated with gentiobiose utilization, consistent with phenotypic observations. Overall, these findings highlight the importance of integrating phenotypic validation with complete genome context to support the identification of candidate structural and regulatory determinants of carbohydrate utilization in lactic acid bacteria. KEY POINTS: • Phenotype microarrays reveal metabolic traits of interest in isolated strains. • Regulatory context is key to understanding carbohydrate metabolism differences. • Basis of subspecies-dependent cellobiose metabolism in L. paracasei is provided.

Carbohydrate Metabolism↗

Earliest changes in the left ventricular transcriptome postmyocardial infarction.

We report a genome-wide survey of early responses of the mouse heart transcriptome to acute myocardial infarction (AMI). For three regions of the left ventricle (LV), namely, ischemic/infarcted tissue (IF), the surviving LV free wall (FW), and the interventricular septum (IVS), 36,899 transcripts were assayed at six time points from 15 min to 48 h post-AMI in both AMI and sham surgery mice. For each transcript, temporal expression patterns were systematically compared between AMI and sham groups, which identified 515 AMI-responsive genes in IF tissue, 35 in the FW, 7 in the IVS, with three genes induced in all three regions. Using the literature, we assigned functional annotations to all 519 nonredundant AMI-induced genes and present two testable models for central signaling pathways induced early post-AMI. First, the early induction of 15 genes involved in assembly and activation of the activator protein-1 (AP-1) family of transcription factors implicates AP-1 as a dominant regulator of earliest post-ischemic molecular events. Second, dramatic increases in transcripts for arginase 1 (ARG1), the enzymes of polyamine biosynthesis, and protein inhibitor of nitric oxide synthase (NOS) activity indicate that NO production may be regulated, in part, by inhibition of NOS and coordinate depletion of the NOS substrate, L: -arginine. ARG1: was the single-most highly induced transcript in the database (121-fold in IF region) and its induction in heart has not been previously reported.

Acute Disease↗

Differential expression of genes related to HFE and iron status in mouse duodenal epithelium.

Iron absorption, distribution, use, and storage are thought to be tightly regulated since altered iron stores may lead to cellular damage and disease. HFE, the hereditary hemochromatosis gene product, is expressed in the crypts of the duodenum, but the molecular mechanism by which it contributes to the inhibition of iron absorption is still unknown. In this study we aimed to identify transcriptional profiles in the duodenal epithelium of Hfe(-/-) mice. We used dedicated microarrays to compare gene expression among the duodenum of Hfe(-/-) mice, induced iron overload mice, and control mice. We found 151 differentially expressed genes and unknown sequences between Hfe(-/-) mice and normal littermates. Gene profiling revealed a gene subset more specific for Hfe inactivation. The functional annotation of upregulated genes highlighted that mucus production and cell maintenance may account for the influence of Hfe on epithelium integrity and luminal iron uptake.

Animals↗

Radiation induces different changes in expression profiles of normal rectal tissue compared with rectal carcinoma.

PURPOSE: Radiotherapy is a very effective adjuvant treatment for rectal cancer with little side effects. Its killing effect on tumor cells seems to be more profound than the effect on normal tissue. The molecular events caused by irradiation are mainly analyzed in in vitro and animal models; investigations on human material are rare. In the current study, we analyzed the effects of irradiation on gene expression in normal and tumor tissue of rectal cancer patients. METHODS AND MATERIALS: Normal and carcinoma tissue of patients from a randomized clinical trial of the benefits of preoperative radiotherapy were analyzed using the Affymetrix Human Cancer Gene Chip. Preoperative radiotherapy was given within 5 days prior to surgery. Results for normal tissue and tumor were compared to investigate the radiation-related differences between normal and tumor cells. We clustered the differentially expressed genes based on their functional annotation. Results were compared with immunohistochemical and literature data. RESULTS: The majority of the investigated cancer-related genes remained unchanged by irradiation (92% in tumor tissue and 93% in normal tissue). The differentially expressed genes varied between tumor and normal tissue except for maspin and IL-8. Both in tumor and normal tissue, differentially expressed genes were present related to cell signaling and cycle control, apoptosis and cell survival and tissue response and repair. However, the spectrum of affected genes was totally different. CONCLUSION: Pre-existing differences in gene expression between normal tissue and tumor tissue might explain the differences in their responses to radiation. This change in response may explain the clinical beneficial effect of radiotherapy on tumor cells (low local recurrence rate) and the less severe effects on normal tissue (minor side effects).

Apoptosis↗

A genetic signal at 8q12.3 modulates GGT levels via the Runx1-CYP7B1 axis in female ethnic minorities from Guizhou.

Gamma-glutamyl transferase (GGT) regarded as a biomarker of liver dysfunction or excessive alcohol consumption; however, existing genome-wide association studies (GWAS) have been conducted predominantly in European populations and East Asian populations from Japan and the Taiwan region, with limited investigation in ethnic minorities from Guizhou Province. Previous genetic studies have demonstrated that Guizhou ethnic minorities share an East Asian genetic background while exhibiting specific genetic structures, a pattern that is also confirmed by our principal component analysis (PCA) results. We therefore performed a GWAS in this population and identified a genome-wide significant signal at 8q12.3 in female ethnic minorities from Guizhou. Fine-mapping and functional annotation analyses suggest that a regulatory pathway involving Runt-related transcription factor 1 (Runx1)-Cytochrome P450 family 7 subfamily B member 1 (CYP7B1)-cholesterol-reactive oxygen species (ROS)-glutathione (GSH) may contribute to the regulation of GGT levels. Mendelian randomization (MR) analyses further supported a causal relationship between GGT levels and autoimmune hepatitis (AIH). These findings uncover a genetic mechanism underlying GGT variation at 8q12.3 in female ethnic minorities from Guizhou, implicating a pathway linked to cholesterol metabolism and oxidative stress, and providing potential targets and insights for precision prevention and treatment of related diseases.

Female↗