Search PubMedSearch

Biomedical subjects

Lei Sun

Publications and source records attributed to Lei Sun.

6 recordsLinked to original sources

Assessing Hardy-Weinberg equilibrium in T2T-aligned 1000 genomes project.

Quality control of markers in genome-wide association studies often includes testing for Hardy-Weinberg equilibrium (HWE). However, this is usually implemented in a homogeneous population without stratifying by sex. Previous work indicates sex-based selection at numerous autosomal loci in cohorts with active recruitment. Sex chromosome sequences can also interfere with autosomal SNPs. These motivate a re-examination of HWE in sex-aware analyses. Using the telomere-to-telomere (T2Tv2)-aligned high-coverage whole genome sequencing data from 2,490 individuals in the 1000 Genomes Project, we examined genome-wide sex-specific deviations from HWE across five super-populations. Our analyses were restricted to bi-allelic SNPs with non-missing genotypes and minor allele frequency (MAF) &#x2265;5% in both sexes of the five super-populations. We applied an allele-based framework to quantify both the magnitude and direction of Hardy-Weinberg disequilibrium (HWD), followed by a second-order omnibus meta-analysis that combined HWD results across populations and sexes. At a genome-wide significance threshold of p&#x2009;<&#x2009;5e-8, 0.9% of autosomal SNPs exhibited significant deviations from HWE. The majority of these deviations were associated with genomic features indicative of poor sequence quality. Restricting the analysis to reliable genomic regions substantially reduced the number of signals, yielding 255 autosomal SNPs and one non-pseudoautosomal chromosome X SNP. Among these, 140 autosomal SNPs displayed significant heterogeneity across populations but not across sexes. Notably, eight SNPs within a 15-bp region on chromosome 14q31.3 showed excess heterozygosity in both sexes of the African super-population (AFR). Finally, we developed a multivariate predictor of HWD based on sequence features, providing a practical tool that can be integrated into existing quality control pipelines for whole genome sequencing studies.

Journal Article

The CsTBH-CsROP2 Module Regulates Waterlogging Tolerance via Auxin-Mediated Adventitious Root Formation in Cucumber.

Cucumber (Cucumis sativus L.) requires frequent irrigation due to its shallow root system and high transpiration rate of the aboveground parts. However, it is also prone to waterlogging damage. Therefore, understanding its response to waterlogging is crucial for breeding waterlogging-tolerant varieties. Although Rho of Plants GTPases play well-established roles in regulating development and stress signalling, their functions in plant adaptation to waterlogging stress has yet to be fully elucidated. Here, we identified nine CsROP genes in the cucumber genome, which exhibit evolutionary diversification but retain conserved functional domains. Functional analysis revealed that CsROP2 acts as a negative regulator of adventitious root formation. It modulates auxin accumulation in hypocotyl vascular bundles, thereby suppressing adventitious root development and enhancing waterlogging sensitivity. The HD-Zip I transcription factor CsTBH directly binds the CsROP2 promoter and activates its expression. Our study uncovers a CsTBH-CsROP2 module that governs adventitious rooting and waterlogging tolerance by modulating auxin homeostasis. These findings provide new insights into the crosstalk between developmental programmes and stress signalling pathways and offer potential genetic targets for improving stress resilience in cucumber and other crops.

CsROP2

Modtector: ultra-fast modification signal mining on mapped sequencing reads.

SUMMARY: Existing tools for RNA epitranscriptomic modification and structural signal analysis are often fragmented, inefficiency, and limited to single signal types. We developed Modtector, an unified tool for extracting mutation and reverse-transcription stop signals from aligned sequencing reads. By using a "count-then-correct" strategy, Modtector reduces computational complexity and enables efficient dual-signal analysis. It achieves multi-fold speedups on large-genome and high-coverage datasets, including completing HEK293 22G data analysis in 5&#x2009;minutes, and show strong scalability on single-cell datasets with speedups exceeding 50-fold. AVAILABILITY: The source code is available at GitHub (https://github.com/TongZhou2017/modtector) and Crates.io (https://crates.io/crates/modtector). The archived source-code snapshot used in this study is available at Zenodo (DOI: 10.5281/zenodo.20967747), corresponding to GitHub commit 7c60e9d. Workflow examples, datasets, and analysis scripts are available at Zenodo (DOI: 10.5281/zenodo.17316476 and 10.5281/zenodo.18523297).

Humans

Deciphering the genetic background of an industrial 2-ketogluconic acid-producing strain Pseudomonas plecoglossicida JUIM01 using whole-genome sequencing.

2-Ketogluconic acid (2KGA) is an important precursor for the food antioxidant erythorbic acid, currently produced via microbial fermentation using Pseudomonas species. To facilitate the genetic improvement of production strains, the complete genome of an industrial 2KGA producer P. plecoglossicida JUIM01 was sequenced and analyzed. The genome consists of a 5.13-Mb circular chromosome with a GC content of 63.58%, encoding 4,517 predicted proteins. Comprehensive functional annotation identified a putative global regulatory network comprising 75 core regulators, which were classified into six functionally cooperative modules, potentially governing the strain's metabolism and environmental adaptability.&#xa0;We further delineated the genetic determinants hypothetically linked to efficient 2KGA synthesis, including glucose metabolism, fatty acid metabolism, and the oxidative phosphorylation system.&#xa0;These outputs could provide the genomic resource for elucidating high productivity and robustness, and rationally engineering the high-performance chassis cells toward robust 2KGA production.

P. plecoglossicida

Accurate identification of abnormal ploidy using an artificial intelligence model in preimplantation genetic testing.

STUDY QUESTION: Can ultra-low-coverage whole-genome sequencing (ulc-WGS) accurately identify abnormal ploidy during preimplantation genetic testing (PGT)? SUMMARY ANSWER: The artificial intelligence (AI)-based PGT-Plus model demonstrates high accuracy in ploidy detection, offering a cost-effective solution that enhances clinical utility of PGT. WHAT IS KNOWN ALREADY: The predominant PGT for aneuploidy can identify chromosomal aneuploidies but cannot determine ploidy status. Transferring embryos with ploidy abnormalities can result in miscarriage and molar pregnancy. On the other hand, in ART, fertilization is assessed by morphological pronuclear assessment at the zygote stage. However, it has a low specificity in the prediction of abnormal ploidy status and embryos deemed abnormally fertilized can yield healthy pregnancies. Accurately identified abnormal ploidy in PGT-A can resolve current limitations and expand the utility range of PGT-A. Several studies have identified ploidy abnormalities; however, they were mainly based on single-nucleotide polymorphism (SNP) arrays or needed to combine additional targeted-next-generation sequencing (NGS) information. Studies based on ulc-WGS remain scarce. STUDY DESIGN SIZE DURATION: The study consisted of two stages: methodology establishment and validation. An AI model, named PGT-Plus, was developed using 653 samples with known ploidy status, which was further validated using 792 different ploidy status samples. In the clinical application stage, the approach was used to analyse the ploidy status of 19&#x2009;103 normally fertilized PGT blastocysts and 140 single pronucleus (1PN)-derived blastocysts collected between May 2022 and December 2023. All blastocysts were tested using trophectoderm biopsy and NGS. PARTICIPANTS/MATERIALS SETTING METHODS: The methodology is based on the ulc-WGS data. First, based on samples with known ploidy status: the heterozygosity rate of high-frequency biallelic SNPs, the likelihood ratio (LLR) of alleles was calculated under different assumptions ('both parental homologs' [BPH] from a single parent, 'single parental homolog' [SPH] from each parent, disomy, and monosomy) by leveraging allele frequencies and linkage disequilibrium (LD) measured in the 1000 genomes project database. Twenty-three continuous candidate features derived from heterozygosity rates and LLRs of chromosomes or selected windows were included to establish the ploidy prediction AI model. Gini importance analysis and multicollinearity mitigation was performed for feature selection, then the performance of Random Forest (RF), Support Vector Machine (SVM), and Logistic Regression for modelling was compared. Subsequently, the parameter optimization was performed based on the RF model. Ploidy constitution concordance was evaluated in known ploidy status samples. The frequency of abnormal ploidy in normal fertilized PGT blastocysts and 1PN-derived blastocysts (including conventional IVF and ICSI) was evaluated. MAIN RESULTS AND THE ROLE OF CHANCE: Eleven features were collected for model architecture compared to SVM and Logistic Regression; RF achieved superior performance for ploidy detection. The AI model achieved an AUC of 1 for genome-wide-uniparental diploidy (GW-UPD), 1 for triploidy, and 0.99 for diploidy. For the 792 validation samples, 99.5% of samples were successfully detected using the AI model, and the model showed 100% accuracy for ploidy classification. In the clinical application stage, out of 19&#x2009;103 PGT samples, 19&#x2009;069 were successfully analysed using the model, with 110 (0.57%) identified as having abnormal ploidy embryos. Among these, 12.7% (14/110) were identified as GW-UPD, and 87.3% (96/110) were triploid. Among 5563 diploid blastocysts transferred, 3478 clinical pregnancies were achieved. Subsequent ploidy analysis was performed for 217 spontaneous abortion and 935 prenatal diagnostic samples, and no abnormal ploidy was identified. Furthermore, of the 140 1PN embryos tested, 40 (28.6%) exhibited GW-UPD, 3 (2.1%) exhibited triploidy, and 97 (69.3%) were determined to be biparental and normally fertilized. Among the 97 biparental embryos, 46 were diploid, 11 were mosaic, and 40 were aneuploid. In terms of the insemination pattern, the percentage of abnormal ploidy in ICSI was significantly higher than in conventional IVF (P&#x2009;<&#x2009;0.01, 37.1% vs. 2.9%, respectively). With full informed consent, 20 patients without euploidy from normal fertilization chose 1PN-derived biparental and diploid blastocysts to transfer, resulting in 10 clinical pregnancies and 9 ongoing pregnancies. LARGE-SCALE DATA: N/A. LIMITATIONS REASONS FOR CAUTION: Some rare ploidy abnormalities, such as polyploidy with an equal number of identical sets of chromosomes and ploidy mosaicism cannot be accurately identified. Moreover, the origin of abnormal ploidy was not identified due to the unavailability of DNA from both parents. WIDER IMPLICATIONS OF THE FINDINGS: The PGT-Plus AI model provides a ploidy evaluation method based on the conventional PGT-A data and integrates directly into standard PGT-A workflows. Clinical utility results suggest that the model is a valuable tool for identifying embryos with abnormal ploidy in PGT-A and rescuing normal diploid embryos from abnormally fertilized embryos. These findings demonstrate that PGT-Plus significantly enhances the diagnostic accuracy of PGT. STUDY FUNDING/COMPETING INTERESTS: This study was supported by grants from Major Scientific Program of CITIC Group (No. 2023ZXKYB34100, to Ge.L.), Hunan Provincial Grant for Innovative Province Construction (2019SK4012), Hunan Xiangjiang New District (Changsha High-tech Zone) key core technology research project in 2023, and Science Foundation of Hunan Province (Grant 2023JJ30422). All authors declared no conflicts of interest..

artificial intelligence

Histone modifications in the regulation of erythropoiesis.

INTRODUCTION: The pathogenesis of anemia and other erythroid dysphasia are mains poorly understood, primarily due to limited knowledge about the differentiation processes and regulatory mechanisms governing erythropoiesis. Erythropoiesis is a highly complex and precise biological process, that can be categorized into three distinct stages: early erythropoiesis, terminal erythroid differentiation, and reticulocyte maturation, and this complex process is tightly controlled by multiple regulatory factors. Emerging evidence highlights the crucial role of epigenetic modifications, particularly histone modifications, in regulating erythropoiesis. Methylation and acetylation are two common modification forms that affect genome accessibility by altering the state of chromatin, thereby regulating gene expression during erythropoiesis. DISCUSSION: This review systematically examines the roles of histone methylation and acetylation, along with their respective regulatory enzymes, in regulating the development and differentiation of hematopoietic stem/progenitor cells (HSPCs) and erythroid progenitors. Furthermore, we discuss the involvement of these histone modifications in erythroid-specific developmental processes, including hemoglobin switching, chromatin condensation, and enucleation.Conclusions This review summarizes the current understanding of the role of histone modifications in erythropoiesis based on existing research, as a foundation for further research the mechanisms of epigenetic regulatory in erythropoiesis.

Erythropoiesis