Search PubMedSearch

Biomedical subjects

Donghui Yan

Publications and source records attributed to Donghui Yan.

2 recordsLinked to original sources

Seq2Saccharide: Discovering Oligosaccharides and Aminoglycosides Natural Products by Integrating Computational Mass Spectrometry and Genome Mining.

Natural oligosaccharides and aminoglycosides are important sources of new drug candidates, especially in the development of antibiotics. In the past, discovering novel saccharides has been time-consuming and costly. However, the rapid expansion of high-throughput data, including genomic and mass spectrometry data sets, has greatly increased opportunities for natural saccharide discovery. Yet, due to the complex biosynthesis pathways of saccharides, no existing method can predict their structures with high precision. To address this, we introduce Seq2Saccharide, a tool designed to automate saccharide natural product discovery by integrating both genomic and mass spectrometry data. To enhance accuracy, Seq2Saccharide predicts hundreds or thousands of putative structures for each gene cluster. The correct structure is then identified from these predictions using a mass spectral search. Benchmarks against saccharides in the MiBIG database show that Seq2Saccharide outperforms existing methods in predicting the structure of saccharides. Furthermore, mass spectrometry analysis indicates that the variable search module can correct mispredictions from genome mining. By searching genomic and mass spectrometry data of microbial strains, Seq2Saccharide correctly identified the biosynthetic gene cluster for the polysaccharide oligosaccharide trestatin B.

Aminoglycosides

Biobank-wide association scan identifies risk factors for late-onset Alzheimer's disease and endophenotypes.

Rich data from large biobanks, coupled with increasingly accessible association statistics from genome-wide association studies (GWAS), provide great opportunities to dissect the complex relationships among human traits and diseases. We introduce BADGERS, a powerful method to perform polygenic score-based biobank-wide association scans. Compared to traditional approaches, BADGERS uses GWAS summary statistics as input and does not require multiple traits to be measured in the same cohort. We applied BADGERS to two independent datasets for late-onset Alzheimer's disease (AD; n=61,212). Among 1738 traits in the UK biobank, we identified 48 significant associations for AD. Family history, high cholesterol, and numerous traits related to intelligence and education showed strong and independent associations with AD. Furthermore, we identified 41 significant associations for a variety of AD endophenotypes. While family history and high cholesterol were strongly associated with AD subgroups and pathologies, only intelligence and education-related traits predicted pre-clinical cognitive phenotypes. These results provide novel insights into the distinct biological processes underlying various risk factors for AD.

Alzheimer Disease