Search PubMedSearch

Biomedical subjects

Xiaojing Gao

Publications and source records attributed to Xiaojing Gao.

2 recordsLinked to original sources

OmniExtract: an automatic data extraction tool based on large language model and prompt engineering.

Extracting structured information from documents or scientific papers is crucial for data sharing and retrieval. Recent advances in large language models (LLMs) have demonstrated strong capabilities in language understanding, and a number of LLM-based tools have been developed for extraction-oriented tasks. However, it's still difficult to find a universal and user-friendly tool for various practical extraction tasks. To address this challenge, we propose OmniExtract, an automatic data extraction tool with user-friendly configuration files that can adapt to various data extraction tasks. OmniExtract employs a prompt optimization method to refine task-specific prompts and achieve high extraction performance. It also supports comprehensive data extraction from both documents and tables, making it applicable to a broad range of data sources. Evaluation results show that OmniExtract obtains a high accuracy ~90% for three datasets. Furthermore, two additional data extraction applications of OmniExtract in real-world scenarios have been presented, achieving an accuracy of 92.21% and ~90% precision and recall, respectively. Specifically, OmniExtract can handle tabular files of various sizes and formats, and achieve over 99% precision and recall on table information extraction tasks. The data reliability performance shows that OmniExtract is a valuable tool for database updating. An online testing service is available at https://ngdc.cncb.ac.cn/omniextract/. The service can be deployed locally with the code in https://github.com/wyb39/OmniExtract.

Large Language Models

Molecular Signature of Prediabetes With High-Risk of Diabetes Revealed by Deep Plasma Proteome.

AIMS: Prediabetes is biologically heterogeneous, but molecular subtypes linked to diabetes progression remain poorly defined. We aimed to identify plasma proteome-based subtypes of impaired fasting glucose (IFG), characterise their molecular features and assess their association with future diabetes risk. MATERIALS AND METHODS: We quantified 2584 plasma proteins using liquid chromatography-mass spectrometry in 538 IFG participants from a prospective discovery cohort (Nutrition and Health of Aging Population in China, NHAPC). Proteomic subtypes were defined by consensus clustering, linked to longitudinal changes in insulin sensitivity and incident type 2 diabetes mellitus (T2DM), which were further validated in an independent Shanghai Brain Aging Study (SBAS) cohort. RESULTS: Two reproducible IFG molecular subtypes based on plasma proteomics were identified. The high-risk subtype showed higher incident diabetes and a greater 6-year decline in insulin sensitivity and was characterised by enrichment of glycolysis/gluconeogenesis, insulin signalling and neutrophil degranulation, together with a dyslipidemic lipidomic profile indicating co-dysregulation of glucose and lipid homeostasis. The low-risk subtype demonstrated a higher complement cascade and high-density lipoprotein particle remodelling signature. In the high-risk subtype, key proteins and lipids showed stronger associations with longitudinal declines in insulin sensitivity, including PPBP, PGK1 and ALDOA, as well as PE-P 18:0/20:3 and PE-P 18:1/20:3. CONCLUSIONS: Proteome-based molecular subtyping stratifies IFG individuals with similar fasting glucose levels but distinct biology and future diabetes risk, supporting earlier and more targeted prevention.

Humans