Search PubMedSearch

Biomedical subjects

Zhigang Meng

Publications and source records attributed to Zhigang Meng.

2 recordsLinked to original sources

Large language models in bioinformatics: a comprehensive survey.

The emergence of foundation models with trillion-level parameters has redefined the landscape of artificial intelligence. Various fields are developing their own large-scale models, which can solve many problems within the field and improve work efficiency. Biological large-scale models are a cross-disciplinary research field that combines mathematics, computer science, and biology, aiming to simulate and understand the structure, function, and dynamic changes of biological systems through the establishment of complex computational models. This field covers multiple levels such as biological pathways, population dynamics, protein folding, etc., providing us with tools for deep exploration of the mysteries of life and applications in medicine, ecology, and other fields. This article reviews the background and research status of biological large-scale models, and discusses future directions. Large language models (LLMs) and other large-scale foundation models have rapidly advanced in recent years, enabling powerful representation learning and generation across text, sequences, and multimodal data. In bioinformatics and biomedicine, these models are increasingly used to analyze genomic sequences, infer protein properties and structures, support drug discovery, and integrate heterogeneous biomedical evidence. This survey reviews the basic principles of LLMs and summarizes representative applications in (i) gene and genome sequence analysis, (ii) protein structure and function prediction, and (iii) drug design, including virtual screening and personalized medicine. We also discuss emerging multi-model modeling approaches, as well as key challenges such as data quality and privacy, interpretability, generalization to new organisms and tasks, and responsible deployment in health-related settings. Finally, we outline future directions for developing reliable, scalable, and explainable bioinformatics foundation models.

bioinformatics

Selection of GhTT2-A07 promoter enhances fiber quality in improved cotton varieties.

Modern cultivated cotton fibers are predominantly white with enhanced quality compared to their wild ancestors. However, the molecular mechanisms and evolutionary drivers linking fiber color to quality remain least focused. In this study, we identified FQC1 (Fiber Quality and Color 1), a major quantitative trait locus (QTL) on chromosome A07 that concurrently regulates both fiber quality and pigmentation. Through map-based cloning, we revealed that Gossypium hirsutum TRANSPARENT TESTA2-A07 (GhTT2-A07), an R2R3-MYB transcription factor, resides within this locus. GhTT2-A07 modulates fiber development by directly activating genes in the general phenylpropanoid pathway, thereby promoting the metabolic flux toward downstream secondary metabolites. Variations in the GhTT2-A07 promoter led to its reduced expression in modern white cotton cultivars. This down-regulation suppresses the accumulation of S/G/H-type lignin monomers and proanthocyanidins, resulting in altered secondary cell wall composition and ultimately enhancing the quality of mature white fibers. Population genetic analyses further indicate that the white-fiber allele GhTT2-A07W has been fixed in modern breeding genotypes, underscoring the impact of artificial selection during cotton domestication. Overall, our study elucidates the biochemical and molecular mechanisms underlying fiber quality and pigmentation in cotton, clarifies the selection criteria for high-quality white fibers in modern cultivars, and provides a theoretical basis for future targeted genetic improvement of cotton fibers.

Alleles