Search PubMedSearch

SEARCH · Search PubMed

Results for “crop breeding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Translating functional molecular knowledge into crop-breeding success.

Historical plant breeding, which optimizes phenotypes through selective crossing guided by phenotypic evaluation and molecular markers, is limited by evolutionary constraints that hinder rapid crop improvement. A new paradigm, precision breeding, circumvents these limitations by targeting genetic variants through functional molecular knowledge. To generate this knowledge at scale, sequence-based deep learning leverages high-quality genome sequence data to predict variant effects at base-pair resolution. When linked to agronomically important traits, these predictions enable breeders to prioritize variants for precision selection or editing. Although it is still in the early stages of development, we foresee three key applications for this approach: introgressing genes from distant breeding pools, purging deleterious mutations and designing new plant ideotypes. Looking ahead, refined computational models will facilitate targeted editing and the systematic redesign of complex physiological processes to address emerging breeding goals under shifting environmental conditions.

Crops, Agricultural

AI-integrated digital breeding for crop improvement.

Crop breeding increasingly depends on the effective integration and interpretation of large, heterogeneous datasets spanning genomic, phenotypic, multi-omics, and environmental layers. Conventional breeding approaches are often insufficient to capture the complex relationships among these data or to support timely selection decisions. Digital breeding can help address this limitation by complementing field experimentation, mixed models, and genomic prediction with the integration of biological data and computational prediction throughout the breeding process. In particular, the rapid advancement of artificial intelligence (AI) has improved the analysis of high-dimensional datasets and broadened its application to trait prediction, selection, and breeding design. Here, we review recent developments in AI-enabled digital breeding, encompassing genomic, phenomic, and multi-omics data generation and analysis, predictive modeling, explainable and generative AI, and data-driven breeding decision support. We further discuss emerging AI applications, their current contributions to crop research and breeding, and the major considerations affecting their reliable and practical implementation. Collectively, this review provides a structured understanding of the roles of AI across the digital breeding process and offers guidance for future methodological development and practical application in crop improvement.

artificial intelligence

The potential of considering photosynthesis parameters in crop yield breeding by genomic prediction.

To meet the growing demand for agricultural products, optimizing photosynthesis is a promising strategy to improve crop yields. Phenotypic variance in photosynthesis has been observed within or between species. To explore the potential of integrating photosynthetic parameters into crop breeding programs, we explored the genetic variation in photosynthesis by assessing photosynthesis-related parameters across plant development in 631 barley recombinant inbred lines (RILs) from eight HvDRR subpopulations under field conditions. The genetic complexity of these parameters was resolved by analyses of bi-parental and multi-parental quantitative trait loci (QTLs). Finally, we examined the merit of integrating photosynthesis-related parameters in genomic prediction of yield and its components. Significant genotypic variations of the photosynthesis-related parameters were found among the RILs, with their heritability ranging from 0.38 to 0.54. The multiple QTLs and dynamic QTLs for photosynthesis observed across different developmental stages underlined the complexity of the genetics of photosynthesis in barley. The considerably higher percentage of phenotypic variance explained for genomic prediction than multi-parental QTL analysis illustrates that the photosynthesis-related parameters are inherited in a more complex way than classical agronomic traits. Notably, the prediction ability for yield was increased by integrating the photosynthesis-related parameters of some developmental stages into genomic prediction models. Thus, our results suggest a novel perspective on increasing the efficiency of crop breeding programs by integrating photosynthesis-related parameters into prediction models.

Photosynthesis

PlantPan: A comprehensive multi-species plant pan-genome database.

The pan-genome represents the complete genomic diversity of specific species, serving as a valuable resource for studying species evolution, crop domestication, and guiding crop breeding and improvement. While there are several single-species-specific plant pan-genome databases, the availability of multi-species pan-genome databases is limited. Additionally, variations in methods and data types used for plant pan-genome analysis across different databases hinder the comparison and integration of pan-genome information from various projects at multi-species or single-species levels. To tackle this challenge, we introduce PlantPan, a comprehensive database housing the results of pan-genome analysis for 195 genomes from 11 plant species. PlantPan aims to provide extensive information, including gene-centric and sequence-centric pan-genome information, graph-based pan-genome, pan-genome openness profiles, gene functions and its variation characteristics, homologous genes, and gene clusters across different species. Statistically, PlantPan incorporates 9 163 011 genes, 694 191 gene clusters, 526 973 370 genome variations, and 1 616 089 non-redundant genome variation groups at the species level, 33 455,098 genome synteny, and 177 827 non-redundant genome synteny groups at the species level. Regarding functional genes, PlantPan contains 5 222 720 genes related to transcription factors, 395 247 literature-reported resistance genes, 455 748 predicted microbial/disease resistance genes, and 1 612 112 genes related to molecular pathways. In summary, PlantPan is a vital platform for advancing the application of pan-genomes in molecular breeding for crops and evolutionary research for plants.

Genome, Plant

A comparative study highlights superiority of LSTM in crop genomic prediction.

We systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods, and found LSTM suitable for capturing additive and epistatic effects. Genomic prediction (GP) has been developed as an important method supporting crop breeding. By utilizing the phenotype values result from GP, breeders could make decisions in the seedling stage that consequently benefit for cost saving. In recent years, machine learning emerged as an efficient technology to solve modeling problems in many fields, including crop breeding. However, numerous modeling approaches have hindered the application of GP since breeders struggle to choose. Therefore, a comprehensively methodological research with guiding significance is extremely necessary. In the present study, we systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods. As for genomic feature processing, we found feature selection (SNP filtering approach) performed better than feature extraction (PCA method). Specifically, the feature relationship dependent methods (GBLUP, RNN, and LSTM) as well as DNN architecture showed superior performance with feature selection. Marker density analysis showed positive correlation with prediction accuracy in a limited threshold. Comparison on effect of population size demonstrated a positive correlation between trait genetic complexity and the optimal population size required. By testing fifteen modeling methods, we found LSTM network displayed superior performance, achieving the highest average STScore (0.967) across six datasets. Further research using all cell states or the latest cell states of LSTM inputs demonstrated its architecture particularly adept with capturing additive and epistatic QTL effects among SNPs. In conclusion, our findings provide basic principles for implementing GP in breeding project to maximize prediction accuracy while maintaining cost-effectiveness.

Plant Breeding

Deep learning-based annotation of plant abiotic stress resistance genes for crops.

The declining costs of DNA sequencing have expanded genomic data, crucial for understanding plant abiotic stress responses and crop improvement. However, accurate gene annotation remains challenging. To address this limitation, we propose the PASRGA, a deep learning approach that leverages transfer learning and contrastive learning to annotate genes related to drought, salt, cold, and UV resistance. PASRGA achieves high F1-scores, area under the receiver operating characteristic (AUROC), area under the precision-recall curve (AUPRC), and Matthews correlation coefficient (MCC) in annotating stress resistance genes, significantly outperforming the general protein annotation model CLEAN, the plant phosphatase gene annotation model PF-NET, the top-ranked model in the CAFA5 challenge NetGO 4.0, and four traditional machine learning methods. Its effectiveness was further validated with a salt stress treatment experiment in Eutrema salsugineum. To facilitate crop breeding practices, we utilized PASRGA to annotate the genomes of 17 major crops. To improve accessibility and utility, we incorporated both manually curated and PASRGA-predicted gene data, together with the PASRGA tool, into the PlantASRG database (https://bioinfor.nefu.edu.cn/PlantASRG/). This comprehensive resource aims to support crop breeding initiatives and ensure food security.

Crops, Agricultural

PSIA: A Comprehensive Knowledgebase of Plant Self-incompatibility.

Self-incompatibility (SI) is an important genetic mechanism in angiosperms that prevents inbreeding and promotes outcrossing, with significant implications for crop breeding, including genetic diversity, hybrid seed production, and yield optimization. In eudicots, SI is typically governed by a single S-locus containing tightly linked pistil and pollen S-determinant genes. Despite major advances in SI research, a centralized, comprehensive resource for SI-related genomic data remains lacking. To address this gap, we developed the Plant Self-Incompatibility Atlas (PSIA), a systematically curated knowledgebase providing an extensive compilation of plant SI, including genomic resources for SI species, S gene annotations, molecular mechanisms, phylogenetic relationships, and comparative genomic analyses. The current release of PSIA includes over 500 genome assemblies from 469 SI species. Using known S genes as queries, we manually identified and rigorously curated 3700 S genes. PSIA provides detailed S-locus information from assembled genomes of SI species and offers an interactive platform for browsing, BLAST searches, S gene analysis, and data retrieval. Additionally, PSIA serves as a unique platform for comparative genomic studies of S-loci, facilitating exploration of the dynamic processes underlying the origin, loss, and regain of SI. As a comprehensive and user-friendly resource, PSIA will greatly advance our understanding of angiosperm SI and serve as a valuable tool for crop breeding and hybrid seed production. PSIA is freely available at http://www.plantsi.cn.

Self-Incompatibility in Flowering Plants

Discovery and Engineering of a Rat Endogenous Retrovirus Reverse Transcriptase for Efficient Prime Editing.

CRISPR-based prime editors (PEs) install precise edits into genomic DNA without generating double-strand breaks. Their editing efficiency is highly dependent on reverse transcriptases (RTs), but efficient RT candidates remain limited. Here, we identified 19 novel active RTs by screening 558 candidates. Among them, RERV-RT, derived from Rattus norvegicus, exhibited the highest activity. Through structure-guided engineering and deep mutational scanning, we developed an optimized variant, enRERV-RT, which outperforms conventional M-MLV-RT-based PE systems by 1.20-fold in mammalian and plant cells, and by 1.88-fold at hard-to-edit loci, while enabling precise multiplex editing of functionally relevant genes. Additionally, we developed a high-throughput platform, TRAP-seq-PE, to systematically evaluate prime editor performance. Across diverse mutation types, we found that PE systems based on enRERV-RT exhibited higher editing efficiencies than those based on M-MLV-RT. Collectively, our work establishes a versatile, high-efficiency PE system, thereby facilitating advances in clinical gene therapy and precise crop breeding.

Animals

Reconstruction of ancestral plant genomes for inter-crop translational research.

We present Ancestral Genome Reconstruction (AGR), an exploratory framework for the automated inference of "paleogenomes" from large-scale comparative datasets. By analyzing 84 extant angiosperm species, we reconstructed 10 key ancestral angiosperm genomes millions of years old. These reconstructed ancestors were instrumental in (1) estimating when angiosperms emerged, when major botanical families originated, and when shared ancestral whole-genome duplication events occurred; and (2) tracing the evolutionary trajectories of ancestral chromosomes and genes, especially those that may have driven the emergence of key life-history traits (e.g., woody vs. herbaceous, aquatic vs. terrestrial, C3 vs. C4, and symbiotic root-nodulating vs. non-nodulating species). We demonstrated that these paleogenomes serve as tractable backbones for inter-crop translational research. Through an open-access web tool, OrthoViewer, we identified orthologs that have retained the same ancestral genomic context, favoring the identification of genes associated with "phenologs"- orthologous genes across species driving analogous phenotypes, traits, or processes-exemplified by FUWA for yield components, FLC for flowering time, and DDM1 for DNA methylation. Taken together, this study provides a testable paleogenomic workflow, opening novel avenues for integrating evolutionary genomics data into modern climate-smart crop breeding and supporting the agroecological transition.

Genome, Plant

Cis-regulatory elements: systematic identification and horticultural applications.

Cis-regulatory elements (CREs) are the genetic DNA fragments bound by transcription factors (TFs). CREs function as molecular switches that precisely modulate the dosage and spatiotemporal patterns of gene expression. The systematic identification of CREs not only facilitates the annotation of the functional non-coding genome but also provides essential insights into the architecture of gene regulatory networks and sheds light on an accurate selection of the target sites for genetic engineering of crops. In this review, we summarize the current high-throughput methodologies used for identifying CREs, illustrate the associations between CREs and agronomic traits in horticultural crops, and discuss how CREs can be exploited to facilitate crop breeding.

Breeding

Chromosome-scale genomes and population resequencing resolve subgenome diversity and halophyte adaptation in Salicornia.

Amid escalating water scarcity and groundwater depletion, halophytes such as Salicornia (Amaranthaceae) represent valuable models for extreme salt tolerance and hold promise for saltwater-based agriculture. Here, we show chromosome-scale genome assemblies for six Salicornia species, revealing four distinct subgenomes, reconciling our assemblies with two existing reference genomes (S. ramosissima UK and S. europaea China), correcting chromosome numbering and orientation. Comparative analyses across ploidy levels demonstrate genome expansion in North American lineages driven by Gypsy retrotransposons, and lineage-specific expansions of two gene families implicated in stress metabolism. Phylogenetic and population-structure analyses of a global resequencing panel of 318 accessions resolve interspecific relationships and establish curated germplasm collections for future crop breeding. Genetic analyses uncover a contrasting population-genetic signal on chromosome 6A between two species, highlighting an OSCA calcium-permeable channel gene as a candidate locus for osmotic adaptation. Together, these resources establish a genomic framework for Salicornia that supports evolutionary studies of halophyte adaptation and crop development.

Chenopodiaceae

Transposable element-driven expansion of enhancer RNA repertoires underlies regulatory innovation and polyploid adaptation in cereal crops.

Cereal genomes have undergone repeated polyploidization and transposable element (TE) proliferation, collectively generating complex regulatory landscapes. However, the evolutionary trajectories and functional implications of these landscapes remain largely unexplored. Using chromatin-bound RNA sequencing across seven cereal species, we systematically mapped 45,952 regulatory element transcripts (RETs), including 32,867 distal RETs corresponding to enhancer RNAs (eRNAs). Our analysis revealed that 56% of lineage-specific eRNAs originated from TE expansions, indicating that TEs serve as major reservoirs of species-specific regulatory innovation in cereals. Notably, we identified remarkable conservation in defense-related functions, root-specific expression, and TE-derived origins of eRNAs across both ancient and recent evolutionary layers of Triticeae, suggesting recurrent recruitment of TE-derived, root-associated regulatory elements throughout Triticeae evolution. Furthermore, we found that young eRNA pairs in hexaploid wheat with high sequence similarity, many originating from RLG_famc8.3 and DTC_famc4.3, exhibited pronounced root specificity and coordinated expression, suggesting targeted amplification and refinement of successful ancestral regulatory strategies established after Triticeae divergence. To facilitate community access, we developed Cereal-eRNAdb (http://bioinfo.cemps.ac.cn/Cereal-eRNAdb/), a comprehensive database integrating 69,426 eRNAs with functional annotations across 296 samples. Our findings suggest that TE-mediated innovation of root-specific eRNAs may contribute to Triticeae adaptation and provide a foundational resource for exploiting regulatory variation in cereal crop breeding.

Enhancer RNAs

Dynamic shading of chloroplasts for enhanced photosynthetic efficiency.

Coping with high light represents a major challenge for plants in nature. Under high light, 1O2 can induce MBS1 to form a low-dynamic condensate, which can effectively shade chloroplasts to avoid photodamage. This mechanism can be used to support breeding crops for both high photoprotection and high photosynthetic light use efficiency.

Photosynthesis

Integrated metabolomics, transcriptional, and physicochemical analysis reveals key metabolites and genes associated with somatic embryogenesis in Phyllostachys pubescens.

Phyllostachys pubescens (Moso bamboo) is a significant perennial crop species that provides valuable nutritional and industrial uses, as well as carbon sequestration. Due to its remarkable growth rate, bamboo offers an ideal system for studying organogenesis, particularly in monocots. Somatic embryogenesis (SE) serves as a useful technique for crop breeding and improvement. SE in moso bamboo (Phyllostachys pubescens) remains challenging due to limited knowledge of its transcriptional and metabolomic reprogramming. To address this, we optimized callus initiation (MS + 18.1 µM 2,4-D + 8.5 µM picloram), callus proliferation (MS + 12.5 µM 2,4-D + 8.5 µM picloram), and somatic embryogenesis (MS + 1.1 µM 2,4-D + 3.3 µM metatopolin), using nodal segments as explants. UHPLC-Q-TOF-MS-based metabolite profiling revealed distinct biochemical trajectories across developmental stages of P. pubescens. NEC (non-embryogenic callus) was enriched in flavonoids, alkaloids, and saponins, while in-vitro shoots showed flavonoids and glycosides enrichment, and ex-vitro shoots showed high accumulation of glycosides and terpenoids. In contrast, EC (embryogenic callus) showed elevated levels of fatty acid derivatives (α-ESA, 26-Methyl Nigranoate), phytoalexins (Wyerone acid), sesquiterpene (Alpha-santalal, Beta-guaiene), flavonoid glycosides, and plant hormones (Cis-Zeatin, Gibberellin A45), indicating a metabolically active state supporting somatic embryogenesis. Similarly, genes and transcription factors controlling cell differentiation and embryogenesis were upregulated during SE. This study provides a comprehensive resource to facilitate future genomic and genetic investigations aimed at deciphering the molecular basis of organogenesis and advancing research on somatic embryogenesis in bamboo.

Plant Somatic Embryogenesis Techniques

A robust biotechnology induces artificial genomic duplication via transient RNAi-mediated suppression of OSD1 in rice.

Ploidy manipulation is a crucial strategy for generating germplasm in crop breeding. However, artificial genomic duplication, often induced by colchicine treatment, is associated with toxicity and unpredictability. Although mutations in OSD1 have shown promise for inducing genomic duplication, the instability of ploidy across generations limits their practical application. In this study, we developed a Plant Polyploidization via Gene Interference (PPGI) system that utilizes transient RNAi-mediated suppression of OSD1 to efficiently induce artificial genomic duplication, demonstrating obvious potential for producing autotetraploids. We first validated this system by successfully generating PPGI-induced autotetraploid plants from the Taichung65 cultivar. These PPGI-induced plants exhibited notable differences from Taichung65 but resembled the existing Taichung65-4x line obtained through colchicine treatment. Haplotype analysis indicated that the OSD1 RNAi fragment is conserved across 2,908 rice cultivars. Consequently, we employed the same PPGI vector to develop autotetraploid lines from various germplasms, including another japonica cultivar, seven indica cultivars, and one Oryza rufipogon line. The probability of genomic duplication achieved by our PPGI method was higher than that obtained by colchicine treatment. Typically, autotetraploid lines exhibit severe sterility in the first generation following polyploidization. Leveraging fertile neo-tetraploid rice and the PPGI system, we designed and verified two strategies to directly induce fertile autotetraploid germplasms in the first generation, thereby substantially shortening the breeding cycle. Our method provides a universal, efficient, and non-toxic approach for inducing autotetraploid rice germplasms and contributes to enriching fertile autotetraploid rice germplasm resources.

OSD1

Efficient homologous replacement and deletion of large genomic fragments through template-jumping prime editing in rice.

Homologous replacement of genomic sequences with large DNA fragments (> 100 bp) holds great potential for crop breeding, yet an efficient method to achieve such edits is lacking in plants. Here, in rice, we developed template-jumping prime editing (TJ-PE), a recently reported PE strategy for large targeted insertion, as an efficient tool for homologous replacement with DNA fragments ranging from dozens to hundreds of base pairs, and using TJ-PE, we replaced genomic fragments of up to 340 bp with homologous fragments of the same length. In addition, our TJ-PE tool also enabled precise deletion of 944- to 2024-bp fragments in rice, with efficiencies of up to 34.6% for c. 2000-bp precise deletions. Collectively, this study expands the editing scope of PE in rice and establishes TJ-PE as a generalist tool for precise deletion and replacement of large DNA fragments.

Oryza

Tobamoviruses: Advances in Molecular Biology, Host Interactions and Integrated Disease Management.

Tobamoviruses (viruses in the genus Tobamovirus, family Virgaviridae) lead to major yield losses in economically important crops around the world. In this review, we go beyond the canonical gene expression framework by integrating recent discoveries of reverse open reading frames (rORFs) on the negative-strand RNA. These rORFs have only been experimentally validated in cucumber green mottle mosaic virus (CGMMV), with predicted sequence-conserved homologs across a subset of the genus, including TMV, ToBRFV, and PMMoV. However, they are not universally present in all tobamoviruses. We systematically dissect the infection cycle-from disassembly and replication to cell-to-cell and systemic movement-with an emphasis on the host factors hijacked at each stage. We synthesize current understanding of plant antiviral immunity, focusing on RNA silencing and NLR receptor-mediated resistance as two pillars of defense, along with the transcription factors and microRNAs that orchestrate these responses. We critically evaluate the experimental evidence for both plant defenses and viral counter-strategies, noting that many mechanistic models derive from limited model systems. We further characterize host genetic resistance and susceptibility factors applicable to crop breeding. These resources include dominant NLR and non-NLR resistance, as well as recessive resistance derived from modified host susceptibility genes. We address how viral mutations, recombination and fitness trade-offs undermine resistance durability. We then evaluate their practical deployment through conventional breeding, the exploitation of quantitative resistance, and genome editing, and outline associated agronomic drawbacks and regulatory constraints. Using ToBRFV as a case study, we analyze its epidemiological traits and assess the current arsenal of surveillance tools, from field diagnostics to remote sensing. Finally, we survey management strategies across a spectrum of maturity. Some approaches, including sanitation protocols and conventionally bred resistant cultivars, have proven effective under field conditions. The first dsRNA-based biopesticide has recently been registered in China, while other biological control agents and low-risk chemical approaches remain largely at the experimental stage. We also discuss the bottlenecks that impede lab-to-field transition and highlight promising solutions such as precision breeding and evolution-oriented cultivar deployment. By bridging molecular virology, epidemiology, and integrated disease management, this review provides a critical, bench-to-field framework for the sustainable control of tobamoviruses.

TMV