Search PubMedSearch

SEARCH · Search PubMed

Results for “coat protein”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The complete genomic sequence of a novel member of the genus Caulimovirus isolated from Dregea volubilis.

A novel caulimovirus was identified from diseased leaves of Dregea volubilis exhibiting yellowing and vein-associated chlorosis in Yuanjiang County, Yunnan Province, China. The virus was tentatively named Dregea volubilis caulimovirus 1 (DVCaV1). The complete genome sequence of DVCaV1, determined by de novo assembly of high-throughput sequencing data, comprises 8,160 bp of circular double-stranded DNA containing two intergenic regions and seven open reading frames (ORFs). These ORFs encode (in order) a movement protein (MP), an aphid transmission factor (ATF), a virion-associated protein (VAP), a coat protein (CP), a polymerase polyprotein (Pol, containing protease, reverse transcriptase, and RNase H domains), a transactivator/viroplasmin (TAV) protein, and a hypothetical protein of unknown function. Sequence comparisons revealed the highest nucleotide similarity with strawberry vein banding virus (SVBV; NC_001725). Phylogenetic analysis confirmed DVCaV1 as a member of the genus Caulimovirus, with SVBV as its closest known relative. According to current ICTV species demarcation criteria for the genus Caulimovirus (host range and > 20% nucleotide sequence divergence in the polymerase region), DVCaV1 represents a novel species. This is, to our knowledge, the first report of a caulimovirus detected in naturally symptomatic Dregea volubilis.

Genome, Viral

The proxiome of a plant viral protein with dual targeting to mitochondria and chloroplasts revealed MAPK cascade and splicing components as proviral factors.

The coat protein (CP) of the melon necrotic spot virus (MNSV) is a multifunctional factor localized in the chloroplast, mitochondria, and cytoplasm, playing a critical role in overcoming plant defenses such as RNA silencing (RNAi) and the necrotic hypersensitive response. However, the molecular mechanisms through which CP interferes with plant defenses remain unclear. Identifying viral-host interactors can reveal how viruses exploit fundamental cellular processes and help elucidate viral survival strategies. Here, we employed a TurboID-based proximity labeling approach to identify interactors of both the wild-type MNSV CP and a cytoplasmic CP mutant lacking the dual transit peptide (ΔNtCP). Of the interactors, eight were selected for silencing. Notably, silencing MAP4K SIK1 and NbMAP3Kε1 kinases, and a splicing factor homolog NbSMU2 significantly reduced MNSV accumulation, suggesting a proviral role for these proteins in plants. Yeast two-hybrid and bimolecular fluorescence complementation assays confirmed the CP and ΔNtCP interaction with NbSMU2 and NbMAP3Kε1 but not with NbSIK1, which interacted with NbMAP3Kε1. These findings open up new possibilities for exploring how MNSV CP might modulate gene expression and MAPK, thereby facilitating MNSV infection.

Chloroplasts

Infectious Clone Development of Zucchini Green Mottle Mosaic Virus Infecting Medicinal Plant Trichosanthes kirilowii and Establishment of a Serological Assay System.

Trichosanthes kirilowii has long been cultivated for application in traditional Chinese medicine. In this study, we identified two isolates of zucchini green mottle mosaic virus (ZGMMV; species Tobamovirus cucurbitae) from T. kirilowii plants. We determined the complete genome sequences of the ZGMMV isolates named ZGMMV-GL-1 and ZGMMV-GL-2. Each ZGMMV genome was 6,517 nucleotides in length, with only a single nucleotide variation detected between two sequences. Sequence analysis revealed that the ZGMMV isolates from this study shared 88.07 to 91.62% nucleotide identity with five other ZGMMV isolates deposited in GenBank. Phylogenetic analysis indicated that ZGMMV isolates can be clustered into two distinct groups; our two isolates shared the highest sequence similarity with the ZGMMV isolate from Nanning (GenBank accession number MF066176) and clustered within Group II. The coat protein (CP) gene was cloned from ZGMMV-infected T. kirilowii samples, and the CPZGMMV was expressed using the pET28(a) vector. Specific polyclonal antiserum CPZGMMV was generated by immunizing rabbits with the purified protein, and its sensitivity was determined to be satisfactory. Leveraging the high accuracy and sensitivity of the CPZGMMV antiserum, we developed a rapid, precise, and scalable diagnostic method for ZGMMV. We then constructed the full-length cDNA clones (ZGMMV-GL-1 and ZGMMV-GL-2). Additionally, the ZGMMV cDNA infectious clones from T. kirilowii were also able to infect Nicotiana benthamiana and Cucumis sativus systemically, inducing rough-textured and curled leaves in N. benthamiana and mosaic symptoms in C. sativus and T. kirilowii. In this study, we produced an antiserum against the ZGMMV CP and developed a sensitive, rapid, and reliable diagnostic assay, which lays a technical foundation for the detection and monitoring of ZGMMV. Therefore, the establishment of the ZGMMV infectious clone facilitates further research on viral protein functions, plant-pathogen interactions, and the formulation of effective ZGMMV management strategies.

Nicotiana benthamiana

Systematic mapping of insertion-tolerant regions enables capsid engineering of an infectious RNA phage.

RNA phages are attractive platforms for the design of programmable bioparticles, but their development has been constrained by limited knowledge of genomic sites that can tolerate sequence insertion. Here, we combined MuA transposase-mediated in vitro insertion mutagenesis with our established reverse genetics systems to systematically identify insertion-tolerant regions (ITRs) in the RNA phages MS2 and PP7. Screening of 4,555 MS2 and 2,228 PP7 random insertion clones identified 29 and 26 non-redundant ITRs, respectively. We further analyzed and compared these ITRs in the context of RNA genome organization and virion architecture. Both phages contained ITRs within the maturation protein, whereas only PP7 tolerated insertions within the coat protein (CP). On the basis of structural location and plaque-forming capacity, an ITR situated between Gly74 and Glu75 (GGC^GAG) in the PP7 CP was selected for further study. Infectious phage particles generated from complementary DNA clones retained the 15-bp insertion at both the RNA and protein levels. Engineered PP7 phages carrying an Arg-Gly-Asp motif inserted into the CP at this ITR displayed enhanced in vivo clearance in a Drosophila model, despite having in vitro stability comparable to that of the wild type. These findings provide the first example of CP engineering in an infectious RNA phage and establish a framework for engineering RNA phages for biological and biotechnological applications.IMPORTANCEA major obstacle to developing RNA phages as synthetic biology platforms is the lack of design principles for genomic insertion. Here, we address this limitation by establishing a mutagenesis-and-recovery workflow that systematically identifies insertion-tolerant regions (ITRs) in the RNA phages MS2 and PP7. The resulting maps reveal distinct structural constraints in the two phages and enable rational engineering of a peptide-display site in the PP7 capsid. Using this approach, we generated an engineered infectious phage with a modified capsid, thereby providing the first demonstration of capsid engineering in an infectious RNA phage, to our knowledge. This study lays the groundwork for the rational design of live RNA phage virions as tractable and engineerable scaffolds for future biological and biotechnological applications.

Animals

Functional overlap between the mammalian Sar1a and Sar1b paralogs in vivo.

Proteins carrying a signal peptide and/or a transmembrane domain enter the intracellular secretory pathway at the endoplasmic reticulum (ER) and are transported to the Golgi apparatus via COPII vesicles or tubules. SAR1 initiates COPII coat assembly by recruiting other coat proteins to the ER membrane. Mammalian genomes encode two SAR1 paralogs, SAR1A and SAR1B. While these paralogs exhibit ~90% amino acid sequence identity, it is unknown whether they perform distinct or overlapping functions in vivo. We now report that genetic inactivation of Sar1a in mice results in lethality during midembryogenesis. We also confirm previous reports that complete deficiency of murine Sar1b results in perinatal lethality. In contrast, we demonstrate that deletion of Sar1b restricted to hepatocytes is compatible with survival, though resulting in hypocholesterolemia that can be rescued by adenovirus-mediated overexpression of either SAR1A or SAR1B. To further examine the in vivo function of these two paralogs, we genetically engineered mice with the Sar1a coding sequence replacing that of Sar1b at the endogenous Sar1b locus. Mice homozygous for this allele survive to adulthood and are phenotypically normal, demonstrating complete or near-complete overlap in function between the two SAR1 protein paralogs in mice. These data also suggest upregulation of SAR1A gene expression as a potential approach for the treatment of SAR1B deficiency (chylomicron retention disease) in humans.

Animals

Functional overlap between the mammalian Sar1a and Sar1b paralogs in vivo.

Proteins carrying a signal peptide and/or a transmembrane domain enter the intracellular secretory pathway at the endoplasmic reticulum (ER) and are transported to the Golgi apparatus via COPII vesicles or tubules. SAR1 initiates COPII coat assembly by recruiting other coat proteins to the ER membrane. Mammalian genomes encode two SAR1 paralogs, SAR1A and SAR1B. While these paralogs exhibit ~90% amino acid sequence identity, it is unknown whether they perform distinct or overlapping functions in vivo. We now report that genetic inactivation of Sar1a in mice results in lethality during mid-embryogenesis. We also confirm previous reports that complete deficiency of murine Sar1b results in perinatal lethality. In contrast, we demonstrate that deletion of Sar1b restricted to hepatocytes is compatible with survival, though resulting in hypocholesterolemia that can be rescued by adenovirus-mediated overexpression of either SAR1A or SAR1B. To further examine the in vivo function of these 2 paralogs, we genetically engineered mice with the Sar1a coding sequence replacing that of Sar1b at the endogenous Sar1b locus. Mice homozygous for this allele survive to adulthood and are phenotypically normal, demonstrating complete or near-complete overlap in function between the two SAR1 protein paralogs in mice. These data also suggest upregulation of SAR1A gene expression as a potential approach for the treatment of SAR1B deficiency (chylomicron retention disease) in humans.

Preprint

The first complete genome sequence of Ammi majus latent virus from the new natural host culantro.

A potyvirus (isolate AMLV-CQ) infecting culantro (Eryngium foetidum L.) imported from Vietnam was identified by RT-PCR. The complete genome sequence of AMLV-CQ was determined to be 9,549 nucleotides in length. It contains a large open reading frame encoding a 3,082-amino-acid putative polyprotein, flanked by 5´ and 3´ untranslated regions (UTRs) of 77 and 226 nt, respectively. AMLV-CQ is closely related to five other completely sequenced potyviruses, sharing 68-69% nucleotide and 69-70% amino acid sequence identity. However, the coat protein (CP) gene shares 89% nucleotide and 93% amino acid sequence identity with that of a partially sequenced potyvirus, Ammi majus latent virus (isolate AMLV-WF17). These results suggest that AMLV-CQ and AMLV-WF17 are isolates of the same species. To our knowledge, this is the first report of a complete genome sequence of an AMLV isolate, and culantro was identified as a new natural host for this virus. In addition, a one-step RT-PCR assay was developed that provides a rapid, robust, and highly sensitive approach for the detection of AMLV.

Eryngium

Cucurbit Leaf Crumple Virus: An Important Pathogen of Cucurbit and Snap Bean Crops.

TAXONOMY: Cucurbit leaf crumple virus (CuLCrV); Begomovirus cucurbitae; Geminiviridae; Geplafuvirales. GEOGRAPHICAL DISTRIBUTION: The presence of CuLCrV is exclusively limited to North America, mainly Mexico and the United States. PHYSICAL PROPERTIES: CuLCrV is a bipartite begomovirus comprising two circular single-stranded DNA molecules (DNA-A and DNA-B), encapsidated within geminate icosahedral particles. GENOME AND ORGANIZATION: CuLCrV possesses a bipartite genome of DNA-A (2632 nucleotides) and DNA-B (2600 nucleotides). DNA-A contains five open reading frames (ORFs): AV1 (coat protein), AC1 (replication-associated protein), AC2 (transcriptional activator protein), AC3 (replication enhancer protein) and AC4. DNA-B contains two ORFs: BV1 (nuclear shuttle protein) and BC1 (movement protein). TRANSMISSION: CuLCrV is transmitted by the sweetpotato whitefly, Bemisia tabaci, in a persistent, circulative and non-propagative manner. HOSTS: CuLCrV primarily infects crop members of the Cucurbitaceae and snap bean (Phaseolus vulgaris, Fabaceae). Multiple weed species belonging to Brassicaceae, Convolvulaceae, Cucurbitaceae and Verbenaceae act as persistent virus reservoir hosts. SYMPTOMS: Symptom expression varies with host and infection timing. In cucurbits, infection induces leaf crumpling, thickening and downward curling of leaves, with green streaks and distortion of fruits. In snap bean, symptoms include leaf distortion, chlorosis and malformed pods. CONTROL: No commercial cultivars with resistance to CuLCrV are available for cucurbit crops, although some resistance has been reported in snap bean cultivars. Therefore, management relies primarily on integrated disease management.

Plant Diseases

Grains, trade and war in the multimodal transmission of Rice yellow mottle virus: An historical and phylogeographical retrospective.

Rice yellow mottle virus (RYMV) is a major pathogen of rice in Africa. RYMV has a narrow host range limited to rice and a few related poaceae species. We explore the links between the spread of RYMV in East Africa and rice history since the second half of the 19th century. The phylogeography of RYMV in East Africa was reconstructed from coat protein gene sequences (ORF4) of 335 isolates sampled over two million square kilometers between 1966 and 2020. Dispersal patterns obtained from ORF2a and ORF2b, and full-length sequences converged to the same scenario. The following imprints of rice cultivation on RYMV epidemiology were unveiled. RYMV emerged in the middle of the 19th century in the Eastern Arc Mountains where slash-and-burn rice cultivation was practiced. Several spillovers from wild hosts to cultivated rice occurred. RYMV was then rapidly introduced into the nearby large rice growing Kilombero valley and Morogoro region. Harvested seeds are contaminated by debris of virus infected plants that subsist after threshing and winnowing. Long-distance dispersal of RYMV is consistent (i) with rice introduction along the caravan routes from the Indian Ocean Coast to Lake Victoria in the second half of the 19th century, (ii) seed movement from East Africa to West Africa at the end of the 19th century, from Lake Victoria to the north of Ethiopia in the second half of the 20th century and to Madagascar at the end of the 20th century, (iii) and, unexpectedly, with rice transport at the end of the First World War as a troop staple food from the Kilombero valley towards the South of Lake Malawi. Overall, RYMV dispersal was associated to a broad range of human activities, some unsuspected. Consequently, RYMV has a wide dispersal capacity. Its dispersal metrics estimated from phylogeographic reconstructions are similar to those of highly mobile zoonotic viruses.

Oryza

Multidimensional Protein Corona Analysis Toward Predictive Nano-Bio Interface Design.

Nanoparticles entering biological fluids are rapidly coated by proteins and other biomolecules, converting their synthetic surfaces into biologically active nano-bio interfaces. These coronas regulate colloidal stability, immune recognition, cellular uptake, biodistribution, pharmacokinetics, cargo delivery, and toxicity. Yet a protein list obtained by mass spectrometry captures only part of this interface. Corona identity and function are also shaped by protein organization, binding stability, exchange dynamics, conformational changes, and molecular accessibility. Here, we discuss recent progress in protein corona isolation and analysis from a question-oriented analytical perspective, with emphasis on how centrifugation, magnetic recovery, affinity- or chemistry-enabled capture, chromatography, filtration, and field-flow fractionation (FFF) influence the fidelity, integrity, and comparability of recovered coronas. We then examine how proteomic profiling can be integrated with binding measurements, interfacial structural analysis and functional validation to distinguish descriptive corona signatures from biologically meaningful mechanisms. We further consider how biofluid composition, disease state, tissue interfaces and cellular environments remodel corona identity, presentation, and bioactivity. Finally, we argue that standardized reporting, computational modeling, and AI-enabled approaches are essential for converting protein corona datasets into reproducible and predictive knowledge that can guide the design of drug delivery systems and precision nanomedicines.

Protein Corona

CRISPR/Cas9 screenings reveal the role of STX1A and CDK1 in Cathepsin G entering and killing colorectal cancer cells.

Neutrophils are the major populations of white blood cells and have been reported to facilitate cancer metastasis. Meanwhile, emerging evidence has recently suggested the anti-cancer role of neutrophils. Our previous study revealed that CB-839 and 5-FU-treated colorectal cancer (CRC) tumors recruited neutrophils and induced neutrophil extracellular traps (NETs). Cathepsin G (CTSG), which is released during NET formation, enters CRC cells through the receptor for advanced glycation end products (RAGE) and cleaves 14-3-3ε to promote apoptosis. However, the detailed mechanism underlying CTSG's anti-tumor function remains less studied. In this study, we report that CTSG enters CRC cells through RAGE-mediated endocytosis. Knocking out RAGE or inhibiting endocytosis blocks CTSG from entering CRC cells and attenuates CTSG-induced apoptosis. Furthermore, the clathrin coat assembly complex and SNARE proteins were enriched in an arrayed CRISPR/Cas9 screening targeting human membrane trafficking genes. Knocking out SNARE protein STX1A prevents the spread of CTSG in CRC cells and the induction of cleaved PARP. A pooled genome-wide CRISPR/Cas9 screening further identifies the role of CDK1 in the NET-induced killing of CRC cells. Inhibiting CDK1 protected CRC cells from killing by CTSG. Our study reveals novel mechanisms by which CTSG enters and kills CRC cells.

CDK1

Treponema pallidum fibronectin-binding proteins.

Putative adhesins were predicted by computer analysis of the Treponema pallidum genome. Two treponemal proteins, Tp0155 and Tp0483, demonstrated specific attachment to fibronectin, blocked bacterial adherence to fibronectin-coated slides, and supported attachment of fibronectin-producing mammalian cells. These results suggest Tp0155 and Tp0483 are fibronectin-binding proteins mediating T. pallidum-host interactions.

Adhesins, Bacterial

Role of functional genes for seed vigor related traits through genome-wide association mapping in finger millet (Eleusine coracana L. Gaertn.).

Finger millet (Eleusine coracana (L.) Gaertn.) is a calcium-rich, nutritious and resilient crop that thrives even in harsh environmental conditions. In such ecologies, seed longevity and seedling vigor are crucial for sustainable crop production amid climate change. The current study explores the genetics of accelerated aging on seed longevity traits across 221 diverse accessions of finger millet through genome-wide association approach (GWAS). A significant variation was identified in germination percentage, germination rate indices, mean germination time, seedling vigor indices and dry weight upon aging treatment. GWAS model from 11,832 high-quality SNPs identified through Genotyping-by-Sequencing (GBS) approach produced 491 marker-trait associations (MTAs) for 27 traits, of which 54 were FDR-corrected. A pleiotropic SNP, FM_SNP_9478 identified on chromosome 7B was associated with the traits viz., germination after aging, germination index after aging and their relative measures. Functional annotation revealed DET1 and expansin-A2 influenced seed coat integrity, critical for germination and aging resilience. Probable protein phosphatase 2C3 and piezo-type ion channels contributed to mechanical sensing and stress adaptation in seeds. Beta-amylase and acetyl-CoA carboxylase 2 were identified for seed metabolism and stress response. These insights lay the framework for targeted breeding efforts to improve seed quality and resilience under diverse production conditions.

Eleusine

Cryo-electron tomography reveals coupled flavivirus replication, budding and maturation.

Flaviviruses replicate their genomes in replication organelles (ROs) formed as bud-like invaginations on the endoplasmic reticulum (ER) membrane, which also functions as the site for virion assembly. While this localization is well established, it is not known to what extent viral membrane remodeling, genome replication, virion assembly, and maturation are coordinated. Here, we imaged tick-borne flavivirus replication in human cells using cryo-electron tomography. We find that the RO membrane bud is shaped by a combination of a curvature-establishing coat and the pressure from intraluminal template RNA. A protein complex at the RO base extends to an adjacent membrane, where immature virions bud. Naturally occurring furin site variants determine whether virions mature in the immediate vicinity of ROs. We further visualize replication in mouse brain tissue by cryo-electron tomography. Taken together, these findings reveal a close spatial coupling of flavivirus genome replication, budding, and maturation.

Journal Article

Oocyte surface proteins EGG-1 and EGG-2 are required for eggshell integrity in Caenorhabditis elegans.

Metazoan eggs are surrounded by a specialized coat of extracellular matrix that mediates sperm-egg interactions. This coat is rapidly remodeled after fertilization to form a barrier that prevents polyspermy, protects against environmental insults, and provides structural support to the developing embryo. In C. elegans several oocyte surface proteins have been identified that mediate these events. However, whether two of these proteins, EGG-1 and EGG-2, are required for fertilization or downstream events has been unclear. Here, we address this question using more recent advances in genome editing tools through the creation of egg-1 egg-2 deletions of the endogenous loci. We found that egg-1 egg-2 oocytes are fertilization competent and form rudimentary eggshells. While the integrity of the egg-1 egg-2 eggshells are compromised and often rupture within the uterus, surprisingly, some embryos are capable of undergoing several rounds of cell division. Overall, our findings demonstrate that EGG-1 and EGG-2 are not required for fertilization but are involved in post-fertilization processes.

Caenorhabditis elegans

Mechanistic roles of GmSWEET10a/b and GmSUT1 in the oil-protein balance in soybean mature seeds at transcriptional and metabolic levels.

Previous investigations indicated that the soybean (Glycine max) SUGARS WILL EVENTUALLY BE EXPORTED TRANSPORTER10a/b (GmSWEET10a/b) genes promote oil accumulation, while inhibiting protein accumulation in seeds. To clarify the mechanisms modulated by GmSWEET10a/b in mediating the oil and protein accumulations in soybean seeds, an integrated comparative multiomics was conducted using the double gmsweet10a,b mutant and wild-type (WT) embryos. Spatial metabolomic analysis revealed that gmsweet10a,b embryos were surrounded by a sugar-reduced seed coat and experienced a sugar-starvation state in embryonic tissues in vivo. The decreased sugar content in the gmsweet10a,b embryos reduced the availability of carbon skeletons required for oil synthesis and was associated with decreased expression levels of genes involved in sucrose metabolism, fatty acid biosynthesis, and triacylglycerol assembly. Meanwhile, the expression of genes encoding storage protein was induced in gmsweet10a,b embryos, when compared with WT. These changes resulted in decreased oil content and increased protein content in gmsweet10a,b embryos versus WT. In vitro sugar-starvation assay also supported the suppression of fatty acid biosynthesis and the enhanced storage protein accumulation in developmental embryo under sugar-starved conditions. Furthermore, the knockout of SUCROSE TRANSPORTER 1 (GmSUT1), which was upregulated in gmsweet10a,b embryos, significantly decreased the sugar level, resulting in lower oil content but higher protein content in gmsut1 embryos than WT ones. Our findings provided a mechanistic understanding of the modulation of sugar transport between seed coat to embryo by both GmSWEET10a/b and GmSUT1, which plays a pivotal role in balancing oil and protein accumulations in soybean mature seeds.

Seeds

Proteomic insights into the immunomodulatory effects of Ca/Sr co-doped sol-gel coatings for titanium implants.

Ionic functionalization of biomaterial coatings has emerged as a powerful strategy to regulate early host responses at the implant interface. However, how combined Ca/Sr incorporation governs the adsorbed proteome and downstream immune signaling remains poorly understood. This study analyses, employing in vitro tests and proteomics, the effect of adding Sr and Ca to Si-based coatings designed to bioactivate Ti implants. Hybrid Si-based coatings were synthesized by the sol-gel route with a fixed Ca content (0.5 wt%) and increasing Sr contents (0.5, 1.0, 1.5 wt%), and their physicochemical properties, ion release kinetics, and hydrolytic stability were characterized. The coatings remained highly crosslinked despite Ca/Sr incorporation, whereas the highest Sr content increased hydrolytic degradation to around 70% after 56 days. Proteomic analysis identified 183 adsorbed proteins, of which 56 were differentially adsorbed on Ca/Sr-coatings, mainly associated with immune and coagulation pathways. In vitro, RAW 264.7 showed increased gene expression of TNF-α and TGF-β; with an enhanced TNF-α secretion by the addition of Ca and Sr. In parallel, MC3T3-E1 indicated that Ca/Sr-coatings were not cytotoxic and did not impair cell proliferation. However, ALP activity was reduced in the co-doped groups, indicating that the immunomodulatory effects induced by Ca/Sr incorporation were not accompanied by enhanced early osteogenic differentiation. The Ca/Sr combination induced alterations in the adsorption of immune-related proteins, which correlated with the in vitro findings. The deeper insight into how Ca/Sr mixtures modulate protein adsorption on biomaterial surfaces may be key to understanding the immunomodulatory capacity of these bioactive cations.

Animals

Analysis of deep-resequencing data of 984 soybean accessions reveals structural variations underlying agronomic traits.

Genomic structural variants (SVs) are major sources of genetic variation and have profound impacts on phenotypic traits. However, their functional effects remain largely unexplored in soybean. Here, we resequence 940 soybean accessions. Together with 44 publicly available datasets, we identify 602,281 SVs. Using a graph-based genome, we detect an additional 58,760 presence/absence variations (PAVs) that broadly affect gene expression. Population genomic analyses reveal that SVs serve as a core driving force for soybean domestication and improvement. Integrating SVs with QTLs for oil and protein content, and performing GWAS on 27 traits, we identify key functional SVs. These include transposable element insertions altering seed coat color, multiple insertions within a cytochrome P450 gene modifying flower and hypocotyl color, and a GmMATE1 deletion enhancing seed size. Together, our study establishes a comprehensive SV map of soybean, offering a valuable resource for dissecting the genetic basis of complex traits to accelerate molecular breeding.

Glycine max