Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Relational genome analysis using reference libraries and hybridisation fingerprinting.

The genomes of eukaryotic organisms are studied by an integrated approach based on hybridisation techniques. For this purpose, a reference library system has been set up, with a wide range of clone libraries made accessible to probe hybridisation as high density filter grids. Many different library types made from a variety of organisms can thus be analysed in a highly parallel process; hence, the amount of work per individual clone is minimised. In addition, information produced on one analysis level instantly assists in the characterisation process on another level. Genetic, physical and transcriptional mapping information and partial sequencing data are obtained for the individual library clones and are cross-referenced toward a comprehensive molecular understanding of genome structure and organisation, of encoded functions and their regulation. The order of genomic clones is established by hybridisation fingerprinting procedures. On these physical maps, the location of transcripts is determined. Complementary, partial sequence information is produced from corresponding cDNAs by hybridising short oligonucleotides, which will lead to the identification of regions of sequence conservation and the constitution of a gene inventory. The hybridisation analysis of the cDNA clones, and the genomic clones as well, could potentially be expanded toward a determination of (nearly) the complete sequence. The accumulated data set will provide the means to direct large-scale sequencing of the DNA, or might even make the sequence analysis of large genomic regions a redundant undertaking due to the already collected information.

Animals↗

Unraveling the genomic blueprint of the Indian black soldier fly: From genome assembly to evolutionary insights.

The black soldier fly (BSF) (Hermetia illucens) has been renowned for its sustainable bioconversion capabilities, resulting in smart protein production with wide applications in animal feed, bioenergy, and biofertilizer. However, the genetic mechanisms underlying efficient bioconversion and productivity remain poorly understood. To advance strain-specific applications and strengthen genetic resource availability, we present the whole genome sequencing (WGS) data for an Indian isolate of black soldier fly. The assembled genome was 1.46 Gb with a scaffold N50 of 172.7 Mb, and a GC content of 42.6%. Furthermore, 64.17% of genomic sequences were masked as repeated, and 14,317 protein-coding sequences were identified. Variant analysis against the reference genome identified 34.44 million variants (∼33.25 million SNPs and ∼ 1.18 million INDELs), with the majority (99.3%) classified as MODIFIER, 0.54% as LOW impact, 0.14% as MODERATE, and only 0.003% as HIGH impact. Comparative genomic analysis with other related species revealed expansions of gene families in BSF associated with Immune effector (Antimicrobial peptides (AMPs), Lysozymes, and Peptidoglycan Recognition Protein (PGRP) and Detoxification (cytochrome P450 enzymes). Notably, AMPs in the Indian isolate showed enhanced copy number variation in defensin (27) and PGRP (40) compared to reference BSF, suggesting potential regional adaptations to pathogen exposure. Collectively, this genomic data provides an improved resource for evolutionary studies, functional genomics, and targeted genetic improvement of BSF for sustainable bioconversion applications.

Comparative genomics↗

Diagnosing missed cases of spinal muscular atrophy in genome, exome, and panel sequencing data sets.

PURPOSE: We set out to develop a publicly available tool that could accurately diagnose spinal muscular atrophy (SMA) in exome, genome, or panel sequencing data sets aligned to a GRCh37, GRCh38, or T2T reference genome. METHODS: The SMA Finder algorithm detects the most common genetic causes of SMA by evaluating reads that overlap the c.840 position of the SMN1 and SMN2 paralogs. It uses these reads to determine whether an individual most likely has 0 functional copies of SMN1. RESULTS: We developed SMA Finder and evaluated it on 16,626 exomes and 3911 genomes from the Broad Institute Center for Mendelian Genomics, 1157 exomes and 8762 panel samples from Tartu University Hospital, and 198,868 exomes and 198,868 genomes from the UK Biobank. SMA Finder's false-positive rate was below 1 in 200,000 samples, its positive predictive value was greater than 96%, and its true-positive rate was 29 out of 29. Most of these SMA diagnoses had initially been clinically misdiagnosed as limb-girdle muscular dystrophy. CONCLUSION: Our extensive evaluation of SMA Finder on exome, genome, and panel sequencing samples found it to have nearly 100% accuracy and demonstrated its ability to reduce diagnostic delays, particularly in individuals with milder subtypes of SMA. Given this accuracy, the common misdiagnoses identified here, the widespread availability of clinical confirmatory testing for SMA, and the existence of treatment options, we propose that it is time to add SMN1 to the American College of Medical Genetics list of genes with reportable secondary findings after genome and exome sequencing.

Humans↗

Highly Contiguous Is Not Chromosomally Accurate: Integrated Cytogenetic and Genomic Mapping in Two Turtle Genome.

High-quality genome assemblies are essential for robust research across biological and medical fields. Assembly errors can have far-reaching consequences for downstream analyses, including gene annotation and the inference of synteny. In contrast to the rapid growth of genomic data volume, there is a notable lag in the integration of chromosome-level assemblies with cytogenetic data. We conducted the first direct genome-to-genome comparison, integrating comparative chromosome painting, the alignment of chromosome-specific probes to available genome assemblies, and synteny-based comparison of independent chromosome-level assemblies of the loggerhead sea turtle (Caretta caretta, 2n = 56) and the red-eared slider (Trachemys scripta elegans, 2n = 50). Using two independent sets of flow-sorted chromosome-specific probes in cross-species hybridizations, together with the sequencing and mapping of chromosome-derived DNA libraries, we assigned assembled scaffolds to all physical chromosomes of both species. In C. caretta, chromosomal assignments and genome-wide synteny were fully consistent with the published assembly, except for the reduced sizes of two microchromosome scaffolds, which we attribute to under-representation of repetitive DNA. In contrast, in T. s. elegans, cytogenetic validation of the assemblies revealed a false rearrangement compared to a missed one. Our results show that even highly contiguous vertebrate genome assemblies can misrepresent chromosome structure. When cytogenetic analyses reveal such inaccuracies, updated reference genomes should be generated for widely studied species to enable accurate inference of karyotype evolution and downstream comparative genomic analyses.

FISH↗

'PePApipe': A complete bioinformatics analysis pipeline for African Swine Fever Virus genome.

African Swine Fever Virus (ASFV) is of high concern in porcine livestock across the world due to both the high mortality rates and the trade restrictions imposed on affected regions. The viral genome is large and complex, and genomic analysis is essential for tracing its origin and evolution. Although several bioinformatics tools exist for genome assembly and analysis, no single platform integrates all necessary steps in an accessible and systematic way. In this study the authors developed 'PePApipe', a custom-built, user-friendly pipeline that enables rapid, complete, and efficient ASFV genome analysis. It is specifically designed for laboratory professionals with limited bioinformatics experience, requiring only basic command-line knowledge. Starting from raw sequencing data, PePApipe integrates thirteen software tools into one automated workflow, covering quality control and pre-processing of raw reads, de novo genome assembly and variant calling. Programmed in Python, it can be executed locally through bash scripts, or using a Slurm protocol for batch processing of multiple samples. The main outputs are the ASFV consensus genome sequence and a file listing its putative variants compared to the selected reference genome. PePApipe classifies generated files into structured folders and produces intermediate files that can be used as inputs for further or parallel analyses; users can also enable or disable specific steps in each particular case. This pipeline is adaptable and complementary to downstream steps such as viral genome annotation or genome visualization. By consolidating all stages of viral genome analysis into a single automated workflow, PePApipe reduces the likelihood of user error, and enhances reproducibility and efficiency. This user-friendly pipeline facilitates the transition from sequencing to assembly and downstream analysis of viral genomes, ensuring a fast and reliable response to molecular analysis demands. Finally, the pipeline can be easily adapted to the study of other viral species, expanding its application in infectious diseases surveillance.

African Swine Fever Virus↗

Mining of important genetic loci and evaluation of genetic effects for growth traits in Baicheng You Chicken.

The Baicheng You Chicken is a precious indigenous breed in Xinjiang, China, prized for its strong disease and stress resistance and superior meat quality. However, the lack of scientific breeding and conservation has led to poor production performance, particularly in growth traits. In this study, we collected phenotypic and whole-genome resequencing data from 1,535 18-week-old Baicheng You Chickens (180 males and 1,355 females). After stringent quality control (SNP call rate > 95%, minor allele frequency > 1%), we constructed the breed's first comprehensive SNP-based genome-wide variation map, which comprised 2,020,743 high-quality SNPs across the genome. The filtered SNPs had high mapping quality (99.73% mapped to the bGalGal1.mat.broiler.GRCg7b reference genome, Q30 = 93.26%) and a reasonable Ti/Tv ratio (2.596), guaranteeing the reliability of subsequent analyses. We estimated genetic effects (SNP-based heritability and phenotypic variance explained (PVE) by individual loci) via the restricted maximum likelihood (REML) method, and performed a genome-wide association study (GWAS) using a mixed linear model (MLM) - with sex as a fixed effect and principal components to correct for population stratification - to identify significant loci and their effect sizes (Beta). All eight growth traits showed moderate to high heritability: body weight (BW) had the highest heritability (0.86±0.11), while chest width (CW, 0.41±0.08) and body slanting length (BSL, 0.43±0.09) were the lowest; keel length (KL), chest girth (CG), pelvic width (PW), chest depth (CD) and shank length (SL) had heritabilities of 0.50±0.09, 0.46±0.09, 0.54±0.09, 0.67±0.10 and 0.74±0.10, respectively. GWAS identified 145 significant SNPs, with a maximum Beta value of 0.39 and PVE ranging from 1.25% to 6.25%. We annotated 22 candidate genes, with TAPT1, IGF2BP1, ADGRB3, LDB2, NCAPG and LCORL as key candidates. These quantifiable genetic markers and effect estimates provide direct targets for marker-assisted selection (MAS) and valuable resources for future genomic selection (GS) programs, offering a practical approach to improve the breed's slow growth while preserving its unique meat quality.

Baicheng You Chicken↗

Gene identification in novel eukaryotic genomes by self-training algorithm.

Finding new protein-coding genes is one of the most important goals of eukaryotic genome sequencing projects. However, genomic organization of novel eukaryotic genomes is diverse and ab initio gene finding tools tuned up for previously studied species are rarely suitable for efficacious gene hunting in DNA sequences of a new genome. Gene identification methods based on cDNA and expressed sequence tag (EST) mapping to genomic DNA or those using alignments to closely related genomes rely either on existence of abundant cDNA and EST data and/or availability on reference genomes. Conventional statistical ab initio methods require large training sets of validated genes for estimating gene model parameters. In practice, neither one of these types of data may be available in sufficient amount until rather late stages of the novel genome sequencing. Nevertheless, we have shown that gene finding in eukaryotic genomes could be carried out in parallel with statistical models estimation directly from yet anonymous genomic DNA. The suggested method of parallelization of gene prediction with the model parameters estimation follows the path of the iterative Viterbi training. Rounds of genomic sequence labeling into coding and non-coding regions are followed by the rounds of model parameters estimation. Several dynamically changing restrictions on the possible range of model parameters are added to filter out fluctuations in the initial steps of the algorithm that could redirect the iteration process away from the biologically relevant point in parameter space. Tests on well-studied eukaryotic genomes have shown that the new method performs comparably or better than conventional methods where the supervised model training precedes the gene prediction step. Several novel genomes have been analyzed and biologically interesting findings are discussed. Thus, a self-training algorithm that had been assumed feasible only for prokaryotic genomes has now been developed for ab initio eukaryotic gene identification.

Algorithms↗

Comparative genomics approaches to identify genomic regions associated with the antimicrobial activity of Pseudomonas protegens PBL3.

The environmental bacterium Pseudomonas protegens PBL3 has antagonistic activity against the plant pathogenic bacterium Burkholderia glumae, an important pathogen in rice. The antimicrobial activity of P. protegens PBL3 was found in the bacteria-free secreted fraction (secretome), but the specific molecules, as well as the genetic basis of that activity, have not been identified. In this study, we integrated genomic information with antimicrobial assays on P. protegens PBL3 and additional six Pseudomonas spp. strains, to identify putative genomic regions in P. protegens PBL3 associated with antimicrobial activity. We hypothesized that Pseudomonas spp. strains with antimicrobial activity against B. glumae have conserved genes with P. protegens PBL3 that are absent in strains lacking activity. Comparative genomics analyses with anvi'o and progressiveMauve, and using P. protegens PBL3 as the reference genome, revealed 188 genes uniquely present in antimicrobial-producing strains. Seven of those genes were annotated as biosynthetic gene clusters predicted to encode secondary metabolites; additional genes were grouped into 25 contiguous clusters with functions annotated as secretion, signal transduction, regulation, transport/efflux, carbohydrate metabolism and one with an additional uncharacterized function. Altogether, this study uncovered a complex and multi-functional network of candidate genes, suggesting that the antimicrobial activity in P. protegens PBL3 is not limited to biosynthetic pathways but also involves additional regulatory, metabolic and export modules to synthesize and deploy antimicrobials.

Pseudomonas↗

Pangenomes aid accurate detection of large insertions and deletions from targeted sequencing: the case of cardiomyopathies.

BACKGROUND: Gene panels represent a widely used strategy for genetic testing in a vast range of Mendelian disorders. While this approach aids reliable bioinformatic detection of short coding variants, it often fails to detect many larger variants. Recent studies have recommended the adoption of pangenome references (as opposed to linear reference genomes like GRCh38) to augment detection of large variants from targeted sequencing, potentially providing diagnostic laboratories with the possibility to streamline diagnostic work-ups and reduce costs. METHODS: Here, we analyze 1969 cardiomyopathy cases and 1805 controls sequenced with the Illumina Trusight Cardio panel using a pangenome-based workflow (GRAF) and five conventional orthogonal methodologies (GATK HaplotypeCaller, GATK-gCNV, ExomeDepth, Manta and Lumpy-SV) to detect variants ≥ 20 bp in size. RESULTS: Following lab-based variant validation by means of PCR and Sanger sequencing, we show that GRAF conjugates higher precision and recall (F1 score 0.86) compared with other methods (F1 0-0.57) in detecting potentially pathogenic variants ≥ 20 bp from short-read panel data. Results were complemented by a comparison of the tools' performance in detecting ground truth variants on reference sample HG002 from Genome In A Bottle, which confirmed GRAF to outperform other tools also on exome sequencing (F1 0.97 vs. 0-0.94). Notably, in the HG002 benchmark dataset, GRAF also showed slightly improved performance compared to GATK HaplotypeCaller in the identification of small variants (1-19 bp; F1 0.975 vs. 0.968). CONCLUSIONS: Our results indicate that pangenome-based workflows aid improved detection of large variants from targeted sequencing data in the clinical context and suggest that they may contribute to more unified variant detection frameworks for all-size genetic variants in the future.

Humans↗

Saccharomyces cerevisiae S288C genome annotation: a working hypothesis.

The S. cerevisiae genome is the most well-characterized eukaryotic genome and one of the simplest in terms of identifying open reading frames (ORFs), yet its primary annotation has been updated continually in the decade since its initial release in 1996 (Goffeau et al., 1996). The Saccharomyces Genome Database (SGD; www.yeastgenome.org) (Hirschman et al., 2006), the community-designated repository for this reference genome, strives to ensure that the S. cerevisiae annotation is as accurate and useful as possible. At SGD, the S. cerevisiae genome sequence and annotation are treated as a working hypothesis, which must be repeatedly tested and refined. In this paper, in celebration of the tenth anniversary of the completion of the S. cerevisiae genome sequence, we discuss the ways in which the S. cerevisiae sequence and annotation have changed, consider the multiple sources of experimental and comparative data on which these changes are based, and describe our methods for evaluating, incorporating and documenting these new data.

Base Sequence↗

Burkholderia arboris bacteremia initially identified as Burkholderia cepacia complex: a genome-based case report.

We report a bloodstream Burkholderia arboris isolate from a 75-year-old man without cystic fibrosis. The organism was recovered from both aerobic bottles of two separately collected blood-culture sets and was initially assigned to the Burkholderia cepacia complex (Bcc) by matrix-assisted laser desorption ionization-time-of-flight mass spectrometry. Whole-genome sequencing yielded three circular chromosomes and one circular plasmid. DFAST_QC identified B. arboris as the only type-strain match above the species threshold, with an average nucleotide identity of 99.48%; the next-highest match was B. seminalis at 93.33%. Multilocus sequence typing identified ST-2575, and ResFinder detected no acquired antimicrobial resistance genes. The patient improved after 14 days of meropenem therapy without recurrent B. arboris bacteremia. This report adds a clinically supported bloodstream infection, a complete genome resource, and detailed susceptibility data, while illustrating the importance of up-to-date reference genomes for species-level interpretation of unusual Bcc isolates.

Humans↗

A high-quality chromosome-level genome assembly of apple of Peru (Nicandra physalodes).

Nicandra physalodes, a member of the Solanaceae family, is known for its medicinal potential and strong natural insect-repellent properties, which are mainly attributed to its bioactive withanolides and alkaloids. Despite its ecological and pharmacological significance, genomic information for this species has remained limited. Here, we generated a chromosome-level reference genome for N. physalodes based on PacBio high-fidelity (HiFi) long-read sequencing and Hi-C scaffolding. The assembled genome is 933.97 Mb in size, with a contig N50 of 87.37 Mb, and 99.95% (933.54 Mb) of the sequences anchored to 10 pseudochromosomes. Repetitive elements account for 73.06% of the genome, and 27,925 protein-coding genes were predicted, 97.81% of which were functionally annotated. This genomic resource provides a valuable foundation for investigating the genetic basis of specialized metabolite biosynthesis, insect resistance, and environmental adaptation in N. physalodes, as well as for comparative studies within the Solanaceae family.

Genome, Plant↗

Genome analysis of Thinopyrum intermedium and Thinopyrum ponticum using genomic in situ hybridization.

Genomic in situ hybridization (GISH) using genomic DNA probes from Thinopyrum elongatum (Host) D.R. Dewey (genome E, 2n = 14), Thinopyrum bessarabicum (Savul. & Rayss) A. Löve (genome J, 2n = 14), and Pseudoroegneria strigosa (M. Bieb.) A. Löve (genome S, 2n = 14), was used to examine the genomic constitution of Thinopyrum intermedium (Host) Barkworth & D.R. Dewey (2n = 6x = 42) and Thinopyrum ponticum (Podp.) Barkworth & D.R. Dewey (2n = 10x = 70). Evidence from GISH indicated that hexaploid Th, intermedium contained the J, Js, and S genomes, in which the J genome was related to the E genome of Th. elongatum and the J genome of Th. bessarabicum. The S genome was homologous to the S genome of Ps. strigosa, while the Js genome referred to modified J- or E-type chromosomes distinguished by the presence of S genome specific sequences close to the centromere. Decaploid Th. ponticum had only the two basic genomes J and Js. The Js genome present in Th. intermedium and Th. ponticum was homologous with E or J genomes, but was quite distinct at centromeric regions, which can strongly hybridize with the S genome DNA probe. Based on GISH results, the genomic formula of Th. intermedium was redesignated JJsS and that of Th. ponticum was redesignated JJJJsJs. The finding of a close relationship among S, J, and Js genomes provides valuable markers for molecular cytogenetic analyses using S genome DNA probes to monitor the transfer of useful traits from Th. intermedium and Th. ponticum to wheat.

Centromere↗

Genotypic analysis at multiple loci across Kaposi's sarcoma herpesvirus (KSHV) DNA molecules: clustering patterns, novel variants and chimerism.

BACKGROUND: The genomes of human Kaposi's sarcoma-associated herpesvirus (KSHV) display several levels of DNA sequence heterogeneity and subgrouping that show distinctive clustering patterns in related human populations. The four major subtype patterns for the hypervariable ORF-K1 protein correlate closely with the principal diasporas resulting from the migration of modern humans out of East Africa and suggest that KSHV is an ancient human virus that is transmitted primarily in a familial fashion with consequent very low recombination rates. However, chimeric genomes have also been detected, especially with regard to the presence of P versus M alleles of the ORF-K15 gene. OBJECTIVES: To understand further the genetic organization and evolutionary history of KSHV, especially with regard to possible new subtypes, recombinant genomes, constant region loci and clustering in particular ethnic groups or among classic versus epidemic cases in the same geographic area. STUDY DESIGN: Direct PCR DNA sequencing was carried out on the ORF-K1 and ORF-K15 genes at the extreme left and right hand sides, as well as on six other internal loci of diagnostic samples collected from 70 new KSHV-positive patients in Israel, South Korea, Sicily, Scandinavia, Brazil, Uganda, South Africa and the US. RESULTS AND CONCLUSIONS: Our overall results from more than 135 KSHV genomes from many different human population groups now provides evidence for seven distinct subtypes of KSHV genomes (referred to as A/P, B/P, C/P, D/P, M, N and Q). However, the two most closely related subtypes (A/P and C/P) are only differentiated at the LHS side of the genome, and the three most distantly related forms (M, N and Q) appear to exist only as small chimeric segments that are remnants from the RHS of more ancient forms of the virus. By analyzing multiple conserved loci across the B subtype genomes that predominate in sub-Saharan Africa, we can also now recognize three to four distinct B genome subgroups with varying patterns of inter and intratypic mosaicism. Analysis of classic KS genomes from Israel has revealed that the ORF-K1 clade referred to as A1' predominates in Ashkenazi Jewish immigrants from Russia, whereas C2 and C6 variants predominate in North African Sephardi Jews. A variety of chimeric genomes containing C2 or C3 ORF-K1 genes are disseminated among classic KS cases throughout Europe and Asia including Israel, Sicily, Scandinavia, South Korea, and Taiwan. Comparison of the genomes from classic versus AIDS-associated KSHV in the US indicates that it was derived originally by reactivation and spread of a subset of the endogenous viruses carried by descendants of immigrants from endemic areas of Northern and Eastern Europe, the Mediterranean and sub-Saharan Africa.

Acquired Immunodeficiency Syndrome↗

Whole-genome re-sequencing.

DNA sequencing can be used to gain important information on genes, genetic variation and gene function for biological and medical studies. The growing collection of publicly available reference genome sequences will underpin a new era of whole genome re-sequencing, but sequencing costs need to fall and throughput needs to rise by several orders of magnitude. Novel technologies are being developed to meet this need by generating massive amounts of sequence that can be aligned to the reference sequence. The challenge is to maintain the high standards of accuracy and completeness that are hallmarks of the previous genome projects. One or more new sequencing technologies are expected to become the mainstay of future research, and to make DNA sequencing centre stage as a routine tool in genetic research in the coming years.

Genome, Human↗

Metatranscriptomic analysis of viral sequences associated with Culex nigripalpus at an Alabama aquaculture site.

Mosquitoes associated with aquaculture habitats can harbor diverse viruses, yet the viromes of many locally abundant species remain poorly characterized. At an aquaculture-associated site in Auburn, Alabama, we surveyed mosquito populations and found Culex nigripalpus to be the dominant species collected. To characterize viruses associated with this mosquito, we performed RNA-seq on pooled female Cx. nigripalpus and compared complementary bioinformatic workflows for viral detection and genome recovery. One workflow removed host-associated reads by mapping to the closest available mosquito reference genome prior to assembly, whereas a second workflow used fully de novo assembly and viral database annotation. Additional protein-level filtering, cross-workflow comparison, and comparison of Trinity and rnaSPAdes assemblies were used to prioritize well-supported viral candidates. Across the original analyses, 16 submitted accessions corresponding to 12 collapsed virus/name groups were recovered, including Merida virus, Hubei mosquito virus 5, Zhejiang mosquito virus, Hubei virga-like virus 3, Rinkaby virus, Elemess virus, Qingnian mosquito virus, Serbia narna-like virus 2, XiangYun narna-levi-like virus 8, Ecclesville picorna-like virus, and baculovirus-like fragments. Several candidates were supported across multiple workflows, while others were recovered only under specific analytical conditions, indicating that candidate recovery was influenced by assembly and filtering choices. Selected viral contigs were independently supported by RT-PCR amplification. Overall, these results provide a first characterization of viral sequences associated with Cx. nigripalpus from an Alabama aquaculture-associated site and show that comparison across assembly and filtering strategies helped prioritize the most consistently supported viral candidates.

Animals↗

Chromosome-scale genomes and population resequencing resolve subgenome diversity and halophyte adaptation in Salicornia.

Amid escalating water scarcity and groundwater depletion, halophytes such as Salicornia (Amaranthaceae) represent valuable models for extreme salt tolerance and hold promise for saltwater-based agriculture. Here, we show chromosome-scale genome assemblies for six Salicornia species, revealing four distinct subgenomes, reconciling our assemblies with two existing reference genomes (S. ramosissima UK and S. europaea China), correcting chromosome numbering and orientation. Comparative analyses across ploidy levels demonstrate genome expansion in North American lineages driven by Gypsy retrotransposons, and lineage-specific expansions of two gene families implicated in stress metabolism. Phylogenetic and population-structure analyses of a global resequencing panel of 318 accessions resolve interspecific relationships and establish curated germplasm collections for future crop breeding. Genetic analyses uncover a contrasting population-genetic signal on chromosome 6A between two species, highlighting an OSCA calcium-permeable channel gene as a candidate locus for osmotic adaptation. Together, these resources establish a genomic framework for Salicornia that supports evolutionary studies of halophyte adaptation and crop development.

Chenopodiaceae↗

Long-read low-pass sequencing enhances variant detection in a peanut MAGIC population.

Accurate genotyping accelerates crop improvement, yet long-read sequencing remains underused in breeding due to cost. We present a scalable long-read low-pass (LRLP) sequencing framework for high-throughput variant discovery and trait mapping. Using PacBio HiFi reads in an allotetraploid peanut (Arachis hypogaea; AABB, 2n = 4x = 40) MAGIC population, we generated both LRLP and short-read low-pass (SRLP) data. At comparable depths, LRLP achieved substantially greater whole-genome and gene-space coverage than SRLP. Data were analyzed using both a single-reference genome and an 18-parent pangenome graph constructed with KhufuPan, a new tool for graph-based genotyping. Across analytical approaches, LRLP consistently identified more SNPs, indels (2-1,000 bp), and structural variants (>1 kb) than SRLP, improving genotype resolution and selection accuracy, particularly for large structural variants. By reducing cost barriers and increasing variant discovery in complex genomes, LRLP provides a practical path for deploying advanced genomics in under-resourced and orphan crops critical to global food security.

Arachis↗