Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Nucleic acid hybridization for detection of cell culture-amplified adenovirus.

A number of recombinant plasmids containing genomic segments of adenovirus were constructed. Seven cloned probes, as well as total adenovirus type 2 (Ad2) and Ad16 genomic DNA, were tested by a nucleic acid hybridization technique for sensitivity and specificity in detecting adenoviruses in infected cells. Adenovirus DNA was spotted onto a nitrocellulose filter and hybridized with 32P-labeled DNA probes. The probes, total Ad2 genomic DNA, and plasmid pAd2-H (containing the hexon gene from Ad2 DNA) all detected 10 reference serotypes of five genomic subgroups (A through E) with similar sensitivities. However, plasmid pAd2-H required less preparation time than did total Ad2 DNA. Probes pAd2-F (containing the fiber gene from Ad2) and pAd16-BD (containing the BamHI D fragment from Ad16) hybridized only with reference serotypes from the homologous subgroups (C and B, respectively). Of 101 patient isolates amplified in cells, pAd2-H detected 100% of all isolates from both the homologous and the heterologous subgroups. The detection rates for pAd2-F were 100% (subgroup C) and 3.6% (subgroups A, B, and D), and those for pAd16-BD were 100% (subgroup B) and 9.4% (subgroups A, C, and D). A commercial biotinylated product (Pathogene II) was also included in this study for comparison.

Adenoviruses, Human↗

The Human Genome Project--an overview.

The human genome sequence will underpin human biology and medicine in the next century, providing a single, essential reference to all genetic information. The international program to determine the complete DNA sequence (3,000 million bases) is well underway. As of January 2000, 50% of the sequence is available in the public domain. A comprehensive working draft is expected this year, and the entire sequence is projected to be finished in 2003. DNA sequencing is carried out on mapped, overlapping bacterial clones of 150-200 kb. The working draft comprises assembled unfinished sequence and is released immediately in the public domain. The draft sequence of each clone is then completed, by closing any remaining gaps and resolving any ambiguities, before the entire sequence is checked, annotated, and submitted to the public databases. The sequence of each clone is finished to an accuracy of >99.99%. The availability of a reference sequence of the genome provides the basis for studying the nature of sequence variation, particularly single nucleotide polymorphisms (SNPs), in human populations. SNP typing is a powerful tool for genetic analysis, and will enable us to uncover the association of loci at specific sites in the genome with many disease traits. SNPs occur at a frequency of approximately 1 SNP/kb throughout the genome when the sequence of any two individuals is compared. Programs to detect and map SNPs in the human genome are underway with the aim of establishing a SNP map of the genome during the next two years. The human genome sequence will provide a complete description of all the genes. Annotation of the sequence with the gene structures is achieved by a combination of computational analysis (predictive and homology-based) and experimental confirmation by cDNA sequencing. Detecting homologies between newly defined gene products and proteins of known function helps to postulate biochemical functions for them, which can then be tested. Establishing the association of specific genes with disease phenotypes by mutation screening, particularly for monogenic disorders, provides further assistance in defining the functions of some gene products, as well as helping to establish the cause of the disease. As our knowledge of gene sequences and sequence variation in populations increases, we will pinpoint more and more of the genes and proteins that are important in common, complex diseases. A more detailed understanding of the function of the human genome will be achieved as we identify sequences that control gene expression. Given the availability of gene sequences, the expression status of genes in particular tissues can be monitored in parallel. By comparing corresponding genomic sequences in different species (for example: man, mouse, chicken, and zebrafish), regions that have been highly conserved during evolution can be identified, many of which reflect conserved functions such as gene regulation. These approaches promise to greatly accelerate our interpretation of the human genome sequence.

Human Genome Project↗

Evolution of hominoid mitochondrial DNA with special reference to the silent substitution rate over the genome.

Focusing on the synonymous substitution rate, we carried out detailed sequence analyses of hominoid mitochondrial (mt) DNAs of ca. 5-kb length. Owing to the outnumbered transitions and strong biases in the base compositions, synonymous substitutions in mtDNA reach rapidly a rather low saturation level. The extent of the compositional biases differs from gene to gene. Such changes in base compositions, even if small, can bring about considerable variation in observed synonymous differences and may result in the region-dependent estimate of the synonymous substitution rate. We demonstrate that such a region dependency is due to a failure to take proper account of heterogeneous compositional biases from gene to gene but that the actual synonymous substitution rate is rather uniform. The synonymous substitution rate thus estimated is 2.37 +/- 0.11 x 10(-8) per site per year and comparable to the overall rate for the noncoding region. On the other hand, the rate of nonsynonymous substitutions differs considerably from gene to gene, as expected under the neutral theory of molecular evolution. The lowest rate is 0.8 x 10(-9) per site per year for COI and the highest rate is 4.5 x 10(-9) for ATPase 8, the degree of functional constraints (measured by the ratio of the nonsynonymous to the synonymous substitution rate) being 0.03 and 0.19, respectively. Transfer RNA (tRNA) genes also show variability in the base contents and thus in the nucleotide differences. The average rate for 11 tRNAs contained in the 5-kb region is 3.9 x 10(-9) per site per year. The nucleotide substitutions in the genome suggest that the transition rate is about 17 times faster than the transversion rate.

Adenosine Triphosphatases↗

The sequence of rice chromosomes 11 and 12, rich in disease resistance genes and recent gene duplications.

BACKGROUND: Rice is an important staple food and, with the smallest cereal genome, serves as a reference species for studies on the evolution of cereals and other grasses. Therefore, decoding its entire genome will be a prerequisite for applied and basic research on this species and all other cereals. RESULTS: We have determined and analyzed the complete sequences of two of its chromosomes, 11 and 12, which total 55.9 Mb (14.3% of the entire genome length), based on a set of overlapping clones. A total of 5,993 non-transposable element related genes are present on these chromosomes. Among them are 289 disease resistance-like and 28 defense-response genes, a higher proportion of these categories than on any other rice chromosome. A three-Mb segment on both chromosomes resulted from a duplication 7.7 million years ago (mya), the most recent large-scale duplication in the rice genome. Paralogous gene copies within this segmental duplication can be aligned with genomic assemblies from sorghum and maize. Although these gene copies are preserved on both chromosomes, their expression patterns have diverged. When the gene order of rice chromosomes 11 and 12 was compared to wheat gene loci, significant synteny between these orthologous regions was detected, illustrating the presence of conserved genes alternating with recently evolved genes. CONCLUSION: Because the resistance and defense response genes, enriched on these chromosomes relative to the whole genome, also occur in clusters, they provide a preferred target for breeding durable disease resistance in rice and the isolation of their allelic variants. The recent duplication of a large chromosomal segment coupled with the high density of disease resistance gene clusters makes this the most recently evolved part of the rice genome. Based on syntenic alignments of these chromosomes, rice chromosome 11 and 12 do not appear to have resulted from a single whole-genome duplication event as previously suggested.

Chromosome Mapping↗

Pairwise graph edit distance characterizes the impact of the construction method on pangenome graphs.

MOTIVATION: Pangenome variation graphs are an increasingly used tool to perform genome analysis, aiming to replace a linear reference in a wide variety of genomic analyses. The construction of a variation graph from a collection of chromosome-size genome sequences is a difficult task that is generally addressed using a number of heuristics. The question that arises is to what extent the construction method influences the resulting graph, and the characterization of variability. RESULTS: We aim to characterize the differences between variation graphs derived from the same set of genomes with a metric which expresses and pinpoint differences. We designed a pairwise variation graph comparison algorithm, which establishes an edit distance between variation graphs, threading the genomes through both graphs. We applied our method to pangenome graphs built from yeast and human chromosome collections, and demonstrate that our method effectively characterizes discordances between pangenome graph construction methods and scales to real datasets. AVAILABILITY AND IMPLEMENTATION: pancat compare is published as free Rust software under the AGPL3.0 open source license. Source code and documentation are available at https://github.com/dubssieg/rs-pancat-compare. Snapshot available on Software Heritage at swh:1:dir:61acda8ba3dac1709ed60530147d3871831be629.

Algorithms↗

A unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification.

The rapid growth of genomic sequencing demands fast, accurate, and scalable analysis methods. In viral genomic classification, expanding labeled reference collections can make supervised models costly to update and dependent on fixed label sets, motivating retrieval-based genomic classification as a simpler, more flexible alternative. We present a unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification across three viral classification tasks: hepatitis C virus (HCV) genotyping, COVID-19 discrimination, and human papillomavirus (HPV) genotyping. We compare standard sequence encodings (one-hot, k-mers, FCGR) with dense embeddings (dna2vec, DNABERT). For each representation, we evaluate supervised classifiers (Random Forest, Decision Tree, XGBoost) and retrieval-based classification, where sequence vectors are indexed with FAISS and labels are assigned via similarity-weighted k-NN. Furthermore, we benchmark multiple FAISS index types (Flat, IVF, HNSW, IVFPQ, OPQ) to characterize accuracy-speed-memory trade-offs at scale. The results show that XGBoost and retrieval using Flat or IVF indexes achieve strong classification performance under different computational profiles. Compressed indexes such as IVFPQ and OPQ substantially reduce memory usage, although their accuracy loss depends on the dataset and representation. Overall, supervised XGBoost provides a favorable accuracy-size trade-off, while retrieval-based classification remains competitive and allows labeled reference sequences to be incorporated without retraining a global classifier. This benchmark provides practical guidance for selecting sequence representations, classifiers, and vector-search indexes under different accuracy, memory, and update requirements.

Genome, Viral↗

Expressive genomic hybridisation: gene expression profiling at the cytogenetic level.

AIMS: To describe a cytogenetic technique suitable for the rapid assessment of global gene expression that is based on comparative genomic hybridisation (CGH), and to use it to understand the relation between genetic amplifications and gene expression. METHODS: Whereas traditional CGH uses DNA as test and reference in hybridisations, expressive genomic hybridisation (EGH) uses globally amplified mRNA as test and normal DNA as reference. EGH is a rapid and powerful tool for localising and studying global gene expression profiles and correlating them with loci of genetic amplifications using traditional CGH. RESULTS: EGH was used to correlate genetic amplifications detected by CGH with the expression profile of two independent cell lines-Colo320 and T47D. Although many amplifications resulted in overexpression, other amplifications were partially or completely silenced at the cytogenetic level. CONCLUSION: This technique will assist in the analysis of overexpressed genes within amplicons and could resolve a controversial issue in cancer cytogenetics; namely, the relation between genetic amplifications and overexpression.

Cell Line↗

Validation of mixed-genome microarrays as a method for genetic discrimination.

Comparative genomic hybridizations have been used to examine genetic relationships among bacteria. The microarrays used in these experiments may have open reading frames from one or more reference strains (whole-genome microarrays), or they may be composed of random DNA fragments from a large number of strains (mixed-genome microarrays [MGMs]). In this work both experimental and virtual arrays are analyzed to assess the validity of genetic inferences from these experiments with a focus on MGMs. Empirical data are analyzed from an Enterococcus MGM, while a virtual MGM is constructed in silico using sequenced genomes (Streptococcus). On average, a small MGM is capable of correctly deriving phylogenetic relationships between seven species of Enterococcus with accuracies of 100% (n=100 probes) and 95% (n=46 probes); more probes are required for intraspecific differentiation. Compared to multilocus sequence methods and whole-genome microarrays, MGMs provide additional discrimination between closely related strains and offer the possibility of identifying unique strain or lineage markers. Representational bias can have mixed effects. Microarrays composed of probes from a single genome can be used to derive phylogenetic relationships, although branch length can be exaggerated for the reference strain. We describe a case where disproportional representation of different strains used to construct an MGM can result in inaccurate phylogenetic inferences, and we illustrate an algorithm that is capable of correcting this type of bias. The bias correction algorithm automatically provides bootstrap confidence values and can provide multiple bias-corrected trees with high confidence values.

Algorithms↗

[Construction of standard human transcript dataset based on RefSeq and human genome sequence database].

The NCBI Reference Sequence (RefSeq) database aimed to provide a biologically non-redundant collection of DNA, RNA, and protein sequences and to promote the research on genes and proteins of human beings and other species. However, because of widely distributed polymorphisms and different quality control of experiments in individual laboratories, there are potential problems need to be identified in the RefSeq database. Regarding which, we herein define the concept, standard transcript, based on the Central Dogmas of Biology that each standard transcript should be perfectly mapped to the standard genomic DNA sequence at the exon level. A large scale analysis for mapping all of the RefSeq records of human being (2005-4-18) to the officially released human genome sequence database (2005-4-20) was further performed using BLAT, Sim4 and a homemade program, EIparser, which was especially designed for this purpose. The standard transcripts based on the RefSeq database were obtained according to the alignment with standard human genome database. There are 9,771 RefSeq records of human being labeled with "NM_" and "NR_" could be perfectly mapped to human genome sequences, while other 10,943 records could be considered as standard transcripts after reasonable revision by comparing with the genome sequences according to all of the three methods. Moreover, the left 203 unrevisable records and 2,676 inconsistent records reported by the above programs could not be considered as standard transcripts and should be checked critically before using because of potential errors in them. Our study has thus provided a reference standard dataset of human beings with high quality for further bioinformatic and experimental analysis such as polymorphism and mutation of human genes. The reference standard dataset based on above criteria could be retrieved from http://biocompute.bmi.ac.cn/transcriptome/index.htm.

Databases, Genetic↗

A robust approach for the quantitation of viral concentration in an adenoviral vector-based human immunodeficiency virus vaccine by real-time quantitative polymerase chain reaction.

A real-time quantitative polymerase chain reaction (PCR)-based method was developed to measure the concentration of recombinant adenoviral vector genomes in purified virus bulks and final container samples of monovalent and multivalent human immunodeficiency virus (HIV) adenoviral vector vaccine candidates. This method, referred to as the genome quantitation assay (GQA), was optimized through a rigorous approach for evaluating PCR detection chemistries, designing a robust assay format, and establishing a properly calibrated reference standard. In addition, the use of a simplified lysis procedure, automated liquid transfer system, and parallel-line data analysis contribute to an accurate, precise, reliable, and high-throughput assay procedure that can be used for process monitoring, final formulation, and release of vaccine products. A variance component analysis study indicated that the GQA typically produces results with an interassay precision of less than 10% relative standard deviation (RSD), allowing generation of final results (average of three runs) with associated interassay precision of 6% RSD or less. The precision, accuracy, specificity, and robustness of the GQA demonstrate its utility for analytical characterization of a wide variety of viral vector- and DNA plasmid- based vaccines or gene therapy products. In addition, we also evaluated the Adenovirus Reference Standard generated by the Adenovirus Reference Material Working Group in the GQA to provide a common point-of-reference for our analytical method.

AIDS Vaccines↗

Reverse sample genome probing, a new technique for identification of bacteria in environmental samples by DNA hybridization, and its application to the identification of sulfate-reducing bacteria in oil field samples.

A novel method for the identification of bacteria in environmental samples by DNA hybridization is presented. It is based on the fact that, even within a genus, the genomes of different bacteria may have little overall sequence homology. This allows the use of the labeled genomic DNA of a given bacterium (referred to as a "standard") to probe for its presence and that of bacteria with highly homologous genomes in total DNA obtained from an environmental sample. Alternatively, total DNA extracted from the sample can be labeled and used to probe filters on which denatured chromosomal DNA from relevant bacterial standards has been spotted. The latter technique is referred to as reverse sample genome probing, since it is the reverse of the usual practice of deriving probes from reference bacteria for analyzing a DNA sample. Reverse sample genome probing allows identification of bacteria in a sample in a single step once a master filter with suitable standards has been developed. Application of reverse sample genome probing to the identification of sulfate-reducing bacteria in 31 samples obtained primarily from oil fields in the province of Alberta has indicated that there are at least 20 genotypically different sulfate-reducing bacteria in these samples.

Journal Article↗

An integrated in-silico approach for drug target identification in human pathogen Shigella dysenteriae.

Shigella dysenteriae, is a Gram-negative bacterium that emerged as the second most significant cause of bacillary dysentery. Antibiotic treatment is vital in lowering Shigella infection rates, yet the growing global resistance to broad-spectrum antibiotics poses a significant challenge. The persistent multidrug resistance of S. dysenteriae complicates its management and control. Hence, there is an urgent requirement to discover novel therapeutic targets and potent medications to prevent and treat this disease. Therefore, the integration of bioinformatics methods such as subtractive and comparative analysis provides a pathway to compute the pan-genome of S. dysenteriae. In our study, we analysed a dataset comprising 27 whole genomes. The S. dysenteriae strain SD197 was used as the reference for determining the core genome. Initially, our focus was directed towards the identification of the proteome of the core genome. Moreover, several filters were applied to the core genome, including assessments for non-host homology, protein essentiality, and virulence, in order to prioritize potential drug targets. Among these targets were Integration host factor subunit alpha and Tyrosine recombinase XerC. Furthermore, four drug-like compounds showing potential inhibitory effects against both target proteins were identified. Subsequently, molecular docking analysis was conducted involving these targets and the compounds. This initial study provides the list of novel targets against S. dysenteriae. Conclusively, future in vitro investigations could validate our in-silico findings and uncover potential therapeutic drugs for combating bacillary dysentery infection.

Shigella dysenteriae↗

A vision of how low-coverage sequence data should contribute to genetic evaluation in the future.

Low-coverage sequencing refers to sequencing DNA of individuals to a low depth of coverage (e.g., 0.5X) and imputing that sequence to a genomic sequence based on reference haplotypes from individuals sequenced to a high depth of coverage (e.g., ≥10X). It has been proposed as an alternative to genotyping by Single-nucleotide polymorphisms (SNP) arrays. At least one commercial product based on it is available for agricultural species. Concerns limiting adoption in its current form are: 1) the cost of storing the huge volume of data it generates and 2) whether that additional data will result in improved accuracy of genetic evaluation. This work envisions future implementation of low-coverage sequencing to reduce storage costs and enhance genetic evaluations by leveraging the additional information in the full sequence of the pangenome to account for more genetic variation. We propose addressing the storage issue by representing genomic sequence of an individual in a pair of haplotype arrays with each element pointing to an enumerated haplotype of the sequence within one of approximately 50,000 defined genome segments. Assuming 60 million genomic variants, the infrastructure required to translate the identifier of any enumerated haplotype into its genomic sequence would require less than 10 gigabytes of binary storage. Each haplotype array element would require 2 bytes, so the marginal binary storage required to represent the genomic sequence of an individual would be about 200 kilobytes (KB), similar to the genotypes from a SNP array with 200,000 markers. This assumes no pedigree and no ambiguity of the imputation, though the latter is unrealistic. Strategies to minimize, and when necessary, to manage and efficiently represent ambiguity are proposed. The genomic sequence of an individual could be stored in about 1 KB (binary) if both parents have unambiguous sequences stored as described above. The proposed system for representing the pangenome includes algorithms for read mapping and imputation intended to leverage all known genetic variation in the target population. It is also designed to use sequencing reads generated for imputing the genomic sequence of new individuals to identify unrecognized mutations, crossovers, and structural variants, thus continuously improving the genome representation, especially if widespread use of low-coverage sequencing in livestock industries is realized. This could make improved genetic merit and management of livestock feasible without computational burden.

Animals↗

Vertebrate genome sequencing: building a backbone for comparative genomics.

The human genome sequence provides a reference point from which we can compare ourselves with other organisms. Interspecies comparison is a powerful tool for inferring function from genomic sequence and could ultimately lead to the discovery of what makes humans unique. To date, most comparative sequencing has focused on pair-wise comparisons between human and a limited number of other vertebrates, such as mouse. Targeted approaches now exist for mapping and sequencing vertebrate bacterial artificial chromosomes (BACs) from numerous species, allowing rapid and detailed molecular and phylogenetic investigation of multi-megabase loci. Such targeted sequencing is complementary to current whole-genome sequencing projects, and would benefit greatly from the creation of BAC libraries from a diverse range of vertebrates.

Animals↗

SilkDB: a knowledgebase for silkworm biology and genomics.

The Silkworm Knowledgebase (SilkDB) is a web-based repository for the curation, integration and study of silkworm genetic and genomic data. With the recent accomplishment of a approximately 6X draft genome sequence of the domestic silkworm (Bombyx mori), SilkDB provides an integrated representation of the large-scale, genome-wide sequence assembly, cDNAs, clusters of expressed sequence tags (ESTs), transposable elements (TEs), mutants, single nucleotide polymorphisms (SNPs) and functional annotations of genes with assignments to InterPro domains and Gene Ontology (GO) terms. SilkDB also hosts a set of ESTs from Bombyx mandarina, a wild progenitor of B.mori, and a collection of genes from other Lepidoptera. Comparative analysis results between the domestic and wild silkworm, between B.mori and other Lepidoptera, and between B.mori and the two sequenced insects, fruitfly and mosquito, are displayed by using B.mori genome sequence as a reference framework. Designed as a basic platform, SilkDB strives to provide a comprehensive knowledgebase about the silkworm and present the silkworm genome and related information in systematic and graphical ways for the convenience of in-depth comparative studies. SilkDB is publicly accessible at http://silkworm.genomics.org.cn.

Animals↗

Genomic Analysis and Clinical Correlation of Non-Small Cell Lung Cancer with Special Reference to Brain Metastasis.

BACKGROUND: Next-generation sequencing (NGS) has improved genomic analysis depth in precision oncology. This study analyzed genomic biomarker testing in stage IV NSCLC, focusing on brain metastasis and clinicopathological correlations. OBJECTIVE: To study molecular markers and clinicopathological correlations in stage IV NSCLC patients, with and without brain metastasis. METHODS: A total of 169 stage IV NSCLC patients were studied from April 2023 to May 2025. Demographic data, clinical presentations, and mutation analyses were assessed using NGS on tissue blocks or liquid biopsies. RESULTS: Among 169 patients, 41.42% (n = 70) had brain metastasis (NSCLC-BM), while 58.58% (n = 99) had no brain metastasis (mNSCLC). Median ages were 51.5 and 56 years, respectively. Adenocarcinoma comprised 95.27% (n = 161) of cases. The cerebral hemisphere was the most common intracranial metastatic site, while skeletal involvement was the most common extracranial site. Headache was the predominant neurological symptom. EGFR mutations were the most common overall. EGFR > TP53 > ALK > other mutations were observed in NSCLC-BM, while EGFR > TP53 > KRAS > other mutations were seen in mNSCLC. Mutation analysis stratified by smoking history (χ²(1) = 1.347, p = 0.245) and sex (χ²(1) = 0.0302, p = 0.862) was not statistically significant. The benefit of gefitinib plus chemotherapy in EGFR exon 19 and exon 21 L858R mutations was greater in mNSCLC (log-rank χ²(1) = 10.813, p = 0.001) than in NSCLC-BM (log-rank χ²(1) = 3.100, p = 0.078). Median survival was 11 months (95% CI: 7.506-14.494) for NSCLC-BM versus 21 months (95% CI: 8.365-33.635) for mNSCLC, with a statistically significant difference (log-rank χ²(1) = 8.639, p = 0.003). CONCLUSION: NSCLC-BM showed higher genomic biomarker enrichment (80% vs. 68.68%) but poorer outcomes than mNSCLC. EGFR was the most common targetable mutation, followed by ALK in NSCLC-BM and KRAS in mNSCLC.

Humans↗

A chromosome-scale assembly for the genome of southern corn rootworm, Diabrotica undecimpunctata.

Diabrotica undecimpunctata ssp. howardi, the southern corn rootworm or eastern 12-spotted cucumber beetle, is a generalist insect herbivore that causes damage and yield loss to several crops in North America including maize. Unresolved phylogenetic relationships within and among D. undecimpunctata subspecies are impacting current quarantine policies. We report the chromosome-level haploid genome assembly, icDiaUnde3, constructed using HiFi and Hi-C read data from a single male D. undecimpunctata collected and identified as subspecies howardi based on geographic location and morphology. The primary 1.74 Gbp assembly is scaffolded into 11 chromosome-length scaffolds representing 9 autosomes, a single X chromosome and a supernumerary (B) chromosome (scaffold N50 = 162.8 Mb and L50 = 5). Ab initio and evidence-based structural reference sequence (RefSeq) annotations predicted 18,959 protein-coding genes, in which 99.2% of the 1,367 Benchmark Universal Single-Copy Orthologs from Insecta were complete. Repeat elements occupy 1.26 Gbp (72.33%) of the icDiaUnde3 assembly, with nearly 36% predicted to be retroelements. Alignment of whole chromosomes from icDiaUnde3 with those previously assembled from Diabrotica spp. predicted 2 and 6 autosomal inversions with D. balteata and D. virgifera virgifera, respectively. The mitochondrial genome had an annotated gene order and orientation conserved among beetles. The icDiaUnde3 reference genome assembly is a vital resource for taxonomic, comparative, and functional studies to enhance sustainable crop production.

agriculture↗

Revisiting the genome assembly of Lupinus species reveals differential diploidization after a shared whole-genome duplication.

Accurate genome assemblies are essential for comparative genomics, yet Hi-C-guided scaffolding can introduce structural errors that misrepresent chromosome architecture and bias evolutionary inferences. Here, we identified pervasive scaffolding errors-including artificial fusions, internal inversions, and incomplete contig mounting-in 2 previously published Lupinus genomes (L. cosentinii and L. digitatus) using a segmentation method based on long terminal repeat (LTR) retrotransposon density. We reassembled both genomes, producing chromosome-level references of 472.7 Mb (16 chromosomes) and 427.2 Mb (21 chromosomes), with BUSCO completeness >98.5%. Synteny validation and reapplication of LTR profiling confirmed that all prior errors were resolved. Using these corrected genomes together with 4 additional Lupinus species and 2 outgroup legumes, we investigated postpolyploid evolution. Synonymous substitution rate (Ks) analysis revealed a genus-specific whole-genome duplication (WGD) event (Ks = 0.17) shared by all 6 Lupinus species. The proportion of WGD-derived genes varied markedly, from 60% in L. digitatus to only 36% in L. mutabilis, indicating differential diploidization. While all species retained a core set of WGD duplicates enriched in cytoskeleton organization, ion transport, and defense responses, each exhibited lineage-specific functional trajectories: cell wall modification in L. cosentinii and L. digitatus, nitrogen metabolism in L. albus and L. angustifolius, flower development in L. luteus, and stress/lipid metabolism in L. mutabilis. Our corrected assemblies provide optimal references for Lupinus comparative genomics, and our findings demonstrate that a shared WGD event can lead to both conserved and highly divergent postpolyploid fates, likely underpinning adaptive diversification within the genus.

Lupinus↗