Search PubMedSearch

SEARCH · Search PubMed

Results for “genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Longitudinal whole-genome analysis of bluetongue virus identifies conserved serotype-specific genomes and distinct genomic constellations within a Colorado sheep flock (2021-2023).

Bluetongue virus (BTV) is a segmented double-stranded RNA virus of ruminants transmitted by Culicoides spp. biting midges. Although the genome consists of ten segments, classification into serotypes is primarily based on genome segment 2. However, reassortment among genomic segments is a major driver of BTV evolution and diversity. This study used longitudinal whole-genome sequencing to characterize BTV genomes collected from 2021 to 2023 within a single sheep flock in Colorado, where multiple serotypes co-circulate. Whole-genome sequences were generated from fourteen blood samples representing four serotypes: BTV-6, -11, -13, and -17. Longitudinal sampling identified multiple BTV serotypes within individual sheep across consecutive years. Tanglegram analysis comparing segment phylogenies to the segment 2 tree demonstrated incongruent topologies across all genomic segments, suggestive of reassortment or the circulation of distinct genomic constellations. Nucleotide-level comparisons revealed high sequence homology among same-serotype samples from the same year, while the greatest genetic divergence was observed among BTV-17 genomes collected in different years. Additionally, all BTV-13 genomes contained a previously undescribed nonsynonymous substitution in segment 10 predicted to extend the encoded protein by three amino acids. Together, these findings demonstrate that highly conserved BTV genomes and distinct genomic constellations can be detected at the flock level across multiple years. This longitudinal whole-genome approach reveals the genetic complexity of endemic BTV populations, including novel variants and genomic patterns consistent with reassortment that are lost with conventional serotyped-based approaches, highlighting the need to integrate whole-genome characterization into endemic BTV monitoring programs.

Animals

Chromosome-scale genome assembly and genomic prediction of essential oil compounds in Atractylodes lancea for genomics-assisted breeding.

Atractylodes lancea rhizomes are used as crude drugs. Essential oil compounds, including atractylodin, hinesol, β-eudesmol, and atractylon, are key determinants of crude drug quality. Conventional breeding of A. lancea is difficult because of its perennial growth. In this study, a chromosome-scale reference genome of A. lancea (4.79 Gb) was generated, and genome-wide association studies (GWAS) and genomic predictions of essential oil compounds were conducted to explore the potential for genome-assisted breeding. Genotyping of 480 lines using double-digest restriction-site-associated DNA-sequencing yielded 29,136 high-quality SNPs. All the compounds showed high genomic heritability (h2 = 0.758-0.915), indicating strong genetic control. Despite the high genomic heritability, GWAS detected only one weak association with atractylon and no significant loci for the three compounds. However, genomic prediction achieved moderate to high accuracy across multiple models, particularly the ridge regression, genomic best linear unbiased prediction, and Bayesian approaches. The prediction accuracy, measured as the Pearson correlation coefficient between the observed and predicted values, exceeded 0.6 for all four essential oil compounds. These results demonstrate the efficacy of genomic selection for improving essential oil compound levels in A. lancea and provide a foundation for genome-assisted breeding of medicinal plants with long breeding cycles.

Atractylodes lancea

The SARS-CoV-2 Integrated Genomic Epidemiology Database (IGED): Linking viral genomes with patient-level metadata to advance statewide genomic surveillance in California.

In July 2021, the California Code of Regulations Title 17 required all laboratories performing SARS‑CoV‑2 whole genome sequencing (WGS) to report their sequencing results to the California Department of Public Health (CDPH). These viral genomic data and patient metadata were compiled into the Integrated Genomic Epidemiology Database (IGED). Linking anonymized viral sequences with patient‑level information enabled monitoring of infectiousness, pathogenicity, transmission dynamics, evolution, and vaccine evasion among emerging SARS‑CoV‑2 lineages. Laboratories performing SARS-CoV-2 WGS transmitted sequencing results to CDPH through Electronic Laboratory Reporting (ELR) and non-ELR pathways. CDPH applied uniform reporting requirements but allowed flexibility in specific data formats to accommodate diverse data systems. To preserve data quality and interoperability across heterogeneous sources, CDPH implemented standardization, validation, and deduplication protocols. Snowflake, a cloud‑based data storage and analytics platform, and Posit Connect, a cloud deployment and automation platform, supported the management, processing, and integration of data within the IGED. The IGED established links between SARS‑CoV‑2 WGS data and epidemiologic metadata for 801,418 sequences, representing 81.7% of all sequences reported in California. Lineages reported to the IGED showed strong concordance with lineage proportions in GISAID. Sequences reported to the IGED had average turnaround times longer than one month, and the majority of sequencing was performed in Southern California and Los Angeles. The IGED enhanced genomic surveillance through predictive modeling and monitoring concerning evolutionary trends such as recombination and saltations in persistent infections. Development of the IGED highlighted the need for standardized data requirements, sustained funding for sequencing, incentives for data submission, and interdisciplinary collaboration to build an effective genomic surveillance system. This framework for linking genomic and epidemiologic data has not only generated critical insights for SARS‑CoV‑2 but also provided the foundation for CDPH and other public health organizations to develop similar IGED‑like systems for other priority pathogens as genomic surveillance expands.

Journal Article

Genome Report: De novo genome assembly of the greater Bermuda land snail, Poecilozonites bermudensis (Mollusca: Gastropoda), confirms ancestral genome duplication.

Poecilozonites bermudensis, the greater Bermuda land snail, is a critically endangered species and one of only two extant members in its genus. These snails are one of Bermuda's few endemic animal clades and their rich fossil record was the basis for the punctuated equilibria model of speciation. Once thought extinct, recent conservation efforts have focused on the recovery of the species, yet no genomic information or other molecular sequences have been available to inform these initiatives. We present a high-quality, annotated genome for P. bermudensis generated using PacBio long read and Omni-C short read sequencing. The resulting assembly is approximately 1.36 Gb with a scaffold N50 of 44.t Mb and 31 chromosome-length scaffolds. Nearly 43 percent of the genome was identified as repeat content. This assembly will serve as a resource for the conservation and study of P. bermudensis, and its only close extant and also critically endangered relative, P. circumfirmatus. Additionally, this genome adds to the growing body of data needed for a more complete understanding of gastropod evolution and for evolutionary processes in general.

Annotation

Complete genomes from a xenic Dolichospermum flosaquae FBCC-A233 culture reveal genome-inferred metabolic asymmetry with associated bacteria.

Cyanobacteria form phycosphere communities with associated bacteria, but genome-resolved resources are needed to formulate testable hypotheses about their metabolic interactions. Here, we reconstructed three complete circular genomes from a unialgal xenic culture, including Dolichospermum flosaquae FBCC-A233 and two associated alphaproteobacterial genomes assigned to Sphingorhabdus sp. and Brevundimonas sp. Genome-wide read mapping and genome-quality assessment supported the three recovered genomes as high-quality circular reconstructions. Comparative genome analysis placed the cyanobacterial genome within the Dolichospermum flosaquae species cluster under the GTDB framework, while the associated bacterial genomes represented Sphingorhabdus sp. and a putative undescribed Brevundimonas species-level lineage. Genome architecture analysis indicated reduced genome size and gene content in Brevundimonas relative to genus-level references although additional metrics did not support a strong conclusion of classical genome streamlining. Selected KEGG module and KO-level reconstructions indicated genome-inferred metabolic asymmetries across the consortium. FBCC-A233 encoded photosynthesis- and nitrogen-related modules and a BioU-mediated de novo biotin biosynthesis route, whereas the associated bacteria lacked complete de novo biotin biosynthesis but retained biotin-dependent carboxylase genes. FBCC-A233 also encoded extensive anaerobic corrinoid biosynthesis potential; however, canonical DMB-containing cobalamin completion, cobamide identity, and complete transporter systems were not resolved. Together, these complete genomes provide a genome-resolved resource for investigating genome-inferred metabolic differentiation and ecological interactions in cyanobacteria-associated bacterial consortia.IMPORTANCEPhycosphere interactions between cyanobacteria and associated bacteria can shape aquatic microbial communities, but many proposed interactions remain difficult to evaluate without genome-resolved resources. This study provides three complete circular genomes from a unialgal xenic Dolichospermum flosaquae culture, capturing the cyanobacterium and two co-maintained bacterial associates. Our analysis identifies genome-inferred metabolic asymmetries, particularly in biotin- and cobamide-related pathways. D. flosaquae FBCC-A233 encoded candidate de novo biotin and corrinoid biosynthesis capacity, whereas the associated bacteria lacked complete de novo pathways but retained cofactor-dependent enzymes. These findings nominate cofactor-related dependencies as experimentally testable hypotheses while emphasizing unresolved uptake, export, cobamide identity, and growth-dependence mechanisms. The complete genomes and KO-level reconstructions generated here provide a resource for future studies of cyanobacteria-associated consortia.

Genome, Bacterial

Comparative genomics of ESKAPE pathogen species: Integrating pan-genome architecture, antimicrobial resistance, and virulence factor repertoires.

BACKGROUND: ESKAPE pathogens are major causes of hospital-acquired infections and are characterized by extensive antimicrobial resistance (AMR) and diverse virulence mechanisms. Although species-specific pan-genome studies have revealed substantial genomic diversity, the relationships among genome plasticity, resistance burden, and virulence remain incompletely understood across the ESKAPE complex. METHODS: We analyzed 120 high-quality genomes representing six single-species ESKAPE groups (20 genomes per species). Genome quality was assessed using CheckM2. Species-specific pan-genomes were constructed with Roary, AMR genes were identified using AMRFinderPlus, and virulence factors were detected against the VFDB database using DIAMOND. AMR genes were mapped to core and accessory genome compartments through integration of Prokka annotations and Roary outputs. Statistical associations were evaluated using Fisher's exact tests and correlation analyses, with false discovery rate correction applied within each test family. Core-genome maximum-likelihood phylogenies were reconstructed to provide an evolutionary framework. RESULTS: Pan-genome sizes ranged from 4720 to 17,272 genes, with Enterobacter and Pseudomonas possessing the largest accessory genomes. Multidrug resistance (MDR; resistance to ≥3 antimicrobial classes) was detected in 93.3% of strains. After false discovery rate correction, AMR genes remained significantly enriched in the accessory genomes of Enterobacter, Enterococcus, Klebsiella, and Staphylococcus, whereas Acinetobacter and Pseudomonas did not show significant enrichment in either genome compartment. Within-species analyses identified significant positive associations between accessory genome size and AMR class burden in Staphylococcus, Enterococcus, and Enterobacter, whereas the moderate Pearson correlation observed in Pseudomonas was not significant after FDR correction. Virulence factor repertoires varied markedly among species, with Pseudomonas exhibiting the highest burden and Enterococcus the lowest. CONCLUSIONS: ESKAPE pathogens display distinct patterns of resistance and virulence. Accessory genome expansion was associated with higher AMR burden in several species, whereas other species showed no significant association between accessory genome size and AMR burden and no significant enrichment of AMR genes in either genome compartment, highlighting the species-specific nature of AMR evolution.

Virulence Factors

Comprehensive genomic and computational insights into Brucella suis: pan-genome analysis, evolutionary perspectives, and in-silico vaccine design.

BACKGROUND: Brucella suis is a zoonotic intracellular pathogen responsible for brucellosis, mainly in swine and humans. Although numerous genome sequences are publicly available, an integrative genomic analysis combining pan-genome architecture, structural organization, evolutionary relationships, and vaccine-associated targets remains limited. RESULTS: In this study, we analyzed 91 publicly available B.suis genomes to characterize their pan-genome composition and genomic structure. The pan-genome exhibited an open configuration, indicating continued genomic diversification. A total of 2,146 core genes were identified, representing conserved functions essential for species maintenance, while the accessory genome reflected strain-level variability. Phylogenetic reconstruction based on single-copy orthologs revealed distinct evolutionary clades among the strains. A complementary phylogenetic analysis of pan-genome gene presence-absence patterns further supported clade differentiation and highlighted variation in accessory gene repertoires. Comparative synteny and genome structural analyses demonstrated largely conserved chromosomal organization with localized rearrangements across strains. Screening of the core proteome identified 64 putative antigenic proteins with predicted surface localization and immunogenic properties. Additionally, resistance-associated determinants related to tetracycline and doxycycline were detected in one genome within the dataset. CONCLUSIONS: This comprehensive genomic analysis defines the pan-genome structure, evolutionary relationships, and genome organization of B.suis. The integration of core and pan-genome-based phylogenies provides complementary insights into strain diversification, while the identified conserved antigenic candidates offer a foundation for future experimental validation and rational vaccine development strategies.

Genome, Bacterial

Reference-Guided Chromosome-Scale Genome Assembly With Insights on Population Genomics of the Atlantic Goliath Grouper (Epinephelus itajara), Islas del Rosario, Colombia.

Epinephelus itajara, commonly known as the Atlantic Goliath grouper, is the largest species among the western North Atlantic groupers and is critically endangered. This species plays a crucial ecological, cultural, and economic role and has been the focus of captive breeding efforts at the Oceanario of the Rosario Islands, Colombia. However, despite its ecological and conservation importance, genomic resources and population genomic data for E. itajara remain scarce, particularly in the Colombian Caribbean. This study presents a reference-guided chromosome-scale genome assembly and an analysis of the population genomic structure of E. itajara using PacBio HiFi sequencing and Illumina technologies. The assembled genome has a total size of 1.12 Gb, with a contig N50 of 42.69 Mb and a scaffold N50 of 46.30 Mb. A total of 22,692 protein-coding genes were identified after masking 46% of the genome, which consists of repetitive elements. Comparative genomic analyses revealed a high degree of collinearity with closely related Epinephelus species and identified E. lanceolatus as the closest relative, supporting recent divergence and conserved genome architecture within the genus. Additionally, a population genomics analysis was conducted using 7706 high-quality SNPs to assess the genomic structure of captive populations. The results revealed four distinct genomic lineages, with moderate genetic differentiation among the sampled individuals. In the Colombian Caribbean, two unique lineages were identified, associated with the localities of Bahía Cispatá and Bahía Barbacoas, suggesting possible geographic isolation. These genomic resources provide valuable tools and new opportunities to better understand the genomic diversity, evolutionary history, and reproductive mechanisms of E. itajara. Moreover, they serve as a foundation for conservation strategies, including selective breeding programs aimed at increasing genomic diversity in captive populations and guiding restoration efforts in its natural habitat.

Epinephelus itajara

Whole genome sequencing and phylogenetic classification accelerate the implementation of respiratory syncytial virus genomic surveillance in Canada: a pilot study.

UNLABELLED: Whole genome sequencing (WGS) has emerged as a powerful tool to facilitate the study of existing and emerging infectious diseases. WGS-based genomic surveillance provides information on the genetic diversity and tracks the evolution of important viral pathogens, including respiratory syncytial virus (RSV). Multiplex tiling polymerase chain reaction (PCR) assays have been used to facilitate sequencing of a variety of pathogens in support of genomics-based surveillance initiatives. We developed, optimized, and implemented multiplex tiling PCR assays for RSVA and RSVB capable of generating near-complete genomes in the majority of contemporaneous specimens tested. A pilot data set comprising 52 RSVA and 37 RSVB genomes derived from Canadian clinical specimens during the 2022-2023 respiratory virus season was used to perform phylogenetic analyses using both near-complete genome and glycoprotein (G) sequences. Overall, the RSV phylogenetic tree built with whole genomes showed identical lineage clusters as compared to the G gene but was more discriminatory. Moreover, the availability of complete genomes enables the identification of a broader range of mutations. For instance, mutations identified in the fusion protein among Canadian isolates tested here, including S377N, K272M, S276N, S211N, S206I, and S209Q, could affect the efficacy of current vaccines or antiviral-based therapeutics. In conclusion, our work reinforces other recent studies demonstrating the utility of multiplex tiling PCR assays to facilitate high-throughput WGS of RSV, which is capable of supporting enhanced genomic surveillance initiatives, as well as the more comprehensive genomic analyses required to inform public health strategies for the development and usage of vaccines and antiviral drugs. IMPORTANCE: We present assays to efficiently sequence genomes of RSVA and RSVB. This enables researchers and public health agencies to acquire high-quality genomic data using rapid and cost-effective approaches. Genomic data-based comparative analysis can be used to conduct surveillance and monitor circulating isolates for efficacy of vaccines and antiviral therapeutics.

Humans

Genomics-enabled dissection of sea wheatgrass genome for advancing wheat genetic resources.

Wheat production is challenged by biotic and abiotic stresses. Alien gene transfer is an effective approach to tackle such challenges. We previously showed that sea wheatgrass (SWG; Thinopyrum junceiforme (2n = 2x = 28; J1J2) is an untapped resource possessing resistance to an array of pests and abiotic stress. However, the transfer of these important traits has been hindered by the lack of genomic resources and a clear picture of its genome constitution. Using multi-color genomic in situ hybridization, we distinguished the SWG sub-genomes and corroborated that the J1 sub-genome is closely related to the E genome of Th. elongatum and the J genome of Th. bessarabicum and the J2 sub-genome to the V genome of Dasypyrum villosum. Meanwhile, we developed a draft SWG genome assembly and 127 SWG-specific DNA markers covering the 14 SWG chromosomes. Screening a population of 466 BC2F1 and BC2F2 individuals, derived from backcrosses of wheat-SWG amphiploid to wheat, by the SWG-specific markers led to selection of 72 plants putatively carrying one or two SWG chromosomes. The genome painting analysis of the 72 plants eventually identified a set of 37 wheat-SWG chromosome addition lines covering all the 14 pairs of SWG chromosomes and two compensating Robertsonian translocations (RobTs). While the wheat-SWG chromosome addition lines and RobTs are invaluable genetic resources for wheat improvement via chromosome engineering, our results showed the power of genome-specific markers in combination with genome painting in dissection of a polyploid genome and implicated the origin of a group of important polyploid grasses.

Triticum

The Rise of Plant Pan-Genomes: From Genome Variation to Predictive Breeding.

Plant pan-genomics is entering a new phase beyond genome variation discovery, requiring a shift from cataloguing genomic diversity toward understanding how variation generates biological function and breeding value. Here, we propose that the future of plant pan-genomics will be shaped by three conceptual transitions. First, structural variation (SV), presence-absence variation (PAV), and haplotype diversity should be interpreted not merely as genomic differences, but as regulatory components that influence gene networks, chromatin organization, and complex traits. Second, the expansion from species-level pan-genomes to genus-level super pan-genomes provides an evolutionary framework for uncovering adaptive genetic modules preserved in wild relatives and overlooked during domestication. Third, integrating pan-genomes with pan-omics, three-dimensional genome analyses, and artificial intelligence will enable the transformation of genomic variation into predictive models for crop improvement. We further propose that the ultimate value of pan-genomes lies not in generating increasingly complete genome collections, but in establishing a mechanistic bridge between genome diversity, biological function, and breeding decisions. This transition will move crop improvement from empirical selection toward rational genome design, where evolutionary diversity can be systematically interpreted, predicted, and engineered.

Journal Article

A complete diploid human genome benchmark for personalized genomics.

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Journal Article

CREAT: A CRISPR-Based Genome Trimming Strategy for Systematic Identification of Dispensable Regions and Rapid Genome Reduction.

The construction of minimal-genome microbes offers an ideal platform for understanding fundamental biological processes and synthetic biology, yet the research is hindered by incomplete lists of essential genes in microbes and by multiple rounds of genome trimming with a trial-and-error nature. To address this, we introduce CREAT (CRISPR-based genome trimming with a multi-homology-arm template)-a streamlined approach that integrates CRISPR-targeted genome cleavage and homology arm walking to classify essential from non-essential genomic subregions, thus providing the basis for predicting essential genes in a given organism. These essential genes were then assembled into synthetic gene cassettes for one-step replacement of the targeted non-deletable genomic regions for further genome trimming. Eight consecutive rounds of CREAT genome trimming achieved a 20.8% reduction in genome size in Saccharolobus islandicus. Furthermore, Cas9-based CREAT genome trimming was developed for Bacillus subtilis and Escherichia coli, with efficiency greatly enhanced by the λ-Red recombinase in the latter. Together, this iterative application of CREAT provides a scalable and generally applicable strategy for rapidly constructing minimal genomes across diverse microorganisms.

CRISPR-Cas Systems

Genomes for Nurses: Understanding and Overcoming Barriers to Nurses Utilizing Genomics.

Background: Genomic testing is an increasingly important technology within pediatric oncology that aids in cancer diagnosis, provides prognostic information, identifies therapeutic targets, and reveals underlying cancer predisposition. However, nurses lack basic knowledge of genomics and have limited self-assurance in using genomic information in their daily practice. This single-institution project was carried out at an academic pediatric cancer hospital in the United States with the aim to explore the barriers to achieving genomics literacy for pediatric oncology nurses. Method: This project assessed barriers to genomic education and preferences for receiving genomics education among pediatric oncology nurses, nurse practitioners, and physician assistants. An electronic survey with demographic questions and 15 genetics-focused questions was developed. The final survey instrument consisted of nine sections and was pilot-tested prior to administration. Data were analyzed using a ranking strategy, and five focus groups were conducted to capture more-nuanced information. The focus group sessions lasted 40 min to 1 hour and were recorded and transcribed. Results: Over 50% of respondents were uncomfortable with or felt unprepared to answer questions from patients and/or family members about genomics. This unease ranked as the top barrier to using genomic information in clinical practice. Discussion: These results reveal that most nurses require additional education to facilitate an understanding of genomics. This project lays the foundation to guide the development of a pediatric cancer genomics curriculum, which will enable the incorporation of genomics into nursing practice.

Humans

Genomic inbreeding coefficients and inbreeding depression of semen production traits at genome-wide and chromosomal levels in Japanese Holstein bulls.

We aimed to estimate inbreeding coefficients and the effects of inbreeding depression on semen production traits at both the genome-wide and chromosomal levels. We utilized pedigree data for 19,921 animals, single nucleotide polymorphism (SNP) data on 5700 Japanese Holstein bulls, and 52,193 semen collection records from 775 bulls. We estimated 4 different inbreeding coefficients, namely a pedigree-based coefficient (FPED) and 3 genomic coefficients derived from SNP data. The genomic coefficients consisted of one based on the genomic relationship matrix (FGRM), one based on runs of homozygosity (ROH), and one based on homozygous-by-descent (HBD) segments (FHBD). These genomic coefficients were estimated at both the genome-wide and chromosomal levels. Furthermore, we investigated the effects of these coefficients on semen production traits: semen volume (VOL), sperm concentration (CON), sperm number (NUM), and sperm motility (MOT). In the genome-wide-level analysis, inbreeding coefficients increased markedly in bulls born after 2009, coinciding with the introduction of genomic selection. Significant inbreeding depression of VOL was found. At the chromosomal level, the inbreeding coefficients for most chromosomes showed a similar trend to the genome-wide metrics, although some (e.g., chr10 and chr20) exhibited a more pronounced trend. Suggestive inbreeding effects were detected on specific chromosomes for all traits (chr1 and chr22 for VOL, chr24 and chr29 for CON, chr1, chr12, and chr27 for NUM, chr10 and chr18 for MOT), including the traits that were not significant at the genome-wide level. Our results highlight that chromosomal-level analysis provides information complementary to whole-genome metrics, offering a more detailed perspective for managing inbreeding effects. To mitigate the adverse effects of inbreeding on semen production traits, future breeding programs would benefit from the control of inbreeding effects on high-risk chromosomal regions.

Genomic inbreeding coefficient

Comparative genomics and phylogenetic analysis of three Malvaceae species on the basis of chloroplast genomes.

INTRODUCTION: The Malvaceae family shows rich species diversity and has substantial economic and medicinal value. However, the frequent interspecific hybridization among members of this family has resulted in confused phylogenetic relationships among the groups, limiting the usefulness of traditional classification methods. METHODS: This study aimed to investigate the phylogenetic relationships among selected taxa of Malvaceae by evaluating 23 chloroplast (CP) genomes, including three newly assembled CP genomes. Among these three genomes, the CP genome of Hibiscus schizopetalus L. was reported for the first time, while the CP genomes of Alcea rosea L. and Hibiscus grewiifolius L., which have been deposited in NCBI, were re-analyzed here alongside newly generated data for comparative purposes. In addition, 20 downloaded CP genomes encompassing 13 genera were analyzed using SNPs in whole CP genomes data. RESULTS: The results showed that the genomes ranged from 160,403 to 161,978 base pairs in length and consisted of small single copies (SSCs) and large single copies (LSCs) separated by two inverted repeat sequences (IRs), forming a typical quadripartite circular structure. The entire genome sequence showed relative conservation across species in terms of structure, GC content, codon usage, and gene composition. The mutation sites were mainly located in the LSC and SSC regions, and the variability in the non-coding regions was higher than that in the coding regions. The nucleotide polymorphism (Pi) analysis identified the non-coding regions such as ndhF-rpl32 and psbZ-trnG as high variable hotspots. A maximum likelihood phylogenetic tree was constructed based on SNPs in whole CP genomes data. The phylogenetic analysis divided these 23 species into five highly supported clades. It also revealed a close sister-group relationship between Abelmoschus and Hibiscus species, suggesting that Hibiscus may have a separate lineage from okra species. DISCUSSION: In conclusion, the increasing availability of CP genome resources will enhance our understanding of the classification and evolutionary patterns of the Malvaceae family. The development of molecular markers will provide important molecular evidence for precise identification and classification revision of plants in this family.

Malvaceae

Combinatorial genome engineering of pseudorabies virus Bartha by developing a reverse genetic system based on three overlapping genomic segments.

INTRODUCTION: The 138-kilobase genome of pseudorabies virus vaccine strain Bartha K61 harbors many nonessential genes for replication and exhibits remarkable capacity for incorporating foreign genes for therapeutic applications. However, the large size of the Bartha genome complicates its efficient engineering. OBJECTIVES: Development of a reverse genetic system for pseudorabies virus Bartha based on three overlapping genomic segments to facilitate multiplex genome engineering. METHODS: The 138-kb genome of Bartha was split into three overlapping segments (42 kb, 43 kb, and 53 kb), each cloned in a bacterial artificial chromosome (BAC) to facilitate genome engineering. The infectious virus was reconstituted by transfecting the 3 genomic fragments released from the BACs into Vero cells in which a complete virus genome was assembled using 2-kb overlaps between adjacent pieces. RESULTS: Employing the reverse genetic system, we individually deleted 15 candidate nonessential genes and confirmed that 10 were dispensable for viral growth in cell culture. Deletion of 7 nonessential genes had no impact on viral growth, whereas UL47 deletion reduced viral growth rate and deletions of UL44, UL47, or US3 resulted in smaller viral plaques. A total of 45 viral genomes with double deletions of nonessential genes were constructed, among which 22 were successfully rescued into infectious virions. Fifteen double-deletion mutant viruses had a viral titer comparable with the wild-type Bartha, while the remaining 7 showed a lower titer. Additionally, expressions of the mNeonGreen reporter gene at nonessential gene loci were evaluated. Cells infected with recombinant viruses carrying mNeonGreen at 8 loci showed strong green fluorescence, whereas those with mNeonGreen at 2 loci exhibited very weak fluorescence. CONCLUSION: The reverse genetic system developed in this study enables rapid and combinatorial engineering of viruses with the large DNA genome, and will accelerate development of large DNA virus-based therapeutics including live-attenuated vaccines, vector vaccines, and oncolytic herpesviruses.

Herpesvirus 1, Suid

MutBERT: probabilistic genome representation improves genomics foundation models.

MOTIVATION: Understanding the genomic foundation of human diversity and disease requires models that effectively capture sequence variation, such as single nucleotide polymorphisms (SNPs). While recent genomic foundation models have scaled to larger datasets and multi-species inputs, they often fail to account for the sparsity and redundancy inherent in human population data, such as those in the 1000 Genomes Project. SNPs are rare in humans, and current masked language models (MLMs) trained directly on whole-genome sequences may struggle to efficiently learn these variations. Additionally, training on the entire dataset without prioritizing regions of genetic variation results in inefficiencies and negligible gains in performance. RESULTS: We present MutBERT, a probabilistic genome-based masked language model that efficiently utilizes SNP information from population-scale genomic data. By representing the entire genome as a probabilistic distribution over observed allele frequencies, MutBERT focuses on informative genomic variations while maintaining computational efficiency. We evaluated MutBERT against DNABERT-2, various versions of Nucleotide Transformer, and modified versions of MutBERT across multiple downstream prediction tasks. MutBERT consistently ranked as one of the top-performing models, demonstrating that this novel representation strategy enables better utilization of biobank-scale genomic data in building pretrained genomic foundation models. AVAILABILITY AND IMPLEMENTATION: https://github.com/ai4nucleome/mutBERT.

Humans