Search PubMedSearch

SEARCH · Search PubMed

Results for “genome size”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

628 records · Page 4Linked to original sources

Assembly and Characterization of the First Complete Mitochondrial Genome of Tussilago farfara L.: Insights into Biological Functions and Phylogenetic Relationships within the Asteraceae Family.

Tussilago farfara L., a member of the Asteraceae family, is an economically valuable species due to its edible and medicinal properties. To elucidate the structural characteristics, genetic mechanisms, and evolutionary pathways of the organelle genomes of T. farfara, we sequenced, assembled, and annotated its mitochondrial genome for the first time. The complete mitochondrial genome of T. farfara spans 306,024 bp and contains 33 mitochondrial protein-coding genes (PCGs), 3 rRNAs, and 22 tRNAs. Analysis of the nucleotide substitution rate and genetic diversity revealed that most mitochondrial genome genes may have undergone purifying selection, indicating a slow evolutionary rate and a relatively conserved genomic structure. We further identified 13 fragments of chloroplast-derived DNA integrated into the mitochondrial genome, evidencing intracellular gene transfer. Collinearity analysis showed that Arctium lappa shares the most extensive mitochondrial homologous sequences and the highest sequence similarity with T. farfara. Phylogenetic analysis based on the mitochondrial genome helped to clarify the evolutionary and taxonomic position of T. farfara within the Asteraceae family. The mitochondrial genome sequence of T. farfara provides a valuable genomic resource for species identification and for evolutionary studies within the Asteraceae family.

Genome, Mitochondrial

Plant species identification by genome skimming across the vascular plant tree of life.

Accurate species identification is essential for biodiversity conservation and sustainable use, yet standard plant DNA barcoding often fails to achieve species-level resolution. We present a large-scale empirical evaluation of genome skimming as a tool to improve plant species discrimination. Using standardised data from 1969 individuals representing 475 species from 32 genera across major lineages of the vascular plant tree of life, we compare conventional plastid + internal transcribed spacer (ITS) barcodes with genome skimming approaches. Standard barcoding using rbcL, matK, trnH-psbA and ITS resolved about half of species (49.3%), with six genera showing <&#x2009;25% species discrimination. By contrast, genome skimming enabled the recovery of complete plastid genomes, yielding 57.6% species discrimination. It also generated sufficient nuclear genomic data for additional resolution from k-mer analysis, achieving 66.8% species discrimination - an average gain of 17.5% over standard barcodes - while eliminating cases of extreme failure (<&#x2009;25% resolution). The recovery of complete plastomes and ribosomal DNAs from genome skims also ensures backward compatibility with existing barcode datasets. Our results demonstrate that genome skimming provides data that substantially improves species-level resolution across diverse plant lineages and offers a scalable, high-throughput approach for building comprehensive reference resources to support global biodiversity initiatives.

DNA Barcoding, Taxonomic

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Artificial Intelligence

Comparative genomic epidemiology of food- and patient-derived diarrheagenic Escherichia coli from sentinel surveillance in Southeast China.

Diarrheagenic Escherichia coli (DEC) remains an important foodborne pathogen, yet long-term comparative genomic surveillance data jointly characterizing food-derived and patient-derived isolates remain limited. This surveillance-based comparative study integrated antimicrobial susceptibility testing and whole-genome sequencing to characterize diarrheagenic Escherichia coli isolates recovered from food and patient sources in Lishui, Southeast China, during 2018-2025, with emphasis on occurrence, resistance profiles, genomic backgrounds, and plasmid replicon-associated features. Antimicrobial susceptibility testing was performed for 258 selected isolates, and whole-genome sequencing was conducted for a curated analytical subset of 204 isolates. The sequenced subset was used for diversity-oriented comparative genomic analysis rather than for unbiased prevalence estimation of the entire DEC collection. EAEC predominated in both sources, although food-associated occurrence was heterogeneous across categories, with the highest recovery rate observed in raw meat. Patient-derived isolates showed a broader overall resistance burden, whereas food-derived isolates retained substantial resistance to tetracycline, chloramphenicol, and florfenicol. Phylogenetic analysis showed partial overlap in genomic backgrounds between food-derived and patient-derived isolates, while representative resistance determinants displayed both broadly distributed and lineage-enriched patterns. Replicon-based plasmid profiling identified 42 plasmid types, including 12 detected in both sources, with IncF-related replicons predominating among these shared profiles. Several food-derived isolates carried multiple plasmid replicon types that were also observed in patient-derived isolates. Overall, food-derived and patient-derived DEC showed partial overlap in genomic backgrounds, resistance determinants, and replicon-defined plasmid profiles within this surveillance setting, while retaining source-associated heterogeneity. These findings should be interpreted as surveillance-based comparative evidence rather than as evidence of direct source attribution or transmission.

Humans

Genomic science and the nurse educator's role: Promoting integration from curriculum to clinical practice.

BACKGROUND: Registered nurses and nurse educators play a critical role in preparing future clinicians to translate genomic discoveries into practice. However, emerging evidence suggests that both groups may lack sufficient knowledge and confidence in genomics, potentially limiting their ability to teach, mentor, and apply genomics in real-world settings. This gap is especially concerning in Aotearoa New Zealand, where the genomic literacy of nurse educators and clinicians remains underexplored. OBJECTIVE: This study aims to: (1) assess nurse educators' genomic literacy and confidence in teaching genomics; and (2) evaluate registered nurses' knowledge and confidence in applying and teaching genomics in clinical practice. DESIGN: Exploratory descriptive qualitative. SETTING: This study was conducted in the greater Auckland area. PARTICIPANTS: A total of 17 participants were recruited using purposive sampling to ensure a diverse range of perspectives across varying levels of teaching experience, disciplinary backgrounds, and exposure to genomic content. METHODS: Data were collected using semi-structured focus group interviews, a method well-suited for generating in-depth discussion and facilitating interaction among participants with shared professional interests. The collected data were analysed using thematic analysis methods. RESULTS: The findings offer insight into the preparedness of New Zealand's nursing workforce to engage with genomic-informed healthcare and inform strategies for integrating genomics into nursing curricula and continuing professional development. Given the interdisciplinary nature of genomic healthcare, these insights may also be relevant to other health professionals-including midwives, pharmacists, and allied health practitioners-who increasingly encounter genomic information in clinical practice and require foundational competencies to support patient care. CONCLUSION: Addressing this educational gap is critical to ensuring that nurses-key facilitators of patient care and public health-are equipped to deliver safe, equitable, and evidence-based genomic healthcare.

Humans

Evaluation of one-step amplicon-based targeted enrichment for SARS-CoV-2 whole-genome sequencing using the Midnight amplicon scheme.

Genomic surveillance proved invaluable during the COVID-19 pandemic for tracking SARS-CoV-2 variants and guiding outbreak responses, underscoring the ongoing need to reduce whole-genome sequencing (WGS) costs and improve workflow efficiency to ensure accessibility in resource limited settings. Here, we evaluated a one-step reverse transcription polymerase chain reaction (RT-PCR) approach using the Midnight V2 primer scheme for targeted amplification of the SARS-CoV-2 genome, assessed its compatibility with Illumina sequencing, and compared its performance to a well-established two-step method. Initially, we determined optimal RT-PCR reaction conditions using the Midnight V2 primer panel for the one-step RT-PCR kit and scaled reaction volumes for both RT-PCR and library preparation. Clinical specimens (n&#x202f;=&#x202f;53) that had undergone routine WGS for surveillance purposes using the established two-step RT-PCR method were compared using the one-step RT-PCR assay. For samples with genome completeness greater than 70%, both methods gave comparable results with similar sequence coverage and 100% concordance for lineage assignment. Further investigation revealed a higher percentage of reads aligning to the SARS-CoV-2 genome with a greater depth of coverage using the one-step method compared to the two-step method. Finally, analysis of scaled one-step and library reaction volumes revealed significant cost savings for samples undergoing WGS. Overall, the results presented here verify the accuracy and reproducibility of one-step targeted amplification and offer an efficient and cost-effective workflow for routine SARS-CoV-2 genomic surveillance.

Humans

Mul-PheG2P: decoupled learning and prediction-space fusion enables robust and interpretable multi-phenotype genomic prediction.

Genomic prediction of multiple phenotypes is crucial in modern plant breeding; however, existing methods struggle with negative transfer and lack interpretability, particularly across high-dimensional small-sample data and diverse species. To address this, we propose Mul-PheG2P, a novel paradigm based on decoupled learning and predictive space fusion. It employs a two-stage design: first training phenotype-specific encoders using genetic data, then decoupling phenotype-specific learning from cross-phenotype aggregation via an interpretable prediction layer. Mul-PheG2P outperforms existing methods across diverse crop datasets, including maize (Zea mays), wheat (Triticum aestivum), and tomato (Solanum lycopersicum). It provides a multi-scale interpretability chain: at the macro level, it quantifies phenotypic contributions via attention-based weighting; at the micro level, Integrated Gradients reveal the genetic basis of predictions. Notably, the model successfully identified the CCT (CONSTANS, CO-like, and TOC) motif regulating photoperiodism and the SQUAMOSA (SQUAMOSA promoter binding protein) promoter for inflorescence development, confirming its ability to capture functional biological mechanisms. These results highlight the high performance and interpretability of Mul-PheG2P, showcasing its value for low-cost, large-scale screening to advance precision breeding.

Phenotype

Genome-guided stage- and tissue-resolved transcriptome analysis of Serrodes campana identifies sex-biased antennal expression and candidate chemosensory-related genes.

Serrodes campana is an erebid moth of ecological and forestry relevance; its larvae are mainly associated with the soapberry tree, Sapindus mukorossi, whereas adults exhibit fruit-piercing behavior. However, stage- and tissue-resolved transcriptomic resources for this species remain limited. Here, using a chromosome-level reference genome, we performed a genome-guided transcriptome analysis of S. campana based on 12 RNA-seq libraries representing major developmental stages and key adult tissues. Global transcriptomic analyses revealed pronounced transcriptional differentiation across developmental stages and tissue types. Tissue-enriched gene sets and functional enrichment analyses identified distinct molecular signatures associated with developmental, sensory, and pheromone-associated tissues. Comparative analysis of female and male antennae further revealed sex-biased expression of several candidate chemosensory-related genes. Among 153 curated chemosensory-related candidate genes, most odorant receptor genes showed strong antennal enrichment, whereas other major chemosensory gene families displayed broader but still tissue-preferential expression patterns. In addition, an exploratory comparison of female terminal abdominal gland tissue and male terminal abdominal coremata revealed divergent expression profiles and highlighted candidate genes potentially associated with pheromone-related physiology, reproduction, and tissue-specific signaling. Together, this study provides the first genome-guided stage- and tissue-resolved transcriptomic resource for S. campana and offers a useful foundation for future studies of chemosensory detection, sex-biased gene expression, and pheromone-associated biology in this species.

Animals

Comparative genomics and full-length transcriptome profiling of wing morphs in Tetrix grossus (Orthoptera: Tetrigidae).

Wing polymorphism represents a paradigmatic dispersal-reproduction trade-off, yet its molecular basis remains uncharacterised in the phylogenetically distant pygmy grasshoppers (Tetrigidae). Here we integrate comparative genomics across ten orthopteran species with full-length transcriptomics of long-winged (FL) and short-winged (FS) Tetrix grossus. OrthoFinder recovered 118 orthogroups specific to T. grossus. Against a backdrop of pronounced gene-family contraction (36 expansions versus 222 contractions; net -186, mirrored at the ancestral Tetrix node, +37/-140), we identified an ancestral, Tetrix-specific expansion of hormone-regulation (12 genes; fold enrichment 7.93) and lipid/carbohydrate-metabolic families organised into syntenic clusters, alongside 513 positively selected genes enriched for integrin-mediated cell adhesion (6 genes), a process relevant to epithelial and appendage morphogenesis. Full-length transcriptomics of one long-winged (FL) and one short-winged (FS) adult female detected 7530 (FL) and 7515 (FS) expressed genes, with 794 FL- and 776 FS-restricted transcriptome-derived SNP-associated genes. The FL morph was enriched for an EGFR/Ras-Rho developmental-patterning axis and neuromuscular flight genes, whereas the FS morph was enriched for insulin/peptide-hormone response and growth-regulatory loci. Overall, we present genomic resources and testable hypotheses concerning the evolution and regulation of wing morphs in Tetrigidae rather than a validated genetic architecture of wing-morph determination.

Animals

Distinct functions of mammalian RAD51 paralogs in genome maintenance.

RAD51 paralogs (RAD51B, RAD51C, RAD51D, XRCC2, and XRCC3) are evolutionarily conserved essential proteins for cell survival and genome maintenance. RAD51 paralogs were originally identified to play a role in homologous recombination-mediated repair of DNA double-strand breaks (DSBs). However, investigations over the last decade have uncovered new roles of RAD51 paralogs beyond DSB repair in replication stress responses, including replication fork progression, fork stability, and its restart. Recent structural studies have not only uncovered the molecular architecture of previously known RAD51 paralog complexes but also identified novel paralog complex assemblies, providing mechanistic insights into their various genome-maintenance functions. Additionally, a role for RAD51 paralogs in resolving R-loops has been identified, and studies with cancer-associated variants suggest that RAD51 paralogs are potential determinants of cancer susceptibility and therapeutic responses. In the present review, we highlight the recently deciphered structures and novel functions of RAD51 paralog complexes and discuss the clinical and therapeutic implications.

Rad51 Recombinase

Genome-wide characterization of heat shock protein genes reveals thermal stress-responsive candidates in Litopenaeus vannamei.

Heat shock proteins (HSPs) are conserved molecular chaperones involved in protein folding, refolding, aggregation prevention, and degradation of damaged proteins. However, the genomic organization and thermal responsiveness of HSP genes in the Pacific white shrimp (Litopenaeus vannamei) remain incompletely understood. Here, we performed a genome-wide analysis of the HSP gene family and examined its phylogenetic relationships, structural features, duplication patterns, sequence variation, interaction networks, and transcriptional responses to acute heat stress. A total of 34 HSP genes were identified and classified into the HSP90, HSP70, HSP40/DNAJ, HSP60, and small HSP families. Phylogenetic, motif, gene structure, synteny, and subcellular localization analyses revealed evolutionary conservation and structural diversification among family members. Three duplicated gene pairs were identified, comprising two segmental duplications and one tandem duplication. All pairs exhibited Ka/Ks ratios below 1, consistent with purifying selection of varying strength. Sequence analysis identified 295 nonsynonymous single-nucleotide polymorphisms, of which 12 were consistently predicted to be deleterious by multiple algorithms. Protein-protein interaction analysis indicated enrichment of protein-folding and cellular stress-response functions. RT-qPCR analysis showed significant induction of HSPA4, HSP90AA1, TRAP1, BiP, and DNAJA1 after 6, 12, and 24&#xa0;h of exposure to 34&#xa0;&#xb0;C, whereas DNAJC3 was significantly induced only at 12&#xa0;h. All six genes reached their highest transcript abundance at 12&#xa0;h. These findings may provide a genomic framework for HSP genes in L. vannamei and identify candidate genes and variants associated with thermal stress responses.

Animals

Gloeotrichia echinulata genomes from the United States are nontoxigenic and likely geosmin producers.

Six Gloeotrichia echinulata genomes derived from planktonic harmful algal blooms (HABs) with similar colonial morphology have been sequenced from lakes in the west and northeast regions of USA, four of them to completion. The c. 7 Mbp genomes exhibit a high level of conservation, with 98-99% pairwise genome-wide average nucleotide identity and high levels of synteny, representing a single species cluster. We observed strong conservation of gene clusters responsible for the synthesis of the secondary metabolites and bioactive peptides that are characteristic of HAB-forming cyanobacteria. All six G. echinulata genomes lack genes for the synthesis of classic cyanotoxins, including microcystin, but possess genes responsible for the synthesis of the taste and odor compound geosmin. Interestingly, the geoA geosmin synthase gene in three genomes is homologous to other cyanobacterial geoA genes, while the other three geoA genes are related to actinomyces geoA. Phylogenomic analysis places the G. echinulata genomes within a clade of benthic Nostocales, reflecting an ecological niche featuring extensive growth on the sediment surface before colonies disperse into the epilimnion for planktonic growth. We identify genes conserved in all six genomes that could represent physiological adaptations supporting active growth on sediments and pelagic recruitment independent of wind-driven mixing: phycoerythrin light harvesting complexes for optimal photosynthesis at depth; gliding motility to access patchy nutrient distributions; and gas vesicles with relatively small GvpC proteins that predict resistance to higher hydrostatic pressure. The strong genomic similarity across geographically distant populations suggests that G. echinulata in the United States is a tightly related non-toxigenic species group with predictable properties relevant to public health and drinking water management.

Cyanobacteria

Genomic determinants underlying biogenic amine detoxification phenotypes in food-associated lactic acid bacteria: Mechanism, evolutionary origin, and relevance to fermented food safety.

Biogenic amines (BAs) are toxic metabolites that accumulate in fermented foods and pose significant food safety concerns. Although several lactic acid bacteria (LAB) have previously been reported to exhibit strain-specific BA-degrading phenotypes, the genetic determinants underlying these activities have remained largely uncharacterized. Here, we analyzed 8251 LAB genomes to validate BA-degrading phenotypes. We predicted five BA-associated genes, including two direct biogenic amine-degrading genes (BADGs), mco and patA, and three polyamine-modifying genes (PMGs), speG, paiA, and bltD. Among BADGs, mco was broadly distributed across LAB and strongly enriched across food-associated niches. patA, organized within a conserved potD-glnB-potABC-patA cassette, is a putative, functionally distinct BADG in LAB, revealing a nitrogen-responsive polyamine uptake-catabolism module. Phylogenomics, phylogenetic reconciliation, and synteny analysis established that all five genes entered the LAB through episodic horizontal gene transfer followed by lineage-specific fixation. GC compositional bias and mobile genetic element association further corroborated the horizontal origin of the two BADGs. Structural analysis confirmed the conservation of catalytic core residues of BADGs across LAB, indicating strong purifying selection. Phenotype-to-genotype correlation with experimentally reported LAB suggested mco as a reliable genomic predictor of degrading phenotype. Integration of degradation and biosynthetic profiles predicted multiple LAB species capable of both synthesizing and degrading BA, along with 1823 genomes with degradation potential but lacking detectable BA biosynthesis genes. This study provides the first large-scale genome framework linking BA-degrading phenotypes with their genetic determinants in LAB and offers a rational basis for selecting BA-detoxifying strains for fermented food applications.

Biogenic Amines

Efficient homologous replacement and deletion of large genomic fragments through template-jumping prime editing in rice.

Homologous replacement of genomic sequences with large DNA fragments (>&#x2009;100&#x2009;bp) holds great potential for crop breeding, yet an efficient method to achieve such edits is lacking in plants. Here, in rice, we developed template-jumping prime editing (TJ-PE), a recently reported PE strategy for large targeted insertion, as an efficient tool for homologous replacement with DNA fragments ranging from dozens to hundreds of base pairs, and using TJ-PE, we replaced genomic fragments of up to 340&#x2009;bp with homologous fragments of the same length. In addition, our TJ-PE tool also enabled precise deletion of 944- to 2024-bp fragments in rice, with efficiencies of up to 34.6% for c. 2000-bp precise deletions. Collectively, this study expands the editing scope of PE in rice and establishes TJ-PE as a generalist tool for precise deletion and replacement of large DNA fragments.

Oryza