Search PubMedSearch

SEARCH · Search PubMed

Results for “Whole genome of bacteria”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

2,998 records · Page 11Linked to original sources

Mul-PheG2P: decoupled learning and prediction-space fusion enables robust and interpretable multi-phenotype genomic prediction.

Genomic prediction of multiple phenotypes is crucial in modern plant breeding; however, existing methods struggle with negative transfer and lack interpretability, particularly across high-dimensional small-sample data and diverse species. To address this, we propose Mul-PheG2P, a novel paradigm based on decoupled learning and predictive space fusion. It employs a two-stage design: first training phenotype-specific encoders using genetic data, then decoupling phenotype-specific learning from cross-phenotype aggregation via an interpretable prediction layer. Mul-PheG2P outperforms existing methods across diverse crop datasets, including maize (Zea mays), wheat (Triticum aestivum), and tomato (Solanum lycopersicum). It provides a multi-scale interpretability chain: at the macro level, it quantifies phenotypic contributions via attention-based weighting; at the micro level, Integrated Gradients reveal the genetic basis of predictions. Notably, the model successfully identified the CCT (CONSTANS, CO-like, and TOC) motif regulating photoperiodism and the SQUAMOSA (SQUAMOSA promoter binding protein) promoter for inflorescence development, confirming its ability to capture functional biological mechanisms. These results highlight the high performance and interpretability of Mul-PheG2P, showcasing its value for low-cost, large-scale screening to advance precision breeding.

Phenotype

Genome-guided stage- and tissue-resolved transcriptome analysis of Serrodes campana identifies sex-biased antennal expression and candidate chemosensory-related genes.

Serrodes campana is an erebid moth of ecological and forestry relevance; its larvae are mainly associated with the soapberry tree, Sapindus mukorossi, whereas adults exhibit fruit-piercing behavior. However, stage- and tissue-resolved transcriptomic resources for this species remain limited. Here, using a chromosome-level reference genome, we performed a genome-guided transcriptome analysis of S. campana based on 12 RNA-seq libraries representing major developmental stages and key adult tissues. Global transcriptomic analyses revealed pronounced transcriptional differentiation across developmental stages and tissue types. Tissue-enriched gene sets and functional enrichment analyses identified distinct molecular signatures associated with developmental, sensory, and pheromone-associated tissues. Comparative analysis of female and male antennae further revealed sex-biased expression of several candidate chemosensory-related genes. Among 153 curated chemosensory-related candidate genes, most odorant receptor genes showed strong antennal enrichment, whereas other major chemosensory gene families displayed broader but still tissue-preferential expression patterns. In addition, an exploratory comparison of female terminal abdominal gland tissue and male terminal abdominal coremata revealed divergent expression profiles and highlighted candidate genes potentially associated with pheromone-related physiology, reproduction, and tissue-specific signaling. Together, this study provides the first genome-guided stage- and tissue-resolved transcriptomic resource for S. campana and offers a useful foundation for future studies of chemosensory detection, sex-biased gene expression, and pheromone-associated biology in this species.

Animals

Comparative genomics and full-length transcriptome profiling of wing morphs in Tetrix grossus (Orthoptera: Tetrigidae).

Wing polymorphism represents a paradigmatic dispersal-reproduction trade-off, yet its molecular basis remains uncharacterised in the phylogenetically distant pygmy grasshoppers (Tetrigidae). Here we integrate comparative genomics across ten orthopteran species with full-length transcriptomics of long-winged (FL) and short-winged (FS) Tetrix grossus. OrthoFinder recovered 118 orthogroups specific to T. grossus. Against a backdrop of pronounced gene-family contraction (36 expansions versus 222 contractions; net -186, mirrored at the ancestral Tetrix node, +37/-140), we identified an ancestral, Tetrix-specific expansion of hormone-regulation (12 genes; fold enrichment 7.93) and lipid/carbohydrate-metabolic families organised into syntenic clusters, alongside 513 positively selected genes enriched for integrin-mediated cell adhesion (6 genes), a process relevant to epithelial and appendage morphogenesis. Full-length transcriptomics of one long-winged (FL) and one short-winged (FS) adult female detected 7530 (FL) and 7515 (FS) expressed genes, with 794 FL- and 776 FS-restricted transcriptome-derived SNP-associated genes. The FL morph was enriched for an EGFR/Ras-Rho developmental-patterning axis and neuromuscular flight genes, whereas the FS morph was enriched for insulin/peptide-hormone response and growth-regulatory loci. Overall, we present genomic resources and testable hypotheses concerning the evolution and regulation of wing morphs in Tetrigidae rather than a validated genetic architecture of wing-morph determination.

Animals

Distinct functions of mammalian RAD51 paralogs in genome maintenance.

RAD51 paralogs (RAD51B, RAD51C, RAD51D, XRCC2, and XRCC3) are evolutionarily conserved essential proteins for cell survival and genome maintenance. RAD51 paralogs were originally identified to play a role in homologous recombination-mediated repair of DNA double-strand breaks (DSBs). However, investigations over the last decade have uncovered new roles of RAD51 paralogs beyond DSB repair in replication stress responses, including replication fork progression, fork stability, and its restart. Recent structural studies have not only uncovered the molecular architecture of previously known RAD51 paralog complexes but also identified novel paralog complex assemblies, providing mechanistic insights into their various genome-maintenance functions. Additionally, a role for RAD51 paralogs in resolving R-loops has been identified, and studies with cancer-associated variants suggest that RAD51 paralogs are potential determinants of cancer susceptibility and therapeutic responses. In the present review, we highlight the recently deciphered structures and novel functions of RAD51 paralog complexes and discuss the clinical and therapeutic implications.

Rad51 Recombinase

Genome-wide characterization of heat shock protein genes reveals thermal stress-responsive candidates in Litopenaeus vannamei.

Heat shock proteins (HSPs) are conserved molecular chaperones involved in protein folding, refolding, aggregation prevention, and degradation of damaged proteins. However, the genomic organization and thermal responsiveness of HSP genes in the Pacific white shrimp (Litopenaeus vannamei) remain incompletely understood. Here, we performed a genome-wide analysis of the HSP gene family and examined its phylogenetic relationships, structural features, duplication patterns, sequence variation, interaction networks, and transcriptional responses to acute heat stress. A total of 34 HSP genes were identified and classified into the HSP90, HSP70, HSP40/DNAJ, HSP60, and small HSP families. Phylogenetic, motif, gene structure, synteny, and subcellular localization analyses revealed evolutionary conservation and structural diversification among family members. Three duplicated gene pairs were identified, comprising two segmental duplications and one tandem duplication. All pairs exhibited Ka/Ks ratios below 1, consistent with purifying selection of varying strength. Sequence analysis identified 295 nonsynonymous single-nucleotide polymorphisms, of which 12 were consistently predicted to be deleterious by multiple algorithms. Protein-protein interaction analysis indicated enrichment of protein-folding and cellular stress-response functions. RT-qPCR analysis showed significant induction of HSPA4, HSP90AA1, TRAP1, BiP, and DNAJA1 after 6, 12, and 24 h of exposure to 34 °C, whereas DNAJC3 was significantly induced only at 12 h. All six genes reached their highest transcript abundance at 12 h. These findings may provide a genomic framework for HSP genes in L. vannamei and identify candidate genes and variants associated with thermal stress responses.

Animals

Gloeotrichia echinulata genomes from the United States are nontoxigenic and likely geosmin producers.

Six Gloeotrichia echinulata genomes derived from planktonic harmful algal blooms (HABs) with similar colonial morphology have been sequenced from lakes in the west and northeast regions of USA, four of them to completion. The c. 7 Mbp genomes exhibit a high level of conservation, with 98-99% pairwise genome-wide average nucleotide identity and high levels of synteny, representing a single species cluster. We observed strong conservation of gene clusters responsible for the synthesis of the secondary metabolites and bioactive peptides that are characteristic of HAB-forming cyanobacteria. All six G. echinulata genomes lack genes for the synthesis of classic cyanotoxins, including microcystin, but possess genes responsible for the synthesis of the taste and odor compound geosmin. Interestingly, the geoA geosmin synthase gene in three genomes is homologous to other cyanobacterial geoA genes, while the other three geoA genes are related to actinomyces geoA. Phylogenomic analysis places the G. echinulata genomes within a clade of benthic Nostocales, reflecting an ecological niche featuring extensive growth on the sediment surface before colonies disperse into the epilimnion for planktonic growth. We identify genes conserved in all six genomes that could represent physiological adaptations supporting active growth on sediments and pelagic recruitment independent of wind-driven mixing: phycoerythrin light harvesting complexes for optimal photosynthesis at depth; gliding motility to access patchy nutrient distributions; and gas vesicles with relatively small GvpC proteins that predict resistance to higher hydrostatic pressure. The strong genomic similarity across geographically distant populations suggests that G. echinulata in the United States is a tightly related non-toxigenic species group with predictable properties relevant to public health and drinking water management.

Cyanobacteria

Efficient homologous replacement and deletion of large genomic fragments through template-jumping prime editing in rice.

Homologous replacement of genomic sequences with large DNA fragments (> 100 bp) holds great potential for crop breeding, yet an efficient method to achieve such edits is lacking in plants. Here, in rice, we developed template-jumping prime editing (TJ-PE), a recently reported PE strategy for large targeted insertion, as an efficient tool for homologous replacement with DNA fragments ranging from dozens to hundreds of base pairs, and using TJ-PE, we replaced genomic fragments of up to 340 bp with homologous fragments of the same length. In addition, our TJ-PE tool also enabled precise deletion of 944- to 2024-bp fragments in rice, with efficiencies of up to 34.6% for c. 2000-bp precise deletions. Collectively, this study expands the editing scope of PE in rice and establishes TJ-PE as a generalist tool for precise deletion and replacement of large DNA fragments.

Oryza

Leveraging traveller genomics for LMIC diarrhoeal disease management.

Diarrhoeal pathogens impose a substantial global health burden, disproportionately affecting low- and middle-income countries (LMICs). However, in these settings, health-seeking behaviours, suboptimal microbiological capacity, and challenges in establishing genomics capacity constrain effective surveillance, including surveillance of antimicrobial resistance (AMR). In contrast, high-income countries routinely generate and share large volumes of diarrhoeal pathogen genomes through established systems, with a significant proportion originating from travellers returning from LMICs. These data reveal strong geographical structuring of lineages and clinically relevant AMR patterns, demonstrating untapped potential to support improvements in geographically granulated surveillance to support antimicrobial treatment recommendations. In this opinion article, we outline the potential to integrate traveller-derived microbial genomic data into LMIC public health decision-making and highlight the scientific, ethical, practical, and governance considerations for implementation.

antimicrobial resistance

Cytonuclear conflict and reticulate evolution in the Morelloid clade (Solanum, Solanaceae): Insights from genome skimming and network Phylogenomics.

The Morelloid clade (black nightshades) is one of the most strongly supported clades within the megadiverse Solanum genus. It comprises 76 globally distributed, non-spiny herbaceous and suffrutescent species. While often erroneously considered poisonous weeds, several species are economically important as orphan crops. The clade is closely related to tomato and potato but, due to a lack of focused breeding efforts, remains a putative reservoir of genetic diversity for crop improvement. Despite this potential, we lack fundamental knowledge on the evolution of the Morelloid clade. The group includes polyploid species with unknown parental origins-likely reflecting reticulate processes such as hybridization, introgression, and associated backcrossing events. Prior analyses have been unable to disentangle these processes, leaving the mechanisms underlying reticulate evolution in the Morelloid clade poorly understood. Here, we use genome skimming to produce a well-supported maximum likelihood plastid phylogeny from complete circularized plastomes and a coalescent-based species tree from combined Angiosperms353 and conserved ortholog set nuclear markers. Our dataset, composed of previously published data and deep genome skimming from herbarium samples, spans 26 Morelloid species. To investigate phylogenetic discordance, we used a nuclear phylogenetic network, multispecies coalescent simulations, a fused rooted nuclear chloroplast tree, and quantification of nuclear gene tree concordance. We show that incongruence between nuclear and plastid trees is pervasive and cannot be explained by incomplete lineage sorting alone. Instead, our results demonstrate that events consistent with repeated chloroplast capture have shaped the reticulate evolutionary history of the clade, especially among African polyploid and Pan-American diploid lineages.

Phylogeny

Ramu stunt virus genome reveals previously unreported segments and nucleocapsid domain duplication in Mechlorovirus.

Ramu stunt virus (RmSV), a member of the genus Mechlorovirus within the family Phenuiviridae, was previously described as a six-segmented RNA virus infecting sugarcane. In this study, we re-examined type material and additional isolates using high-throughput sequencing and RT-PCR validation, revealing that RmSV possesses a nine-segmented genome, making it the largest reported in the Phenuiviridae. This expanded architecture includes duplicated RNA segments (RNA 2a and RNA 2b) encoding nucleocapsid-like proteins and two novel segments (RNA 7 and RNA 8). Comparative analysis showed that RNA 2a and 2b share about 84% amino acid identity, while RNA 5 encodes a third nucleocapsid homolog, indicating unprecedented domain redundancy. Structural modeling confirmed that all three nucleocapsid proteins maintain a conserved fold despite low sequence identity, with electrostatic mapping suggesting differential RNA-binding potential. Additionally, RNA 6 encodes a hypothetical protein structurally similar to the rice stripe virus disease-specific S-protein, implicating a role in symptom development. Transcript abundance analysis revealed RNA 6 as the most highly expressed segment across isolates. These findings revise the genomic composition of RmSV, highlight mechanisms of genome plasticity and adaptive evolution in plant-infecting bunyaviruses, and underscore practical implications for diagnostic assay design, resistance breeding, and biosecurity surveillance.

Genome, Viral

Cost-Effectiveness and the Economics of Genomic Testing and Molecularly Matched Therapies.

Cost-effectiveness analysis of precision oncology can help guide value-driven care. Next-generation sequencing is increasingly cost-efficient over single gene testing because diagnostic algorithms require multiple individual gene tests to determine biomarker status. Matched targeted therapy is often not cost-effective due to the high cost associated with drug treatment. However, genomic profiling can promote cost-effective care by identifying patients who are unlikely to benefit from therapy. Additional applications of genomic profiling such as universal testing for hereditary cancer syndromes and germline testing in patients with cancer may represent cost-effective approaches compared with traditional history-based diagnostic methods.

Humans

Revealing the Shared Genetic Architecture of Metabolic Dysfunction-Associated Steatotic Liver Disease-Related Traits Through Genomic Structural Equation Modeling.

Although individual traits related to metabolic dysfunction-associated steatotic liver disease (MASLD) have been investigated through large-scale genome-wide association studies (GWASs), the shared genetic susceptibility across these traits remains unclear. We therefore conducted a multivariate GWAS of key MASLD-related traits to elucidate their common genetic architecture. We applied genomic structural equation modeling to model a latent genetic factor (MASLD-F) underlying genetically correlated MASLD-related traits, leveraging their GWAS-derived genetic correlations. We then performed functional annotations, including fine-mapping, transcriptome-wide association study, and cell- and tissue-type-specific enrichment analyses, and conducted Mendelian randomization analyses to identify modifiable risk factors. Our multivariate MASLD-F GWAS identified 50 independent variants across 48 genomic loci. Transcriptomic imputation identified several MASLD-F-associated genes, including ARNTL, NPC1, BTBD10, VDAC2, TSKU, SFMBT1, and ABHD17C. We observed significant enrichment of MASLD-F-related genetic signals predominantly in brain tissues, pancreatic islets, and the adrenal gland. Additionally, six modifiable risk factors and four modifiable protective factors for MASLD-F were identified. These findings reveal a complex shared genetic architecture underlying MASLD components, thereby expanding our understanding of disease pathogenesis and providing novel insights for precision medicine and public health interventions.

Humans

Clinical and endocrine correlates of genetic etiologies in severe hypospadias: Study from 34 patients.

OBJECTIVE: Hypospadias is a prevalent congenital anomaly (0.3%-1.0%); however, severe hypospadias (defined as proximal cases with the meatus at the penoscrotal junction, scrotum, or perineum) is a rare and clinically challenging entity with a multifactorial etiology. This study aimed to characterize the interrelationships among the clinical, endocrine, and genetic profiles in children with severe hypospadias. MATERIALS AND METHODS: We conducted a comprehensive analysis of 34 male patients with severe hypospadias. Preoperative hormone levels were measured using two methods: chemiluminescent immunoassay for luteinizing hormone and follicle-stimulating hormone, and liquid chromatography-tandem mass spectrometry for testosterone (T), dihydrotestosterone (DHT), dehydroepiandrosterone (DHEA), 17α-hydroxyprogesterone (17α-OHP), and other steroids. Genetic analysis was conducted via whole exome sequencing. RESULTS: The diagnostic yield of clinically relevant genetic variants (including pathogenic and likely pathogenic, and variants of uncertain significance) in our cohort was 41.2% (14/34) of patients. Patients carrying these variants exhibited a more complex phenotypic profile compared to non-carriers, including a significantly higher rate of patients with ≥3 associated malformations and a greater prevalence of cryptorchidism. Furthermore, the group with clinically relevant variants showed selective elevations in adrenal-derived precursors, specifically 17α-OHP and DHEA. Correlation analysis revealed significant positive associations of both 17α-OHP levels and the T/DHT ratio with the number of associated malformations. CONCLUSION: This study reveals significant genetic heterogeneity in patients with severe hypospadias. Those carrying genetic variants was associated with more severe clinical phenotypes, while certain endocrine variations, including the elevation of adrenal-derived hormones, were also observed in this cohort.

Humans