Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Genome-guided stage- and tissue-resolved transcriptome analysis of Serrodes campana identifies sex-biased antennal expression and candidate chemosensory-related genes.

Serrodes campana is an erebid moth of ecological and forestry relevance; its larvae are mainly associated with the soapberry tree, Sapindus mukorossi, whereas adults exhibit fruit-piercing behavior. However, stage- and tissue-resolved transcriptomic resources for this species remain limited. Here, using a chromosome-level reference genome, we performed a genome-guided transcriptome analysis of S. campana based on 12 RNA-seq libraries representing major developmental stages and key adult tissues. Global transcriptomic analyses revealed pronounced transcriptional differentiation across developmental stages and tissue types. Tissue-enriched gene sets and functional enrichment analyses identified distinct molecular signatures associated with developmental, sensory, and pheromone-associated tissues. Comparative analysis of female and male antennae further revealed sex-biased expression of several candidate chemosensory-related genes. Among 153 curated chemosensory-related candidate genes, most odorant receptor genes showed strong antennal enrichment, whereas other major chemosensory gene families displayed broader but still tissue-preferential expression patterns. In addition, an exploratory comparison of female terminal abdominal gland tissue and male terminal abdominal coremata revealed divergent expression profiles and highlighted candidate genes potentially associated with pheromone-related physiology, reproduction, and tissue-specific signaling. Together, this study provides the first genome-guided stage- and tissue-resolved transcriptomic resource for S. campana and offers a useful foundation for future studies of chemosensory detection, sex-biased gene expression, and pheromone-associated biology in this species.

Animals↗

A gap-free, telomere-to-telomere chromosome-scale genome assembly of the mangrove red snapper, Lutjanus argentimaculatus.

The mangrove red snapper (Lutjanus argentimaculatus) is a commercially important marine fish species in the Indo-Pacific region. Despite its significant economic value for aquaculture, existing genomic resources remain fragmented, limiting the advancement of molecular breeding and functional genomic studies. Here, we present a gap-free, telomere-to-telomere (T2T) genome assembly of L. argentimaculatus, generated using a hybrid approach combining PacBio HiFi, Oxford Nanopore ultra-long reads and Hi-C technology. The resulting assembly comprises exactly 24 scaffolds spanning 1.03 Gb, perfectly matching the haploid chromosome number with a contig N50 of 46.17 Mb. Notably, this assembly resolves all physical gaps present in previous versions, achieving a BUSCO completeness score of 98.2%. Comprehensive genome annotation successfully predicted 23,167 protein-coding genes. Among these, 22,067 genes (95.25%) were functionally annotated across major public databases, including eggNOG, InterPro, and Swiss-Prot. Furthermore, structural analysis successfully identified 19 telomeres and 20 centromeres, validating the chromosomal integrity. This high-fidelity, gap-free reference genome provides a robust foundation for comparative genomics, population genetics, and the genetic improvement of Lutjanidae species.

Animals↗

A genome-wide coverage-based pipeline for the identification of host-derived candidate DNA biomarkers from cell-free blood.

We have created a new data-analysis pipeline for the discovery of host-specific candidate DNA biomarkers derived from sequencing data of cell-free blood. Unlike approaches that rely on specific molecular or genetic signatures, our method leverages the coverage distribution of cell-free DNA sequences mapped to a reference genome, applying statistical analyses to identify informative short genomic regions for biomarker discovery. The pipeline is applicable to diverse diseases and can be used to analyze cell-free DNA sequences from plasma or serum to identify candidate biomarkers that are characteristic of disease states in mammals. Core functionalities were developed in Java and integrated with open-source software tools for the preprocessing of raw sequencing data, complemented by Python scripts for the machine-learning analysis and statistical validation. The pipeline is designed for HPC use and users can access the pipeline through a Galaxy workflow, which offers a user-friendly web interface for input selection prior to execution and analysis progress monitoring. Performance tests, carried out using duplicate sets of COVID-19 samples and controls, showed linear scalability of execution time with an increasing dataset size, as well as a substantial reduction in execution time through parallelized computation, whereby each HPC node is used to process the data of one chromosome. Further statistical tests confirmed the quality of the pipeline's results by showing that the set of identified candidate biomarkers remained stable across varying dataset sizes.

Biomarkers↗

Distinct evolutionary trajectories of subgenomic centromeres in polyploid wheat.

BACKGROUND: Centromeres are crucial for precise chromosome segregation and maintaining genome stability during cell division. However, their evolutionary dynamics, particularly in polyploid organisms with complex genomic architectures, remain largely enigmatic. Allopolyploid wheat, with its well-defined hierarchical ploidy series and recent polyploidization history, serves as an excellent model to explore centromere evolution. RESULTS: In this study, we perform a systematic comparative analysis of centromeres in common wheat and its corresponding ancestral species, utilizing the latest comprehensive reference genome assembly available. Our findings reveal that wheat centromeres predominantly consist of five types of centromeric-specific retrotransposon elements (CRWs), with CRW1 and CRW2 being the most prevalent. We identify distinct evolutionary trajectories in the functional centromeres of each subgenome, characterized by variations in copy number, insertion age, and CRW composition. By utilizing CENH3-ChIP data across various ploidy levels, we uncover a series of CRW invasion events that have shaped the evolution of AA subgenome centromeres. Conversely, the evolutionary process of the DD subgenome centromeres involves their expansion from diploid to hexaploid wheat, facilitating adaptation to a larger genomic context. Integration of complete einkorn centromere assemblies and Aegilops tauschii pan-genomes further revealed subgenome-specific centromere evolutionary trajectories. By inclusion of synthetic hexaploid from S2-S3 generations, alongside 2x/6 × natural accessions, we demonstrate that DD subgenome centromere expansion represents a gradual evolutionary process rather than an immediate response to polyploidization. CONCLUSIONS: Our study provides a comprehensive landscape of centromere adaptation, evolution, and maturation, along with insights into how retrotransposon invasions drive centromere evolution in polyploid wheat.

Centromere↗

Extensive editing of both hepatitis B virus DNA strands by APOBEC3 cytidine deaminases in vitro and in vivo.

Because the replication of hepatitis B virus (HBV) proceeds via an obligatory reverse transcription step in the viral capsid, cDNA is potentially vulnerable to editing by cytidine deaminases of the APOBEC3 family. To date only two edited HBV genomes, referred to as G --> A hypermutants, have been described in vivo. Recent work suggested that HBV replication was indeed restricted by APOBEC3G but by a mechanism other than editing. The issue of restriction has been explored by using a sensitive PCR method allowing differential amplification of AT-rich DNA. G --> A hypermutated HBV genomes were recovered from transfection experiments involving APOBEC3B, -3C, -3F, and -3G indicating that all four enzymes were able to extensively deaminate cytidine residues in minus-strand DNA. Unexpectedly, three of the four enzymes (APOBEC3B, -3F, and -3G) deaminated HBV plus-strand DNA as well. From the serum of two of four patients with high viremia, G --> A hypermutated genomes were recovered at a frequency of approximately 10(-4), indicating that they are, albeit relatively rare, part of the natural cycle of HBV infection. These findings suggest that human APOBEC3 enzymes can impact HBV replication via cytidine deamination.

APOBEC-3G Deaminase↗

Genome-Resolved Functional Profiling of Osteoporosis-Associated Gut Bacteria Highlights Putative Metabolic and Immunogenic Signatures of the Gut-Bone Axis.

The gut microbiota has emerged as a potential regulator of bone metabolism, but the genome-encoded functional repertoire of osteoporosis-associated gut bacteria remains insufficiently characterized. This study performed in silico functional profiling of gut bacterial taxa associated with osteoporosis, low bone mineral density, or comparator bone-related phenotypes. Twenty candidate taxa were selected from evidence in the human microbiome and represented by 26 curated bacterial reference genomes. Genome-wide annotations were used to map predicted gut-bone axis signatures, carbohydrate-active enzyme (CAZyme) repertoires, selected Kyoto Encyclopedia of Genes and Genomes pathways, and gutSMASH-predicted metabolic gene clusters. Functional burdens were normalized as hits per 1000 annotated proteins and integrated into metabolic, immunogenic, CAZyme, KEGG, and metabolic gene cluster profiles. Twelve predicted gut-bone axis signatures were identified, comprising 3337 primary candidate protein hits and a strict high-confidence subset of 2497 hits. Dominant signatures included vitamin B12/cobalamin metabolism, folate/one-carbon metabolism, peptidoglycan/cell-wall biosynthesis, and short-chain fatty acid-related functions. Dialister invisus, Dialister succinatiphilus, Megamonas funiformis, and Megamonas hypermegale showed the strongest normalized predicted gut-bone axis signal. These hypothesis-generating findings prioritize microbial metabolic and immunogenic features for future metagenomic, metabolomic, and experimental validation studies.

Osteoporosis↗

PNA targeting the PBS and A-loop sequences of HIV-1 genome destabilizes packaged tRNA3(Lys) in the virions and inhibits HIV-1 replication.

During assembly of the HIV-1 virions, cellular tRNA(Lys)(3) is packaged into the virion particles and is utilized as a primer for the initiation of reverse transcription. The 3'-terminal 18 nucleotides of the cellular tRNA(Lys)(3) are complementary to nucleotides 183-201 of the viral RNA genome, referred to as the primer binding sequence (PBS). Additional sequences (A-Loop) upstream of the PBS are essential for tRNA primer selection. We report here that a PNA targeted to PBS and A-Loop sequence (PNA(PBS)) exhibits high specificity for its target sequence and prevents tRNA(Lys)(3) priming on the viral genome. We also demonstrate that PNA(PBS) is able to invade the duplex region of the tRNA(Lys)(3)-viral RNA complex and destabilize the priming process, thereby inhibiting the in vitro initiation of reverse transcription. The endogenously packaged tRNA(Lys)(3) bound to the PBS region of the viral RNA genome in the HIV-1 virion is efficiently competed out by PNA(PBS), resulting in near complete inhibition of initiation of endogenous reverse transcription. Examination of the effect of PNA(PBS) on HIV-1 production in CEM cells infected with pseudotyped HIV-1 virions carrying luciferase reporter exhibited dramatic reduction of HIV-1 replication by nearly 99%. Analysis of the mechanism of PNA(PBS)-mediated inhibition indicated that PNA(PBS) interferes at the step of reverse transcription. These findings suggest the antiviral efficacy of PNA(PBS) in blocking the process of HIV-1 replication.

Base Sequence↗

Differential methylation of genes and retrotransposons facilitates shotgun sequencing of the maize genome.

The genomes of higher plants and animals are highly differentiated, and are composed of a relatively small number of genes and a large fraction of repetitive DNA. The bulk of this repetitive DNA constitutes transposable, and especially retrotransposable, elements. It has been hypothesized that most of these elements are heavily methylated relative to genes, but the evidence for this is controversial. We show here that repeat sequences in maize are largely excluded from genomic shotgun libraries by the selection of an appropriate host strain because of their sensitivity to bacterial restriction-modification systems. In contrast, unmethylated genic regions are preserved in these genetically filtered libraries if the insert size is less than the average size of genes. The representation of unique maize sequences not found in plant reference genomes is also greatly enriched. This demonstrates that repeats, and not genes, are the primary targets of methylation in maize. The use of restrictive libraries in genome shotgun sequencing in plant genomes should allow significant representation of genes, reducing the number of reactions required.

Cloning, Molecular↗

Telomere-to-telomere genome of Phoebe chekiangensis reveals that age-dependent CHG hypomethylation promotes floral transition via MADS-box gene activation.

Phoebe species are renowned for their highly valuable 'golden thread' timber; however, their protracted juvenile phase presents a significant obstacle to mechanistic investigations of floral induction. Phoebe chekiangensis, a rare early-flowering representative within this genus, provides a unique model system for dissecting the vegetative-to-reproductive phase transition. Nevertheless, the absence of a high-quality reference genome has severely hindered molecular insights into its developmental regulation. Here, we present the first telomere-to-telomere (T2T) genome assembly for P. chekiangensis, comprising two completely gap-free haplotypes with contig N50 values exceeding 65 Mb, base-level quality scores >36, and Long Terminal Repeat Assembly Index scores surpassing the gold standard threshold of 20. Approximately 29 000 genes were annotated per haplotype, supported by a BUSCO completeness score of >97%. Age-resolved transcriptomic landscapes identified two MADS-box transcription factors, PcMADS5 (AP1-like) and PcMADS19.1 (SOC1-like), as core activators of the floral transition. Both genes triggered precocious flowering when ectopically expressed in Arabidopsis thaliana. Whole-genome bisulfite sequencing revealed a progressive, age-dependent decline in CHG (where H is A, C, or T) DNA methylation, which was particularly pronounced at the PcMADS19.1 locus. Notably, DML1/2, which mediate active DNA demethylation, were coordinately upregulated during the onset of reproductive growth. Chemical demethylation using 5-azacytidine further diminished CHG methylation and selectively enhanced PcMADS19.1 expression, confirming a causal relationship between CHG hypomethylation and transcriptional activation. This work delivers the first chromosome-scale T2T genome within the genus Phoebe and uncovers CHG demethylation as a previously unrecognized epigenetic switch governing reproductive competence in woody perennials.

Journal Article↗

Wastewater-based sequencing of respiratory syncytial virus to investigate lineage dynamics and antigenic site mutations: a retrospective genomic epidemiology study.

BACKGROUND: Respiratory syncytial virus (RSV) infections pose a substantial health burden, particularly for clinically vulnerable populations such as infants and older adults. Although novel immunoprophylactic interventions show promise in providing protection, many countries may not have robust surveillance systems to monitor circulating RSV lineages and detect mutations that might reduce the effectiveness of these new interventions. We aimed to assess the diversity and temporal dynamics of circulating RSV lineages in urban populations through amplicon-based sequencing and analysis of wastewater extracts. METHODS: In this prospective observational wastewater-based genomic surveillance study, 32 raw influent 24-h composite samples were collected during the 2022-23 and 2023-24 RSV seasons from both Zurich and Geneva, Switzerland. We applied an RSV subtype-specific amplicon-based sequencing approach to obtain RSV-A and RSV-B sequences from all 64 samples. Mutations relative to reference genomes were identified at positions with read depth above 30. Relative abundances of RSV lineages were estimated from frequencies of lineage-signature mutations, present in greater than 90% of publicly available sequences of that lineage. FINDINGS: Relative abundances of RSV-B (2022-23) and RSV-A (2023-24) lineages were estimated over the two RSV seasons. During the 2022-23 season, the RSV-B B.D.E.1 lineage prevailed in both cities. In the 2023-24 season, multiple RSV-A lineages cocirculated, including A.D.1, A.D.3, A.D.5, and their sub-lineages. Identification and frequency estimation of mutations showed low-frequency, non-synonymous mutations in antigenic sites on the fusion gene of both RSV-A and RSV-B, some of which have not been reported in clinical sequences. The primary outcome was identification and relative abundance of RSV lineages in wastewater samples. INTERPRETATION: These findings show the potential of wastewater-based genomic surveillance to identify and track circulating RSV lineages and clinically relevant mutations. As novel RSV immunoprophylaxis measures are introduced in upcoming RSV seasons, wastewater-derived genomic RSV data provide a valuable baseline for understanding RSV diversity and future viral evolution under increased immunological pressure. FUNDING: This study was funded by the Swiss National Science Foundation and in part by the National Institute Of Allergy And Infectious Diseases of the National Institutes of Health. Funding for sample collection and processing was provided by the Swiss Federal Office of Public Health.

Humans↗

Partial mitochondrial genome sequences of Ostrinia nubilalis and Ostrinia furnicalis.

Contiguous 14,535 and 14,536 nt near complete mitochondrial genome sequences respectively were obtained for Ostrinia nubilalis and Ostrinia furnicalis. Mitochondrial gene order was identical to that observed from Bombyx. Sequences comparatively showed 186 substitutions (1.3% sequence divergence), 170 CDS substitutions (131 at 3(rd) codon positions), and an excess of transition mutation likely resulting by purifying selection (d(N)/d(S) = omega congruent with 0.15). Overall substitution rates were significantly higher at 4-fold (5.2%) compared to 2-fold degenerate codons (2.6%). These are the 3(rd) and 4(th) lepidopteran mitochondrial genome reference sequences in GenBank and useful for comparative mitochondrial studies.

Animals↗

Robustness of metabolic map reconstruction.

With the ever increasing amount of genomic data available, the interest for generating biochemical pathways has grown tremendously. So far, mainly complete genomes have been used to reconstruct the biochemical pathways and their associated interactions. However, a large number of low coverage genomes, as well as other sources of partial genomic data, are currently available for many organisms. In order to be able to use incomplete data for metabolic reconstruction, the inherent properties of this procedure need to be investigated. In this short note, we describe the robustness and predictive power of metabolic reconstructions using partial information from Schizosaccharomyces pombe. We also discuss the implications of the results on reference genome projects as well as other large-scale sequencing data.

Chromosome Mapping↗

The genetic map and comparative analysis with the physical map of Trypanosoma brucei.

Trypanosoma brucei is the causative agent of African sleeping sickness in humans and contributes to the debilitating disease 'Nagana' in cattle. To date we know little about the genes that determine drug resistance, host specificity, pathogenesis and virulence in these parasites. The availability of the complete genome sequence and the ability of the parasite to undergo genetic exchange have allowed genetic investigations into this parasite and here we report the first genetic map of T.brucei for the genome reference stock TREU 927, comprising of 182 markers and 11 major linkage groups, that correspond to the 11 previously identified chromosomes. The genetic map provides 90% probability of a marker being 11 cM from any given locus. Its comparison to the available physical map has revealed the average physical size of a recombination unit to be 15.6 Kb/cM. The genetic map coupled with the genome sequence and the ability to undertake crosses presents a new approach to identifying genes relevant to the disease and its prevention in this important pathogen through forward genetic analysis and positional cloning.

Animals↗

HSDSnake: a user-friendly SnakeMake pipeline for analysis of duplicate genes in eukaryotic genomes.

SUMMARY: Gene duplication is a well-known driver of molecular evolution-it acts as a source of genetic novelty, thereby providing the raw substrate for organismal adaption. However, detecting different types of gene duplicates and comparing them in sequence datasets can be difficult. Existing tools can identify and classify gene duplicates that have arisen by various processes, but have limitations; for example, some do not have a user-friendly workflow and can include many intermediate steps requiring manual adjustments of parameters and/or are not maintained for the benefit of research community members. Here, we have developed HSDSnake, a user-friendly SnakeMake pipeline that can detect and classify gene duplications into five categories: dispersed, proximal, tandem, transposed, and whole genome. It also curates and evaluates the highly similar gene duplicates (HSDs) in each gene duplication category with reliance on both sequence similarity and conserved domains. Lastly, the detected gene duplicates can be visualized within a KEGG functional pathway framework and the substitution rates (Ka, Ks, and their Ka/Ks ratio) can be analyzed for all the duplicate gene pairs. We demonstrate HSDSnake's capabilities by analyzing two reference genomes directly downloaded from NCBI and provide detailed instructions for each step. AVAILABILITY AND IMPLEMENTATION: The HSDSnake pipeline uses SnakeMake and Conda to run and install dependencies. The distribution version is available online at GitHub: https://github.com/zx0223winner/HSDSnake and the archived version at Zenodo is https://doi.org/10.5281/zenodo.15521945.

Software↗

ATAC-seq in Emerging Model Organisms: Challenges and Strategies.

The Assay for Transposase-Accessible Chromatin with sequencing (ATAC-seq) is a versatile and widely utilized method for identifying potential regulatory regions, such as promoters and enhancers, within a genome. ATAC-seq has been successfully applied to a wide range of established and emerging model organisms. However, implementing this method in emerging model systems, such as arthropods, can be challenging due to several factors that influence data quality. These factors include the availability of a sufficient amount and quality of tissue or cells, the need for species- and tissue-specific protocol optimization, the completeness and accuracy of the reference genome, and the quality of the genome annotation. In this article, we emphasize the key steps in the ATAC-seq protocol that, based on our experience, have the greatest impact on data quality when adapting this method for emerging model organisms. Specifically, we discuss the importance of nuclei isolation, the incubation conditions of the Tn5 transposase, and PCR amplification of the library. Furthermore, we outline essential quality checkpoints during the bioinformatic analysis of ATAC-seq data to assist in assessing data integrity and consistency. Given that many emerging model organisms may not be readily available in laboratory cultures, we also emphasize the importance of evaluating how different preservation methods affect ATAC-seq data quality. Based on examples in one spider and one ant species, we demonstrate that replication and thorough quality controls at all steps of the protocol and data analysis are essential to assess the usability of ATAC-seq data. Our data highlights the importance of isolating the right number of intact nuclei, as well as ensuring optimal amplification conditions during library preparation to obtain good-quality sequence data for downstream analyses. We recommend using fresh tissue samples if possible because we show that direct cryopreservation of the tissue may affect chromatin integrity. This effect could be avoided or reduced by preserving the homogenate in cell culture medium. Overall, we explain the ATAC-seq protocol and downstream analyses in detail and give step-by-step advice to researchers who are new to the field and want to implement this method. With careful planning and validation, ATAC-seq can reveal the regulatory landscape of a genome and aid in identifying elements that govern gene expression.

Animals↗

Reduced replication of human immunodeficiency virus type 1 mutants that use reverse transcription primers other than the natural tRNA(3Lys).

Replication of the human immunodeficiency virus type 1 (HIV-1) and other retroviruses involves reverse transcription of the viral RNA genome into a double-stranded DNA. This reaction is primed by the cellular tRNA(3Lys) molecule, which binds to a complementary sequence in the viral genome, referred to as the primer-binding site (PBS). In order to study the specificity of primer usage, we constructed a set of HIV-1 mutants with altered PBS sites corresponding to other tRNA species (tRNA(Ile), tRNA(1,2Lys), tRNA(Phe), tRNA(Pro), tRNA(Trp)). These mutant viruses were able to replicate, although with delayed replication kinetics compared with wild-type HIV-1. Identification of the tRNA species associated with the genomic RNA demonstrated binding of tRNAs complementary to the new PBS sites. However, the occupancy of the mutant PBS sites by these new primers was reduced and correlated well with the replication potential of the mutant viruses. These results suggest that the PBS sequence is not sufficient for annealing of the tRNA primer. Upon prolonged culturing, all mutants reverted to the wild-type PBS(3Lys) sequence. Minor sequence changes in the nucleotides flanking the PBS site indicate that these reversions resulted from annealing of the wild-type tRNA(3Lys) primer onto the mutant PBS sites, followed by copying of part of the tRNA(3Lys) sequence during reverse transcription. Furthermore, the reversion efficiency of the different PBS mutants was found to correlate with their tRNA(Lys)3 binding capacity. A remarkable reversion pathway was observed for the PBSPro variant (PBSPro-->PBSIle-->PBSwt). This pathway can be explained by efficient base pairing of tRNA(Ile) to PBSPro, followed by annealing of tRNA(3Lys) onto the PBSIle intermediate. These results demonstrate that HIV-1 is dedicated to the tRNA(3Lys) primer and that factors other than the PBS sequence determine the selective primer usage of this retrovirus.

Base Sequence↗

Haplotype-aware long-read error correction.

Error correction of long reads is an important initial step in genome assembly workflows. For organisms with ploidy greater than one, it is important to preserve haplotype-specific variation during read correction. This challenge has driven the development of several haplotype-aware correction methods. However, existing methods are based on either ad-hoc heuristics or deep learning approaches. In this paper, we introduce a rigorous formulation for this problem. Our approach builds on the minimum error correction framework used in reference-based haplotype phasing. We prove that the proposed formulation for error correction of reads in de novo context, i.e., without using a reference genome, is NP-hard. To make our exact algorithm scale to large datasets, we introduce practical heuristics. Experiments using PacBio HiFi sequencing datasets from human and plant genomes show that our approach achieves accuracy comparable to state-of-the-art methods. Implementation: https://github.com/at-cg/HALE .

Clustering↗

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software↗