Search PubMedSearch

SEARCH · Search PubMed

Results for “genome streamlining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Adaptive deletion of functional duplicate genes in Drosophila.

Gene deletion is traditionally viewed as a nonadaptive mechanism that eliminates functional redundancy, yet emerging evidence indicates that it disproportionately affects tissue-specific duplicates with unique functions. Here, we test whether gene deletion preferentially removes weakly constrained, degenerating duplicates or instead eliminates functionally active duplicates through an adaptive process. To identify the evolutionary and functional factors that determine which duplicates are lost, we systematically analyzed 100 gene deletion events in Drosophila by integrating sequence, expression, interaction, and structural data. We uncovered a strong bias toward the loss of younger child copies among functionally unique duplicates, whereas no such bias was observed for redundant duplicates. Contrary to expectations under relaxed constraint, deleted functionally unique genes evolve more slowly, show higher expression, engage in more protein-protein interactions, and do not exhibit elevated structural divergence or intrinsic disorder relative to redundant duplicates. When compared with single-copy genes, deleted functionally unique genes display similar evolutionary rates, slightly lower expression, greater network connectivity, comparable structural divergence, and lower intrinsic disorder. These patterns suggest that deletion frequently affects functionally active rather than degenerate genes. Collectively, our results support the hypothesis that gene deletion in Drosophila can represent an adaptive process acting on transiently functional duplicates, potentially driven by either genome streamlining or context-dependent deleterious effects.

evolution

Complete genomes from a xenic Dolichospermum flosaquae FBCC-A233 culture reveal genome-inferred metabolic asymmetry with associated bacteria.

Cyanobacteria form phycosphere communities with associated bacteria, but genome-resolved resources are needed to formulate testable hypotheses about their metabolic interactions. Here, we reconstructed three complete circular genomes from a unialgal xenic culture, including Dolichospermum flosaquae FBCC-A233 and two associated alphaproteobacterial genomes assigned to Sphingorhabdus sp. and Brevundimonas sp. Genome-wide read mapping and genome-quality assessment supported the three recovered genomes as high-quality circular reconstructions. Comparative genome analysis placed the cyanobacterial genome within the Dolichospermum flosaquae species cluster under the GTDB framework, while the associated bacterial genomes represented Sphingorhabdus sp. and a putative undescribed Brevundimonas species-level lineage. Genome architecture analysis indicated reduced genome size and gene content in Brevundimonas relative to genus-level references although additional metrics did not support a strong conclusion of classical genome streamlining. Selected KEGG module and KO-level reconstructions indicated genome-inferred metabolic asymmetries across the consortium. FBCC-A233 encoded photosynthesis- and nitrogen-related modules and a BioU-mediated de novo biotin biosynthesis route, whereas the associated bacteria lacked complete de novo biotin biosynthesis but retained biotin-dependent carboxylase genes. FBCC-A233 also encoded extensive anaerobic corrinoid biosynthesis potential; however, canonical DMB-containing cobalamin completion, cobamide identity, and complete transporter systems were not resolved. Together, these complete genomes provide a genome-resolved resource for investigating genome-inferred metabolic differentiation and ecological interactions in cyanobacteria-associated bacterial consortia.IMPORTANCEPhycosphere interactions between cyanobacteria and associated bacteria can shape aquatic microbial communities, but many proposed interactions remain difficult to evaluate without genome-resolved resources. This study provides three complete circular genomes from a unialgal xenic Dolichospermum flosaquae culture, capturing the cyanobacterium and two co-maintained bacterial associates. Our analysis identifies genome-inferred metabolic asymmetries, particularly in biotin- and cobamide-related pathways. D. flosaquae FBCC-A233 encoded candidate de novo biotin and corrinoid biosynthesis capacity, whereas the associated bacteria lacked complete de novo pathways but retained cofactor-dependent enzymes. These findings nominate cofactor-related dependencies as experimentally testable hypotheses while emphasizing unresolved uptake, export, cobamide identity, and growth-dependence mechanisms. The complete genomes and KO-level reconstructions generated here provide a resource for future studies of cyanobacteria-associated consortia.

Genome, Bacterial

Fructophilic lactic acid bacteria as a window into multi-scale convergent evolution.

Fructophilic lactic acid bacteria (FLAB) are a group of lactic acid bacteria with unique growth characteristics, that is, poor growth on glucose. Their growth is enhanced in the presence of fructose or external electron acceptors. These organisms inhabit fructose-rich environments such as flowers, fruits, and pollinating insects, particularly honey bees. Apilactobacillus spp. and Fructobacillus spp. are representatives of FLAB, although they belong to phylogenetically distant clades. These organisms commonly possess markedly small genomes with a low number of coding DNA sequences. Furthermore, their genomes are characterized by a markedly reduced number of genes involved in carbohydrate transport and metabolism. Genome reduction in FLAB reflects convergent adaptation to fructose-rich environments rather than general genome streamlining. The two distinct FLAB genera, Fructobacillus and Apilactobacillus, independently lost more than 100 genes in statistically similar orders. In contrast, genes involved in carbohydrate and amino acid metabolism exhibited reversed orders of loss between the two genera. Furthermore, FLAB genomes lack an intact bifunctional alcohol/aldehyde dehydrogenase gene (adhE), which causes their poor growth on glucose. A comparative genomic study suggested the evolutionary process underlying adhE gene decay during adaptation to the fructose-rich environments, including pollinating insects. In conclusion, FLAB represent a unique example of habitat-driven convergent reductive evolution that can be investigated across multiple biological scales - from individual genes to whole genomes - in the diverse LAB group with a wide range of habitats, and partially share the fructophilic evolution with eukaryotic yeasts found in fructose-rich habitats.

Fructose

MAFin: motif detection in multiple alignment files.

MOTIVATION: Whole Genome and Proteome Alignments, represented by the multiple alignment file format, have become a standard approach in comparative genomics and proteomics. These often require identifying conserved motifs, which is crucial for understanding functional and evolutionary relationships. However, current approaches lack a direct method for motif detection within MAF files. We present MAFin, a novel tool that enables efficient motif detection and conservation analysis in MAF files to address this gap, streamlining genomic and proteomic research. RESULTS: We developed MAFin, the first motif detection tool for Multiple Alignment Format files. MAFin enables the multithreaded search of conserved motifs using three approaches: (i) using user-specified k-mers to search the sequences. (ii) with regular expressions, in which case one or more patterns are searched, and (iii) with predefined Position Weight Matrices. Once the motif has been found, MAFin detects the motif instances and calculates the conservation across the aligned sequences. MAFin also calculates a conservation percentage, which provides information about the conservation levels of each motif across the aligned sequences, based on the number of matches relative to the length of the motif. A set of statistics enables the interpretation of each motif's conservation level, and the detected motifs are exported in JSON and CSV files for downstream analyses. AVAILABILITY AND IMPLEMENTATION: MAFin is offered as a Python package under the GPL license as a multi-platform application and is available at: https://github.com/Georgakopoulos-Soares-lab/MAFin.

Software

A streamlined base editor engineering strategy to reduce bystander editing.

Base editing (BE) can permanently correct over half of known human pathogenic genetic variants without requiring a repair template, thus serving as a promising therapeutic tool to treat a broad spectrum of genetic diseases. However, the broad activity windows of current base editors pose a major challenge to their therapeutic application. Here, we show that integrating a naturally occurring oligonucleotide binding module into the deaminase active center of TadA-8e, a highly active deoxyadenosine deaminase, enhances its editing specificity. When conjugated with a Cas9 nickase or alternative PAM Cas9 variants, the engineered TadA variant-TadA-NW1-consistently achieves robust A-to-G editing efficiencies within an editing window consisting of four nucleotides, substantially narrower than the 10-bp editing window of the TadA-8e-derived ABEs. Moreover, compared to ABE8e, ABE-NW1 shows significantly decreased Cas9-dependent and -independent off-target activity while maintaining similar on-target editing efficiency. Further, TadA-NW1 can be reprogrammed to perform desired cytidine deamination and adenine transversion within a restricted editing window. Finally, in a cystic fibrosis (CF) cell model, ABE-NW1 outperforms existing ABEs in accurately and efficiently correcting the CFTR W1282X variant, one of the most common CF-causing mutations. In all, we engineered a suite of base editors with refined activity windows, enabling more precise base editing. Importantly, this study presents a streamlined genome editor re-engineering strategy to accelerate the development of therapeutic base editing.

Gene Editing

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans

Genome-scale insights into metabolic streamlining and photosynthetic energy balance in the extremophile green alga Picocystis salinarum (Picocystophyceae, Chlorophyta).

Picocystis salinarum is an early-diverging chlorophyte and the sole described member of the Picocystophyceae, frequently dominating hypersaline and alkaline lakes despite extreme physicochemical constraints. To elucidate the genomic foundations of its ecological success, we generated a fully annotated, chromosome-scale nuclear genome assembly of the type strain originally isolated from a saline pond in San Francisco Bay. The 18.5-Mb genome comprises 30 chromosomal assemblies, exhibits clear diploidy, and contains multiple copies of intact Ty3/Gypsy and Ty1/Copia long terminal repeat retrotransposons encoding polyproteins with atypical accessory domains. Phylogenomic analyses reveal strong affinity with the Nephroselmidophyceae. Comparative analyses reveal extensive metabolic streamlining, including the absence of a queuosine salvage pathway, the 2-methylcitrate cycle, β-oxidation of propionate, and branched-chain amino acid catabolism, traits retained in several marine prasinophyte lineages. In contrast, the genome preserves multiple ancestral bacterial derived systems. Notably, P. salinarum features a complete chloroplast NADH dehydrogenase-like complex, including all membrane, electron binding, and assembly components, a configuration not previously reported in sequenced chlorophyte algae. This retention implies substantial capacity for cyclic electron flow and chlororespiration, processes expected to be critical in chronically low-light and chemically extreme environments. The genome further reveals a distinctive biochemical CO2-concentrating mechanism centered on plastid-targeted phosphoenolpyruvate carboxykinase, complete plastid peptidoglycan biosynthetic and remodeling pathways, and partial retention of lipid-A-related machinery. Conversely, P. salinarum lacks canonical non-photochemical quenching proteins while retaining xanthophyll-cycle enzymes that support slower photoprotective responses. Together, these features define a coordinated genomic architecture that underpins the specialization of P. salinarum to hypersaline, alkaline, and persistently low-light ecosystems.

3‐deoxy‐D‐manno‐octulo

Robust and highly efficient transformation method for a minimal mycoplasma cell.

UNLABELLED: Mycoplasmas have been widely investigated for their pathogenicity, as well as for genomics and synthetic biology. Conventionally, transformation of mycoplasmas was not highly efficient, and due to the low transformation efficiency, large amounts of DNA and recipient cells were required for that purpose. Here, we report a robust and highly efficient transformation method for the minimal cell JCVI-syn3B, which was created through streamlining the genome of Mycoplasma mycoides. When the growth states of JCVI-syn3B were examined in detail by focusing on such factors as pH, color, absorbance, colony forming unit, and transformation efficiency, it was found that the growth phase after the lag phase can be divided into three distinct phases, of which the highest transformation efficiency was observed during the early exponential growth phase. Notably, the transformation efficiency of up to 4.4 × 10-2 transformants per cell per microgram of plasmid DNA was obtained. A method to obtain several hundred to several thousand transformants with less than 0.2 mL of culture with approximately 1 × 107-108 cells and 10 ng of plasmid DNA was developed. Moreover, a transformation method using a frozen stock of transformation-ready cells was established. These procedures and information could simplify and enhance the transformation process of minimal cells, facilitating advanced genetic engineering and biological research using minimal cells. IMPORTANCE: Mycoplasmas are parasitic and pathogenic bacteria for many animals. They are also useful bacteria to understand the cellular process of life and for bioengineering because of their simple metabolism, small genomes, and cultivability. Genetic manipulation is crucial for these purposes, but transformation efficiency in mycoplasmas is typically quite low. Here, we report a highly efficient transformation method for the minimal genome mycoplasma JCVI-syn3B. Using this method, transformants can be obtained with only 10 ng of plasmid DNA, which is around one-thousandth of the amount required for traditional mycoplasma transformations. Moreover, a convenient method using frozen stocks of transformation-ready cells was established. These improved methods play a crucial role in further studies using minimal cells.

Transformation, Bacterial

Functional unknomics of the SAR11 clade reveal hidden genetic potential underlying adaptation to bottom-up and top-down pressures.

UNLABELLED: A substantial fraction of the genes in bacteria lack detectable sequence similarity to genes with known functions. These functionally uncharacterized genes-collectively referred to as the "unknome"-represent a largely unexplored genetic repertoire harboring insights into marine bacterial ecology. In this study, we explored the function of the unknome of the SAR11 clade, the most abundant bacterial lineage in the ocean, with a particular focus on genes that provide insight into its ecology. Based on the Clusters of Orthologous Genes and Kyoto Encyclopedia of Genes and Genomes classifications, approximately 56% of SAR11 ortholog groups were classified as members of the unknome. Among the SAR11 unknome, we successfully inferred the functions of 57 ortholog groups that are conserved in the SAR11 clade by protein structure similarity searches and genomic context analyses. These ortholog groups include putative transporter components, supporting the current ecological understanding that the SAR11 clade is specialized in substrate uptake to adapt to oligotrophic marine environments. Furthermore, structural analysis indicated that the DUF2237-containing protein, enriched in marine environments, may interact with purine nucleotide-containing compounds. This may suggest the existence of unique nucleotide utilization mechanisms in marine bacteria. In addition, we identified candidate viral defense systems within the unknome, indicating that diverse defense systems are present in at least one-third of cultured SAR11 strains. The presence of these defense systems, even within streamlined SAR11 genomes, suggests that they confer significant ecological advantages. Our analyses provide insights into the genetic basis of bottom-up processes (adaptation to oligotrophic environments) and top-down processes (antiviral defense strategy) contributing to ecological success. IMPORTANCE: Many microbial genes have no experimentally established function, limiting our ability to explain how microorganisms adapt to their environments. We examined this uncharacterized gene space, or "unknome" in SAR11, the most abundant bacterial clade in the ocean, by integrating evolutionary conservation, genomic context, predicted protein structure, and environmental distribution. This approach enabled us to prioritize components of the SAR11 unknome, including a core unknome conserved across the clade and genes enriched in specific lineages, and to identify several candidates with possible ecological roles in nutrient acquisition and defense against viruses. Our results suggest that the SAR11 unknome contains important clues to the ecological success of SAR11 rather than merely reflecting incomplete annotation or gene-prediction artifacts. Our study highlights the potential value of unknome analysis for identifying ecologically relevant genes in environmental microorganisms.

Pelagibacterales

Acquisition and erosion of toxin-antitoxin systems in bacterial chromosomes.

Toxin-antitoxin systems (TAs) are widespread in bacterial genomes. Yet, their integration, persistence, and impact in chromosome dynamics remain unclear. Here, we identified 80 type II TAs in the single chromosome of Photorhabdus laumondii TT01, 50 of which were experimentally validated. Comparative analysis across the Photorhabdus genus revealed a highly heterogeneous distribution, with TAs frequently clustering within discrete genomic regions, either alone or associated with cointegrate-forming transposases and integrases. TAs rarely clustered with other putative defense systems and are preferentially associated with different types of recombinases, suggesting distinct pathways of acquisition for the two types of functions. Functional analyses showed that most validated TAs display addictive properties and stabilize plasmids. These addictive TAs are preferentially located in genomic regions characterized by high gene turnover, consistent with recent acquisition events. Despite their plasmid-stabilizing capacity, TAs do not promote long-term conservation of their immediate chromosomal neighborhoods. Instead, we observed frequent TA loss, either through complete deletion or toxin pseudogenization, indicating relaxed selection for their persistence in bacterial lineages. We propose a stepwise model for TA evolution in bacterial chromosomes: initial acquisition mediated by mobile genetic elements, preferential integration into permissive genomic regions, subsequent genetic streamlining of linked loci, and progressive gene loss. The short-lasting linkage between TAs and their genomic neighborhoods is consistent with the view that TA modules can behave as autonomous, selfish genetic elements.

Journal Article

Private detection of relatives in forensic genomics using homomorphic encryption.

BACKGROUND: Forensic analysis heavily relies on DNA analysis techniques, notably autosomal Single Nucleotide Polymorphisms (SNPs), to expedite the identification of unknown suspects through genomic database searches. However, the uniqueness of an individual's genome sequence designates it as Personal Identifiable Information (PII), subjecting it to stringent privacy regulations that can impede data access and analysis, as well as restrict the parties allowed to handle the data. Homomorphic Encryption (HE) emerges as a promising solution, enabling the execution of complex functions on encrypted data without the need for decryption. HE not only permits the processing of PII as soon as it is collected and encrypted, such as at a crime scene, but also expands the potential for data processing by multiple entities and artificial intelligence services. METHODS: This study introduces HE-based privacy-preserving methods for SNP DNA analysis, offering a means to compute kinship scores for a set of genome queries while meticulously preserving data privacy. We present three distinct approaches, including one unsupervised and two supervised methods, all of which demonstrated exceptional performance in the iDASH 2023 Track 1 competition. RESULTS: Our HE-based methods can rapidly predict 400 kinship scores from an encrypted database containing 2000 entries within seconds, capitalizing on advanced technologies like Intel AVX vector extensions, Intel HEXL, and Microsoft SEAL HE libraries. Crucially, all three methods achieve remarkable accuracy levels (ranging from 96% to 100%), as evaluated by the auROC score metric, while maintaining robust 128-bit security. These findings underscore the transformative potential of HE in both safeguarding genomic data privacy and streamlining precise DNA analysis. CONCLUSIONS: Results demonstrate that HE-based solutions can be computationally practical to protect genomic privacy during screening of candidate matches for further genealogy analysis in Forensic Genetic Genealogy (FGG).

Humans

A streamlined protocol for small-scale protoplast generation and CRISPR/Cpf1-mediated genome editing in Fusarium oxysporum.

Fusarium oxysporum is a significant threat to agriculture and One Health, requiring advanced molecular tools for functional genomic analyses and biological control agent development. Existing gene-editing methods are hampered by costly protoplast preparation protocols and by CRISPR-Cas9 limitations, such as restricted protospacer adjacent motif (PAM) sequences and complex guide RNA requirements. We engineered an efficient CRISPR/Cpf1 system that overcomes these issues through three main innovations: small-scale protoplast generation using filter column-based methods that greatly reduce enzyme consumption while simplifying workflows, a CRISPR/Cpf1 system with shorter guide RNA design and staggered DNA cleavage to promote homologous recombination, and minimal homology arm strategies that significantly decrease cloning complexity. Extensive validation confirms successful gene targeting with molecular verification and functional analysis via standardized pathogenicity assays. This integrated platform offers affordable, accessible tools for systematic F. oxysporum research, enhancing fundamental understanding of plant-pathogen interactions and supporting high-throughput screening vital for agricultural biotechnology and biological agent development.

CRISPR/Cpf1

Single-nucleotide transcription start sites profiling via Nascent Strand-Specific RNA sequencing uncovers IFN-γ-induced promoter dynamics.

Transcriptional regulation is a highly dynamic process in which nascent RNAs provide the most immediate readout of transcriptional activity. Precise mapping of transcription start sites (TSSs) is therefore critical for understanding promoter architecture and gene regulation, yet remains technically challenging. Here, we introduce Nascent Strand-Specific RNA sequencing (NSS-seq), a robust and streamlined method for genome-wide profiling of the capped 5' ends of nascent RNAs. By directly capturing transcription initiation events, NSS-seq overcomes the temporal delay inherent to conventional RNA-seq and enables time-resolved interrogation of transcriptional dynamics. Applied to interferon-γ (IFN-γ)-stimulation, NSS-seq uncovers previously unrecognized IFN-γ-responsive genes and transient transcription factor activation patterns underlying interferon-mediated tumor-suppressive functions. Together, NSS-seq provides a cost-effective and technically accessible platform for dissecting promoter-level regulatory dynamics during cellular responses.

Promoter Regions, Genetic

CREAT: A CRISPR-Based Genome Trimming Strategy for Systematic Identification of Dispensable Regions and Rapid Genome Reduction.

The construction of minimal-genome microbes offers an ideal platform for understanding fundamental biological processes and synthetic biology, yet the research is hindered by incomplete lists of essential genes in microbes and by multiple rounds of genome trimming with a trial-and-error nature. To address this, we introduce CREAT (CRISPR-based genome trimming with a multi-homology-arm template)-a streamlined approach that integrates CRISPR-targeted genome cleavage and homology arm walking to classify essential from non-essential genomic subregions, thus providing the basis for predicting essential genes in a given organism. These essential genes were then assembled into synthetic gene cassettes for one-step replacement of the targeted non-deletable genomic regions for further genome trimming. Eight consecutive rounds of CREAT genome trimming achieved a 20.8% reduction in genome size in Saccharolobus islandicus. Furthermore, Cas9-based CREAT genome trimming was developed for Bacillus subtilis and Escherichia coli, with efficiency greatly enhanced by the λ-Red recombinase in the latter. Together, this iterative application of CREAT provides a scalable and generally applicable strategy for rapidly constructing minimal genomes across diverse microorganisms.

CRISPR-Cas Systems

ChemGenXplore: an interactive tool for exploring and analysing chemical genomic data.

MOTIVATION: Chemical genomics is a powerful high-throughput approach to systematically link phenotypes to genotypes. However, the vast datasets generated remain challenging to explore due to the lack of integrated, interactive tools for visualization and analysis. Existing workflows often require multiple independent software tools, limiting data accessibility and collaboration. Therefore, we created a user-friendly platform that enables efficient exploration and sharing of chemical genomics data. RESULTS: We developed ChemGenXplore, a web-based Shiny application designed to streamline the visualization and analysis of chemical genomic screens. It offers two primary functionalities: one for exploring pre-implemented datasets and another for analysing user-uploaded datasets. ChemGenXplore enables users to visualize phenotypic profiles, assess gene-gene and condition-condition correlations, perform GO and KEGG enrichment analysis, and generate customizable, interactive heatmaps. To further support collaborative research, ChemGenXplore also facilitates the comparative analysis of chemical genomic and other omics datasets. By consolidating these features into a single interactive and accessible tool, ChemGenXplore facilitates data sharing, enhances reproducibility, and promotes collaboration within the research community. AVAILABILITY AND IMPLEMENTATION: ChemGenXplore is freely accessible as a web application at https://chemgenxplore.kaust.edu.sa/. Source code and documentation, including instructions for local installation, are provided on GitHub (https://github.com/Hudaahmadd/ChemGenXplore). A Docker image is also available on DockerHub (https://hub.docker.com/r/hudaahmad/chemgenxplore) to ensure reproducibility and simplify installation.

Software

insilicoSV: a flexible grammar-based framework for structural variant simulation and placement.

SUMMARY: Structural variants (SVs) are key drivers of genetic variation and disease in the genome. Their discovery remains challenging, however, in large part due to the scarcity of validated SV callsets and comprehensive benchmarks, which are essential for method development and evaluation. The growing number of data-driven learning-based approaches for SV discovery, in particular, requires large, diverse, and well-balanced training datasets to achieve reliable performance. To address this need, SV simulation has served as a key tool for assessing method performance and training SV models. However, existing SV simulators only support a fixed and limited set of SV classes and do not provide fine-grained control over the placement of SVs within specific contexts of the genome. Here we present insilicoSV, a versatile framework for SV simulation, which models SVs using a simple and flexible grammar, allowing users to easily define standard and custom arbitrary genome rearrangements, as well as encode genome placement constraints. This design allows insilicoSV to naturally support new and bespoke SV types, such as the complex rearrangements of cancer genomes. In addition to grammar-based modeling, insilicoSV provides built-in support for 26 predefined SV types, placement of user-provided SVs, small variant simulation, streamlined workflows for the simulation of genome evolution and genome mixtures, read simulation, alignment, and visualization. These features enable the creation of comprehensive genomic datasets for a variety of downstream applications, such as in-depth benchmarking of alignment and variant calling methods, as well as training of data-driven learning-based approaches for SV detection. AVAILABILITY AND IMPLEMENTATION: insilicoSV is available under the MIT license at https://github.com/PopicLab/insilicoSV and https://doi.org/10.5281/zenodo.17402009.

Software

An open-source clinical bioinformatics pipeline for real-world NGS implementation: translating genomic variants into actionable treatment strategies in oncology.

BACKGROUND: Next-Generation Sequencing (NGS) has become a cornerstone technology in clinical practice, yet its adoption presents significant challenges. Physicians and oncologists must manage vast amounts of genome-scale data and transform it into actionable insights for complex decision-making. While commercial systems exist to synthesize data from NGS experiments into clinical reports, many are hindered by limitations such as closed-source designs that restrict transparency and customization. Additionally, some fail to leverage publicly available genomic databases, missing opportunities to integrate valuable external data. Furthermore, the rigidity of many tools in accommodating diverse NGS panels limits their applicability across varied clinical scenarios. METHODS: To address these limitations, we developed OncoReport, an open-source tool that generates comprehensive reports from NGS analyses. By integrating publicly accessible databases, OncoReport provides a robust, user-friendly environment equipped with essential tools for NGS analysis. This design aims to enhance data interpretation and support informed clinical decision-making. RESULTS: Rigorous testing has demonstrated OncoReport’s effectiveness in producing detailed, actionable reports that are clear and easy to use. By automating key aspects of the workflow, the tool significantly reduces manual effort and expedites the synthesis and interpretation of NGS results, making genomic insights more accessible to clinicians. CONCLUSION: OncoReport offers a transparent, flexible, and efficient framework for clinicians to analyze and apply genomic data in patient care. By streamlining workflows and leveraging open-source principles, it empowers healthcare professionals to make informed, data-driven decisions. OncoReport is freely available at https://oncoreport.atlas.dmi.unict.it, with source code and issue tracking on GitHub: https://github.com/knowmics-lab/oncoreport .

Humans

An inducer-independent, single-plasmid CRISPR-Cas9 system for genome editing in Bacillus species.

Advances in molecular biology tools are essential for streamlining and accelerating genetic engineering of cells across industrial and academic applications. While CRISPR-Cas improves genome editing efficiency, current systems have limitations and are often host specific, which restricts their versatility. This study describes a versatile CRISPR-Cas9 system for genome editing in industrially relevant Bacillus species. By adapting the well-established pJOE8999 vector-based CRISPR-Cas9 genome editing system, we constructed an inducer-independent, broad-host-range genome editing system. It maintains the benefits of low toxicity to the target cell and the cloning host as well as the ease to use of a single-plasmid CRISPR-Cas9 system. We utilized the constitutive Sigma70-type promoter from the conserved veg gene of Bacillus, to develop and test the suitability of promoter variants of different strengths for Cas9 expression. Successful gene deletions in three different Bacillus species demonstrated the versatility of the modified system for this industrially important genus. This was further confirmed by the integration of a reporter gene fusion and the introduction of a single point mutation in the genome of Bacillus licheniformis. This one-step CRISPR-based transformation protocol developed in this study enables fast genome editing workflows with minimal hands-on time. KEY POINTS: • Editing and screening of promoter variants for balanced Cas9 expression in Bacillus. • Development of a versatile inducer-independent, single-plasmid CRISPR-Cas-based system. • Verification of the modified CRISPR-based system for genome editing in different Bacilli.

CRISPR-Cas Systems