Search PubMedSearch

SEARCH · Search PubMed

Results for “genome streamlining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Molecular characterization of archival adrenal tumor tissue from patients with ACTH-independent Cushing syndrome.

Cushing syndrome represents a multitude of signs and symptoms associated with long-term and excessive exposure to glucocorticoids. Solitary cortisol-producing adenomas (CPAs) account for most cases of ACTH-independent Cushing syndrome (CS). Technological advances in next-generation sequencing have significantly increased our understanding about the genetic landscape of CPAs. However, the conventional approach utilizes fresh/frozen tissue samples, which are not routinely available for most clinical adrenal adenoma specimens. This coupled with the fact that CS is relatively rare reduces the accessibility to CPAs for research. In order to circumvent this issue, our group recently developed a sequencing strategy that allowed the use of formalin-fixed paraffin-embedded (FFPE) CPA samples for mutation analysis. Our streamlined approach includes the visualization and genomic DNA (gDNA) capture of the cortisol-producing regions in the tumor using immunohistochemistry (IHC)-guided techniques followed by targeted and/or whole-exome sequencing analysis. This approach has the advantage of using both prospective and retrospective CPA cohorts since FFPE pathologic specimens are routinely banked. This review discusses this advanced approach using IHC-guided gDNA capture of pathologic tissue followed by NGS as a preferred method for mutational analysis of CPAs.

Humans

The European Health Data Space and the Secondary Use of Sensitive Health Data.

INTRODUCTION: The European Health Data Space (EHDS) is one of the European Union's most ambitious data-governance projects. It aims to create a common framework through which electronic health data can be accessed and reused across Member States for care, research, innovation, policy, and public-interest purposes. Its practical viability depends not only on digital infrastructure, but also on legal, ethical, and organisational harmonisation, particularly for genetic and genomic data. METHODS: This paper examines the EHDS with emphasis on the secondary use of health data. It reviews the EHDS institutional architecture, discusses Finland's Findata as a national model for structured access, and analyses challenges for data holders and data donors, including interoperability, governance burdens, privacy protection, residual re-identification risk, and genomic-data sensitivity. RESULTS: A cross-border cancer-genomics case study shows that the EHDS can streamline data discovery and the routing of access requests, but does not by itself eliminate legal fragmentation, heterogeneous ethics review, and consent-related barriers. DISCUSSION: Effective implementation will require harmonisation beyond infrastructure, including clearer consent standards, more consistent ethics procedures, interoperable metadata, and proportionate safeguards for genomic data.

Electronic Health Records

CRISPRessoSea: streamlined analysis and comparison of pooled amplicon CRISPR screens.

BACKGROUND: CRISPR genome editing enables precise modification of genomic targets but may also induce unintended edits at off-target sites with similar sequences. Pooled amplicon sequencing can assess on- and off-target editing across many samples, yet analyzing, aggregating, and visualizing results from multiple pooled experiments remains challenging. Tools to simplify and standardize these analyses are needed to provide reproducible and comparable interpretation of editing data. RESULTS: We developed CRISPRessoSea, a software package that processes, compares, and visualizes genome editing rates from pooled amplicon sequencing experiments. The tool provides standardized workflows for analyzing editing across multiple targets and samples, supports both nuclease- and base-editing modalities, and generates clear, data-rich summaries suitable for downstream interpretation. CONCLUSIONS: CRISPRessoSea facilitates reproducible, scalable analysis of CRISPR editing outcomes across diverse experimental designs, enabling more efficient and transparent assessment of genome editing specificity. The software is freely available at https://github.com/clementlab/CRISPRessoSea .

Software

VirDetector: a bioinformatic pipeline for virus surveillance using nanopore sequencing.

SUMMARY: Virus surveillance programmes are designed to counter the growing threat of viral outbreaks to human health. Nanopore sequencing, in particular, has proven to be suitable for this purpose, as it is readily available and provides rapid results. However, as special bioinformatic programs are required to extract the relevant information from the sequencing data, applications are needed that allow users without extensive bioinformatics knowledge to carry out the relevant analysis steps. We present VirDetector, a bioinformatic pipeline for virus surveillance using nanopore sequencing. The pipeline automatically installs all required programs and databases and allows all its steps to be executed with a single console command. After preprocessing the samples, including the possibility for basecalling, the pipeline classifies each sample taxonomically and reconstructs the viral consensus genomes, which are then used in phylogenetic analyses. This streamlined workflow provides a user-friendly and efficient solution for monitoring viral pathogens. AVAILABILITY AND IMPLEMENTATION: VirDetector is freely available at https://github.com/NLKaiser/VirDetector and https://zenodo.org/records/14637302 (10.5281/zenodo.14637302).

Nanopore Sequencing

Impact of precision oncology research in pediatric poor prognosis cancer: patient, parent and healthcare provider perspectives.

BACKGROUND: Comprehensive genomic analyses are increasingly accessible to children, adolescents and young adults (AYAs) with poor prognosis cancers. Challenges and successes of pediatric precision oncology studies from the perspectives of AYA patients, parents and healthcare providers (HCPs) are poorly described. METHODS: Between March 2021 and May 2023, we interviewed AYA patients (12-21 years), parents and HCPs who participated in pediatric precision oncology studies for poor prognosis cancers in British Columbia. Interviews followed an investigator-developed semi-structured topic guide. Data were coded inductively and deductively by one qualitative researcher and one trainee, supported by two additional team members. Analytic themes were established using qualitative thematic analysis. RESULTS: We interviewed 9 AYAs, 10 parents, and 17 HCPs. We identified five analytic themes: importance of clear communication of study information between patients, families and multidisciplinary HCPs; a need to support disclosure, understanding and clinical integration of research results; barriers to accessing innovative therapy and mitigation strategies; approaches to managing parent, patient and HCP hopes and expectations; personal challenges and stressors related to participation. CONCLUSIONS: We highlight unmet needs and offer practical considerations for integrating precision oncology into clinical practice. Considerations include educating and supporting oncologists through genomics results disclosure, increasing engagement with multidisciplinary HCPs, streamlining access to study information and results, coordinating efforts to clinically validate results and access therapies, and establishing real-world outcome data to inform clinical decision-making. Implementation of these strategies will optimize care for patients and families who are navigating poor prognosis cancers.

Humans

Measuring Cell Dimensions in Fission Yeast Using Machine Learning.

In fission yeast (Schizosaccharomyces pombe), cell length is a crucial indicator of cell cycle progression. Microscopy screens that examine the effect of agents or genotypes suspected of altering genomic or metabolic stability and thus cell size are crucial for studying disruptions to cell cycle dynamics. This method is based on using an automated cell segmentation algorithm to measure S. pombe cells imaged by brightfield (BF) microscopy methods. PhotoPhenosizer (PP) is a machine learning-based tool designed for automated cell measuring and dimensional analysis of morphology frequency distributions. Integration of this method into large-scale pipelines for tracking cell dimension change streamlines morphological measurements, which facilitates the examination of cellular responses to genomic and metabolic stresses. In this protocol, we use PP to observe the effect of genomic instability on cell size dynamics over a 12-day chronological lifespan assay. Our results show that relative to wild-type cells, a replication stress mutant shows larger cells during chronological aging in excess glucose media. Our results are consistent with activation of checkpoints that regulate cell morphology in response to DNA damage. This method's application highlights the relevance of its incorporation in experimental routines that require large-scale image processing and its adoption by users with routine needs in S. pombe molecular research projects.

Schizosaccharomyces

Pretraining improves prediction of genomic datasets across species.

MOTIVATION: Recent studies suggest that deep neural network models trained on thousands of human genomic datasets can accurately predict genomic features, including gene expression and chromatin accessibility. However, training these models is computation- and time-intensive, and datasets of comparable size do not exist for most other organisms. RESULTS: Here, we identify modifications to an existing state-of-the-art model that improve model accuracy while reducing training time and computational cost. Using this streamlined model architecture, we investigate the ability of models pretrained on human genomic datasets to transfer performance to a variety of different tasks. Models pretrained on human data but fine-tuned on genomic datasets from diverse tissues and species achieved significantly higher prediction accuracy while significantly reducing training time compared to models trained from scratch, with Pearson correlation coefficients between experimental results and predictions as high as 0.8. Further, we found that including excessive training tasks decreased model performance and that this decrease could be partially but not completely rescued by fine-tuning. Thus, simplifying model architecture, applying pretrained models, and carefully considering the number of training tasks may be effective and economical techniques for building new models across data types, tissues, and species. AVAILABILITY AND IMPLEMENTATION: Code is available on GitHub and Figshare: https://github.com/optimizedlearning/genomicsML, https://doi.org/10.6084/m9.figshare.31796116.

Genomics

Managing workflow executions with WESkit.

SUMMARY: In biomedical research, managing computational workflows across numerous projects-with varying parameters, tools, and environments-creates major challenges in scalability, reproducibility, and collaboration. Here, we present WESkit, an implementation of the Global Alliance for Genomics and Health (GA4GH) Workflow Execution Service (WES) interface, designed to streamline the execution, monitoring, and documentation of data processing workflows. It addresses the complexities involved in managing numerous executions with varying parameters across diverse research projects. Supporting both Snakemake and Nextflow, the system enables consistent automation and centralized monitoring, which benefits research groups aiming for long-term reproducibility and scalable collaboration. Its suitability for larger teams and service units is further enhanced by seamless integration into cloud environments, contributing to the GA4GH cloud framework. AVAILABILITY AND IMPLEMENTATION: The software WESkit is available under MIT license at the GitLab repository (https://gitlab.com/one-touch-pipeline/weskit). The WESkit main repository is archived at Software Heritage (https://archive.softwareheritage.org/browse/origin/directory/?origin_url=https://gitlab.com/one-touch-pipeline/weskit/api.git) and can be found using "one-touch-pipeline/weskit" term in the search section.

Workflow

One Plasmid Is All You Need: Genome Editing in Escherichia coli Using Endogenous TnpB and Endogenous Recombination System.

Escherichia coli (E. coli) is a key workhorse of biotechnology. Commonly used CRISPR-Cas9 systems for E. coli genome editing are complex and impose metabolic stress on the host, creating demand for more streamlined strategies. Recent studies identified the IS605 transposon-associated TnpB as a programmable RNA-guided (ωRNA) DNA endonuclease, prompting us to explore whether endogenous TnpB in E. coli (EcoTnpB) could be harnessed for genome editing. Biochemical and cellular analyses demonstrated that EcoTnpB efficiently cleaves both chromosomal and plasmid DNA at custom-specified sites in a TAM-dependent manner. Interestingly, E. coli possesses an endogenous recombination machinery capable of repairing EcoTnpB-induced DNA double-strand breaks (DSBs), challenging the long-held view that bacteria lack efficient homologous recombination systems. Based on these findings, we established a single-plasmid editing system (SPEED) in which genome editing is achieved by simply providing ωRNA and a homologous recombination template. By utilizing endogenous EcoTnpB together with the host HR pathway, this system enabled inducible and seamless genome editing at multiple genomic loci in BL21 (DE3), with editing efficiencies ranging from approximately 29% to 56%. Our results demonstrate for the first time that endogenous TnpB can be harnessed for genome editing and may hold potential for broader applications, such as species-specific antimicrobial development.

Escherichia coli

RAPID-DASH: Fast and Efficient Assembly of Guide RNA Arrays for Multiplexed CRISPR-Cas9 Applications.

Guide RNA (gRNA) arrays can enable targeting multiple genomic loci simultaneously using CRISPR-Cas9. In this study, we present a streamlined and efficient method to rapidly construct gRNA arrays with up to 10 gRNA units in a single day. We demonstrate that gRNA arrays maintain robust functional activity across all positions, and can incorporate libraries of gRNAs, combining scalability and multiplexing. Our approach will streamline combinatorial perturbation research by enabling the economical and rapid construction, testing, and iteration of gRNA arrays. To facilitate the adaptation of this approach, we have made a web tool to design oligo sequences necessary to assemble gRNA arrays.

CRISPR-Cas9

Scalable medium-density genotyping platforms for cultivar identification, pedigree authentication, marker-assisted and genomic selection, and other applications in strawberry.

A broad spectrum of high-density genotyping approaches, including single-nucleotide polymorphism (SNP) arrays, genotyping-by-sequencing, and whole-genome reduced-representation sequencing, have been shown to perform well in strawberry (Fragaria × ananassa), despite the inherent complexity of the octoploid genome. While these approaches are effective, their routine deployment in breeding programs can be constrained by cost, computational requirements, and workflow complexity. In parallel, many breeding programs continue to rely on locus-specific assays for marker-assisted selection, resulting in fragmented and inefficient genotyping strategies. Here, we describe medium-density amplicon-based genotyping platforms for strawberry designed to provide cost-effective, turnkey solutions that integrate markers used for marker-assisted selection with genome-wide markers suitable for genomic prediction in a single laboratory assay. These platforms were developed by targeting 1,650 or 4,811 target SNPs via amplicon sequencing, and are interoperable with existing high-density genotyping resources, including a widely used 50K SNP array, thereby facilitating data integration across platforms. We benchmarked their performance relative to the 50K SNP array across breeding-relevant applications, including identity and purity testing, pedigree authentication, marker-assisted selection, and genomic selection, and further evaluated the feasibility of genotype imputation to enhance genome-wide information content. Across analyses, the 1,650- and 4,811-amplicon platforms produced results comparable to higher-density platforms while substantially reducing genotyping cost and analytical overhead. This work demonstrates that targeted amplicon-based genotyping can support efficient, scalable, and integrated genome-informed breeding, enabling the routine application of both marker-assisted and genomic selection within strawberry breeding workflows. Open-source R workflows are provided to support streamlined analyses in breeding contexts.

Fragaria

Pre-Meta: priors-augmented retrieval for LLM-based metadata generation.

MOTIVATION: While high-throughput sequencing technologies have dramatically accelerated genomic data generation, the manual processes required for dataset annotation and metadata creation impede the efficient discovery and publication of these resources across disparate public repositories. Large language models (LLMs) have the potential to streamline dataset profiling and discovery. However, their current limitations in generalizing across specialized knowledge domains, particularly in fields such as biomedical genomics, prevent them from fully realizing this potential. This article presents Pre-Meta, an LLM-agnostic and domain-independent data annotation pipeline with an enriched retrieval procedure that leverages related priors-such as pre-generated metadata tags and ontologies-as auxiliary information to improve the accuracy of automated metadata generation. RESULTS: Validated using five selected metadata fields sampled across 1500 papers, the Pre-Meta assisted annotation experiment-without finetuning and prompt optimization-demonstrates a systemic improvement in the annotation task: shown through a 23%, 72%, and 75% accuracy gain from conventional RAG adoptions of GPT-4o mini, Llama 8B, and Mistral 7B respectively. AVAILABILITY AND IMPLEMENTATION: The code, data access, and scripts are available at: https://github.com/SINTEF-SE/LLMDap.

Metadata

Cre-loaded integrase-defective lentiviral vectors for targeted cassette exchange in CHO cells.

Genome-modifying enzymes, such as recombinases and CRISPR-associated nucleases, enable targeted gene insertion when delivered transiently to minimize off-target effects. Precise genome engineering requires controlled enzyme activity, as well as efficient donor DNA transfer. Integrase-defective lentiviral vectors (IDLVs) provide a promising platform for transient episomal DNA transfer; however, their integration efficiency depends on complementary genome-targeting strategies. Here, we engineered Cre-loaded IDLVs (Cre-IDLVs) that co-package lentiviral vector genomes together with bioactive Cre recombinase. Cre was inserted into the Gag region of an integrase-defective gag-pol construct, allowing for efficient encapsidation and protease-mediated release during virion maturation without compromising the viral titer. The resulting particles carried donor cassettes flanked by heterospecific loxP sites. When applied to CHO founder cells harboring compatible genomic loxP landing pads, Cre-IDLVs efficiently mediated recombination-mediated cassette exchange, producing the highest number of G418-resistant colonies among the plasmid ratios tested. Genomic PCR and sequencing confirmed precise locus-specific insertion without detectable random integration in the analyzed clones. These findings establish Cre-IDLVs as a streamlined dual-delivery platform that couples transient recombinase activity with episomal donor DNA transfer. This hybrid lentiviral strategy provides a programmable approach for controlled and site-specific genome modification in mammalian cells.

Integrases

DeepGeSeq: deep learning library for genomic sequence modeling and analysis.

MOTIVATION: Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. RESULTS: By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq's versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/JiaqiLi1024/DeepGeSeq.

Deep Learning

A universal, high-quality, and high-yield DNA purification method for mycobacteria, including Mycobacterium tuberculosis: large-scale assessment of the chloroform-bead method.

UNLABELLED: Genomic analysis of mycobacteria has become increasingly crucial for understanding drug-resistance mechanisms, molecular epidemiology, and pathogenesis. However, efficient extraction of high-molecular-weight genomic DNA from these organisms remains challenging because of their thick mycolic acid-rich cell walls. In this study, we report the chloroform-bead method, a universal DNA extraction protocol that combines chemical and mechanical disruptions to overcome these challenges. Multi-laboratory evaluation (16 sites) demonstrated the chloroform-bead method's superiority over conventional methods for Mycobacterium tuberculosis (DNA yield: 17.9 vs 1.9 &#xb5;g, purity A260/A230: 1.86 vs 1.22, both P < 0.001). Single-facility assessment extended these findings to >32 nontuberculous mycobacterial species (n = 1,058), showing performance comparable to M. tuberculosis (n = 1,000), with both achieving median yields of 22.2 &#xb5;g DNA and consistent quality metrics. The chloroform-bead method significantly reduced the processing time from 2 to 3 days to 2 h while ensuring complete sample sterilization, eliminating the need for species-specific optimization. This streamlined and universally applicable protocol represents a practical advancement in mycobacterial DNA extraction methodology, ideal for high-throughput genomic studies and routine clinical diagnostics. IMPORTANCE: Mycobacterial genomics is crucial for understanding pathogenesis and drug resistance; however, DNA extraction remains a significant challenge because of its unique cell wall. Traditional methods rely on enzymatic treatments, resulting in complex and time-consuming protocols with variable results. The chloroform-bead method introduces a paradigm shift by chemically and mechanically disrupting the mycolic acid layer and eliminating the need for enzymatic treatment. This standardized approach ensures consistent, high-quality DNA extraction across diverse mycobacterial species, thereby enhancing research capabilities and clinical applications.

Chloroform

Pangenomes aid accurate detection of large insertions and deletions from targeted sequencing: the case of cardiomyopathies.

BACKGROUND: Gene panels represent a widely used strategy for genetic testing in a vast range of Mendelian disorders. While this approach aids reliable bioinformatic detection of short coding variants, it often fails to detect many larger variants. Recent studies have recommended the adoption of pangenome references (as opposed to linear reference genomes like GRCh38) to augment detection of large variants from targeted sequencing, potentially providing diagnostic laboratories with the possibility to streamline diagnostic work-ups and reduce costs. METHODS: Here, we analyze 1969 cardiomyopathy cases and 1805 controls sequenced with the Illumina Trusight Cardio panel using a pangenome-based workflow (GRAF) and five conventional orthogonal methodologies (GATK HaplotypeCaller, GATK-gCNV, ExomeDepth, Manta and Lumpy-SV) to detect variants &#x2265;&#x2009;20&#xa0;bp in size. RESULTS: Following lab-based variant validation by means of PCR and Sanger sequencing, we show that GRAF conjugates higher precision and recall (F1 score 0.86) compared with other methods (F1 0-0.57) in detecting potentially pathogenic variants &#x2265;&#x2009;20&#xa0;bp from short-read panel data. Results were complemented by a comparison of the tools' performance in detecting ground truth variants on reference sample HG002 from Genome In A Bottle, which confirmed GRAF to outperform other tools also on exome sequencing (F1 0.97 vs. 0-0.94). Notably, in the HG002 benchmark dataset, GRAF also showed slightly improved performance compared to GATK HaplotypeCaller in the identification of small variants (1-19&#xa0;bp; F1 0.975 vs. 0.968). CONCLUSIONS: Our results indicate that pangenome-based workflows aid improved detection of large variants from targeted sequencing data in the clinical context and suggest that they may contribute to more unified variant detection frameworks for all-size genetic variants in the future.

Humans

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software

CAGEcleaner: reducing genomic redundancy in gene cluster mining.

SUMMARY: Mining homologous biosynthetic gene clusters (BGCs) typically involves searching colocalised genes against large genomic databases. However, the high degree of genomic redundancy in these databases often propagates into the resulting hit sets, complicating downstream analyses and visualization. To address this challenge, we present CAGEcleaner, a Python-based pipeline with auxiliary bash scripts designed to reduce redundancy in gene cluster hit sets by dereplicating the genomes that host these hits. CAGEcleaner integrates seamlessly with widely used gene cluster mining tools, such as cblaster and CAGECAT, enabling efficient filtering and streamlining BGC discovery workflows. AVAILABILITY AND IMPLEMENTATION: Source code and documentation is hosted at GitHub (https://github.com/LucoDevro/CAGEcleaner) and Zenodo (https://doi.org/10.5281/zenodo.14726119) under an MIT license. For accessibility, CAGEcleaner is installable from Bioconda (https://anaconda.org/bioconda/cagecleaner) and PyPi (https://pypi.org/project/cagecleaner/), and is also available as a Docker image from DockerHub (https://hub.docker.com/r/lucodevro/cagecleaner).

Software