Search PubMedSearch

SEARCH · Search PubMed

Results for “software tools”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Ten quick tips for spatial transcriptomics analysis.

Spatial transcriptomics (ST) enables genome-wide gene expression profiling while retaining spatial context within tissue sections. Since the foundational work by Ståhl et al. in 2016, the field has expanded rapidly, with diverse platforms now spanning sequencing-based (e.g., Visium, Visium HD, Slide-seq, Stereo-seq, and Seq-Scope) and imaging-based (e.g., MERFISH, Xenium, and CosMx SMI) approaches. The breadth of platforms, data structures, and computational tools, however, can be daunting for newcomers. Here, we present ten quick tips spanning the entire ST research workflow: whether ST suits a given biological question, how to select a platform aligned with study objectives, how to understand and process ST data, and which software tools to employ for analysis and visualization. We further discuss interpreting spatial patterns in biological context, integrating complementary modalities such as single-cell RNA sequencing and spatial proteomics, and leveraging public datasets and sharing results. Finally, we highlight current limitations of ST, particularly the challenge of reconstructing three-dimensional tissue architecture from serial tissue sections. This review provides biologists, bioinformaticians, and clinician-scientists with a concise, platform-neutral roadmap for incorporating ST into research, from experimental design to biological discovery.

Spatial Transcriptomics

GBRAP: A Comprehensive Database and Tool for Exploring Genomic Diversity Across All Domains of Life.

Evolutionary studies require extensive examination of genomic information across all domains of life. Despite the availability of a large number of genomes through GenBank, the effective visualization or comparison of the information they contain is challenging due to many reasons, including their size. We introduce genome-based retrieval and analysis parser, a comprehensive software tool to analyze genome files, and an online database housing an extensive collection of carefully curated, high-quality genome statistics for all the organisms available in the RefSeq database of National Center for Biotechnology Information. Users can either directly search, or select from precategorized groups, the organisms of their choice and retrieve data, and the output is generated as tables containing more than 200 columns of useful genomic information (base counts, GC content, Shannon entropy, codon usage, etc.) separately calculated for different genomic elements (e.g. coding sequences, introns, transfer RNA, ribosomal RNA, noncoding RNA, etc.). The data are independently displayed (if applicable) for each chromosomal, mitochondrial, plastid, or plasmid sequence. All the data can be visualized on the database or downloaded as comma-separated value or Excel files. The genome-based retrieval and analysis parser database is free to access without any registration and is publicly available at http://tacclab.org/gbrap/.

Software

FUSE-PhyloTree: linking functions and sequence conservation modules of a protein family through phylogenomic analysis.

SUMMARY: FUSE-PhyloTree is a phylogenomic analysis software for identifying local sequence conservation associated with the different functions of a multi-functional (e.g. paralogous or multi-domain) protein family. FUSE-PhyloTree introduces an original approach that combines advanced sequence analysis with phylogenetic methods. First, local sequence conservation modules within the family are identified using partial local multiple sequence alignment. Next, the evolution of the detected modules and known protein functions is inferred within the family's phylogenetic tree using three-level phylogenetic reconciliation and ancestral state reconstruction. As a result, FUSE-PhyloTree provides a gene tree annotated with both predicted sequence modules and ancestral gene functions, enabling the association of functions with specific sequence regions based on their co-emergence. AVAILABILITY AND IMPLEMENTATION: FUSE-PhyloTree is provided as Docker and Singularity images including all the required software tools. Images, source code, test data, and documentation are available at https://github.com/OcMalde/fuse-phylotree and https://zenodo.org/records/15855068.

Phylogeny

Complete genome sequence and genomic characterization of the probiotic Limosilactobacillus reuteri PSC102.

BACKGROUND: Gut microbiota are potential sources of probiotics and play an essential role in maintaining intestinal health. Limosilactobacillus reuteri PSC102 (L. reuteri PSC102), which was isolated from the feces of healthy pigs, exhibited health-beneficial properties. AIM: We aimed to conduct a whole-genome sequencing analysis of L. reuteri PSC102 to determine its molecular characteristics as a probiotic strain. METHODS: Limosilactobacillus reuteri PSC102 cells were cultured in De Man-Rogosa-Sharpe medium, followed by DNA extraction for genomic analysis using the PacBio-Illumina sequencing platform. The EzBioCloud software was used to perform gene assembly, and the genes were interpreted by the National Center for Biotechnology Information (NCBI) and the Glimmer program. Core and pan-genomic analyses were performed to assess the extent of functional conservation in the genomic sequence. Moreover, the NCBI database and the Basic Local Alignment Search Tool software were used to identify antimicrobial resistance genes and virulence factors. RESULTS: Limosilactobacillus reuteri PSC102 consists of a single circular chromosome with 2,048,626 bp, a guanine- cytosine of 38.9%, 18 rRNA genes, and 69 tRNA genes. Among the 1,846 protein-coding sequences, genes associated with probiotic characteristics were identified, including genes involved in host-microbe interactions, stress tolerance, biogenesis, and defense mechanisms. Furthermore, the genome of L. reuteri PSC102 comprises 2,446 pan-genome and 1,222 core-genome orthologous gene clusters. A total of 74 unique genes were identified in L. reuteri PSC102 genome. These genes mostly encode proteins potentially involved in the transport and metabolism of amino acids and carbohydrates. Moreover, antibacterial resistance genes and virulence factors were absent in L. reuteri PSC102. CONCLUSION: The results of the molecular insight into L. reuteri PSC102 corroborates its use as a probiotic in humans and other animals.

Limosilactobacillus reuteri

Vcfexpress: flexible, rapid user-expressions to filter and format VCFs.

MOTIVATION: Variant call format (VCF) files are the standard output format for various software tools that identify genetic variation from DNA sequencing experiments. Downstream analyses require the ability to query, filter, and modify them simply and efficiently. Several tools are available to perform these operations from the command line, including BCFTools, vembrane, slivar, and others. RESULTS: Here, we introduce vcfexpress, a new, high-performance toolset for the analysis of VCF files, written in the Rust programming language. It is nearly as fast as BCFTools, but adds functionality to execute user expressions in the lua programming language for precise filtering and reporting of variants from a VCF or BCF file. We demonstrate performance and flexibility by comparing vcfexpress to other tools using the vembrane benchmark. AVAILABILITY AND IMPLEMENTATION: vcfexpress is available under the MIT license at https://github.com/brentp/vcfexpress with code used for the manuscript deposited in https://doi.org/10.5281/zenodo.14756838.

Software

1-Mb resolution array-based comparative genomic hybridization using a BAC clone set optimized for cancer gene analysis.

Array-based comparative genomic hybridization (aCGH) is a recently developed tool for genome-wide determination of DNA copy number alterations. This technology has tremendous potential for disease-gene discovery in cancer and developmental disorders as well as numerous other applications. However, widespread utilization of a CGH has been limited by the lack of well characterized, high-resolution clone sets optimized for consistent performance in aCGH assays and specifically designed analytic software. We have assembled a set of approximately 4100 publicly available human bacterial artificial chromosome (BAC) clones evenly spaced at approximately 1-Mb resolution across the genome, which includes direct coverage of approximately 400 known cancer genes. This aCGH-optimized clone set was compiled from five existing sets, experimentally refined, and supplemented for higher resolution and enhancing mapping capabilities. This clone set is associated with a public online resource containing detailed clone mapping data, protocols for the construction and use of arrays, and a suite of analytical software tools designed specifically for aCGH analysis. These resources should greatly facilitate the use of aCGH in gene discovery.

Cell Line, Tumor

A Step-by-Step Guide to Sequencing and Assembly of Complete Bacterial Genomes Using the Oxford Nanopore MinION.

The Oxford Nanopore (ONT) MinION enables sequencing of longer DNA/RNA fragments compared to other sequencers, such as Illumina, etc. This nanopore method provides distinct advantages for generating complete genome assemblies from microorganisms. Specifically, the R9.4 flow cells used for MinION sequencing have much lower error rates compared with earlier versions of the ONT platform. Coupled with base calling using Dorado software, higher-quality long reads can now be generated for complete bacterial genome assembly. In this chapter, we describe a detailed MinION method to assemble a complete genome from a microorganism, polish the final assembly, and evaluate the genome quality using various software tools. Because of the low cost for MinION sequencing, this platform could be an asset for virtually any laboratory interested in generating complete genomes from microorganisms.

Genome, Bacterial

Rapid prototyping of interactive software for automated instrumentation in rehabilitative therapy.

Rapid prototyping is a quick, efficient way to evaluate new electrical instruments. A hardware interface between a computer and an instrument can be simulated in software. The authors demonstrate the technique of rapid prototyping by developing interfaces for two therapeutic strength-testing devices and an electromagnetic tracker/digitizer. The LabVIEW rapid prototyping software tool was used to create "virtual instruments," which combine to form interfaces. The functions and connections of each virtual instrument are described.

Equipment Design

'PePApipe': A complete bioinformatics analysis pipeline for African Swine Fever Virus genome.

African Swine Fever Virus (ASFV) is of high concern in porcine livestock across the world due to both the high mortality rates and the trade restrictions imposed on affected regions. The viral genome is large and complex, and genomic analysis is essential for tracing its origin and evolution. Although several bioinformatics tools exist for genome assembly and analysis, no single platform integrates all necessary steps in an accessible and systematic way. In this study the authors developed 'PePApipe', a custom-built, user-friendly pipeline that enables rapid, complete, and efficient ASFV genome analysis. It is specifically designed for laboratory professionals with limited bioinformatics experience, requiring only basic command-line knowledge. Starting from raw sequencing data, PePApipe integrates thirteen software tools into one automated workflow, covering quality control and pre-processing of raw reads, de novo genome assembly and variant calling. Programmed in Python, it can be executed locally through bash scripts, or using a Slurm protocol for batch processing of multiple samples. The main outputs are the ASFV consensus genome sequence and a file listing its putative variants compared to the selected reference genome. PePApipe classifies generated files into structured folders and produces intermediate files that can be used as inputs for further or parallel analyses; users can also enable or disable specific steps in each particular case. This pipeline is adaptable and complementary to downstream steps such as viral genome annotation or genome visualization. By consolidating all stages of viral genome analysis into a single automated workflow, PePApipe reduces the likelihood of user error, and enhances reproducibility and efficiency. This user-friendly pipeline facilitates the transition from sequencing to assembly and downstream analysis of viral genomes, ensuring a fast and reliable response to molecular analysis demands. Finally, the pipeline can be easily adapted to the study of other viral species, expanding its application in infectious diseases surveillance.

African Swine Fever Virus

moiraine: an R package to construct reproducible pipelines for the application and comparison of multi-omics integration methods.

MOTIVATION: In the past decades, many statistical methods for integrating multi-omics data have been developed. They have been implemented into software tools, which differ widely in their programming choices, such as the format required for data input, or the format of the generated integration results. This lack of standards renders cumbersome and time-intensive the application and comparison of different integration tools to the same multi-omics dataset. RESULTS: We have developed the moiraine R package for constructing reproducible multi-omics integration pipelines, which enables users to apply one or more statistical methods for multi-omics integration to their own multi-omics dataset. moiraine facilitates the preprocessing of the omics datasets and automates their formatting for the integration step. It simplifies the interpretation and evaluation of the integration results through the construction of visualizations in which metadata about samples and features can easily be included. Crucially, it enables the comparison of results obtained with different integration tools, allowing users to assess the robustness of their results. AVAILABILITY AND IMPLEMENTATION: The moiraine R package is publicly available at https://github.com/Plant-Food-Research-Open/moiraine; an archival snapshot of the package is available on Zenodo at https://doi.org/10.5281/zenodo.17172718. A detailed tutorial is available at https://plant-food-research-open.github.io/moiraine-manual/.

Software

G4SNVHunter: An R/Bioconductor Package for Evaluating SNV-Induced Disruption of G-Quadruplex Structures Leveraging the G4Hunter Algorithm.

G-quadruplexes (G4s) are nucleic acid secondary structures with important regulatory functions. Single-nucleotide variants (SNVs), one of the most common forms of genetic variation, can potentially impact the formation of G4 structures if they occur within G4 regions. However, there is currently a lack of software tools specifically designed to assess such effects. Here, we present an R/Bioconductor package named G4SNVHunter, which enables rapid detection of variants that may disrupt G4 structures. This tool, based on the core principles of the G4Hunter algorithm, can provide precise quantitative assessment of the propensity for G4 formation within genomic sequences. Specialized experimental methods can then be designed based on the results provided by G4SNVHunter to further verify the specific functions of the affected G4 structures, facilitating deeper insights into the biological impacts of genetic variants from the perspective of G4 structures. To showcase the functionality of the G4SNVHunter package, we analyzed the Neandertal and Denisovan archaic introgressed variants detected by the Sprime software, and identified approximately 5,800 variants located within G4 regions, among which around 230 may impair G4 structure formation propensity. The source code for the G4SNVHunter package has been publicly released under the MIT license at https://github.com/rongxinzh/G4SNVHunter and https://bioconductor.org/packages/devel/bioc/html/G4SNVHunter.html.

G-Quadruplexes

Construction of simple pathways and simple cycles in ecosystems.

We present software tools for overcoming the problem of combinatorics in the enumeration of simple pathways and simple cycles in a first flow-through analysis of carbon transfer in large ecosystems. Rather than search through the very large number of potential routes in a reasonably sized ecosystem for the relatively small number of actual routes, our main algorithm performs an efficient rule-based construction of the actual routes. The enumeration of the unique pathways becomes tractable in terms of CPU time, which increases linearly with ecosystem size and connectedness. Networks of up to 80 entities can be evaluated using our software.

Algorithms

A picture communicator for symbol users and/or speech-impaired people.

There are several approaches to producing communication aids for people with disabilities. The system described here takes the approach of utilizing as much mainstream hardware as possible, and adapting it with some modular software tools which have been designed to facilitate the building of customized symbol communication systems. The target audience are clients who are symbol users and/or have a speech impairment. The system provides several levels of screens, each of which can contain grids of scalable icons. A number of input methods are provided including keyboard, switch, mouse and touch-screen. Digitized and/or text to speech synthesis can be used for reinforcement of selections and for communication. The structure of the system is discussed and initial feedback from the first field trials is presented.

Communication Devices for People with Disabilities

Do the benefits outweigh the costs of PACS? The results of an International Workshop on Technology Assessment of PACS.

On May 26-27, 1991, an International Workshop on the Technology Assessment of PACS was held at Enkhuizen, The Netherlands. During this workshop 35 experts in the field, from 13 different countries, discussed amongst others the required functionality of PACS, diagnostic aspects and the quality of care and organizational aspects. The key question was whether, when and how PACS is feasible, both from a financial and a clinical point of view. Data which were collected with the aid of the software tool CAPACITY formed the starting point of this meeting. This paper gives an outline of the discussions during this workshop. The main conclusion is that more clinical research is needed, to get a better insight into the costs and the clinical benefits of PACS. Because of the high costs of the PACS technology, international cooperation in this field is requested. It is recommended that the CAPACITY project, which is set up to stimulate the international dialogue and data exchange on PACS, is continued.

Cost-Benefit Analysis

A dedicated caller for DUX4 rearrangements from whole-genome sequencing data.

Rearrangements involving the DUX4 gene (DUX4-r) define a subtype of paediatric and adult acute lymphoblastic leukaemia (ALL) with a favourable outcome. Currently, there is no 'standard of care' diagnostic method for their confident identification. Here, we present an open-source software tool designed to detect DUX4-r from short-read, whole-genome sequencing (WGS) data. Evaluation on a cohort of 210 paediatric ALL cases showed that our method detects all known, as well as previously unidentified, cases of IGH::DUX4 and rearrangements with other partner genes. These findings demonstrate the possibility of robustly detecting DUX4-r using WGS in the routine clinical setting.

Humans

A genome-wide coverage-based pipeline for the identification of host-derived candidate DNA biomarkers from cell-free blood.

We have created a new data-analysis pipeline for the discovery of host-specific candidate DNA biomarkers derived from sequencing data of cell-free blood. Unlike approaches that rely on specific molecular or genetic signatures, our method leverages the coverage distribution of cell-free DNA sequences mapped to a reference genome, applying statistical analyses to identify informative short genomic regions for biomarker discovery. The pipeline is applicable to diverse diseases and can be used to analyze cell-free DNA sequences from plasma or serum to identify candidate biomarkers that are characteristic of disease states in mammals. Core functionalities were developed in Java and integrated with open-source software tools for the preprocessing of raw sequencing data, complemented by Python scripts for the machine-learning analysis and statistical validation. The pipeline is designed for HPC use and users can access the pipeline through a Galaxy workflow, which offers a user-friendly web interface for input selection prior to execution and analysis progress monitoring. Performance tests, carried out using duplicate sets of COVID-19 samples and controls, showed linear scalability of execution time with an increasing dataset size, as well as a substantial reduction in execution time through parallelized computation, whereby each HPC node is used to process the data of one chromosome. Further statistical tests confirmed the quality of the pipeline's results by showing that the set of identified candidate biomarkers remained stable across varying dataset sizes.

Biomarkers

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score = 0.67-0.90) in genus diversity and showed a high correlation (rSpearman = 0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics

Exchange of Veterans Affairs medical data using national and local networks.

Remote data exchange is extremely useful to a number of medical applications. It requires an infrastructure including systems, network and software tools. With such an infrastructure, existing local applications can be extended to serve national needs. There are many approaches to providing remote data exchange. Selection of an approach for an application requires balancing of various factors, including the need for rapid interactive access to data and ad hoc queries, the adequacy of access to predefined data sets, the need for an integrated view of the data, the ability to provide adequate security protection, the amount of data required, and the time frame in which data is required. The applications described here demonstrate new ways that the VA is reaping benefits from its infrastructure and its compatible integrated hospital information systems located at its facilities. The needs that have been met are also needs of private hospitals. However, in many cases the infrastructure to allow data exchange is not present. The VA's experiences may serve to establish the benefits that can be obtained by all hospitals.

CD-ROM