Search PubMedSearch

SEARCH · Search PubMed

Results for “software tools”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Phasis: a software tool for register-resolved discovery of plant phased small RNA loci.

Plant PHAS locus discovery remains challenging because phasiRNA-producing loci must be distinguished from other sRNA-producing regions with high abundance or apparent periodicity. This problem is especially acute for reproductive 24-PHAS loci, which occur within genomes that also produce abundant 24-nt siRNAs from nonPHAS regions. We present Phasis, an open-source Python software tool for plant PHAS-locus discovery from small RNA sequencing data. Phasis combines statistical evidence for phased accumulation with locus-level features and a Register-Resolved Locus Interpretation Layer that evaluates whether candidate loci show coherent phased architecture. Across diverse plant datasets, Phasis recovered validated or annotated 21- and 24-PHAS loci with a strong balance between call-level precision and reference-locus recall, and generally outperformed PhaseTank and ShortStack in matched benchmark analyses. The register-resolved interpretation layer reduced unsupported calls by separating coherent phased loci from ambiguous sRNA-producing regions. In maize dcl5 mutant libraries, Phasis showed strong depletion of 24-PHAS recovery, supporting DCL5-dependent recovery of reproductive 24-PHAS signal. Together, these results support Phasis as a biologically interpretable tool for large-scale discovery of plant DCL-dependent phasiRNA loci.

bioinformatics

GRable Version 1.0: A Software Tool for Site-Specific Glycoform Analysis With Improved MS1-Based Glycopeptide Detection With Parallel Clustering and Confidence Evaluation With MS2 Information.

High-throughput intact glycopeptide analysis is crucial for elucidating the physiological and pathological status of the glycans attached to each glycoprotein. Mass spectrometry-based glycoproteomic methods are challenging because of the diversity and heterogeneity of glycan structures. Therefore, we developed an MS1-based site-specific glycoform analysis method named "Glycan heterogeneity-based Relational IDentification of Glycopeptide signals on Elution profile (Glyco-RIDGE)" for a more comprehensive analysis. This method detects glycopeptide signals as a cluster based on the mass and chromatographic properties of glycopeptides and then searches for each combination of core peptides and glycan compositions by matching their mass and retention time differences. Here, we developed a novel browser-based software named GRable for semi-automated Glyco-RIDGE analysis with significant improvements in glycopeptide detection algorithms, including "parallel clustering." This unique function improved the comprehensiveness of glycopeptide detection and allowed the analysis to focus on specific glycan structures, such as pauci-mannose. The other notable improvement is evaluating the "confidence level" of the GRable results, especially using MS2 information. This function facilitated reduced misassignment of the core peptide and glycan composition and improved the interpretation of the results. Additional improved points of the algorithms are "correction function" for accurate monoisotopic peak picking; one-to-one correspondence of clusters and core peptides even for multiply sialylated glycopeptides; and "inter-cluster analysis" function for understanding the reason for detected but unmatched clusters. The significance of these improvements was demonstrated using purified and crude glycoprotein samples, showing that GRable allowed site-specific glycoform analysis of intact sialylated glycoproteins on a large-scale and in-depth. Therefore, this software will help us analyze the status and changes in glycans to obtain biological and clinical insights into protein glycosylation by complementing the comprehensiveness of MS2-based glycoproteomics. GRable can be freely run online using a web browser via the GlyCosmos Portal (https://glycosmos.org/grable).

Glycopeptides

Generating multiple alignments on a pangenomic scale.

MOTIVATION: Since novel long read sequencing technologies allow for de novo assembly of many individuals of a species, high-quality assemblies are becoming widely available. For example, the recently published draft human pangenome reference was based on assemblies composed of contigs. There is an urgent need for a software-tool that is able to generate a multiple alignment of genomes of the same species because current multiple sequence alignment programs cannot deal with such a volume of data. RESULTS: We show that the combination of a well-known anchor-based method with the technique of prefix-free parsing yields an approach that is able to generate multiple alignments on a pangenomic scale, provided that large-scale structural variants are rare. Furthermore, experiments with real world data show that our software tool PANgenomic Anchor-based Multiple Alignment significantly outperforms current state-of-the art programs. AVAILABILITY AND IMPLEMENTATION: Source code is available at: https://gitlab.com/qwerzuiop/panama, archived at swh:1:dir:e90c9f664995acca9063245cabdd97549cf39694.

Software

vcfsim: flexible simulation of all-sites VCFs with missing data.

BACKGROUND |: VCFs are the most widely used data format for encoding genetic variation. By design, standard VCFs do not include data from sites where all individuals are homozygous for the reference allele ("invariant sites") and thus do not differentiate these from sites where data are completely missing. However, missing data are a key feature of biological datasets across all domains of genomics, and many recent studies have shown that missing data can introduce a variety of statistical biases in the estimation of key population genetic parameters. A solution to this limitation is to include invariant sites in a standard VCF, creating an "all-sites VCF", exposing missing and invariant sites explicitly. One hurdle to the wider adoption of all-sites VCFs is a reliable parameterized simulation framework for generating biologically realistic all-sites VCFs. RESULTS |: Here, we introduce an open-source command line tool, vcfsim, that interfaces with the popular coalescent simulation platform msprime and provides convenience functions for simulating all-sites VCFs with variable levels of ploidy and missing data. We show that the post-processed VCFs generated using vcfsim align precisely with population genetic expectations (i.e. are statistically identical to raw msprime output), accurately introduce missing data, and permit the simulation of data with varying ploidy levels, including the simulation of intraindividual ploidy variation (e.g. heterogametic sex chromosomes) and population structures. CONCLUSIONS |: Our results vcfsim is a useful and easy-to-use tool for the benchmarking of new software tools, performing population genetic inference, training of machine learning models, and the exploration of the effects of missing data in genomics data sets.

Benchmarking

The Hunt Lab Guide to De Novo Peptide Sequence Analysis by Tandem Mass Spectrometry.

Donald Hunt has made seminal contributions to the fields of proteomics, immunology, epigenetics, and glycobiology. The foundation of every important work to come out of the Hunt Laboratory is de novo peptide sequencing. For decades, he taught hundreds of students, postdocs, engineers, and scientists to directly interpret mass spectral data. To honor his legacy and ensure that the art of de novo sequencing is not lost, we have adapted his teaching materials into "The Hunt Lab Guide to De Novo Peptide Sequence Analysis by Tandem Mass Spectrometry". In addition to the de novo sequencing tutorials, we present two freely available software tools that facilitate manual interpretation of mass spectra and validation of search results. The first, "Hunt Lab Peptide Fragment Calculator", calculates precursor and fragment mass-to-charge ratios for any peptide. The second program, "Predator Protein Fragment Calculator", was inspired in part by the fragment calculator developed in the Hunt Lab. Its capabilities are enhanced to facilitate interpretation of mass spectral data derived from intact proteins. We hope that the combination of these educational tools will continue to benefit students and researchers by empowering them to interpret data on their own.

Tandem Mass Spectrometry

ChemGenXplore: an interactive tool for exploring and analysing chemical genomic data.

MOTIVATION: Chemical genomics is a powerful high-throughput approach to systematically link phenotypes to genotypes. However, the vast datasets generated remain challenging to explore due to the lack of integrated, interactive tools for visualization and analysis. Existing workflows often require multiple independent software tools, limiting data accessibility and collaboration. Therefore, we created a user-friendly platform that enables efficient exploration and sharing of chemical genomics data. RESULTS: We developed ChemGenXplore, a web-based Shiny application designed to streamline the visualization and analysis of chemical genomic screens. It offers two primary functionalities: one for exploring pre-implemented datasets and another for analysing user-uploaded datasets. ChemGenXplore enables users to visualize phenotypic profiles, assess gene-gene and condition-condition correlations, perform GO and KEGG enrichment analysis, and generate customizable, interactive heatmaps. To further support collaborative research, ChemGenXplore also facilitates the comparative analysis of chemical genomic and other omics datasets. By consolidating these features into a single interactive and accessible tool, ChemGenXplore facilitates data sharing, enhances reproducibility, and promotes collaboration within the research community. AVAILABILITY AND IMPLEMENTATION: ChemGenXplore is freely accessible as a web application at https://chemgenxplore.kaust.edu.sa/. Source code and documentation, including instructions for local installation, are provided on GitHub (https://github.com/Hudaahmadd/ChemGenXplore). A Docker image is also available on DockerHub (https://hub.docker.com/r/hudaahmad/chemgenxplore) to ensure reproducibility and simplify installation.

Software

A comparison of software for analysis of rare and common short tandem repeat (STR) variation using human genome sequences from clinical and population-based samples.

Short tandem repeat (STR) variation is an often overlooked source of variation between genomes. STRs comprise about 3% of the human genome and are highly polymorphic. Some cause Mendelian disease, and others affect gene expression. Their contribution to common disease is not well-understood, but recent software tools designed to genotype STRs using short read sequencing data will help address this. Here, we compare software that genotypes common STRs and rarer STR expansions genome-wide, with the aim of applying them to population-scale genomes. By using the Genome-In-A-Bottle (GIAB) consortium and 1000 Genomes Project short-read sequencing data, we compare performance in terms of sequence length, depth, computing resources needed, genotyping accuracy and number of STRs genotyped. To ensure broad applicability of our findings, we also measure genotyping performance against a set of genomes from clinical samples with known STR expansions, and a set of STRs commonly used for forensic identification. We find that HipSTR, ExpansionHunter and GangSTR perform well in genotyping common STRs, including the CODIS 13 core STRs used for forensic analysis. GangSTR and ExpansionHunter outperform HipSTR for genotyping call rate and memory usage. ExpansionHunter denovo (EHdn), STRling and GangSTR outperformed STRetch for detecting expanded STRs, and EHdn and STRling used considerably less processor time compared to GangSTR. Analysis on shared genomic sequence data provided by the GIAB consortium allows future performance comparisons of new software approaches on a common set of data, facilitating comparisons and allowing researchers to choose the best software that fulfils their needs.

Humans

map3C: a computational tool for processing multiomic single-cell Hi-C data.

SUMMARY: The emergence of multiomic single-cell Hi-C (scHi-C) methods, which simultaneously profile chromatin conformation and other modalities such as gene expression or DNA methylation, creates tremendous opportunities for studying the genome's structure-function relationships. Existing tools for processing multiomic scHi-C datasets lack certain key functions for downstream bioinformatics analysis. We present map3C, a software tool that incorporates additional key functions. Specifically, we demonstrate that map3C facilitates multiomic scHi-C processing, quality control, and identification of structural variant locations in the genome. AVAILABILITY AND IMPLEMENTATION: map3C is available at https://github.com/luogenomics/map3C and is archived at https://doi.org/10.5281/zenodo.20724719.

Software

Clinical Variant Interpretation with the Integrative Genomics Viewer (IGV) for Molecular Pathologists.

The integrative genomics viewer (IGV) is a pivotal tool in clinical genomics, enabling the visualization and interpretation of complex sequencing data. Bringing clinical knowledge to bear with visual evaluation of sequencing results is the primary means by which molecular pathologists and other professionals assess and finalize cases. A variety of software tools can assist, but their relationship to the underlying data must be understood and applied systematically. This study includes essential background on next-generation sequencing (NGS) data file types (e.g., FASTQ, BAM, VCF) with a discussion of their format and purpose. We then describe features of IGV that derive nuances from these files. We utilize a series of curated practical cases based on clinical vignettes through which the reader will interact with clinical NGS sequencing data using the IGV software to review various types of clinically relevant variants relative to the human reference genome. These clinical vignettes have been curated to describe examples of some of the complexities of interpretation of genomic data, and how utilizing IGV as part of a routine workflow can provide additional interpretive information for variants beyond routine bioinformatic software algorithm variant calls. The visual inspection of genomic variants utilizing the tools within IGV can unmask subtle contextual cues (i.e., variant allele frequency, strand bias, tissue-specific context) that can influence the interpretation of genomic variants. Although this study focuses on using IGV for the detection and interpretation of somatic variants, the provided applications can be extrapolated for use in the germline setting, including analysis of complex variants and detection of mosaicism.

Humans

Comprehensive evaluation of ACMG/AMP-based variant classification tools.

MOTIVATION: The American College of Medical Genetics and Genomics/Association for Molecular Pathology (ACMG/AMP) guidelines represent the gold standard for clinical variant interpretation. Despite the widespread adoption of ACMG/AMP guidelines, a comprehensive comparison of the software tools designed to implement them has been lacking. This represents a significant gap, as clinicians require evidence-based guidance on which tools to use in their practice. RESULTS: We benchmarked four ACMG/AMP-based tools (Franklin, InterVar, TAPES, Genebe) selected from 22 tools, and compared their performance with LIRICAL, a top-performing phenotype-driven tool, using 151 expert-curated datasets from Mendelian disorders. Selection criteria included free availability, VCF compatibility, operational reliability, and not being disease-specific. Our evaluation framework assessed top-N accuracy (N = 1, 5, 10, 20, 50), retention rates, precision, recall, F1 scores, and area under the curve (AUC). Statistical validation employed bootstrap confidence intervals (n = 1000) and Friedman tests. LIRICAL (68.21%) and Franklin (61.59%) demonstrated superior top-10 variant prioritization accuracy in Mendelian disorders, significantly outperforming other tools (P = .0000). Results demonstrate that tools with advanced phenotypic integration significantly outperform those relying primarily on genomic features. AVAILABILITY AND IMPLEMENTATION: All data and source code required to reproduce the findings of this study are openly available in the Code Ocean repository at https://doi.org/10.24433/CO.6562438.v1.

Software

Ten quick tips for spatial transcriptomics analysis.

Spatial transcriptomics (ST) enables genome-wide gene expression profiling while retaining spatial context within tissue sections. Since the foundational work by Ståhl et al. in 2016, the field has expanded rapidly, with diverse platforms now spanning sequencing-based (e.g., Visium, Visium HD, Slide-seq, Stereo-seq, and Seq-Scope) and imaging-based (e.g., MERFISH, Xenium, and CosMx SMI) approaches. The breadth of platforms, data structures, and computational tools, however, can be daunting for newcomers. Here, we present ten quick tips spanning the entire ST research workflow: whether ST suits a given biological question, how to select a platform aligned with study objectives, how to understand and process ST data, and which software tools to employ for analysis and visualization. We further discuss interpreting spatial patterns in biological context, integrating complementary modalities such as single-cell RNA sequencing and spatial proteomics, and leveraging public datasets and sharing results. Finally, we highlight current limitations of ST, particularly the challenge of reconstructing three-dimensional tissue architecture from serial tissue sections. This review provides biologists, bioinformaticians, and clinician-scientists with a concise, platform-neutral roadmap for incorporating ST into research, from experimental design to biological discovery.

Spatial Transcriptomics

GBRAP: A Comprehensive Database and Tool for Exploring Genomic Diversity Across All Domains of Life.

Evolutionary studies require extensive examination of genomic information across all domains of life. Despite the availability of a large number of genomes through GenBank, the effective visualization or comparison of the information they contain is challenging due to many reasons, including their size. We introduce genome-based retrieval and analysis parser, a comprehensive software tool to analyze genome files, and an online database housing an extensive collection of carefully curated, high-quality genome statistics for all the organisms available in the RefSeq database of National Center for Biotechnology Information. Users can either directly search, or select from precategorized groups, the organisms of their choice and retrieve data, and the output is generated as tables containing more than 200 columns of useful genomic information (base counts, GC content, Shannon entropy, codon usage, etc.) separately calculated for different genomic elements (e.g. coding sequences, introns, transfer RNA, ribosomal RNA, noncoding RNA, etc.). The data are independently displayed (if applicable) for each chromosomal, mitochondrial, plastid, or plasmid sequence. All the data can be visualized on the database or downloaded as comma-separated value or Excel files. The genome-based retrieval and analysis parser database is free to access without any registration and is publicly available at http://tacclab.org/gbrap/.

Software

FUSE-PhyloTree: linking functions and sequence conservation modules of a protein family through phylogenomic analysis.

SUMMARY: FUSE-PhyloTree is a phylogenomic analysis software for identifying local sequence conservation associated with the different functions of a multi-functional (e.g. paralogous or multi-domain) protein family. FUSE-PhyloTree introduces an original approach that combines advanced sequence analysis with phylogenetic methods. First, local sequence conservation modules within the family are identified using partial local multiple sequence alignment. Next, the evolution of the detected modules and known protein functions is inferred within the family's phylogenetic tree using three-level phylogenetic reconciliation and ancestral state reconstruction. As a result, FUSE-PhyloTree provides a gene tree annotated with both predicted sequence modules and ancestral gene functions, enabling the association of functions with specific sequence regions based on their co-emergence. AVAILABILITY AND IMPLEMENTATION: FUSE-PhyloTree is provided as Docker and Singularity images including all the required software tools. Images, source code, test data, and documentation are available at https://github.com/OcMalde/fuse-phylotree and https://zenodo.org/records/15855068.

Phylogeny

Complete genome sequence and genomic characterization of the probiotic Limosilactobacillus reuteri PSC102.

BACKGROUND: Gut microbiota are potential sources of probiotics and play an essential role in maintaining intestinal health. Limosilactobacillus reuteri PSC102 (L. reuteri PSC102), which was isolated from the feces of healthy pigs, exhibited health-beneficial properties. AIM: We aimed to conduct a whole-genome sequencing analysis of L. reuteri PSC102 to determine its molecular characteristics as a probiotic strain. METHODS: Limosilactobacillus reuteri PSC102 cells were cultured in De Man-Rogosa-Sharpe medium, followed by DNA extraction for genomic analysis using the PacBio-Illumina sequencing platform. The EzBioCloud software was used to perform gene assembly, and the genes were interpreted by the National Center for Biotechnology Information (NCBI) and the Glimmer program. Core and pan-genomic analyses were performed to assess the extent of functional conservation in the genomic sequence. Moreover, the NCBI database and the Basic Local Alignment Search Tool software were used to identify antimicrobial resistance genes and virulence factors. RESULTS: Limosilactobacillus reuteri PSC102 consists of a single circular chromosome with 2,048,626 bp, a guanine- cytosine of 38.9%, 18 rRNA genes, and 69 tRNA genes. Among the 1,846 protein-coding sequences, genes associated with probiotic characteristics were identified, including genes involved in host-microbe interactions, stress tolerance, biogenesis, and defense mechanisms. Furthermore, the genome of L. reuteri PSC102 comprises 2,446 pan-genome and 1,222 core-genome orthologous gene clusters. A total of 74 unique genes were identified in L. reuteri PSC102 genome. These genes mostly encode proteins potentially involved in the transport and metabolism of amino acids and carbohydrates. Moreover, antibacterial resistance genes and virulence factors were absent in L. reuteri PSC102. CONCLUSION: The results of the molecular insight into L. reuteri PSC102 corroborates its use as a probiotic in humans and other animals.

Limosilactobacillus reuteri

Vcfexpress: flexible, rapid user-expressions to filter and format VCFs.

MOTIVATION: Variant call format (VCF) files are the standard output format for various software tools that identify genetic variation from DNA sequencing experiments. Downstream analyses require the ability to query, filter, and modify them simply and efficiently. Several tools are available to perform these operations from the command line, including BCFTools, vembrane, slivar, and others. RESULTS: Here, we introduce vcfexpress, a new, high-performance toolset for the analysis of VCF files, written in the Rust programming language. It is nearly as fast as BCFTools, but adds functionality to execute user expressions in the lua programming language for precise filtering and reporting of variants from a VCF or BCF file. We demonstrate performance and flexibility by comparing vcfexpress to other tools using the vembrane benchmark. AVAILABILITY AND IMPLEMENTATION: vcfexpress is available under the MIT license at https://github.com/brentp/vcfexpress with code used for the manuscript deposited in https://doi.org/10.5281/zenodo.14756838.

Software

1-Mb resolution array-based comparative genomic hybridization using a BAC clone set optimized for cancer gene analysis.

Array-based comparative genomic hybridization (aCGH) is a recently developed tool for genome-wide determination of DNA copy number alterations. This technology has tremendous potential for disease-gene discovery in cancer and developmental disorders as well as numerous other applications. However, widespread utilization of a CGH has been limited by the lack of well characterized, high-resolution clone sets optimized for consistent performance in aCGH assays and specifically designed analytic software. We have assembled a set of approximately 4100 publicly available human bacterial artificial chromosome (BAC) clones evenly spaced at approximately 1-Mb resolution across the genome, which includes direct coverage of approximately 400 known cancer genes. This aCGH-optimized clone set was compiled from five existing sets, experimentally refined, and supplemented for higher resolution and enhancing mapping capabilities. This clone set is associated with a public online resource containing detailed clone mapping data, protocols for the construction and use of arrays, and a suite of analytical software tools designed specifically for aCGH analysis. These resources should greatly facilitate the use of aCGH in gene discovery.

Cell Line, Tumor

A Step-by-Step Guide to Sequencing and Assembly of Complete Bacterial Genomes Using the Oxford Nanopore MinION.

The Oxford Nanopore (ONT) MinION enables sequencing of longer DNA/RNA fragments compared to other sequencers, such as Illumina, etc. This nanopore method provides distinct advantages for generating complete genome assemblies from microorganisms. Specifically, the R9.4 flow cells used for MinION sequencing have much lower error rates compared with earlier versions of the ONT platform. Coupled with base calling using Dorado software, higher-quality long reads can now be generated for complete bacterial genome assembly. In this chapter, we describe a detailed MinION method to assemble a complete genome from a microorganism, polish the final assembly, and evaluate the genome quality using various software tools. Because of the low cost for MinION sequencing, this platform could be an asset for virtually any laboratory interested in generating complete genomes from microorganisms.

Genome, Bacterial

'PePApipe': A complete bioinformatics analysis pipeline for African Swine Fever Virus genome.

African Swine Fever Virus (ASFV) is of high concern in porcine livestock across the world due to both the high mortality rates and the trade restrictions imposed on affected regions. The viral genome is large and complex, and genomic analysis is essential for tracing its origin and evolution. Although several bioinformatics tools exist for genome assembly and analysis, no single platform integrates all necessary steps in an accessible and systematic way. In this study the authors developed 'PePApipe', a custom-built, user-friendly pipeline that enables rapid, complete, and efficient ASFV genome analysis. It is specifically designed for laboratory professionals with limited bioinformatics experience, requiring only basic command-line knowledge. Starting from raw sequencing data, PePApipe integrates thirteen software tools into one automated workflow, covering quality control and pre-processing of raw reads, de novo genome assembly and variant calling. Programmed in Python, it can be executed locally through bash scripts, or using a Slurm protocol for batch processing of multiple samples. The main outputs are the ASFV consensus genome sequence and a file listing its putative variants compared to the selected reference genome. PePApipe classifies generated files into structured folders and produces intermediate files that can be used as inputs for further or parallel analyses; users can also enable or disable specific steps in each particular case. This pipeline is adaptable and complementary to downstream steps such as viral genome annotation or genome visualization. By consolidating all stages of viral genome analysis into a single automated workflow, PePApipe reduces the likelihood of user error, and enhances reproducibility and efficiency. This user-friendly pipeline facilitates the transition from sequencing to assembly and downstream analysis of viral genomes, ensuring a fast and reliable response to molecular analysis demands. Finally, the pipeline can be easily adapted to the study of other viral species, expanding its application in infectious diseases surveillance.

African Swine Fever Virus