Search PubMedSearch

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Accelerating minimap2 for whole-genome alignment.

SUMMARY: Recent advances in long-read sequencing and genome assembly techniques have enabled the generation of high-quality assemblies, often comprising megabase-scale sequences that span entire chromosomes. This results in longer but fewer sequences per genome, which affects the parallelization efficiency of whole-genome alignment tools. Current methods that assign one thread per query sequence now face suboptimal CPU use and longer runtimes because the processing of fewer sequences leaves many threads idle. We present mm2-plus, a fast and efficient method for whole-genome alignment, built upon the commonly used minimap2 aligner. Our improvements include a fine-grained parallel chaining algorithm and a fast method for differentiating primary and secondary chains. These optimizations accelerate the alignment of human, plant, and primate genomes by 1.6× to 7.2× without compromising accuracy. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/at-cg/mm2-plus and https://doi.org/10.5281/zenodo.18220923.

Sequence Alignment

wgatools: an ultrafast toolkit for manipulating whole-genome alignments.

SUMMARY: With the rapid development of long-read sequencing technologies, the era of individual complete genomes is approaching. We have developed wgatools, a cross-platform, ultrafast toolkit that supports a range of whole-genome alignment formats, offering practical tools for conversion, processing, evaluation, and visualization of alignments, thereby facilitating population-level genome analysis and advancing functional and evolutionary genomics. AVAILABILITY AND IMPLEMENTATION: wgatools supports diverse formats and can process, filter, and statistically evaluate alignments, perform alignment-based variant calling, and visualize alignments both locally and genome-wide. Built with Rust for efficiency and safe memory usage, it ensures fast performance and can handle large datasets consisting of hundreds of genomes. wgatools is published as free software under the MIT open-source license, and its source code is freely available at https://github.com/wjwei-handsome/wgatools and https://zenodo.org/records/14882797.

Software

Assessing Hardy-Weinberg equilibrium in T2T-aligned 1000 genomes project.

Quality control of markers in genome-wide association studies often includes testing for Hardy-Weinberg equilibrium (HWE). However, this is usually implemented in a homogeneous population without stratifying by sex. Previous work indicates sex-based selection at numerous autosomal loci in cohorts with active recruitment. Sex chromosome sequences can also interfere with autosomal SNPs. These motivate a re-examination of HWE in sex-aware analyses. Using the telomere-to-telomere (T2Tv2)-aligned high-coverage whole genome sequencing data from 2,490 individuals in the 1000 Genomes Project, we examined genome-wide sex-specific deviations from HWE across five super-populations. Our analyses were restricted to bi-allelic SNPs with non-missing genotypes and minor allele frequency (MAF) &#x2265;5% in both sexes of the five super-populations. We applied an allele-based framework to quantify both the magnitude and direction of Hardy-Weinberg disequilibrium (HWD), followed by a second-order omnibus meta-analysis that combined HWD results across populations and sexes. At a genome-wide significance threshold of p&#x2009;<&#x2009;5e-8, 0.9% of autosomal SNPs exhibited significant deviations from HWE. The majority of these deviations were associated with genomic features indicative of poor sequence quality. Restricting the analysis to reliable genomic regions substantially reduced the number of signals, yielding 255 autosomal SNPs and one non-pseudoautosomal chromosome X SNP. Among these, 140 autosomal SNPs displayed significant heterogeneity across populations but not across sexes. Notably, eight SNPs within a 15-bp region on chromosome 14q31.3 showed excess heterozygosity in both sexes of the African super-population (AFR). Finally, we developed a multivariate predictor of HWD based on sequence features, providing a practical tool that can be integrated into existing quality control pipelines for whole genome sequencing studies.

Journal Article

A-liner: linear alignment visualizer for genome comparisons.

SUMMARY: A-liner is a flexible command-line tool for linear visualization of genome-scale sequence alignments, supporting outputs from multiple aligners and integrated visualization of annotations, highlights, quantitative tracks, and coordinate scales. It is applicable to a wide range of organisms, from bacteria to large eukaryotic genomes, and facilitates efficient generation of publication-ready comparative genome visualizations. AVAILABILITY AND IMPLEMENTATION: The source code and example output files for a-liner are available in the GitHub repository: https://github.com/mokuno3430/a-liner. A-liner v1.1.0 has been archived on Zenodo at https://doi.org/10.5281/zenodo.19702001.

Software

Genomic erosion in the assessment of species' extinction risk and recovery potential.

Many species are undergoing rapid population declines and environmental deterioration, leading to genomic erosion. Here we define genomic erosion as the loss of genetic diversity, accumulation of deleterious mutations, maladaptation, and introgression, all of which can undermine individual fitness and long-term population viability. Critically, this process continues even after demographic recovery due to a time-lagged impact of genetic drift, which is known as drift debt. Current conservation assessments, such as the International Union for Conservation of Nature Red List, focus on short-term extinction risk and do not capture the long-term consequences of genomic erosion. Likewise, the longer-term assessments of the International Union for Conservation of Nature Green Status may overestimate population recovery by failing to account for the enduring effects of genomic erosion. As genome sequencing becomes increasingly accessible, there is a growing opportunity to quantify genomic erosion and integrate it into conservation planning. Here, we use genomic simulations to illustrate how different genomic metrics are sensitive to the drift debt. We test how ancestral effective population size (Ne) and bottleneck history influence the tempo and severity of genomic erosion. Furthermore, we demonstrate how these dynamics shape genetic load and additive genetic variation, which are key indicators of long-term evolutionary potential. Finally, we present a proof-of-concept for a Genomic Green Status framework that aligns genomic metrics with conservation impact assessments, laying the foundation for genomics-informed strategies to support species recovery.

Extinction, Biological

An alignment-free strategy for circulating tumor DNA detection and tumor fraction estimation from whole-genome sequencing data.

Circulating tumor DNA (ctDNA) is emerging as a promising biomarker for postoperative monitoring of cancer patients. Precise estimation of circulating tumor fraction is crucial for evaluating treatment effects and timely detection of disease recurrence. All current ctDNA detection methods that utilize whole-genome sequencing (WGS) data rely on the reference genome alignment of sequencing reads and often apply separate tools for detecting different variant types. However, various bioinformatic analysis confounders and the application of external variant calling tools could be avoided by analyzing k-mers from unaligned sequencing reads. While k-mer-based methods have successfully been applied for somatic variant validation and detection, the potential of k-mer-based ctDNA detection is unexplored. We have developed a tumor-informed alignment-free ctDNA detection tool called ctDNAmer that detects tumor-specific somatic variation directly from unaligned sequencing data by identifying k-mers unique to the tumor DNA. ctDNAmer detects variant information across the genome by comparing the primary tumor and germline WGS data and accounts for sample-specific germline variability and technical noise in the same framework. We tested the utility of ctDNAmer for tumor fraction estimation on postoperative plasma cfDNA WGS data (mean sequencing depth&#x2009;~&#x2009;28x) from 90 stage III colorectal cancer patients with three years of follow-up. The tumor fraction (TF) estimates agreed with the available clinical information and ctDNA was detected in 77% (17/22) of recurring patients with a median lead time of 8 months compared to radiological imaging. We further validated ctDNAmer's tumor fraction estimates based on a comparison with the mean cfDNA allele frequencies of somatic clonal SNVs identified from aligned primary tumor sequencing data. The TF estimates showed a strong Pearson correlation of 0.897 with the mean allele frequencies and improved ctDNA detection results across samples with an AUC of 0.79 compared to 0.75 if the mean allele frequency of clonal mutations is used.

Circulating Tumor DNA

Lift&Add-rapid and robust addition of new species to alignments of conserved non-coding sequences.

MOTIVATION: Identifying sequence constraint across long evolutionary distances is a powerful method for the discovery of functional genomic sequences, especially putative non-coding elements. Conserved elements have been a mainstay of comparative genomic research, and can be further investigated for species-specific sequence acceleration to dissect the genetic basis of trait evolution. The conclusions of these comparative genomic studies are contingent on the number and range of species included in this phylogenetic analysis. However, while the number of metazoan genomes sequences is increasing rapidly, adding new genomes to existing whole-genome alignments remains computationally expensive. RESULTS: Here, we present a bioinformatic workflow, Lift&Add, that enables conserved elements, coding or non-coding, to be rapidly mapped to new genomes ("Lift") and subsequently be added to pre-existing multiple species alignments ("Add"), thus providing an avenue for easy exploration of these putative functional elements. Focusing here on a group of species that has been largely under-represented in genomic comparisons, the marsupials, we demonstrate the intuition behind this workflow and provide an example comparative genomic analysis that can be performed. IMPLEMENTATION AND AVAILABILITY: Lift&Add is implemented as a series of scripts in Snakemake and bash, which can be downloaded from https://github.com/navyashukladr/Lift_and_Add.

Conserved Sequence

Analysis of intrastrain recombination in herpes simplex virus type 1 strain 17 and herpes simplex virus type 2 strain HG52 using restriction endonuclease sites as unselected markers and temperature-sensitive lesions as selected markers.

The viral and host factors involved in herpes simplex virus (HSV) recombination are little understood. To identify features of the process, recombination in HSV-1 and HSV-2 has been studied by analysing the segregation of unselected markers in the form of restriction endonuclease (RE) sites. By confining parental interactions to only one strain of virus of each serotype, restrictions imposed by non-homology are overcome and differential growth phenotypes can be discounted. The analysis of unselected and selected recombinants using RE sites in conjunction with temperature-sensitive mutations is consistent with (i) HSV being highly recombinogenic, (ii) parental and progeny molecules taking part in the process, (iii) the four genomic isomers participating in recombination, (iv) genome alignment being part of the recombination process and (v) cellular factors in conjunction with genome homology influencing the efficiency of recombination.

Animals

AniAnn's: alignment-free annotation of tandem repeat arrays using fast average nucleotide identity estimates.

MOTIVATION: Satellite DNA has long posed challenges for genome assembly and analysis due to its low sequence complexity and poor mappability. These large heterochromatic arrays of tandem repeats are ubiquitous across eukaryotic genomes, yet remain understudied. Current methods for annotating satellite regions, and other classes of tandem repeat arrays, are limited in their ability to annotate divergent or novel sequences. RESULTS: In this work, we introduce AniAnn's, an algorithm for annotating large blocks of tandemly repeating DNAs. AniAnn's exploits the high Average Nucleotide Identity (ANI) shared between repeat units of the same array to quickly and accurately infer the boundaries of such arrays. We show that AniAnn's improves the annotation of satellites and other tandem repeats within a variety of plant and animal genomes, while requiring only a fraction of the runtime compared to previous approaches. We conclude by exploring several use cases of AniAnn's as a lightweight method for masking repeats prior to whole-genome alignment as well as the de novo annotation and classification of satellite repeats. AVAILABILITY: AniAnn's is open source software and available at github.com/marbl/anianns.

Algorithms

Improving quality control of microbial agri-inputs by confirming strain identity with an easy and low-cost PCR-multiplex: A study case with Azospirillum brasilense.

The first commercial product containing the Azospirillum brasilense elite strains Ab-V5 and Ab-V6 was launched in Brazil in 2009. These strains have demonstrated agronomic efficiency in grasses and in legume co-inoculation, accounting for approximately 43 million doses in 2024. Official identification of these strains is currently performed by rep-PCR, a reliable but time-consuming and laborious method. In this study, a multiplex PCR assay was developed for the simultaneous identification of Ab-V5 and Ab-V6 in a single reaction using strain-specific SNPs. Forward primers were designed so that the terminal nucleotide at the 3' end corresponded to a strain-specific SNP unique to each target strain. To further enhance specificity, artificial mismatches were introduced at the fourth nucleotide from the 3' end of the forward primers. SNPs were identified using Snippy based on genomic alignments between Ab-V5 and Ab-V6 and confirmed by local BLASTn against the genomes of other Azospirillum species. In the multiplex assay, simultaneous and specific amplification of both strains was observed in a single reaction, without non-specific amplification. Primer specificity was also experimentally evaluated against other A. brasilense strains (Ab-V1, Ab-V2, Ab-V4, Ab-V7, Ab-V8, and Sp7T), in silico against bacteria from different genera associated with agricultural inoculants, and in commercial inoculant samples containing Ab-V5 and Ab-V6. The results confirmed the high specificity of the primers for Ab-V5 and Ab-V6 and demonstrated that the assay was capable of identifying the strains in commercial inoculants. This assay facilitates inoculant quality control by enabling strain confirmation using a simple, rapid, and low-cost method.

Azospirillum

Gene-level complexity explains genome-wide variation in the distribution of fitness effects.

The distribution of fitness effects (DFE)-describing how harmful, neutral, or beneficial new mutations are-is central to understanding how populations evolve. Although the DFE varies across genomes and species, it remains unclear which aspects of genomic organization drive this variation. Here, we inferred gene-level selective constraints across the genomes of Mus musculus castaneus, Drosophila melanogaster and Saccharomyces cerevisiae using a combination of population genetics and machine learning trained on diverse gene features. Many gene features were predictive of selective constraint, with conservation, gene structure, and expression being the most informative. These selective constraints delineated gene classes with distinct DFEs. Genes with higher connectivity and expression-features reflecting how many traits a gene influences-experienced stronger and less dispersed deleterious effects with increasing selective constraint. Between species, the rate of adaptation decreased with increasing organismal complexity, whereas across the genome it did not decrease monotonically with selective constraint, but tended to be higher at intermediate levels. While between-species comparisons of DFE parameters were less consistent with predictions of Fisher's geometric model (FGM) based on organismal complexity, variation in DFE parameters across the genome aligned more closely with FGM when complexity was considered at the gene level. Our results suggest that gene-level complexity, captured by genomic feature proxies, provides a more informative definition of complexity for DFE variation than organism-level labels, and highlight the value of using gene features collectively to link genomic architecture, fitness landscapes, and patterns of molecular evolution.

Animals

Alignment-free integration of single-nucleus ATAC-seq across species with sPYce.

Changes in gene regulation largely contribute to differences in cellular identities and phenotypes between species. Single-nucleus assays for transposase-accessible chromatin with sequencing (snATAC-seq) are an efficient strategy to identify putative gene regulatory elements and provide new insight into evolutionary divergence of regulatory programmes. However, no dedicated framework exists to integrate and compare snATAC-seq data across species, while methods designed for single-cell gene expression data have serious limitations. Here we present sPYce, a cross-species snATAC-seq integration method that relies on sequence composition similarities through k-mer histograms of regulatory regions, removing the need for genome alignments to anchor data from different species. sPYce can embed datasets from multiple species into the same mathematical space and permits further downstream analysis steps. We benchmarked sPYce against existing approaches on two publicly available datasets spanning more than 160&#x2009;myr of evolution, showing that it successfully uncovers conserved cellular programmes while preserving biologically relevant species-specific differences. By comparing cerebellar development in mice and opossums, sPYce identifies regulatory divergence in granule cell differentiation programmes, particularly driven by nuclear factor 1. As an easy-to-use, alignment-free cross-species snATAC-seq integration approach, sPYce opens new perspectives to compare gene regulatory evolution across species.

Animals

(Re)imagining the Future of Genetic Counseling: A Reflexive Qualitative Analysis of Sociopolitical Power, Cultural Safety, Systemic Racism, and Comparative Practice in the United Kingdom, Aotearoa New Zealand and, Australia.

Genetic counseling is undergoing a rapid transformation as genomic medicine becomes embedded within mainstream healthcare systems. At the same time, the profession is being challenged to respond to systemic racism, colonial legacies, technological change, and evolving expectations regarding equity and justice. Historically, genetic counseling emerged within twentieth-century medical genetics and was influenced by political, social, scientific, and medical forces that included eugenic ideology, values, and practices. The profession has since evolved substantially toward psychosocial, patient-centered, and non-directive models of care. Contemporary debates regarding "newgenics" or "neugenics" further demonstrate how concerns regarding equity, reproductive ethics, disability, and genomic stratification continue to shape genomic healthcare discourse. This qualitative reflexive practice paper explores how systemic racism, colonial legacy, cultural safety and structural power shape genetic counseling practice in the United Kingdom (UK), Aotearoa New Zealand and Australia, and how these forces continue to reshape the profession's future identity. A reflexive, narrative, and comparative qualitative approach was employed, grounded in the authors' lived professional experiences across UK and Australasian contexts and informed by purposively selected policy, professional and scholarly literature relating to cultural safety, dignity, anti-racism, and Human Rights-Based Decision-Making. Through iterative reflexive dialogue, comparative analysis, and thematic synthesis, four interrelated themes were developed examining sociopolitical context, systemic racism, cultural safety and technologization within contemporary genetic counseling practice. Comparative analysis identified substantial differences in how culturally responsive practice is conceptualized and operationalized across settings. In Aotearoa, cultural safety is strongly shaped by Te Tiriti o Waitangi, bicultural accountability, and M&#x101;ori sovereignty frameworks. In Australia, culturally safer genomic care has increasingly developed through Indigenous-led initiatives and workforce reform, including the Australian Alliance for Indigenous Genomics (ALIGN). In contrast, UK practice remains largely situated within equality, diversity, and inclusion (EDI) frameworks that may insufficiently address systemic racism and structural power within increasingly diverse populations. Reflexive clinical examples demonstrated how inequities may emerge through undocumented patient values, standardized pathways, assumptions regarding autonomy, and misinterpretation of culturally specific communication styles. Re-imagining the future of genetic counseling requires more than just technological advancement. It requires reflexive engagement with dignity, inequity, and the sociopolitical realities of the populations served. These insights re-imagine a culturally grounded, socially responsive future for genetic counseling in an era shaped by genomic mainstreaming, digital transformation, artificial intelligence and workforce reform and one in which the profession remains ethically anchored, relationally attuned, and committed to justice-oriented practice.

Humans

CholeraSeq: a comprehensive genomic pipeline for cholera surveillance and near real-time outbreak investigation.

SUMMARY: Next Generation Sequencing is widely deployed in cholera-endemic regions, yet an end-to-end reproducible pipeline that unifies read QC, filtering, reference mapping, variant calling/annotation, recombination screening, and extraction of parsimony informative sites/variant codons, phylogenetic inference for downstream phylodynamic and epidemiological analyses have been lacking, slowing outbreak investigation and public health response. CholeraSeq is a high-throughput genomics pipeline for cholera genomic surveillance. It ingests consensus genomes, short read sequence data, draft assemblies, and scales seamlessly from local to cloud environments. To accelerate epidemiological context placement of new outbreak strains, we provide a curated ready-to-use core genome alignment compiled from public data, enabling flexible, fast, integration of new samples for outbreak investigations. AVAILABILITY AND IMPLEMENTATION: CholeraSeq is freely available on the GitHub platform https://github.com/CERI-KRISP/CholeraSeq. CholeraSeq is implemented in Nextflow with a modular design building upon the nf-core community standards.

Cholera

Genomic epidemiology of enteropathogenic Escherichia coli in southwestern Nigeria.

BACKGROUND: Enteropathogenic Escherichia coli (EPEC) are etiological agents of diarrhea. We studied the genetic diversity and virulence factors of EPEC in southwestern Nigeria, where this pathotype is rarely characterized. METHODOLOGY/PRINCIPAL FINDINGS: EPEC isolates (n&#x2009;=&#x2009;96) recovered from recent southwestern Nigeria diarrhea case-control studies were whole genome-sequenced using Illumina technology. Genomes were assembled using SPAdes and quality was evaluated using QUAST. Virulencefinder, Ectyper, and ResFinder were used to identify virulence genes, serotypes, and resistance genes. Multilocus sequence typing was done by STtyping. Single nucleotide polymorphisms (SNPs) were called out of whole genome alignment using SNP-sites and a phylogenetic tree was constructed using IQtree. Thirty-nine of the 96(40.6%) EPEC isolates were from diarrhea cases diarrhea. Nine isolates from diarrhea patients and four from healthy controls were typical EPEC, harboring bundle-forming pilus (bfp) genes whilst the rest were atypical EPEC. There were 15 EPEC-EAEC hybrids. Atypical serotypes O71:H19 (16, 16.6%), O108:H21 (6, 6.3%), O157:H39 (5, 5.2%), and O165:H9 (4, 4.2%) were the most prevalent; only 8 (8.3%) isolates belonged to classical EPEC serovars. The largest, ST517 clade harbored multiple siderophore and serine protease autotransporter genes and included an O71:H19 subclade <10 SNPs apart, representing a likely outbreak involving 15 children, four with diarrhea. Likely outbreaks, of typical O119:H6(ST28) and atypical O127:H29(ST7798) were additionally identified. CONCLUSION/SIGNIFICANCE: EPEC circulating in southwestern Nigeria are diverse and differ substantially from well-characterized lineages seen previously elsewhere. EPEC carriage and outbreaks could be commonplace but are largely undetected, hence, unreported, and require genomic surveillance for identification.

Nigeria

Cloning and nucleotide sequence of the ispA gene responsible for farnesyl diphosphate synthase activity in Escherichia coli.

The molecular cloning and the determination of the nucleotide sequence of the ispA gene responsible for farnesyl diphosphate (FPP) synthase [EC 2.5.1.1] activity in Escherichia coli are described. E. coli ispA strains have temperature-sensitive FPP synthase, and the defective gene is located at about min 10 on the chromosome. The wild-type ispA gene was subcloned from a lambda phage clone containing the chromosomal fragment around min 10, picked up from the aligned genomic library of Kohara et al. [Kohara, Y., Akiyama, K., & Isono, K. (1987) Cell 50, 495-508]. The cloned gene was identified as the ispA gene by the recovery and amplification of FPP synthase activity in an ispA strain. A 1,452-nucleotide sequence of the cloned fragment was determined. This sequence specifies two open reading frames, ORF-1 and ORF-2, encoding proteins with the expected molecular weights of 8,951 and 32,158, respectively. A part of the deduced amino acid sequence of ORF-2 showed similarity to the sequences of eucaryotic FPP synthases and of crtE product of a photosynthetic bacterium. The plasmid carrying ORF-2 downstream of the lac promoter complemented the defect of FPP synthase activity of the ispA mutant, showing that the product encoded by ORF-2 is the ispA product. The maxicell analysis indicated that a protein of molecular weight 36,000, approximately consistent with the molecular weight of the deduced ORF-2-encoded protein, is the gene product.

Alkyl and Aryl Transferases

MKMC enables reference-free transcriptomic analysis using k-mer representations.

Traditional RNA-seq analysis depends heavily on genome alignment and gene annotation, limiting its utility in non-model organisms and introducing biases that can obscure regulatory complexity. We present MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer-based statistics to detect biological variation without requiring alignment. MKMC integrates fast k-mer counting, abundance matrix generation, normalization, dimensionality reduction, and differential analysis into a unified workflow. Across diverse datasets, MKMC recapitulates key biological signals-including sex differences in killifish liver-and matches alignment-based pipelines in differential expression analysis and transcriptomic age prediction. Notably, MKMC detects isoform-specific events missed by traditional methods, one of which we validated using in situ hybridization. These results reveal previously hidden isoform-level regulatory events that contribute to sex- and age-associated transcriptional programs. MKMC offers a robust, extensible alternative to alignment-based approaches, enabling transcriptomic discovery across both model and non-model systems. While we focus here on RNA-seq as a primary application, MKMC is broadly applicable to any k-mer-based analysis of next-generation sequencing data.

MKMC